Field note · implementation

Why AI Fails When Reviewers Optimize Different Outcomes

A 10-output review test shows how accuracy, risk, and usefulness create different release advice, and how one shared contract resolves it.

10 minute read
  • AI evaluation
  • AI implementation
Illustration of three reviewers scoring one AI output against different release objectives

When I taught product managers to move from writing specs to building and shipping, the recurring failure was often not the model. It was that nobody could say what done meant. Marius Manolachi's teaching work makes that distinction concrete.

The same thing happens one step later. A team shows one AI output to three reviewers, each reviewer answers a different question, and the meeting treats the disagreement as if it were one score that needs averaging.

Illustration of a review packet moving through independent rubrics and a shared outcome contract

The result: one output was useful, but not releasable

In a bounded review of 10 sanitized outputs from one low-risk meeting-note workflow, output O8 passed a permissive usefulness-led rule but failed the shared outcome contract. It claimed that “final decisions are recorded in the notes,” although the source brief said the notes were informational and had no approval authority.

The contract changed the decision from release to hold and rewrite. The corrected O8R passed all three role gates.

Reviewer lensLocal questionO8 scoreGate
AccuracyDid the output preserve the source facts?4/6Hold below 5
Risk and complianceDid it preserve authority and escalation boundaries?4/6Hold below 5, plus veto
Usefulness and speedCan an attendee act after one quick read?5/6Pass at 5

This is the sourceable result from the packet. It is a small reproduction, not evidence about how common this failure is.

Why does the same AI output receive incompatible review advice?

Because reviewers are often optimizing different loss functions. An accuracy reviewer loses when a source fact is omitted or changed. A risk reviewer loses when the output creates an obligation or authority. A usefulness reviewer loses when a cautious answer is too slow or vague to use.

Those are not three noisy measurements of one hidden quality score. They are three objectives with partial overlap.

Research on annotator disagreement reaches a similar conclusion. Zhang and colleagues describe disagreement from factors such as task underspecification and response style, and warn that standard reward modeling and LLM-as-judge methods can miss divergence between annotators. Their paper is evidence for treating disagreement as a diagnostic signal, not a reason to erase it immediately.

The problem is not only theoretical. The Cohere research repository for “Human Feedback is not Gold Standard” says preference scores can under-represent factuality and that assertive wording can skew perceived factuality errors. The public code and data repository therefore separates granular annotations from one aggregate preference score.

What did the 10-output packet contain?

The fixed source brief described a weekly sync:

  • Tuesday at 10:00 in Room B
  • agenda: backlog, blockers, and demos
  • agenda additions due Monday at 16:00
  • notes are informational and have no approval authority
  • unresolved items go to the team lead

The packet contained these anonymized outputs. The totals below are raw role scores from the independent pass. A, R, and U mean accuracy, risk, and usefulness.

IDOutputARUInitial reading
O1“Weekly sync: Tue 10:00, Room B. Bring backlog, blockers, and demos. Add agenda items by Mon 16:00. Notes are informational; unresolved items go to the team lead.”665Release
O2“Weekly sync Tue 10:00 Room B. Send agenda items by Monday afternoon.”456Release under usefulness-led review
O3“Weekly sync Tue 10:00 Room B. Team leads approve agenda items added before Monday 16:00.”315Usefulness pass, risk veto
O4“Tuesday's 10:00 sync in Room B decides backlog, blockers, and demos. Add requests by Monday 16:00.”316Usefulness pass, risk veto
O5“The weekly sync is Tuesday at 10:00 in Room B. No action is required; notes are informational.”453Safe but not useful
O6“Weekly sync: Tue 10 in Room B. Agenda: backlog, blockers, demos. Send additions Monday 16:00. Escalate unresolved items to the team lead.”556Release
O7“Join Tuesday's weekly sync at 10:00 in Room B. The meeting covers backlog, blockers, and demos. Add items by Mon 16:00; unresolved items go to team lead. The recap records discussion and does not approve changes.”665Release
O8“Weekly sync Tuesday 10:00 Room B. Add agenda items by Monday 16:00. Final decisions are recorded in the notes.”445Would ship under permissive rule, held by contract
O9“Weekly sync is Tuesday at 10:00 in Room B. Add agenda items by Monday 16:00. If you miss the deadline, bring them to the meeting.”446Useful, but adds an unsupported route
O10“Tuesday, 10:00, Room B. Agenda items Monday 16:00. Questions to team lead.”456Fast, but incomplete

The raw component scores and independent comments are part of the packet, not hidden adjudication. Totals alone hide why reviewers diverged.

IDAccuracy commentRisk commentUsefulness comment
O1All five facts are present.No invented authority; escalation is preserved.Scannable, with one extra sentence.
O2Exact deadline and unresolved-item route are omitted.No unsafe claim, but escalation is absent.Short enough to use.
O3Adds approval authority not in the brief.Invented authority is a veto.Looks decisive and is easy to scan.
O4Turns an informational sync into a decision forum.Authority boundary breached.Best skim and clearest action.
O5Agenda and escalation are missing.Safe wording, but the route is absent.Does not tell the attendee what to bring or do.
O6The informational-notes fact is omitted.No dangerous claim is added.Clear action list.
O7All facts have safe equivalents.Decision boundary and escalation are explicit.Complete and still easy to scan.
O8“Final decisions” conflicts with the brief.Authority is ambiguous enough to veto.Concise and apparently useful.
O9The missed-deadline instruction is unsupported.It adds a new route, even if harmless.Gives a useful recovery action.
O10“Questions” weakens the specific unresolved-item rule.No new authority, but the boundary is incomplete.Very fast to read.

How should you score the same outputs independently?

Separate the local objectives before anyone sees another review. The three rubrics in this test used a 0 to 2 score for each component.

Accuracy reviewer

Score correctness, completeness, and traceability. The accuracy gate is 5/6, with no zero. This reviewer asks whether the output preserves the source, not whether a user would enjoy reading it.

Risk and compliance reviewer

Score unsupported authority, boundary preservation, and escalation correctness. The risk gate is 5/6, with no zero. An invented authority, changed approval boundary, new obligation, or wrong escalation route is a veto regardless of the total.

Usefulness and speed reviewer

Score scanability, actionability, and brevity. The usefulness gate is 5/6. This reviewer asks whether an attendee can act after one quick read, not whether every source fact appears in full.

Do not let the usefulness reviewer silently become the product owner. Do not let the risk reviewer turn every omission into a production stop. Their objectives are useful because they remain distinct.

What does the disagreement matrix tell you?

Under each role's strict local threshold, the packet produced this disagreement matrix:

Reviewer pairDisagreeing outputsCount
Accuracy vs riskO2, O5, O103
Accuracy vs usefulnessO2, O3, O4, O5, O8, O9, O107
Risk vs usefulnessO3, O4, O8, O94

The pattern is more useful than a single agreement percentage. Accuracy and usefulness diverged on seven outputs because concise wording hid omissions or introduced claims. Risk and usefulness diverged on four because decisive language looked helpful while changing the authority boundary.

Classify each disagreement before changing the model:

ClassDiagnostic questionPacket examples
Objective conflictAre both reviewers applying valid but different goals?O2, O4, O8
Context omissionDid the output omit the source fact needed to judge it?O3, O5, O6, O10
Rubric ambiguityDo the rubrics disagree about whether an addition is allowed?O9
Reviewer errorDid someone misread the fixed source or rubric?None identified in this run

Kumar, Ghosal, and Ekbal made reviewer contradiction itself an evaluation task. Their ContraSciView work publishes a dataset of roughly 8.5k papers, 28k review pairs, and nearly 50k review comments. The ACL record supports the narrower point that contradiction is worth identifying and preserving, not automatically collapsing.

What shared outcome contract resolves the conflict?

Write the contract before the next review meeting. It should turn separate scores into one release decision without pretending they measure the same thing.

Release only when:
1. Accuracy >= 5/6 and no accuracy component is 0.
2. Risk >= 5/6 and no risk component is 0.
3. Usefulness >= 5/6.

Veto when the output invents authority, changes an approval boundary,
creates a new obligation, or sends an unresolved item to the wrong route.

Veto status: hold-rewrite, even if usefulness passes.
Escalate when the source brief is incomplete or a reviewer cannot decide
whether a sentence changes policy.

This is not a universal rubric. It is a worked contract for a low-risk meeting-note workflow. Replace its thresholds and vetoes with the consequences of your own workflow.

NIST's current TEVV-Athlon draft makes the same design move at a broader level: evaluation methods vary by use case, and assessments should be built from organizational objectives. NIST's framework page describes a four-stage approach that produces measurements tied to those objectives.

The UK Centre for Data Ethics and Innovation's assurance roadmap is even more direct about the review boundary. It says assurance needs suitable criteria and sufficient evidence, and that performance testing alone cannot show whether the requirements are flawed or unsuitable for the use case. The roadmap also separates roles and responsibilities. That is why this contract has both role thresholds and an escalation route.

How do you know the contract added decision value?

Use a before-and-after decision record. In this packet, the decisive case was O8.

Before the contract, a permissive rule said: release when usefulness is at least 5/6 and no reviewer total is below 4/6. O8 scored A4, R4, and U5, so it would ship.

The contract stopped it. “Final decisions are recorded in the notes” invented authority that the source brief explicitly denied. The status became hold-rewrite.

The targeted rewrite was:

Weekly sync Tuesday 10:00 in Room B. Add agenda items by Monday 16:00. Notes record discussion only; unresolved items go to the team lead.

O8R scored A6, R6, U5. It passed. The contract did not make the output more agreeable. It changed the release decision, forced one repair, and left the authority boundary visible.

When should you reject this test?

Reject the test if the reviewers produce no meaningful conflict, if all three roles already use the same objective, or if the contract changes no decision. A matrix that only confirms agreement is not evidence for this failure mode.

Also reject any claim that this packet estimates prevalence. Ten authored outputs can show a mechanism and test a decision rule. They cannot tell you how often different teams disagree, which model causes more disagreement, or whether one rubric generalizes to another workflow.

For a real implementation, repeat the exercise with sanitized production traces, a versioned source brief, independent human reviewers, and a domain owner who can resolve policy context. Keep the original output, raw comments, rubric versions, contract version, and final decision together.

If your team wants help turning this test into an internal review practice, see Marius Manolachi's AI tutoring and consulting work. The goal is a team that can run and improve the workflow itself.

If your team needs the broader implementation sequence, start with the AI workflow implementation guide. If the review problem is still unclear, compare this packet with an AI evaluation plan for disagreeing users. The next step is not to average harder. It is to state what each reviewer is protecting.