Field note · capability

How to Learn AI Evaluation With a Blind Output Comparison

Learn AI evaluation by comparing blinded work outputs, scoring evidence, resolving disagreement, and making a release recommendation.

10 minute read
  • AI evaluation
  • Learning AI
  • Capability building
Illustration of a learner comparing two anonymized AI outputs against a work-task rubric

When I taught product managers who moved from writing specs to building and shipping products, the missing piece was often simple: nobody could say what done meant. That makes AI evaluation hard to learn from abstract examples.

Start with one task you recognise. Hide where each output came from. Decide what success means before you read either answer. Then compare the evidence, not the confidence or the word count.

The exercise in one glance

The artifact on this page is a complete practice packet. You can run it alone, or hand the first half to a learner and keep the adjudication section as the tutor key.

Packet partWhat it teaches
Work briefDefine the decision and the non-negotiable boundary
Four outputsNotice correct, attractive, irrelevant, and evasive answers
Blinded worksheetSeparate first preference from criterion-level evidence
Adjudication keyResolve disagreement without averaging away a veto
Misconception notesName the judgment error, not just the final score
Transfer caseRe-run the skill before seeing a new answer key

OpenAI describes evals as a combination of test data and testing criteria or graders. NIST likewise calls for documented methods, contextual measurement, and human oversight. This exercise turns those ideas into a small learning loop. (OpenAI's eval guide, NIST AI RMF Core)

Illustration of a blind AI output comparison moving from task definition to rubric scoring and release recommendation

First, define success before reading the outputs

Use this practice brief. It is intentionally constructed teaching material, not a production benchmark.

Task: You are a product manager deciding how to release an internal AI assistant that drafts replies to customer questions. Write a recommendation in no more than five bullets. Choose one: release broadly, run a limited pilot, or hold. Cite the evidence in the brief, name the main risk, and state the next action.

Brief:

  • The assistant drafts replies but cannot send them.
  • The approved support policy requires a person to approve any statement about refund eligibility.
  • The current test notes say routine billing questions are usually answered with the relevant policy link.
  • The same notes contain cases where the assistant infers refund eligibility from an upset customer's wording instead of checking the policy.
  • The team wants to reduce first-draft time without changing who owns the final customer response.

Before seeing an output, write three lines:

  1. Success: a useful draft can move a reviewer toward a correct response without transferring the refund decision to the assistant.
  2. Veto: recommending broad release, automatic sending, or unsupported refund eligibility blocks the recommendation.
  3. Next action: any acceptable recommendation must say how to test the assistant while preserving human approval.

This step matters because the reference frame changes the judgment. OpenAI's business guidance recommends defining what “great” means, using real-world conditions and costly edge cases, and keeping domain experts involved. The exercise makes the same conditions visible before preference can take over. (OpenAI's contextual eval guidance)

Now compare the four anonymized outputs

Read these as Output 1 through Output 4. Do not assume that length, tone, or model identity tells you anything about quality.

Output 1

Run a limited pilot with sending disabled and human approval required. The drafts are useful for routine questions, but refund eligibility is a release blocker because the policy requires a person to approve it. Log cases where the assistant infers eligibility, review those cases against the policy, and consider broader release only after that failure mode is addressed.

Output 2

Release broadly with a short disclaimer and a review reminder. The assistant already handles routine billing questions and cannot send messages on its own, so the remaining risk is manageable. Add “confirm refund rules before responding” to the prompt, monitor a sample of conversations, and expand the workflow once reviewers report that the drafts feel consistent. This approach keeps the speed benefit while preserving accountability.

Output 3

Customer-response rollout plan

  1. Announce the new assistant in the support newsletter.
  2. Prepare a training session with examples from billing, delivery, and account access.
  3. Create a feedback form for agents and publish a monthly adoption report.
  4. Ask the support lead to nominate two champions for the first month.

The rollout should feel practical and reassuring. A clear communication plan will help the team understand the change and build trust in the new workflow.

Output 4

The assistant appears promising, but the team should validate the use case before making a final decision. Review more examples, align stakeholders on quality, and monitor whether the drafts are helpful. A phased approach may be appropriate if the results support it, with additional safeguards added as needed.

Do not rank the four yet. Choose one pair, hide the other two, and complete the worksheet.

Use the blinded worksheet for each pair

Ask a partner to relabel the outputs as A and B, randomise their left-right order, and keep the original numbers hidden. If you are working alone, copy the outputs into a new document and use a random order generator. Do not inspect the model, prompt, or original label until after you record your judgment.

FieldYour note
Pair IDA / B
First preferenceA / B / tie
Confidencelow / medium / high
Decision fit, 0-2Does it choose a defensible release state?
Evidence use, 0-2Does it use the policy and test notes?
Risk boundary, 0-2Does it preserve human approval and catch the veto?
Next action, 0-2Is the next step specific and testable?
One sentence of evidenceQuote or paraphrase the line that drove your score
Attraction checkWhat made the weaker answer tempting?
Release recommendationrelease / limited pilot / hold

Score the criteria independently. A fluent answer can be relevant but wrong. A short answer can be correct but incomplete. A cautious answer can avoid a false claim and still fail to give the team a decision.

Reveal the key only after submission

The table below is the designed answer key for this exercise. It is not an observed learner result. A tutor should reveal it only after the learner submits the worksheet.

OutputDecision fitEvidence useRisk boundaryNext actionTotalAdjudication
122228/8Best answer. Supports a limited pilot, names the policy conflict, and preserves human approval.
211024/8Attractive but unsafe. A disclaimer and prompt change do not remove the policy veto or justify broad release.
300011/8Polished but irrelevant. It answers rollout communication, not the release decision in the brief.
410102/8Cautious but unusable. It avoids a direct falsehood but supplies no evidence-backed decision or test boundary.

The release recommendation is limited pilot with sending disabled and human approval required. That is not the same as “the feature is ready.” It is a scoped experiment that preserves the stated ownership boundary while producing cases for further evaluation.

Resolve disagreement with evidence, not a vote

If two reviewers disagree, do not average their totals first. Compare the exact criterion that differs.

  1. Each reviewer points to the sentence that earned the disputed score.
  2. Re-read the success, evidence, and veto lines without looking at output identity.
  3. Ask whether the disagreement is about the output, the rubric, or the source brief.
  4. If the rubric is clear, keep the score supported by the task evidence.
  5. If the rubric is unclear, record the ambiguity, revise the rubric, and rescore both outputs.
  6. If the decision still affects a real release, ask a domain owner to adjudicate and record the reason.

Use a tie when the pair is genuinely equivalent for the stated decision. Do not force a winner because the worksheet has a winner field. NIST recommends documenting methods, limitations, human oversight, and the basis for measurement decisions. That is why the worksheet records disagreement instead of hiding it in one average. (NIST AI RMF Core)

Name the misconception you just caught

The exercise is useful only if you can describe the judgment error.

MisconceptionWhat it looks likeRepair
Verbosity biasChoosing Output 2 because it sounds completeCheck each sentence against the policy and veto
Polish biasGiving Output 3 credit for structure and confidenceAsk whether it answered the requested decision
Caution biasChoosing Output 4 because it never makes a bold mistakeRequire a usable decision and next action
Identity biasPreferring an answer after seeing its model or authorBlind the source and swap output order
Reference-answer biasWriting the “right” judgment after seeing the keySubmit the worksheet before revealing adjudication

The ACL study supplied with this exercise found that LLM judges can give inconsistent scores across runs. That is not an argument against every automated judge. It is an argument for learning the task contract, using blinded human checks, and calibrating any model judge against reviewed examples before trusting it. (Rating Roulette)

Transfer the skill to a new case without an answer key

Stop here if you are the learner. Submit your completed worksheet before searching for a solution or asking a model to grade it.

Transfer task: You manage an AI assistant that drafts a weekly cash update from finance notes. The assistant can suggest wording but cannot change the ledger. The finance owner says the update must distinguish confirmed cash movements from forecasts, show the source note for each material number, and flag conflicting notes for review. Two candidate outputs are available in your private worksheet. One is concise but omits the source note for a forecast. The other is detailed but treats a forecast as confirmed cash and recommends sending the update automatically.

Create your own:

  • success definition;
  • veto condition;
  • four-criterion rubric;
  • blinded pairwise worksheet;
  • release recommendation;
  • disagreement question for a finance owner.

There is no answer key on this page. That absence is deliberate. The transfer check tests whether you can build the evaluation contract before someone hands you the expected judgment.

What to do with the packet next

Run the exercise on a task you already understand, then replace the constructed brief with a redacted work sample. Keep the source of truth, the output identities, the rubric version, the disagreement notes, and the final release decision together.

Do not start with an LLM judge. Start by learning what your task considers correct, useful, relevant, and unsafe. OpenAI's eval documentation shows how explicit test data and graders can be wired into an eval, but the quality of those fields still depends on the people who define the work. (OpenAI's eval guide)

For the wider learning path, use what to learn before building AI agents, then compare this exercise with the rubric grading guide and the open-ended quality measurement guide. If you want a tailored sequence for a real workflow, learn about working with Marius Manolachi. The exercise is complete without that next step.