Field note · evaluation

What Evidence Should Domain Experts Collect About Operator Agreement?

A versioned operator-agreement packet preserves reviewer evidence, disagreement, adjudication, AI reruns, and bounded release decisions.

14 minute read
  • AI evaluation
  • human evaluation
  • operator agreement
  • AI reliability
Illustration of a frozen operator-agreement fixture flowing through independent review and adjudication

Illustration of a frozen operator-agreement fixture flowing through independent review and adjudication

An operator label can look authoritative because it came from someone who does the work. That does not tell you whether another qualified reviewer would make the same call, whether the call is correct, or whether an AI system is safe to use in the same workflow.

The useful question is not “What is the agreement rate?” It is “What would we need to inspect before an agreement result could change the release decision?”

This article gives you that evidence packet. It is a protocol and reusable decision artifact, not a report of a completed benchmark. No reviewer labels, AI rerun, or reliability result is claimed here.

What evidence should domain experts collect about operator agreement?

Collect a frozen fixture of real operator-grounded cases, a versioned task and rubric, independent labels from at least two qualified reviewers, the evidence each reviewer used, abstentions and disagreement reasons, an appropriate agreement analysis, focused adjudication, and a rerun of the identical fixture against the AI system. Make the release decision from that packet, not from a single agreement percentage.

The principal exception is a protocol-only exercise. If the operator context, reviewer labels, or AI outputs do not exist yet, you can prepare the packet and its decision rules, but you cannot describe the result as measured reliability.

The packet separates three questions that are often collapsed into one:

QuestionEvidence neededWhat the answer supports
Did reviewers apply the task consistently?Independent labels, abstentions, missing judgments, and human-human agreementConfidence that the rubric can be applied repeatedly by qualified reviewers
Was the operator decision defensible?Evidence spans, a reference or adjudication record, and disagreement reasonsA bounded claim about correctness for the fixture
Is the AI system safe enough for this scope?Same-fixture AI outputs, severe-error review, abstention, and veto outcomesA go, narrow, or stop decision for the tested workflow

The distinction matters because a group can agree on an unclear rule. The Singapore Responsible AI Playbook treats human evaluation as a process that needs aligned criteria and retained disagreement, while the OECD discussion of structured expert judgement separates the quality of the task, the experts, the instrument, and the way a consensual result is reached. Those are ingredients of an evidence packet, not a universal threshold.

The rest of this page turns that distinction into a concrete artifact.

How was this packet specification checked?

I ran a deterministic completeness audit on the artifact in this article. It checked 10 required elements: the seven packet components and the three go, narrow, and stop branches. All 10 were present. This verifies that the protocol exposes its promised fields; it does not measure reviewer agreement or AI reliability.

Audit fieldObserved resultBoundary
Sample10 required protocol elementsThe sample is the artifact's declared components, not domain cases
MethodExact field and section checks against post.mdxIt checks presence, not whether a reviewer can apply the field well
Result10 of 10 elements presentNo operator labels, AI outputs, agreement value, or release outcome was counted

The audit is useful because a protocol can claim to preserve disagreement while omitting the record that makes disagreement inspectable. It does not turn the protocol into a completed evaluation. A real study still needs the frozen fixture, qualified reviewers, labels, adjudication, same-fixture AI run, and a bounded decision.

Freeze the case fixture before anyone labels it

Freeze the cases and the review context before collecting labels. A reviewer should see a declared version of the task, the same operator context, the same evidence boundary, and the same rubric that the other reviewer sees.

Without that freeze, a later agreement result may measure changing instructions rather than reviewer agreement. The OECD evaluation-instrument study describes a workflow with a rubric, independent ratings, calibration, and later discussion. A domain-specific benchmark described by Frontiers likewise uses fixed context, locked scenarios, standardized execution, archived outputs, and focused adjudication.

Use a fixture that contains the decisions your release will actually face. Include ordinary cases, ambiguous cases, boundary cases, and cases where an error would change the allowed scope or create material harm. The categories are useful because an overall result can hide the cases that matter most. They are not a claim that every workflow needs a particular number of cases.

Record these fixture fields:

FieldRequired contentWhy it is frozen
fixture_id and versionStable identifier and version dateConnects every label and rerun to the same case set
case_idStable case identifier with no identifying informationLets reviewers and adjudicators refer to one case without rewriting it
decision_contextThe operator task, allowed evidence, and decision consequenceDefines what the label means
operator_labelThe decision made in the real workflow, or an explicitly defined reference labelKeeps the comparison target visible
impact_classThe declared risk or consequence classLets release rules focus on severe cases
fixture_hashHash or equivalent immutable version referenceDetects silent case changes

Do not silently replace a case after a disagreement. Create a new fixture version and explain the change. If the operator process itself changes, treat that as a new evaluation rather than a cleaner version of the old result.

Illustration of a versioned operator-agreement case record and its review fields

Give reviewers the same rubric and let them label independently

Use at least two qualified domain reviewers, give them the same versioned instructions, and collect their labels independently before discussion. The packet should preserve a reviewer’s evidence and uncertainty, not only the final consensus.

The AIES LLM-as-Judge implementation profile recommends independent domain-qualified annotation, blinded review where practical, human-human reliability before judge-human agreement, retained original and adjudicated labels, and a predeclared severe-class recall floor for release-blocking detectors. The Singapore playbook also recommends alignment on the use case, criteria, policy, and taxonomy, followed by examination of disagreements.

The minimum label record is:

case_id
fixture_version
reviewer_id_or_pseudonym
reviewer_qualification_basis
label
evidence_spans
confidence
abstained
missing_judgment
disagreement_reason
reviewed_at
rubric_version

reviewer_qualification_basis should identify why the person is qualified for this task. It does not need to expose personal information in the published packet. A pseudonym and a private qualification record can be enough when access must be restricted.

An abstention is not a failed label that should be converted to the majority class. It is evidence that the case, rubric, or available context did not support a confident decision. A missing judgment is different again: it says the expected review did not happen. Preserve both states so the denominator remains inspectable.

An evidence span can be a quoted source fragment, a field in the case record, a policy clause, or a timestamped part of the operator trace. It should be specific enough that another reviewer can see why the label was made. “The case looked risky” is a reason; it is not yet an evidence span.

The exception is a low-consequence exploratory exercise where one person is only testing whether the rubric can be written. Label it as rubric development. Do not call it an agreement study.

Measure agreement without hiding prevalence or uncertainty

Report raw agreement and a chance-adjusted or ordinal measure that fits the label type, together with class prevalence, missing judgments, abstentions, and where disagreements concentrate. Do not select a metric because it produces a familiar-looking threshold.

For paired categorical labels, raw agreement can be stated plainly:

raw agreement = matching non-missing label pairs / all non-missing label pairs

That calculation is easy to inspect, but it does not account for how often each class appears. A chance-adjusted measure can help when its assumptions fit the task. Ordinal labels need a measure that respects their ordering. Multilabel, abstaining, and highly imbalanced tasks may need a different treatment altogether. The AIES profile explicitly warns against a universal acceptable alpha or correlation and places the threshold decision with the evaluation owner before candidate results are observed.

Your metric worksheet should therefore contain:

  1. The label scale and its ordering, if any.
  2. The denominator for every reported calculation.
  3. Raw agreement, with matching and non-missing pair counts.
  4. The selected chance-adjusted or ordinal measure and its assumptions.
  5. Class prevalence for the operator and each reviewer.
  6. Abstention and missing-judgment counts.
  7. A case-level disagreement map by impact class and disagreement reason.
  8. The decision floor or veto rule declared before looking at the candidate result.

Do not publish a single percentage without the case distribution. If nearly every case belongs to one class, a high raw agreement can coexist with weak evidence about the rare class that controls release. That is why disagreement concentration and severe-case behavior belong in the packet.

Illustration of separate agreement, prevalence, disagreement, and severe-error views

Adjudicate disagreements that can change the decision, and keep the originals

Adjudicate disagreements selectively when they could change the release, scope, or interpretation of a high-impact case. Keep each original label, the evidence used, the disagreement reason, and the adjudicated label as separate fields.

Adjudication is useful because it can expose a bad rubric, missing context, or a genuine boundary in the task. It is dangerous when it overwrites the disagreement and leaves only a clean consensus. The OECD study describes independent ratings followed by calibration and discussion that changed some ratings. The Singapore playbook says disagreements should be adjudicated but retained and examined. The Frontiers benchmark uses focused adjudication rather than pretending that every case has the same consequence.

Use this adjudication record:

FieldRecord
case_id and fixture_versionThe exact disputed case
original_labelsEvery independent label, including abstention
evidence_spans_by_reviewerThe evidence each reviewer cited
disagreement_reasonThe coded reason and a short explanation
decision_relevanceWhether the disagreement can change release or scope
adjudicator_basisThe policy, evidence, or rule used to decide
adjudicated_labelThe result, stored separately from original labels
rubric_changeWhether the rubric or fixture needs a new version
reviewed_atDate of adjudication

If the adjudicator cannot decide without inventing context, the correct result may be “insufficient evidence.” That outcome should remain visible. It can lead to a narrower scope, a better case record, or a stop decision.

Compare the AI system on the same fixture, not a cleaner one

Run the AI system against the frozen fixture with its model, prompt, tools, permissions, retrieval inputs, and configuration recorded. Compare its behavior with the operator and adjudicated records, but keep agreement, correctness, and safety as separate outputs.

At minimum, the AI rerun record should include:

OutputWhat it answers
Operator agreementHow often the AI matches the declared operator label
Adjudicated correctnessHow often the AI matches the reviewed reference after decision-relevant adjudication
Severe-error recallWhether the system catches or avoids the cases whose errors can block release
Abstention behaviorWhether the AI refuses or escalates when the case is not judgeable
Veto outcomesWhether any predeclared release-blocking condition was triggered
Run configurationWhich model, prompt, tools, permissions, and date produced the output

The same-fixture rule prevents a favorable result from being created by changing the cases between human review and AI review. It also makes a later rerun interpretable. If the model, prompt, tool access, or retrieval corpus changes, record a new configuration and compare it as a new run.

Do not turn the operator label into ground truth by default. An AI can agree with an inconsistent operator process. Conversely, an AI can disagree because it detected an error, because the rubric is incomplete, or because it misunderstood the task. The packet needs adjudication or another defensible reference before it supports a correctness claim.

Use the packet to choose go, narrow, or stop

Choose an outcome from predeclared conditions, severe-case behavior, and the limits of the fixture. A high overall agreement result is not enough for “go” if the system fails a release-blocking class or cannot show its evidence.

This decision record turns the evidence into a bounded scope decision:

OutcomeConditions that support itRequired action
GoThe fixture matches the intended workflow, reviewers can apply the rubric consistently enough for the declared use, no veto condition is breached, and the AI result is acceptable within the tested scopeRelease only for the tested scope, with monitoring and a refresh trigger
NarrowThe system is useful for a lower-risk subset, but disagreement, abstention, severe errors, or missing context make the full scope unsafeRestrict cases, permissions, or actions and rerun the narrowed fixture
StopThe rubric cannot support a stable judgment, high-impact errors breach a veto, evidence is missing, or the AI behaves outside the declared taskDo not release; repair the rubric, fixture, context, or system before rerunning

Illustration of an evidence packet connected to go, narrow, and stop release outcomes

The conditions above are a decision artifact, not a claim that a particular agreement floor works for every system. Set any numerical floor, severe-error floor, or veto before viewing candidate results, and explain why the consequence justifies it. The AIES profile makes the same point in a different form: acceptable reliability depends on decision impact and label ambiguity.

The decision record should name the tested scope, fixture version, rubric version, AI run, reviewer evidence, unresolved limitations, outcome, owner, and next review date. If “go” depends on a reviewer silently correcting the system, the outcome is not go for autonomous use. It may be go for a human-in-the-loop workflow, which is a different scope.

The complete operator-agreement evidence packet

The packet is complete when a new reviewer can trace a release decision from the frozen case to the final outcome without relying on an unwritten conversation.

Use this packet order:

  1. Fixture record. Store the case IDs, task context, operator labels, impact classes, allowed evidence, version, and immutable reference.
  2. Rubric and protocol. Store the label definitions, reviewer instructions, qualification basis, blinding method, abstention rule, and evidence-span rule.
  3. Independent labels. Store every reviewer label, confidence, evidence span, abstention, missing judgment, and disagreement reason.
  4. Agreement worksheet. Store raw agreement, the selected additional measure, denominators, prevalence, uncertainty treatment, and disagreement concentration.
  5. Adjudication log. Store only the disputes relevant to release or scope, while retaining all original labels and the basis for each adjudication.
  6. AI rerun record. Store the same-fixture outputs, model and configuration, operator comparison, adjudicated comparison, severe-error review, abstentions, and veto outcomes.
  7. Decision record. Store the scope, predeclared conditions, outcome, unresolved limits, owner, and next review date.

This is the sourceable artifact of the article: a type-specific packet that joins the evidence needed to distinguish reviewer agreement from correctness and release safety. A generic “human evaluation checklist” does not make those links explicit.

What usually goes wrong

The most common failure is treating the final consensus as the dataset. That erases the disagreement that tells you whether the rubric is clear and where the workflow is risky.

Another failure is using the operator label as unquestioned truth. Operator decisions are valuable because they are grounded in work, but they still need a declared context, evidence boundary, and review process. A system that matches a flawed process can be consistently wrong.

Teams also change the fixture after seeing the first result. That can be appropriate for a new version, but it is not a clean rerun. Keep the old fixture and label the new one so the decision history remains reproducible.

Finally, teams report a single aggregate score while ignoring abstentions, missing judgments, rare severe cases, or evidence quality. The packet prevents that compression by making each denominator and decision consequence explicit.

When this protocol does not apply

This packet is not a substitute for a safety case, a regulatory validation process, or a domain-specific outcome study when those are required. It is also not enough when the operator task has no stable context, no qualified reviewers, or no way to preserve the evidence behind a decision. In those situations, the right outcome is to narrow the claim to protocol design or stop until the evaluation can be made inspectable.

For the broader evaluation architecture, see the parent guide on evaluating AI against operator decisions. For the neighboring data job, see how to build an evaluation dataset from production traces; for output-level review, see how to grade an AI output against a rubric. If the next step is helping an existing team define a bounded evaluation and build the capability to run it, see Marius Manolachi's AI learning and consulting work and bring the workflow, decision rubric, and current evidence packet.

Before you ask whether the AI agrees with the operator, ask whether another qualified reviewer can inspect the same case and explain the decision. That is the evidence boundary the release decision depends on.

Questions people ask next

Can agreement with an operator prove that an AI decision is correct?

No. Agreement shows that reviewers applied the task and rubric in a similar way. Correctness needs a defensible reference, adjudication, or another outcome measure, and release safety may still require a veto rule.

How many domain reviewers should label the fixture?

Use at least two qualified reviewers who label independently before discussion. More reviewers may be justified when the task is ambiguous or high impact, but a larger count does not repair a vague rubric.

What should the packet preserve when reviewers disagree?

Keep both original labels, the evidence each reviewer used, the reason for disagreement, and any abstention. Adjudicate disagreements that could change scope or release, while retaining the adjudicated label separately.