Field note · capability
What Should an AI Training Transfer Evidence Packet Contain?
A practical evidence packet turns AI training transfer into a reviewable record with cases, a rubric, a trace, a correction, and a decision boundary.

The difficult part of AI training comes after the workshop. A participant can repeat the demonstration and still fail when the input changes, a source conflicts, or the output needs checking.
When I taught product managers who moved from writing specifications to building and shipping, the recurring gap was often an undefined meaning of done, not a missing model trick. A transfer packet makes that meaning visible. It gives a manager something to inspect before reducing routine support.
What does a transfer evidence packet need to prove?
It needs to connect one real task to a new case, a visible quality rule, an independent check, a bounded correction, and a recorded decision. Ten records make that chain inspectable.
I derived the packet from the task-performance emphasis in the OECD's assessment work, the role-sensitive practice and support boundary in the Skills England tools package, the output-checking behaviors in the Skills England benchmark, and Pearson's guidance on independent work and thinking. (OECD task-performance assessment, Skills England benchmark, Skills England tools package, Pearson AI guidance)
I then checked the schema itself on 2026-08-24. The sample was one packet schema, not a group of participants. Each required record was checked for presence and a distinct purpose.
| Required record | The question it answers | Schema audit result |
|---|---|---|
| Task brief | What real task is being tested? | Present |
| Representative cases | What normal and difficult inputs define the task? | Present |
| Withheld fresh case | Can the method travel beyond practice examples? | Present |
| Published rubric | What counts as acceptable, unsafe, incomplete, or escalated? | Present |
| Raw work trace | What did the person actually do and inspect? | Present |
| Independent verification | Who or what checked the result against a source of truth? | Present |
| Bounded ambiguity | What changed or failed during the test? | Present |
| Learner correction | What did the person change without the trainer taking over? | Present |
| Rerun record | Did the correction produce the intended change? | Present |
| Assessor decision | What support can change, and what approval remains? | Present |
| Schema total | Is the packet structurally complete? | 10 of 10 records represented |
That is the original result here: a compact packet contract that another manager can reproduce and audit. It is not evidence that a learner passed. A blank or unrun packet is still only a plan.
The packet belongs beside the broader workplace AI capability guide. This article owns the evidence record that a manager can use inside that capability decision.
How was the packet schema audited?
The audit used a fixed presence check, not a judgment about whether a person was capable. Each of the ten records received present when the schema named the record and its purpose, or missing when it did not. The audit passed only when all ten records were present and the records formed a sequence from task to decision.
The method has three parts:
- Define the task as an occupational activity with a source of truth, an acceptable outcome, and a failure boundary. This follows the OECD's preference for practical task performance over an abstract capability label. (OECD assessment chapter)
- Map the behavior to observable checks. Skills England's benchmark includes checking outputs for accuracy and spotting errors, while its tools package describes role-sensitive practice, professional judgment, and knowing when to seek support. (Skills England benchmark, Skills England tools package)
- Require a trace and a correction path. NIST treats AI measurement and evaluation as context-dependent, and Anthropic separates the task, trial, transcript, grader, and outcome. Those distinctions prevent a polished final answer from standing in for the process that produced it. (NIST measurement and evaluation, Anthropic evaluation definitions)
The result is intentionally small. It checks that the packet can hold the evidence. It does not estimate reliability, compare trainers, or establish a threshold for every role. The schema is a gate for evidence collection, not a substitute for evidence collection.
What belongs in the task and case records?
The task record should describe one real, low-risk activity, and the case records should show enough variation to expose the rule being taught. Use three representative cases for practice and one fresh case withheld until the transfer check. The number is a packet design choice, not a universal sample-size claim.
Write these fields before the session starts:
| Field | What to record | Why it matters |
|---|---|---|
| Task | The input, output, owner, and normal manual path | Keeps the test tied to actual work |
| Source of truth | Policy, record, document, or human authority used to check the result | Separates correctness from fluent wording |
| Acceptable outcome | What must be true for the output to be usable | Gives the rubric something concrete to inspect |
| Failure cost | What happens if the output is wrong, incomplete, late, or overconfident | Sets the review and escalation boundary |
| Permissions | What the AI system and participant may read, change, or send | Prevents an assessment from becoming an uncontrolled production action |
| Representative cases | Ordinary input, missing detail, and conflicting or ambiguous evidence | Tests the behavior around the happy path |
| Withheld fresh case | A new input with the same task shape but different details | Tests transfer rather than recall |
The fresh case should be sealed from the practice discussion. It can be a sanitized internal record or a newly written case that preserves the real task structure without exposing private information. Do not use a case so unfamiliar that it tests a different job. Do not let the trainer explain the answer while the participant is performing the transfer check.
A good case changes one meaningful condition at a time. One case can omit a required fact. Another can contain two sources that disagree. A third can make the cost of an invented answer clear. The fresh case then asks whether the participant applies the same checks without being handed the original vocabulary.

How should the rubric make transfer observable?
The rubric should score observable behavior and keep critical failures as vetoes. Do not grade “confidence” or “overall quality” first. Separate whether the result is faithful, useful, complete, safe, and in the requested form.
Use a small rubric with anchors that another assessor can apply:
| Dimension | Pass evidence | Veto condition |
|---|---|---|
| Fidelity | Claims and extracted fields agree with the source of truth, or uncertainty is marked | Invented fact, unsupported certainty, or source contradiction |
| Decision usefulness | The output states the next action and its owner | No actionable next step when the task requires one |
| Completeness | Required fields and relevant exceptions are present | A missing field changes the decision |
| Risk and uncertainty | The participant identifies what needs review and why | High-impact ambiguity is treated as routine |
| Format and traceability | The output fits the requested form and links back to evidence | The assessor cannot reconstruct the basis of a decision |
The rubric is a measurement instrument for this task, not a general AI score. NIST's measurement guidance makes the same contextual point: the characteristic, task, and method need to fit the system and setting. A finance approval task, a draft internal brief, and a read-only document extraction task should not inherit the same vetoes.
Publish the anchors before the fresh case. Otherwise the assessor can unconsciously move the goalposts after seeing the output. Skills England's guidance supports observable output checking and knowing when to seek support. Pearson's guidance adds a related boundary: the person must demonstrate their own independent work and thinking rather than submit an answer that cannot be explained.
The rubric must also state what does not count as a pass. A participant who gets the final field right by guessing has not shown the same evidence as someone who names the source, spots the ambiguity, and explains the choice. A fluent output can therefore fail while a cautious, incomplete output can earn a correction opportunity.
What should the raw trace and verification record capture?
The raw trace should preserve enough of the interaction to show the participant's decisions, checks, and requests for help. The verification record should be separate enough that the assessor can see what was checked independently.
Capture:
- the task version and case ID;
- the participant's instructions, edits, tool use, and intermediate checks;
- the output shown to the assessor;
- the source or rule used to verify each material field;
- any request for help and the exact help given;
- the assessor's verification notes, including disagreements;
- the permission boundary and whether any external state changed.
Anthropic's evaluation vocabulary is useful here because it distinguishes a trial from its transcript and the final outcome. For a workplace packet, the transcript can be a screen recording, saved prompt and output history, a versioned document, or a structured work log. The format matters less than preserving the sequence and the evidence.
Independent verification does not have to mean a second person checks every word. It can be a deterministic field check, a source comparison, a policy lookup, or a qualified reviewer for a consequential decision. The packet should name who or what performed the check and what remained uncertain.
If the task uses an AI system with tools, record the tool boundary as well. A participant who can draft a recommendation is not automatically authorized to send it, change a customer record, approve money, or alter a production system. The test should observe the intended capability without granting permissions the role does not have.
Why do ambiguity, correction, and rerun matter?
One bounded failure is more informative than several selected successes. The participant should face an ambiguity that the rubric recognizes, choose a safe response, make a defined correction, and rerun the affected case.
Use this sequence:
- Present the case without the trainer's answer.
- Preserve the participant's first output and explanation.
- Mark the ambiguity or failure against the published rubric.
- Ask the participant to state the smallest safe change.
- Let the participant make the change with the agreed permissions.
- Rerun the same case and record what changed.
- Check one neighboring case to see whether the correction created a new error.
The correction should be bounded. It might be a changed instruction, a clearer source-selection rule, a new review condition, or a revised output field. It should not become an open-ended rescue by the trainer. If the trainer types the fix, the packet can record assisted recovery, but it cannot label the correction independent.
The rerun is the point of the packet. It shows whether the person can connect a failure to an intervention and inspect the new result. It also exposes a common trap: a fix can repair one case while breaking another. That is why the neighboring-case check belongs in the record even when it is not a separate scored gate.

The packet should keep the failed first attempt. Deleting it produces a tidy story and weak evidence. A manager needs to see whether the participant recognized the error, whether the correction was proportionate, and whether the final result was independently checked.
What decision should the assessor sign?
The decision record should reduce a specific kind of support for a specific task and preserve approval where the consequences require it. “Ready” is too broad to sign.
Use a decision record with these fields:
| Decision field | Example entry |
|---|---|
| Task covered | Read-only extraction of policy fields into an internal review brief |
| Evidence reviewed | Fresh case, raw trace, rubric, verification notes, correction, rerun, neighboring-case check |
| Support that can reduce | Routine prompting for the defined task and ordinary input shape |
| Support that remains | Review for conflicting sources, missing ownership, and high-impact decisions |
| Permission boundary | No external messages, record changes, payments, or production writes |
| Escalation trigger | Source conflict, missing evidence, policy ambiguity, or a veto condition |
| Assessor decision | Reduce routine support for this bounded task, with a dated recheck |
| Unknowns | Transfer to other roles, future model changes, scale, and long-tail cases |
The decision should be scoped by task, risk, and time. A successful read-only transfer check does not authorize an automated customer communication. A participant who can verify a draft does not automatically own a production release. The packet helps a manager say exactly what changed and what did not.
This is also where the article connects to the buying decision. The AI tutoring versus implementation comparison separates a person's ability to change and rerun a workflow from the existence of a runnable handoff. The packet here is the evidence layer for the first question. For rubric construction, use How to Grade an AI Output Against a Rubric.

What does this packet still not prove?
It does not prove broad job performance, durable retention, model reliability across a population, or safe operation at production scale. It proves only what the preserved task, cases, rubric, trace, verification, correction, rerun, and decision can support.
The schema audit reported 10 of 10 required records represented. That result says the packet has the places where evidence belongs. It says nothing about whether a participant filled those places well. A real run could still reveal that the fresh case is too easy, the rubric misses a critical failure, the verification source is circular, or the assessor cannot reproduce the decision.
The evidence also has freshness risks. The Skills England documents can change. Pearson guidance is tied to its delivery context. OECD's assessment chapter is a methodology source, not a result from this packet. NIST's measurement principles are broad guidance, not a certification. Recheck the sources when the packet is reused for a new role or a new risk boundary.
Keep human approval for work involving money, legal or health consequences, employment decisions, customer commitments, sensitive records, or irreversible external actions. The packet can show that a person knows when to escalate. It cannot make the approval boundary disappear.
If you run the packet, publish the actual observation separately from this schema result: task, sample, cases, raw trace, failure, correction, rerun, decision, and limitations. Until then, the honest claim is narrower and useful: this is a reproducible evidence contract for inspecting AI training transfer, not proof that anyone has passed it.