Field note · capability
How to Diagnose Why AI Training Does Not Transfer
Reproduce an AI training transfer failure with paired cases, trace the first broken condition, repair it, and verify the fix on changed work.

AI training transfer fails when a person can repeat the taught sequence but cannot adapt it when the task, evidence, or authority boundary changes. The useful question is not whether the workshop was clear. It is where the work first stops being independently safe and defensible.
This article uses a worked diagnostic artifact, not a team study. The cases are fictional and low risk. Their purpose is to make the failure trace reproducible without pretending that an authored walkthrough measures training effectiveness.

The sourceable atom is the paired-case transfer failure trace below. It records the observable break, its likely cause, the smallest repair, and the check that can confirm the repair. You can copy the structure into a real assessment, then replace the fictional cases with approved work.
How do you reproduce an AI training transfer failure?
Compare a guided case with an unseen case that changes one meaningful condition. Save the prompt or instructions, source pack, work artifact, verification notes, and handoff decision for both cases. A fluent answer is not enough evidence of transfer.
This paired-task design follows the practical direction of work-task-oriented AI literacy assessment, which tests applied behaviour in context rather than relying only on abstract knowledge or self-report (the work-task assessment paper). The AIL AT WORK project also distinguishes self-assessment from objective assessment, a useful reminder that confidence and observable performance answer different questions (AIL AT WORK instruments).
Use a low-risk task with a clear source of truth. For the worked trace, the task is to turn a small set of sanitized operations notes into a one-page decision brief. The participant may use an AI assistant to organize the notes, but the notes remain the source of truth. No external action is allowed.
- Run the guided case with explicit instructions to cite the notes, mark assumptions, and name the human owner.
- Change the evidence and one operating constraint without changing the requested artifact.
- Remove the step-by-step checklist, but keep the same timebox and source rule.
- Compare the two work trails, not just the final prose.
Here is the worked reproduction trace. It is an authored example, not a record of participants.
| Stage | Changed condition | Observable trace | First failure to investigate |
|---|---|---|---|
| Guided case | Notes describe a recurring report delay. Public or sanitized text is allowed. | The brief cites two notes, marks one uncertainty, and recommends a read-only check. | None in the worked case. The artifact shows the expected evidence trail. |
| Unseen case | A note introduces restricted internal records. A second note requires a named human approval before any response is sent. | The draft repeats the earlier tool plan, does not flag the data boundary, and leaves approval unnamed. | Boundary recognition, before prompt quality or writing quality. |
| Reproduction | Keep the instructions and change only the case cards. Replay the same request once. | The same tool plan appears again because the task-specific boundary is absent from the working checklist. | The training taught a sequence, but not how to derive a safe boundary from new evidence. |
The trace makes a useful distinction. The person can produce the requested shape. The transfer failure is the unchanged decision process, not an inability to write a brief. That distinction determines the repair.
How do you diagnose the first broken condition?
Trace the work from task framing to handoff and stop at the first condition that fails. Do not label the whole learner or the whole training program from one bad final answer. Diagnose the missing observable behaviour.
Use this worksheet after each paired case:
| Condition | Evidence to inspect | Failure signal | Diagnosis |
|---|---|---|---|
| Frame | Goal, scope, constraints, and what the artifact does not decide | The changed case is treated as the old case | The task brief does not expose the decision boundary. |
| Source check | Citations, source IDs, assumptions, and conflicts | The answer sounds certain but cannot point to the changed note | Evidence handling was taught as a final polish step. |
| Tool choice | Tool, input class, permissions, and allowed actions | Restricted material is sent to a tool without approval | The workflow boundary is missing or too abstract. |
| Verification | Corrections, rejected claims, and checks against the source of truth | The model's draft is accepted because it is plausible | Verification has no visible stopping rule. |
| Handoff | Owner, approval point, escalation path, and next check | A recommendation is presented as permission to act | The artifact ends before responsibility is assigned. |
The AICC instrument repository is useful for separating assessment types such as self-report and content assessment, but it does not supply a universal workplace pass score (AICC instrument repository). Treat the worksheet as a local diagnostic artifact. Its value is the evidence it asks you to save, not a number that travels unchanged across roles.
In the worked trace, the first break is tool choice and boundary recognition. The missing approval is a second break in handoff. The polished brief is downstream noise. Fixing the prompt or asking for more fluent prose would leave both defects intact.
The exception is a task where the source of truth itself is unstable. If the notes conflict or are incomplete, a person may be right to stop before choosing a tool. Record that as an unresolved input condition, not as a learner failure. The diagnostic should reveal uncertainty rather than punish someone for refusing to invent certainty.
What is the smallest repair for a transfer failure?
Repair the missing condition in the task and workflow before adding another general lesson. A good repair makes the required behaviour visible, bounded, and testable on the next case.
For the worked trace, the repair is a four-field task card placed before tool use:
| Field | Repair wording for the next case | Why it changes the failure |
|---|---|---|
| Source of truth | List the exact notes that may support the brief. Mark every assumption. | The participant must derive claims from the new evidence. |
| Input boundary | State which records may enter the selected tool and which may not. | The tool choice now depends on the case, not the habit. |
| Action boundary | Separate drafting from sending, editing, approving, or changing a system. | A useful draft cannot silently become authorization. |
| Owner and stop rule | Name the human owner and the condition that requires escalation. | The artifact has a safe endpoint when certainty or authority runs out. |
Run the repair as a short procedure:
- Rewrite the case card so the changed constraint is explicit but not solved for the learner.
- Require one source reference beside each material recommendation.
- Require the learner to state the tool's allowed input and output actions before opening it.
- Add a reviewer check for privacy, authority, and unresolved evidence.
This is not an argument for a larger course by default. The U.S. Department of Labor presents its AI Literacy Framework as a resource for program design, so it can inform the capability topics a role needs without becoming a local pass certificate (DOL AI Literacy Framework). Canada’s AI Readiness Scorecard makes a similar decision-oriented move by asking whether a specific problem is a fit for an AI solution and whether to proceed, pivot, or pause (Canada AI Readiness Scorecard).
If the failure is missing vocabulary or a missing technique, train that narrow skill and repeat the case. If the person knows the rule but the workflow hides it in a policy page, redesign the task card, interface, or approval path. If the boundary cannot be made safe, pause. More instruction cannot grant authority that the workflow does not have.
How do you verify that the repair worked?
Retest with a changed case and no step-by-step rescue. Verification requires the repaired behaviour to appear in the work trail, not just a correct answer after coaching.
Use a second case with a different evidence conflict, input boundary, or approval requirement. Keep the output format stable so the comparison remains useful. Then mark the following checks:
| Verification check | Pass evidence | Fail evidence |
|---|---|---|
| Framing | The learner states what changed and what the brief will not decide. | The old case is copied with new nouns. |
| Evidence | Material claims point to the new source cards or are marked as assumptions. | The answer relies on fluent but untraceable claims. |
| Boundary | The selected tool and input path fit the new permission rule. | Restricted data or an unapproved action remains in the plan. |
| Repair | The learner corrects the reproduced defect without being shown the answer. | The reviewer must point out the same defect again. |
| Handoff | A responsible owner, approval step, and next check are named. | The artifact implies that an AI output is permission to act. |
The verification result is a decision artifact, not a universal score:
| Result | Decision | Follow-up |
|---|---|---|
| All checks pass and the task is low risk | Continue within the stated permissions | Recheck when the task, tool, data, or owner changes. |
| The technique is sound but one observable skill fails | Train or coach that skill | Run another changed case focused on the failed condition. |
| The workflow hides the needed evidence or boundary | Redesign the task or approval path | Retest after the workflow exposes the condition. |
| Privacy, authority, or safety remains unresolved | Pause and escalate | Do not treat more training as permission to proceed. |

The principal exception is high-impact work. A successful low-risk drafting retest does not authorize unsupervised financial, legal, employment, healthcare, customer-send, or production actions. Keep the qualified reviewer and the relevant approval chain in place.
How should the result change the next training decision?
Use the first broken condition to choose among targeted training, workflow redesign, and pause. Do not report a single completion percentage as proof that transfer happened.
The artifact gives a clean handoff:
- Train when the learner cannot perform a required technique even after the task boundary and source material are clear.
- Redesign when the learner could act safely but the workflow hides the source, permission, owner, or stop rule.
- Pause when the authority, privacy, or safety condition is unresolved, or when no qualified reviewer is available.
The wider What to Learn Before Building AI Agents guide provides the collection context for this diagnostic. For the practice loop that should precede a transfer check, see How to Make AI Training Stick in a Small Team. If the defect is in judging an individual output rather than in training transfer, use How to Grade an AI Output Against a Rubric.
The worked trace is deliberately modest. It proves that the artifact can expose a specific failure and connect it to a repair and a retest. It does not prove that a team improved, that one rubric fits every role, or that training caused a later business result. A real assessment should preserve the same evidence trail with approved cases, role-specific reviewers, and explicit limits.