Field note · capability
Workplace AI Task Transfer Assessment
Assess workplace AI task transfer with a five-dimension rubric, a decision record, and one changed condition that reveals whether the learner can adapt safely.

Most AI learning plans measure attendance, hours, or whether someone finished a guided exercise. Those signals tell you that practice happened. They do not tell you whether the learner can recognize a bad assumption when the work changes.
Use the rubric below on one low-risk workplace task, then repeat it with one changed condition. The result is a coaching decision, not a claim about general intelligence or a replacement for professional review.
The broader AI learning capability guide covers the larger curriculum. This page adds a concrete scorecard. If you are designing a team program, pair it with how to make AI training stick in a small team.

The rubric should score decisions, not activity
Score the learner's visible decisions and artifacts, not prompt volume or time spent in a training session. A useful workplace rubric asks whether the person can define the job, use AI selectively, verify what matters, recover from a failed attempt, and finish with an explainable result.
The five dimensions below are a practical synthesis of the OECD AI Capability Indicators, UNESCO's competency-based assessment guidance, the Frontiers AI-literacy rubric, the task-oriented AI-literacy study, and the U.S. Department of Labor framework. Those sources support observable, contextualized performance. The scorecard and its thresholds are this page's artifact.
| Dimension | What the reviewer looks for | Required evidence |
|---|---|---|
| Frame the job | States the user, outcome, constraints, and definition of done | A short task brief with an explicit completion condition |
| Delegate selectively | Gives AI work that is bounded and keeps judgment where context or risk requires it | A delegation note naming the AI task and the human task |
| Check evidence and uncertainty | Traces important claims, marks unknowns, and avoids treating fluent output as proof | Source notes, checks, or an uncertainty log |
| Adapt after failure | Identifies what failed, changes the relevant assumption, and reruns the work | A failed attempt, diagnosis, and repaired attempt |
| Complete or explain independently | Produces the result and can explain why it meets the stated condition | Final artifact plus a concise decision record |
The exception is a task with no meaningful judgment, such as a fully deterministic transformation with an independent machine check. Even there, the framing and completion condition still protect against checking the wrong output.
Use four levels and a safety veto
Use four levels to describe the quality of each dimension. Do not average away a safety failure. A learner can be strong at producing an answer and still need review because they cannot verify its source or recognize an unsafe boundary.
| Level | Observable behavior | Coaching decision |
|---|---|---|
| 1. Guided | Needs the task framed, the checks named, or the next action supplied | Model the missing decision and repeat the same task with support |
| 2. Assisted | Completes the step with a checklist, prompt, or reviewer's correction | Keep the checklist, then remove one support on the next run |
| 3. Reliable | Makes the decision, records evidence, and repairs a normal failure with limited prompting | Run the changed-context task |
| 4. Transferable | Repeats the reasoning on a changed condition and names when to stop or escalate | Increase task variation, not consequence, before widening authority |
Apply a safety veto when the learner invents evidence, ignores a material uncertainty, crosses an access or privacy boundary, or claims completion without checking the definition of done. A veto means “not ready for this workflow boundary,” even if the final prose looks polished.
This distinction follows the evidence base. The OECD describes capability across domains such as problem solving and metacognition, while the DOL framework emphasizes contextualized learning and evaluation of AI outputs. The scorecard turns those broad ideas into evidence a manager can inspect. It does not turn them into a validated numeric scale.
Run one authentic task and one changed case
Run the first task in the learner's real work at low risk, then change one condition that should alter the decision. The changed case is the transfer check. It prevents a familiar procedure from looking like independent problem solving.
- Choose a low-risk task with a real user, a clear output, and a reversible consequence.
- Write the definition of done before the learner starts. Include the source of truth and the human owner.
- Ask the learner to complete the task with the AI tool they normally use. Preserve the prompt, output, edits, checks, and failed attempt if one occurs.
- Score each dimension from 1 to 4. Write one sentence of evidence for every score. A score without an artifact is a judgment, not a record.
- Create a changed case by moving one relevant boundary: a source becomes stale, the audience changes, a required field is missing, or the action becomes irreversible.
- Re-run the task without explaining the expected answer. Score whether the learner notices the changed condition and changes the decision.
- Choose the next coaching action: keep support, remove one support, repeat with a different boundary, or stop the workflow until domain review is present.
UNESCO's guidance supports authentic tasks and transferability, and the task-oriented AI-literacy research uses contextualized work tasks to examine applied capability. The procedure above is an operational adaptation of that guidance, not a claim that these sources validate this exact scorecard.
Give the learner a decision record, not another quiz
Require a small artifact that exposes the decision boundary. A useful record contains the task, chosen action, evidence, uncertainty, rejected alternative, escalation trigger, and final result. It should be short enough to complete during work and concrete enough for another reviewer to check.
Task and user:
Definition of done:
What AI did:
What the human owned:
Evidence checked:
Unknowns or uncertainty:
Failed attempt and diagnosis:
Changed condition:
Chosen action after the change:
Rejected alternative:
Escalation trigger:
Final artifact or link:
Reviewer score: 1 2 3 4
Reviewer evidence:
The artifact separates output quality from decision quality. A correct answer with no evidence may score well on completion and poorly on verification. A cautious answer that names the missing source may score better on uncertainty, even when it cannot finish the task. That is useful information for coaching.
The record also protects the learner from a vague standard. “Use AI responsibly” is not a completion condition. “Draft the internal update from today's board, leave the external date out, and send unresolved ownership to the project owner” is inspectable.
Interpret the score with worked decisions
Use the score to decide what support or boundary comes next. Do not turn the total into a universal readiness label. The two cases below are illustrative applications of the artifact, not observed learner results.
| Illustrative case | First task | Changed case | Decision |
|---|---|---|---|
| Support reply | Learner drafts a FAQ-grounded invoice-location reply for a verified customer | The customer requests a refund and an account change without verified identity | Hold and route to support; the changed boundary triggers a safety veto |
| Project update | Learner summarizes a current internal board with a named owner and review date | The source is stale, the owner is missing, and the update includes an external commitment | Hold publication and escalate; the learner must not fill gaps from an old note |
In the first case, the initial procedure can be correct while the changed case requires a different action. In the second, fluent wording can hide stale evidence and missing ownership. The reviewer is scoring whether the learner changed the decision for the right reason, not whether the response sounded careful.
Know where the rubric stops
Use this rubric for coaching and task design, not as permission to remove review from consequential work. High-stakes decisions need the domain's required controls, accountable ownership, privacy safeguards, access limits, and documented approval even when a learner reaches level 4 on a practice task.
The rubric also cannot tell you whether the chosen task represents the learner's whole role. One changed case tests one boundary. A learner may transfer the skill to a missing-source case and still struggle with conflicting policy, unfamiliar data, or a different audience. Add task variation before adding consequence.
The next useful decision is small: choose one reversible task, write its definition of done, and preserve the decision record. If the learner can complete the answer but cannot show the evidence or changed-case decision, coach that missing step before introducing a more autonomous workflow.
Questions people ask next
Can this rubric prove that someone is ready for every AI task?
No. It gives a bounded coaching signal for one task and one changed case. Use role-specific evaluation, domain review, and formal controls when the work carries material legal, financial, employment, medical, privacy, or safety consequences.
Does a high score mean the learner should work without review?
No. Independent completion means the learner can show the required decisions and evidence for the exercise. It does not remove accountable human approval, access controls, or review requirements for the real workflow.