Field note · capability

Workplace AI Task Transfer Assessment

Assess workplace AI task transfer with a five-dimension rubric, a decision record, and one changed condition that reveals whether the learner can adapt safely.

7 minute read
  • AI learning
  • AI training
  • Workplace capability
  • Problem solving
Illustration of a workplace AI learning progress rubric beside a task brief and transfer check

Most AI learning plans measure attendance, hours, or whether someone finished a guided exercise. Those signals tell you that practice happened. They do not tell you whether the learner can recognize a bad assumption when the work changes.

Use the rubric below on one low-risk workplace task, then repeat it with one changed condition. The result is a coaching decision, not a claim about general intelligence or a replacement for professional review.

The broader AI learning capability guide covers the larger curriculum. This page adds a concrete scorecard. If you are designing a team program, pair it with how to make AI training stick in a small team.

Illustration of a workplace AI learning rubric moving from a task brief to evidence checks and a changed-context transfer card

The rubric should score decisions, not activity

Score the learner's visible decisions and artifacts, not prompt volume or time spent in a training session. A useful workplace rubric asks whether the person can define the job, use AI selectively, verify what matters, recover from a failed attempt, and finish with an explainable result.

The five dimensions below are a practical synthesis of the OECD AI Capability Indicators, UNESCO's competency-based assessment guidance, the Frontiers AI-literacy rubric, the task-oriented AI-literacy study, and the U.S. Department of Labor framework. Those sources support observable, contextualized performance. The scorecard and its thresholds are this page's artifact.

DimensionWhat the reviewer looks forRequired evidence
Frame the jobStates the user, outcome, constraints, and definition of doneA short task brief with an explicit completion condition
Delegate selectivelyGives AI work that is bounded and keeps judgment where context or risk requires itA delegation note naming the AI task and the human task
Check evidence and uncertaintyTraces important claims, marks unknowns, and avoids treating fluent output as proofSource notes, checks, or an uncertainty log
Adapt after failureIdentifies what failed, changes the relevant assumption, and reruns the workA failed attempt, diagnosis, and repaired attempt
Complete or explain independentlyProduces the result and can explain why it meets the stated conditionFinal artifact plus a concise decision record

The exception is a task with no meaningful judgment, such as a fully deterministic transformation with an independent machine check. Even there, the framing and completion condition still protect against checking the wrong output.

Use four levels and a safety veto

Use four levels to describe the quality of each dimension. Do not average away a safety failure. A learner can be strong at producing an answer and still need review because they cannot verify its source or recognize an unsafe boundary.

LevelObservable behaviorCoaching decision
1. GuidedNeeds the task framed, the checks named, or the next action suppliedModel the missing decision and repeat the same task with support
2. AssistedCompletes the step with a checklist, prompt, or reviewer's correctionKeep the checklist, then remove one support on the next run
3. ReliableMakes the decision, records evidence, and repairs a normal failure with limited promptingRun the changed-context task
4. TransferableRepeats the reasoning on a changed condition and names when to stop or escalateIncrease task variation, not consequence, before widening authority

Apply a safety veto when the learner invents evidence, ignores a material uncertainty, crosses an access or privacy boundary, or claims completion without checking the definition of done. A veto means “not ready for this workflow boundary,” even if the final prose looks polished.

This distinction follows the evidence base. The OECD describes capability across domains such as problem solving and metacognition, while the DOL framework emphasizes contextualized learning and evaluation of AI outputs. The scorecard turns those broad ideas into evidence a manager can inspect. It does not turn them into a validated numeric scale.

Run one authentic task and one changed case

Run the first task in the learner's real work at low risk, then change one condition that should alter the decision. The changed case is the transfer check. It prevents a familiar procedure from looking like independent problem solving.

  1. Choose a low-risk task with a real user, a clear output, and a reversible consequence.
  2. Write the definition of done before the learner starts. Include the source of truth and the human owner.
  3. Ask the learner to complete the task with the AI tool they normally use. Preserve the prompt, output, edits, checks, and failed attempt if one occurs.
  4. Score each dimension from 1 to 4. Write one sentence of evidence for every score. A score without an artifact is a judgment, not a record.
  5. Create a changed case by moving one relevant boundary: a source becomes stale, the audience changes, a required field is missing, or the action becomes irreversible.
  6. Re-run the task without explaining the expected answer. Score whether the learner notices the changed condition and changes the decision.
  7. Choose the next coaching action: keep support, remove one support, repeat with a different boundary, or stop the workflow until domain review is present.

UNESCO's guidance supports authentic tasks and transferability, and the task-oriented AI-literacy research uses contextualized work tasks to examine applied capability. The procedure above is an operational adaptation of that guidance, not a claim that these sources validate this exact scorecard.

Give the learner a decision record, not another quiz

Require a small artifact that exposes the decision boundary. A useful record contains the task, chosen action, evidence, uncertainty, rejected alternative, escalation trigger, and final result. It should be short enough to complete during work and concrete enough for another reviewer to check.

Task and user:
Definition of done:
What AI did:
What the human owned:
Evidence checked:
Unknowns or uncertainty:
Failed attempt and diagnosis:
Changed condition:
Chosen action after the change:
Rejected alternative:
Escalation trigger:
Final artifact or link:
Reviewer score: 1  2  3  4
Reviewer evidence:

The artifact separates output quality from decision quality. A correct answer with no evidence may score well on completion and poorly on verification. A cautious answer that names the missing source may score better on uncertainty, even when it cannot finish the task. That is useful information for coaching.

The record also protects the learner from a vague standard. “Use AI responsibly” is not a completion condition. “Draft the internal update from today's board, leave the external date out, and send unresolved ownership to the project owner” is inspectable.

Interpret the score with worked decisions

Use the score to decide what support or boundary comes next. Do not turn the total into a universal readiness label. The two cases below are illustrative applications of the artifact, not observed learner results.

Illustrative caseFirst taskChanged caseDecision
Support replyLearner drafts a FAQ-grounded invoice-location reply for a verified customerThe customer requests a refund and an account change without verified identityHold and route to support; the changed boundary triggers a safety veto
Project updateLearner summarizes a current internal board with a named owner and review dateThe source is stale, the owner is missing, and the update includes an external commitmentHold publication and escalate; the learner must not fill gaps from an old note

In the first case, the initial procedure can be correct while the changed case requires a different action. In the second, fluent wording can hide stale evidence and missing ownership. The reviewer is scoring whether the learner changed the decision for the right reason, not whether the response sounded careful.

Know where the rubric stops

Use this rubric for coaching and task design, not as permission to remove review from consequential work. High-stakes decisions need the domain's required controls, accountable ownership, privacy safeguards, access limits, and documented approval even when a learner reaches level 4 on a practice task.

The rubric also cannot tell you whether the chosen task represents the learner's whole role. One changed case tests one boundary. A learner may transfer the skill to a missing-source case and still struggle with conflicting policy, unfamiliar data, or a different audience. Add task variation before adding consequence.

The next useful decision is small: choose one reversible task, write its definition of done, and preserve the decision record. If the learner can complete the answer but cannot show the evidence or changed-case decision, coach that missing step before introducing a more autonomous workflow.

Questions people ask next

Can this rubric prove that someone is ready for every AI task?

No. It gives a bounded coaching signal for one task and one changed case. Use role-specific evaluation, domain review, and formal controls when the work carries material legal, financial, employment, medical, privacy, or safety consequences.

Does a high score mean the learner should work without review?

No. Independent completion means the learner can show the required decisions and evidence for the exercise. It does not remove accountable human approval, access controls, or review requirements for the real workflow.