Field note · capability
How to Assess Whether AI Training Made a Team Capable
Use a role-specific work sample, failure repair, and transfer case to tell whether AI training created capability rather than attendance or confidence.

After teaching product managers who went from writing specs to building and shipping the product, and automating work around it, I pay attention to a different signal than workshop enthusiasm: can someone make a sound AI-assisted decision when the inputs change? That is the capability question, and it needs a work sample.
The packet below is designed for a product manager. It uses fictional, non-sensitive inputs, so a team can practise without putting customer or employee data into a tool. It also makes a useful distinction visible: a person can complete every field and still fail to show independent judgment.
For the follow-up loop that helps a new method survive after the assessment, see How to Make AI Training Stick in a Small Team.
Observed result from the bounded desk pilot: all three authored sample responses completed the assessment form. Only one met the final capability threshold after artifact, repair, and transfer scores were combined. This was a rubric pilot, not a study of trained employees.

What counts as capability after AI training?
Capability means that a person can frame a task, use AI inside a known boundary, check the result against evidence, repair a failure, explain a trade-off, and adapt the method when the case changes. Attendance, confidence, and a polished first answer are useful signals, but they are not proof of that capability.
The Australian National AI Centre describes capability as something built through practical learning, shared habits, clear expectations, and regular review rather than a one-off event. Its guidance also recommends safe experiments, feedback loops, and clear use boundaries. See the team-capability activity.
Use these three labels in your report:
| Result | What it proves | What it does not prove |
|---|---|---|
| Completion | The participant submitted the requested fields and stayed within the exercise format. | That the reasoning is sound or transferable. |
| Supervised practice | The participant can perform the task with prompts, examples, or an expert nearby. | That they can handle a new failure alone. |
| Demonstrated capability | The participant produces a defensible artifact, repairs a failure, explains the trade-off, and adapts to changed inputs. | That they can act without domain-specific approval in high-risk work. |
This separation matters because transfer is not automatic. A meta-analysis of 89 empirical studies found that transfer relates to learner and work-environment factors, and that measuring training and outcome through the same source can inflate the relationship. Read the transfer review.
Run this role-specific assessment
Give the participant 45 minutes, access to the AI tool your training covered, and the fictional source cards below. Do not give them a model answer. Ask them to save the final brief and a short note showing what they checked independently.
Case: choose a small product change
You are a product manager for an invoicing product. Choose the smallest next experiment for a two-week sprint. You may use AI to organize and challenge your thinking, but every claim in the brief must trace to a source card or be marked as an assumption.
Fictional source cards:
- Three account administrators say opening invoices one by one makes a month-end export slow.
- Two finance leads ask for saved filters so they can review unpaid invoices weekly.
- Two users ask for dark mode, but neither describes a work failure caused by its absence.
- One account administrator reports that the current CSV sometimes drops timezone and account-ID fields.
- Two finance leads say a scheduled month-end export would be useful if the fields were reliable.
Produce a one-page decision brief with these fields:
- The user problem, written without adding facts.
- The recommended two-week experiment.
- Two source-card references supporting the recommendation.
- One unresolved uncertainty and the smallest next check.
- One trade-off, such as speed versus data reliability or narrow scope versus user coverage.
- The AI's role, your independent checks, and the point where a human owns the decision.
This task is intentionally ordinary. A capability check should resemble the decisions a role actually makes, while remaining safe enough to run with sanitized or fictional inputs. The EU AI Act's AI-literacy provision explicitly points to technical knowledge, experience, education, training, and the context in which an AI system is used. Read Article 4.
The National AI Centre makes the same practical move by asking teams to examine roles, tasks, confidence, uncertainty, and support needs before selecting capability actions. Use its training-needs activity.
Score the artifact, repair, and transfer separately
Do not hide the result inside one satisfaction score. Score the submitted brief out of 10, then use repair and transfer as gates.
This rubric evaluates the person's work, not the model's prose. For a separate guide to grading an AI output itself, see How to Grade an AI Output Against a Rubric.
| Artifact criterion | 0 points | 1 point | 2 points |
|---|---|---|---|
| Evidence traceability | Claims have no source cards or invent facts. | Some claims trace, but important support is unclear. | At least two important claims trace to cards and assumptions are marked. |
| Decision quality | Recommendation ignores the task or the strongest risk. | Recommendation is plausible but its scope or rationale is thin. | Recommendation fits the two-week constraint and explains why it wins now. |
| Verification | Treats AI output as correct. | Mentions checking but does not show what was checked. | Names concrete checks and removes or marks unsupported claims. |
| Trade-off reasoning | No trade-off or a slogan. | Names a trade-off without showing its effect. | Explains what is gained, what is sacrificed, and why the choice is acceptable. |
| Ownership and boundary | AI is treated as the decision-maker or sensitive use is assumed. | A human owner is mentioned but the boundary is vague. | AI's role, human ownership, and escalation point are explicit. |
Repair exercise, 3 points
Give the participant this flawed AI-generated brief after they submit the first one:
Build dark mode first because two users asked for it and it will improve retention. The export issue is a one-off. To validate the next version, paste real customer invoices into a public model and ask it to find the missing fields.
Ask them to mark each unsupported or unsafe claim, rewrite the recommendation, and name the smallest safe next check.
Award one point for each of these:
- They reject the retention claim and the “one-off” claim because neither follows from the cards.
- They notice that the input boundary is unsafe and replace real invoices with redacted or fictional examples approved for the tool.
- They repair the next step into a bounded check, such as reproducing the CSV field issue on sanitized records before choosing a larger export feature.
NIST's AI Risk Management Framework is useful here because it separates governing, mapping, measuring, and managing risk, and stresses context, roles, documentation, and human oversight. Use the NIST AI RMF Core as a governance reference. The participant does not need to recite NIST. They need to show the same kind of behavior in the case.
Transfer case, 3 points
Change the inputs and the role. You are now an operations lead deciding whether to improve a weekly internal request report or create a reusable response template. The new cards say:
- Four team leads spend time reconciling the report because two categories are often missing.
- Three coordinators answer the same onboarding question with slightly different instructions.
- One request contains a deadline that must be confirmed by a human before any response is sent.
- The team has no approval to send internal records to an external AI tool.
Ask for the same six-field brief, but do not repeat the invoice examples or tell the participant which risk matters. Award one point for each of these:
- The participant identifies the changed input boundary and keeps records out of an unapproved tool.
- The recommendation changes because the new case has a human-confirmation requirement or a different dominant bottleneck.
- The participant names a new, smallest next check instead of copying the first case's validation step.
Repeated testing can support transfer. In a four-experiment study, repeated testing produced better delayed retention and transfer than repeated studying. See Butler's study. In this packet, the transfer case is the test. It is not an optional extension for people who finish early.
Use thresholds that preserve the exception
For this packet, classify the result as demonstrated capability only when all conditions hold:
- artifact score is at least 8/10;
- repair score is at least 2/3;
- transfer score is at least 2/3;
- no critical boundary failure appears, such as inventing evidence, treating AI as the owner, or sending restricted data to an unapproved tool.
If the participant misses the artifact threshold but passes repair and transfer, call the result supervised practice and give feedback on the artifact. If they score well on the first brief but fail transfer, do not call them incapable. Call the capability unproven in changed conditions and repeat with a new case.
The trade-off is deliberate. A strict all-conditions gate produces fewer passes and more coaching work. A single average score produces easier reporting but can hide a dangerous boundary failure or a person who only learned the example. NIST's framework treats risk management as iterative and context-dependent, so the threshold should rise with the consequences of the task.
What the small pilot exposed
The desk pilot used three authored responses, not learners. Its aim was to see whether the packet and scoring guidance could distinguish completion from capability.
| Sample | What it did | Final score | Failure mode |
|---|---|---|---|
| A | Filled every field, chose dark mode, and repeated the AI's unsupported retention claim. | 4/10, 1/3 repair, 0/3 transfer | Polished completion without evidence discipline. |
| B | Traced the export problem and repaired the unsafe data suggestion, but copied the first case's reasoning into the changed case. | 8/10, 3/3 repair, 1/3 transfer | Correct method in the taught example, weak adaptation. |
| C | Chose a bounded export experiment, marked uncertainty, repaired the flawed brief, and changed the next check for the operations case. | 8/10, 3/3 repair, 2/3 transfer | Met the packet threshold. |
The first rubric was ambiguous in two places. “Uses evidence” allowed a correct recommendation with no traceable support. “Shows transfer” allowed a participant to reuse the same categories without responding to the changed constraint. I revised the criteria to require two source-card references, one explicit unresolved uncertainty, and one changed input that alters the next action. I also clarified that “needs more data” earns no uncertainty point unless the participant names the smallest next check.
The three authored sample responses, in condensed form, were:
- A: “Build dark mode first because two users asked for it; it will improve retention.” It filled the form but made unsupported claims.
- B: “Run a two-week reproduction test for the CSV field problem before committing to a larger export feature.” It traced evidence and repaired the unsafe suggestion, then reused the first case's logic in the transfer case.
- C: “Use sanitized records to reproduce the missing fields, then choose the smallest export reliability fix that the check supports.” It changed the next check when the operations case introduced a different boundary.
That is the useful evidence this page can own: a ready-to-run packet, its sample responses, and a documented revision made after applying the rubric. It is not evidence that three people became capable after training. For a real team, run the packet before and after training with parallel cases, then inspect a later work artifact with a manager or role expert.
When this assessment is not enough
Do not use this packet as the sole release gate for hiring, medical, legal, financial, employment, safety-critical, or production-change decisions. Replace the fictional source cards with an approved, sanitized case only after privacy, security, and authority boundaries are clear. Add the domain expert who owns the consequence.
Also avoid a generic test for a team with very different roles. A product manager, support lead, engineer, and finance analyst may need different artifact criteria even if they use the same model. The sourceable result here is the assessment shape, not a universal score.
If the training has already happened, the next step is small: give each role one fresh case, collect the artifact, run repair and transfer, and report the three labels separately. That tells you what attendance cannot: whether the team can use AI responsibly when the example is no longer doing the thinking for them.
Questions people ask next
Is attendance a useful measure of AI training?
Yes, but only as an exposure or completion measure. Pair it with a fresh work sample, failure repair, and transfer check before calling the team capable.
How many people should take a capability assessment?
Start with the roles and risk boundaries that matter, then assess enough people to expose disagreement and common failure modes. There is no universal pass count; calibrate the task with role experts.
Should the post-training task use a real business task?
Use a real task shape with sanitized or fictional inputs when privacy, security, or legal boundaries are uncertain. Move to live work only after permissions and human review are clear.