Field note · capability
What Evidence Should AI Training Collect for Workplace Readiness?
A five-source audit defines the evidence packet an AI training program should collect before claiming workplace readiness for an individual.

Training completion tells you that someone finished training. It does not tell you whether they can choose a sensible task, check an AI output, or stop when the evidence is thin. That distinction matters when a manager has to decide whether a new capability is ready for ordinary work.
This article reports a bounded source audit, not a participant study. I reviewed five primary sources and turned the gap into a reusable evidence packet. The result is a decision about what to collect before making a workplace-readiness claim.
What did the five-source audit actually observe?
The observed result is a gap across the five sources: each defines or measures one part of the decision, but none supplies one complete individual-readiness packet. Use them to design an assessment, not as proof that a learner can work independently.
| Source | What the source contributes | What it does not establish for an individual learner | Use in a training review |
|---|---|---|---|
| U.S. Department of Labor AI Literacy Framework | Practical and contextual learning, output evaluation, responsible use, and continued learning | A participant-level task result or independent-completion threshold | Turn broad outcomes into observable task criteria DOL (https://www.dol.gov/newsroom/releases/eta/eta20260213) |
| Skills England foundation-skills benchmark | Technical, non-technical, and responsible-use skills, including checking outputs and spotting errors | A repeat-task result showing transfer into ordinary work | Select the verification behaviors to observe Skills England (https://www.gov.uk/government/publications/ai-foundation-skills-for-work/ai-foundation-skills-for-work-benchmark) |
| Alan Turing Institute AI skills framework | Knowledge, skills, aptitudes, behaviors, competencies, and proficiency levels | A scored workplace task with a participant-independent decision rule | Map the rubric to a proficiency description Turing Institute (https://www.turing.ac.uk/skills/collaborate/ai-skills-business-framework) |
| Task-oriented AI literacy paper | Evidence that a contextual scenario can measure applied literacy more closely than generic tests in its US Navy robotics training context | General evidence for ordinary cross-role workplace readiness | Borrow the scenario principle, then test your own population paper (https://arxiv.org/abs/2511.05475) |
| U.S. Census working paper | Firm, business-function, and worker-task diffusion measures | Individual readiness, independent completion, or a training threshold | Keep adoption claims separate from capability claims Census (https://www.census.gov/library/working-papers/2026/adrm/CES-WP-26-25.html) |
That is the sourceable result: the evidence categories are complementary, not interchangeable. A framework can tell you what to teach. An applied task can show what someone did. Adoption data can show where use appears. None of those alone proves that a person can use AI independently at work.

How was the audit conducted?
I sampled five primary sources already selected for this topic: three official competency or workforce sources, one original research paper, and one official statistical working paper. The access date for each source was 2026-08-24.
The unit of analysis was the source, not a learner. For each source, I recorded four questions:
- Does it define a capability or behavior worth teaching?
- Does it require a contextual task rather than a decontextualized quiz?
- Does it report an observable individual result with a scoring rule?
- Does it show independent workplace completion or transfer?
The audit used the source descriptions and claims in the primary documents. It did not pool their populations, convert framework language into scores, or estimate an effect. The result table therefore reports what each source can support and where it stops.
This method is deliberately small. It answers the reader's immediate decision, which is whether a program has enough evidence to make a readiness claim. It does not answer how common readiness is in the workforce.
Why does a capability framework not prove readiness?
A capability framework names a target. Readiness requires an observed performance against that target. The difference is the difference between saying “verify AI output” and showing a learner find a planted ambiguity, explain the check, and change or reject the output.
The Department of Labor framework emphasizes practical, contextual learning and evaluation. Skills England includes checking AI outputs for accuracy and spotting errors. The Turing Institute describes proficiency through skills and behaviors. Those are useful design inputs, but they do not become participant results merely because they are specific. DOL, Skills England, Turing Institute
The exception is a program that uses a framework as the rubric's source. That is sound. The program should still publish the task, scoring rule, observed output, and decision boundary separately.
What should an evidence packet contain?
A decision-ready packet needs five linked artifacts, each answering a different question. If one is missing, label the result as training evidence or a pilot observation, not workplace readiness.
| Artifact | Question it answers | Minimum contents |
|---|---|---|
| Task brief and reference pack | What did the learner have to do? | Work context, allowed tool, source material, ambiguity or trap, time limit, and version |
| Frozen rubric | What counts as good enough? | Observable criteria, critical-failure rules, and who can score without seeing the learner's prompts |
| Learner output | What did the person actually produce? | Original output, relevant prompt or interaction record, and the input version |
| Verification record | Did the person check the result? | Checks performed, evidence used, corrections made, and unresolved uncertainty |
| Stop or escalation record | Did the person know when not to proceed? | Refusal, escalation, or approval decision and the evidence that triggered it |
This packet is the reusable artifact from the audit. It keeps a training team from treating a completion certificate, a prompt transcript, or a polished demo as a substitute for work evidence. Store the versions together so a later reviewer can reproduce the decision.
How should a team score the task without rewarding prompt style?
Score the work and the judgment around it, not how sophisticated the prompt sounds. A participant can write a long prompt and still miss the task's risk. The rubric should make that failure visible.
| Dimension | Pass evidence | Fail or hold evidence |
|---|---|---|
| Task and tool selection | Chooses a task within the stated scope and a tool that can handle the work | Chooses an unsuitable task, tool, or permission level |
| Usable output | Produces an output that meets the task criteria and identifies missing inputs | Produces fluent text that misses a required condition or invents support |
| Independent verification | Checks the output against the reference pack or explicit criteria and records a correction or confirmation | Accepts the answer without a relevant check, or cites a check that did not happen |
| Stop or escalation | Stops, refuses, or asks for review when the reference pack cannot support the next step | Continues through an ambiguity or hides uncertainty |
The scorer should not need the participant's exact prompt wording to judge the four dimensions. Prompt records can help explain a failure, but they are not the outcome. The participant-independent rule also reduces the risk that a facilitator quietly rescues the task and then reports the rescued result.
This is an assessment artifact, not a measured result from the five-source audit. Freeze it before scoring. If the rubric changes after seeing outputs, record the change and treat the earlier and later results as different versions.

When can a program claim workplace readiness?
Make the claim only when the learner completes a representative task with the evidence packet intact, meets the frozen rubric, and shows the same judgment without facilitator prompting. A single polished answer is not enough.
Use this decision rule:
- Ready to test: the task is representative, the tool and reference pack are versioned, and the rubric is frozen.
- Observed capability: the learner produces a usable output, verifies it independently, and records a justified stop or escalation when needed.
- Transfer evidence: the same pattern appears on a second task or in the learner's ordinary work, with the context recorded.
- Bounded claim: the result states the task, population, tool version, date, scoring rule, and what the test cannot generalize to.
If only the first two conditions exist, say “passed a controlled task assessment.” If training completion or self-reported confidence is all you have, say exactly that. Do not call it independent workplace use.
The exception is a low-risk internal practice exercise where the purpose is learning rather than authorization. You can use a lighter packet, but label it as practice evidence and keep it separate from a release or role-readiness decision.
What does adoption data tell you, and what does it not tell you?
Adoption data tells you that firms, functions, or workers are using AI in some way. It does not tell you that a named learner can complete a task without help. The Census working paper is useful precisely because it separates firm use, business functions, and worker-task integration. Census
Keep the two questions separate in a program review:
- Capability question: can this person perform and verify the task under the stated conditions?
- Adoption question: is this task or tool being used in the organization, and under what access and policy conditions?
A high adoption signal can coexist with weak independent judgment. A strong task result can also fail to transfer if the worker lacks access, time, permission, or a suitable workflow. That is why the evidence packet includes context and limits rather than only a score.
What did this audit still not establish?
It did not establish a universal readiness threshold, a predictive relationship, a causal training effect, or a participant result. The sample was five sources, not workplace learners. The task-oriented paper was set in a US Navy robotics training context, so its scenario-based finding should not be generalized to ordinary cross-role work without a new test. The frameworks are guidance and competency models, not validation studies.
The audit also did not compare model versions, prompt volumes, facilitator interventions, or repeat-task performance. Those are the next measurements if a program wants to claim transfer. A future study should freeze the tool and model version, recruit a documented sample across roles, archive anonymized outputs, log facilitator prompts, and report bounded results with failure cases.
The honest unknown is the one a manager most wants answered: how much evidence is enough for a particular role? This article gives the packet and the decision boundary. It does not pretend that five documents can set that threshold.

How should a capability program use the result now?
Start with one low-risk task that already exists in the team's work. Assemble the brief, reference pack, and rubric before training. Run the task without facilitator rescue, preserve the output and verification record, then repeat it after the intervention. Review the stop or escalation decision separately from fluency.
That sequence gives a manager a concrete next action without pretending the result is more general than it is. It also gives a learner feedback they can use: which part failed, what evidence was missing, and what decision should have stopped the run.
Marius Manolachi's service is built around making existing people capable of building AI products on their own work. If your next step is turning a real work task into a safe practice and assessment packet, learn more about AI tutoring and consulting. For the broader capability-building sequence, start with AI capability building. For a concrete scoring artifact, see how to grade an AI output against a rubric.