Field note · capability

What Should an AI Learner Demonstrate Before Working Without Supervision

Use an unseen transfer task, source checks, and explicit vetoes to decide whether an AI learner can continue alone, needs tutoring, or must escalate.

10 minute read
  • AI learning
  • Team capability
  • AI training
Illustration of an AI learner completing an independence audit before working alone

Course completion is not independence. A learner can produce a polished answer while still missing the source conflict, the privacy boundary, or the moment when somebody else must decide.

Before I reduce supervision, I want to see the learner handle a task they have not rehearsed step by step. They should make a useful attempt, show how they checked it, and stop cleanly when the task moves beyond their authority.

What should an AI learner demonstrate before working without supervision?

They should demonstrate independent framing, bounded workflow choice, source-based verification, clear documentation, and safe escalation on an unseen, low-risk task. The test should use a local rubric and explicit vetoes. It should not be treated as a universal certification score, and it should not remove qualified review from high-risk work.

Here is the worked result from the self-contained exercise behind this guide. It is an author-run artifact dated 2026-08-23, not a learner study.

CaseResultDecision
Guided practice11/12, no veto. The output preserved the uncertain metric, used a read-only workflow, and named the human approver.Continue alone for this bounded drafting task, with owner signoff for the pilot.
Unseen transfer5/12, with a privacy and authority veto. The output copied raw customer notes into an unapproved tool and did not block auto-send.Escalate to a qualified reviewer. Do not run the proposed workflow.

The useful finding is not the number 11 or 5. It is that the unseen case changed the decision. An output that looked independent in guided practice still had to prove that the method transferred when the data and risk boundary changed.

Illustration of a guided practice case flowing into an unseen transfer test and a continue-alone, tutor, or escalate decision

What counts as evidence of independent AI work?

Evidence is a visible work trail, not confidence or a fluent final answer. For one bounded task, the learner should leave an artifact that shows the goal, sources, assumptions, checks, corrections, and next owner.

The OECD and European Commission describe AI literacy as knowledge, skills, and attitudes that help learners understand AI, critically evaluate outputs, use it ethically and creatively, and make informed decisions about risks. That points to a stronger test than “can this person prompt a model?” (OECD and European Commission framework)

I use a simple distinction:

What happenedWhat it proves
The learner watched a demonstrationExposure to a method
The learner produced a good result with step-by-step helpAssisted performance
The learner produced, checked, explained, and transferred the workBounded evidence of independent capability

When I taught product managers who moved from writing specifications to building and shipping products, the recurring difficulty was often agreeing on what “done” meant. That is a teaching observation, not a measured result. It matters here because an undefined finish line lets an AI output look complete when nobody has decided what must be true. (how I help people learn AI)

The audit therefore asks for an artifact that another reviewer can inspect without replaying the entire learning session.

Which capabilities should the independence audit test?

Test six capabilities. They are concrete enough to observe and broad enough to transfer across tools.

CapabilityWhat the learner must showWhat to save
Frame the taskStates the decision, scope, constraints, and what the output is not decidingDated task brief
Choose the workflowUses AI for a bounded draft or analysis and keeps irreversible action outside the workflowShort workflow description
Preserve sourcesSeparates source facts, uncertainty, assumptions, and conflictsSource table or annotations
Verify the outputChecks claims against the source of truth and records a correction or rejectionVerification notes
Document the workLeaves a usable brief with tool use, assumptions, evidence, and next ownerFinal artifact and decision record
Escalate safelyRecognises privacy, authority, uncertainty, or risk boundaries and names the next reviewerEscalation note

These dimensions map cleanly to the major frameworks without pretending that any framework defines a workplace pass score. Stanford's framework uses functional, ethical, rhetorical, and pedagogical domains. Its progressive objectives include analyzing AI limits, critiquing outputs, choosing an appropriate level of collaboration, and keeping human agency in the work. (Stanford's AI literacy framework)

How do you run a reproducible transfer test?

Run one guided case, then change the task and remove the checklist. Keep the timebox and evidence requirements fixed. The second case is the decision point.

1. Write the task brief before opening the tool

Use a low-risk task with a reviewable output. State the role, goal, allowed tools, timebox, source-of-truth rule, privacy boundary, and escalation path.

For the exercise used here:

  • The role is an operations analyst preparing a one-page decision brief.
  • The learner may use one AI assistant, an editor, a calculator, and the supplied source pack.
  • No private or personal data may be pasted into an AI assistant.
  • The supplied notes and policy excerpt outrank the model.
  • The total timebox is 50 minutes, split between guided practice, unseen transfer, and the decision record.
  • No external action is allowed during the exercise.

The U.S. Department of Labor calls its AI Literacy Framework voluntary guidance for program design. It is designed to flex across audiences and roles, and it addresses prerequisites such as digital literacy. That is why this packet defines its own task conditions instead of presenting the rubric as a government standard. (U.S. Department of Labor release, accessible framework notice)

2. Give one guided practice case

The guided case teaches the shape of a good attempt without giving the answer to the transfer case.

In this packet, the source pack says a team spends about six hours per week turning meeting notes into action lists. A policy excerpt allows public or sanitised text, requires human review before distribution, and forbids personal data in unapproved tools. A metric note says the six-hour estimate came from one manager and has not been measured across a month.

The learner must recommend a two-week, read-only experiment. A good output keeps the estimate qualified, states the input boundary, requires human review, and names the owner of the final decision.

3. Give an unseen transfer case

Change the task, evidence, and risk boundary. Do not change the rubric.

The unseen case asks the learner to assess AI-assisted replies to customer feedback. The notes include customer email addresses, a policy allowing sanitised examples only, and a warning that the sample is too small to infer satisfaction. The brief asks whether the team should auto-send replies, but names no accountable approver.

The safe response is a sanitised, draft-only test. The learner must preserve the small-sample limitation, escalate the missing approver, and block any send action. Uploading the raw notes to an unapproved tool is a privacy veto.

ETS describes its progression as hypothesized and connects assessment to scaffolded instruction. Its task-design principles include relevance, reducing access barriers, and giving learners room to advance. A guided case followed by an unseen case is my application of that logic, not a procedure prescribed by ETS. (ETS AI-literacy report)

4. Score observable behaviour, not polish

Score each dimension 0, 1, or 2. A zero means absent or unsafe, one means partial or tutor-dependent, and two means clear, correct, and evidenced.

For this exercise only, 10/12 or higher with no dimension below 1 and no veto is a continue-alone signal for the bounded low-risk task. A score of 7-9/12 with no veto means targeted tutoring and another transfer case. A score of 0-6/12, or any veto, means qualified review or escalation.

Those thresholds are local decision rules. The cited frameworks establish capability themes, not a universal professional pass score.

What does a worked scoring sheet reveal?

The unseen output in this exercise recommended drafting customer replies but pasted raw notes, including email addresses, into an unapproved AI tool. It preserved the sample-size caveat, but it did not separate drafting from sending or identify an approver.

DimensionScoreEvidence
Frame the task1Identified reply drafting but did not separate drafting from sending.
Choose the workflow1Recognised draft-only in part, but used an unsafe input path.
Preserve sources1Kept the sample-size caveat but failed to carry the policy boundary into execution.
Verify the output1Checked one limitation but left no source-by-source verification.
Document the work1Produced a usable recommendation but omitted the accountable approver.
Escalate safely0Did not stop at the privacy or authority boundary.
Total5/12The privacy veto overrides the numeric score.

This is the kind of failure a completion certificate misses. The writing is plausible. The decision is unsafe. The right repair is not “try a better prompt.” It is to stop, remove the personal data, define who owns the decision, and repeat the transfer case with draft-only permissions.

How should a manager act on the result?

Use the result to choose the next intervention, not to label the person.

ResultNext moveWhat to test next
Continue alonePermit the bounded, low-risk task within stated permissionsA later transfer case and periodic review when the workflow changes
Targeted tutoringTeach the weakest dimension, then repeat a changed caseThe same capability under less prompting
Qualified review or escalationPause the workflow and involve the accountable owner or specialistWhether the risk, privacy, or authority boundary has been resolved

Australia's National AI Centre uses a team-readiness activity to surface gaps in skills, communication, oversight, and culture. It asks teams to identify where staff must review outputs, who owns final decisions, which tasks need stronger oversight, and where escalation pathways are unclear. That supports treating an independence decision as an operating agreement, not a private learner score. (National AI Centre readiness activity)

The principal exception is risk. A learner can pass a low-risk drafting case and still require supervision for a financial approval, employment decision, healthcare recommendation, legal conclusion, customer-facing send, production change, or task involving restricted data. Independence is always scoped to a task, permission set, and review chain.

If the team has not yet built a practice loop, How to Make AI Training Stick in a Small Team covers the recurring task, saved artifact, and next-use trigger that should come before this audit. If the learner needs help using AI during practice, start with How to Use AI to Learn a Technical Skill. The parent collection for this capability work is What to Learn Before Building AI Agents.

What should you do before removing supervision?

Give the learner a new task, a source pack, a timebox, and a safe boundary. Ask for the artifact and the decision record, not just the final answer. Continue alone only when the learner can show the work, verify it, and stop at the boundary without being prompted.

That is a small test. It is also a much stronger handoff than course completion.

If you are designing this kind of capability transfer for your own work, Marius Manolachi helps people build AI capability around the tasks they already own. The audit is complete without that next step, but it gives a tutoring or consulting conversation something concrete to inspect.

Questions people ask next

Does passing an independence audit mean the learner is fully competent?

No. It shows bounded evidence for one task family under stated conditions. Repeat the transfer test when the role, data, tools, authority, or risk changes, and keep qualified review for high-impact work.

Can one AI learner score be used for every work task?

No. The rubric can be reused, but the source pack, vetoes, reviewer, authority boundary, and pass rule should be adapted to the task. A low-risk drafting pass does not authorize unsupervised financial, legal, employment, healthcare, or production decisions.