Field note · capability

How to Run a Changed-Task Rehearsal After an AI Workshop

A bounded field note and worksheet for checking whether an AI workshop transfers to a changed task through explanation, prediction, verification, recovery, and judgment.

13 minute read
  • AI capability
  • AI training
  • workplace AI
Illustration of a changed-task rehearsal after an AI workshop

Most AI workshops end when the facilitator's example works. That shows that a learner can follow a path with help. It does not show that the learner can recognize the task, check the output, or recover when the next case is different.

I use a changed-task rehearsal for that decision. It is a small, low-risk exercise after the workshop in which the learner must explain the work before opening the AI tool, predict what could go wrong, inspect the result, recover from one controlled failure, and apply the same judgment to a new case. The point is not to create a pass rate. The point is to make the missing evidence visible.

Illustration of a changed-task rehearsal after an AI workshop

The result is a rehearsal, not a better workshop demo

The evidence supports one narrow conclusion: a facilitator should inspect a chain of decisions after the demonstration, not use demonstration smoothness as a proxy for transfer. The chain is useful because each step exposes a different failure.

CheckObservable resultDecision it supportsEvidence boundary
ExplainThe learner states the task, acceptance condition, and stop condition before AI use.Continue with a changed task if the work is understood.It does not prove the learner will repeat the judgment later.
PredictThe learner names an expected output or a likely failure before seeing the response.Compare expectation with the actual result.It does not measure model quality by itself.
VerifyThe learner points to a source, rule, test, or human check that could catch an error.Treat verification as part of the work, not an optional final glance.It does not establish a universal verification standard.
RecoverThe learner stops, revises, escalates, or seeks approval after one controlled failure.Check whether the person can operate within a safe boundary.It does not estimate recovery performance in production.
TransferThe learner handles a new case with one relevant condition changed.Decide whether the session produced evidence worth following up.It is local evidence about one learner and task.

This is the article's sourceable atom: the Changed-Task Rehearsal Sheet, with an explicit rule against turning one rehearsal into a population claim. The AI capability-building guide gives the wider capability context. This page supplies the post-workshop test.

The primary studies support the caution. A randomized study of 164 law students compared no GenAI access, optional access, and access with brief targeted training. The trained-access condition was associated with higher adoption and better issue-spotting performance than untrained access, but that does not answer whether a professional can transfer the practice to a changed work task. Chen and Bao report the study and its conditions.

Ask for an explanation before the AI is opened

Start with the work, not the prompt. Ask the learner to describe the task, the acceptable output, the first reason to stop, and who owns the final decision.

When I taught product managers who went from writing specs to building and shipping the product, the failure was often an undefined “done,” not the model. That observation changes the first rehearsal question. “What prompt would you use?” is too early. “What would make this output acceptable?” reveals whether the learner can name the work contract.

Use a low-risk task the learner already knows. It can be a draft, a classification, a comparison, or a small analysis. Keep the data fictional or non-sensitive if the task is being run in a group. Do not make the exercise depend on private workplace records.

Record the learner's explanation before showing a model response. If the explanation is vague, stop there and repair the task definition. A polished AI answer cannot compensate for an undefined acceptance condition.

The exception is a genuinely exploratory task with no agreed output yet. In that case, score the learner on whether they can propose a next question, a safe boundary, and a way to learn what “good” means. Do not score an exploratory answer as if it were a finished deliverable.

Ask for a prediction before showing the output

Prediction separates following a recipe from forming a testable expectation. Before the AI responds, ask what the output will contain and where it might fail.

The prediction can be short. It might identify a missing field, a likely ambiguity, a source the model cannot access, or a case that needs human approval. The important part is that the learner commits to a check before seeing the answer.

This matters because a working demo can conceal skipped evaluation. In my Udemy teaching, I have seen learners start at frameworks, skip evaluation, and confuse a working demo with a release. The approved Udemy record covers 109,753 students across four courses and 23,929 reviews, but those totals are teaching scale, not a study sample or a failure rate. The course profile is the source for the scale.

Score the prediction for specificity, not for being correct. A wrong but concrete prediction gives the learner and facilitator something to compare. A vague prediction such as “the model might hallucinate” does not.

The exception is a task where the learner cannot reasonably anticipate the content. Ask for a prediction about the process instead: what must be checked, what would trigger a stop, and what evidence would change the decision.

Illustration of a changed-task rehearsal with prediction and verification checks

Score verification as a work step

Verification passes only when the learner names what was checked, how it was checked, and what would happen if the check failed. Reading the output once is not enough.

Give the learner a response containing one planted, low-risk defect that the task's normal check could find. Do not use a hidden trick that tests trivia. The defect should connect to the work contract, such as a missing requirement, an unsupported claim, an incorrect field, or a calculation that does not reconcile.

The learner's verification record can use this compact format:

Claim or field checked:
Evidence or test used:
Observed result:
Action if the check fails:
Owner of the final decision:

Learning-centred GenAI research gives a useful external comparison. Zhao and colleagues describe task-focused use that includes planning, checking, and revision, then report prospective associations with later self-regulated learning and academic functioning in a student panel. That is relevant support for checking and revision as observable behaviours. It is not evidence that this rehearsal predicts professional workplace transfer. The panel study and its limits are available in the full article.

Use a separate score for verification omission. A learner can produce a good answer by luck and still fail the work requirement if they cannot show how they checked it. Conversely, an imperfect answer with a sound check may reveal a repairable model or task problem.

Add one controlled failure and an approval boundary

Recovery passes when the learner notices the failure, chooses a safe next action, and identifies when another person must decide. Do not simulate a dramatic incident. A small, visible defect is enough.

The controlled failure might be a stale input, a missing source, an ambiguous instruction, or a response that exceeds the task's authority. Ask the learner to say what they would do next before they retry. If the right action is “ask a person,” let that be a valid result. The rehearsal is testing judgment, not independence at any cost.

This check reflects a product constraint I face while building TryUncle, an AI agent that watches the screen and annotates it live. Latency and human approval are part of the product boundary. They cannot be left for a final discussion after the happy path works. TryUncle is the first-party reference for that product context.

A recovery result should record four things: the failure signal, the learner's first action, the approval boundary, and the final disposition. “The learner tried again” is incomplete. The facilitator needs to know whether the retry was justified, whether the state had changed, and who could stop the process.

The exception is a task with no meaningful failure injection. Use an uncertainty card instead. Give the learner a missing input or a conflicting instruction and score whether they pause, ask for clarification, or state an assumption before continuing.

Change one condition in the transfer task

The changed task should preserve the underlying judgment while changing one relevant condition. Change the input shape, priority, exception, audience, or source. Do not change everything at once.

The Orange workshop provides the right starting point for this design. It began with the work attendees already did, rather than with a generic agent demonstration. The workshop evidence is documented by Natalia Melniciuc. Start with familiar work, then alter one condition that matters to the decision.

Familiar taskChanged conditionTransfer question
Classify a set of requestsOne request is ambiguousDoes the learner ask for clarification or invent a label?
Draft a short responseThe source is missing one key factDoes the learner mark the gap before writing?
Compare two optionsThe priority changes from speed to auditabilityDoes the recommendation change for a stated reason?
Summarize a documentOne section conflicts with anotherDoes the learner surface the conflict and assign an owner?

The changed task is not a surprise exam. Tell the learner which dimension changed. The test is whether they can reuse the decision process, not whether they can guess the facilitator's trick.

A six-month randomized field experiment across 66 firms and 7,137 knowledge workers found changes mainly in independently adjustable work patterns. It did not identify which learning practice caused sustained independent use. That is why this rehearsal is a local diagnostic for a training decision, not a replacement for workplace telemetry or a causal study. The HBS working paper states the study scope.

Illustration of a changed-task transfer rehearsal with one altered work condition

Use the Changed-Task Rehearsal Sheet

The sheet is complete when every row has an observation and a next decision. A facilitator can run it in 20 to 30 minutes after a workshop, but that time range is a suggested format, not a measured result.

FieldPrompt to the learnerEvidence to retainStop or follow-up rule
Task contractWhat is the task, what counts as acceptable, and when do you stop?A one-sentence task definition and acceptance conditionStop if “done” cannot be stated without the facilitator supplying it.
PredictionWhat do you expect the AI to do, and where could it fail?A dated prediction before the output is shownFollow up if the prediction is only a generic warning.
VerificationWhat did you check, using which source, rule, or test?The check, result, and any defect foundDo not pass on output quality alone when the check is missing.
RecoveryWhat changed after the controlled failure, and who can approve the next action?Failure signal, first action, approval boundary, dispositionHold if the learner retries without explaining why.
Changed taskWhat condition changed, and what judgment stayed the same?Transfer decision with reasonRepeat with a smaller change if the task changed too much.

This is a procedure, not a personality test. Keep the learner's exact task response, the facilitator's observation, and the interpretation in separate columns. If the result is disputed, another reviewer should be able to see what happened without reconstructing the session from memory.

Use the sheet with How to Make AI Training Stick in a Small Team when you are planning reinforcement after the rehearsal. That page addresses the persistence problem. This one addresses the evidence you should collect at the handoff from workshop to practice.

Interpret the result without overclaiming

Treat a rehearsal as a decision aid. It can tell you what to reinforce for this learner on this task. It cannot tell you that one practice predicts independent AI use across a company.

Rehearsal patternAppropriate decisionInappropriate conclusion
Strong explanation, weak verificationTeach source checking and release criteria next.The learner lacks AI ability in general.
Strong verification, weak recoveryRehearse stop, escalation, and approval boundaries.The model or workshop alone caused the failure.
Strong familiar task, weak changed taskReduce the change and repeat with one condition.The learner only knows how to prompt.
Weak explanation from the startRepair task definition and domain context before another AI lesson.More prompting will solve the problem.
Strong across all five checks onceRecord a positive local result and schedule another observation.Training transfer is proven.

The European study of more than 36,600 workers across 35 countries examines adoption in relation to exposure, skills, job content, employee influence, digitalisation, and training provision. That context can help an L&D leader ask better rollout questions, but adoption correlates are not a changed-task rehearsal result. Henseke's study is the primary source for that scope.

Do not publish a percentage from a single workshop. Do not rank facilitators from one learner's result. If you need a rate, define a sample, preregister the scoring rules, retain failures, and run a real study. This field note does not contain one.

What we still do not know

We do not know whether the five checks predict later independent workplace use. We do not know how many changed-task rehearsals are needed before a result is stable. We do not know whether explanation, prediction, verification, recovery, or transfer carries the most information in a given role.

We also do not know how much the result depends on domain knowledge, task design, facilitator skill, or the model used in the session. The primary workplace research does not answer those questions, and the locked observations in this article cannot answer them either.

That uncertainty is useful. It prevents the rehearsal from becoming another completion badge. The next responsible step is to repeat the same sheet across a defined set of low-risk tasks, retain the raw observations, and report the limits with the result.

Illustration of a facilitator recording a changed-task rehearsal result with an explicit evidence boundary

Run the rehearsal after your next workshop

Choose one familiar, low-risk task. Write the acceptance and stop conditions before the session. Ask for the learner's prediction before showing an AI response. Add one visible defect, record the recovery decision, then change one condition and run the task again.

At the end, make one of three decisions: reinforce a missing check, repeat with a smaller task change, or schedule another observation. Keep the result attached to the task and learner context. If you want help turning the sheet into a capability program, learn about Marius Manolachi's AI tutoring and consulting work.

The goal is simple. Leave the workshop with evidence about what the learner can judge, not just evidence that the facilitator's example worked.