Field note · commercial

How to Design an AI Tutoring Engagement Around a Team Decision

Design AI tutoring around one decision the team must make alone, then test transfer with a baseline, exit test, and follow-up check.

9 minute read
  • AI tutoring
  • team capability
Illustration of a team reviewing evidence before making an AI pilot decision independently

The usual tutoring brief starts with hours: two workshops, four coaching calls, or a month of office hours. That makes the purchase easy to describe and the outcome hard to inspect.

The better starting point is one decision. What must the team decide without the tutor, using which evidence, under which human control?

This is a narrow child of AI consulting and tutoring decisions. For the broader question of making training stick, see how to make AI training stick in a small team.

Illustration of a baseline decision moving through tutoring, an independent exit test, and a transfer check

Start with the decision the team must own

Design the engagement around a decision with a named owner, a hard boundary, and a reversible outcome. Do not begin with a curriculum of prompts or tools.

The CDC evaluation guidance separates learning from learning transfer and recommends assessing both when possible. The Department of Labor's AI Literacy Framework is likewise designed to adapt across roles and contexts. Together, those sources point to a practical commercial rule: define capability at the level of work the team must perform, not at the level of content the tutor plans to cover.

Use this brief before you sell or buy tutoring:

FieldWhat to record
Decision ownerThe person accountable for the final call
Decision boundaryWhat the team may approve, reject, or defer, and what remains out of scope
InputsThe evidence cards, records, policies, and assumptions available to the team
Evidence standardWhat must be cited or verified before a decision can pass
Permitted AI roleWhat the AI may explain, retrieve, or draft
Human overrideWho can pause, reject, or escalate, and for which triggers
Exit testThe independent task, including what changes from practice
Transfer checkThe date, new workplace case, and record to inspect later

If you cannot fill these fields, you have a topic for tutoring, not yet an engagement design.

The worked simulation: a read-only pilot decision

The following is a simulated case, not a client result. It is deliberately small enough to reproduce.

The decision owner is an operations manager. The team must decide whether to approve a read-only pilot that drafts weekly inventory-exception summaries. The pilot cannot write to a system, send a message, or make the final decision. The manager may approve, reject, or request evidence.

The six evidence cards cover sample quality, missing fields, the data boundary, review capacity, policy constraints, and the expected operating change. The evidence standard is strict: cite the material cards, name a material risk, define a human control, and state what would reverse the decision.

The practice fixture is:

CardPractice valueChanged in independent re-run
E1 sample quality8 of 10 summaries identify the exception category; 2 miss supplier ETAUnchanged
E2 missing fields3 of 10 records lack supplier ETAUnchanged
E3 data boundarySupplier IDs and internal SKU; no customer contact dataOne record includes customer contact data
E4 review capacity18 summaries expected per week; manager can review 30Capacity falls to 12 per week
E5 policyRead-only, human approval, no outbound communicationUnchanged
E6 operating changeSummaries prioritize follow-up; no automatic actionUnchanged

These are simulation inputs, not observations from a real organization.

The human and AI boundary is explicit. The tutor may explain the rubric, ask questions, and expose missing evidence. It may not rank the options or supply the decision. The operations manager owns the call and must record any override or pause.

NIST's AI RMF calls for documented roles and responsibilities, defined human oversight, assessed operator proficiency, and an initial go/no-go decision after context is mapped. Its human-AI interaction guidance also warns that roles vary by configuration and that override rationale can be useful to collect. The simulation turns those governance ideas into fields a buyer can inspect. (NIST AI RMF Core; NIST Appendix C)

Run a baseline before tutoring

The baseline should be a real work sample with no tutor assistance. It tells you whether the team lacks knowledge, judgment, evidence discipline, or a control boundary.

Simulated baseline output: “Go. The sample looks accurate enough and the summaries should save the team time. The operations manager can spot-check the results each week.”

The baseline scores 4/10 under the rubric below. It names a decision and a reviewer, but it does not cite the evidence cards, apply the read-only boundary, test review capacity, define a pause rule, or state what would reverse the decision. This is a fixture label, not a measured learner result.

The failure is useful because it sounds plausible. “Accurate enough” is treated as a release gate even though the team has not shown that its human control can operate under the stated conditions.

Tutor the reasoning, not the answer

The intervention should change what the team can do alone. A compact sequence is:

  1. Ask the team to rewrite the decision boundary in one sentence.
  2. Make the team classify each evidence card as capability, risk, control, or unknown.
  3. Require a human-AI responsibility split, including the person who can pause the pilot.
  4. Introduce one failure case and ask what the team would record.
  5. Ask for the evidence that would change the decision.
  6. Remove the tutor and rerun the task with changed evidence.

This matches the DOL framework's emphasis on role- and context-sensitive design, hands-on work, and judgment about AI outputs. It also fits the OECD's description of AI literacy as knowledge, skills, and attitudes that support critical evaluation and informed decisions about opportunities and risks.

The tutor should not polish the team's baseline answer until it looks good. That would measure session performance. The point is to make the reasoning visible, practice the missing move, then test whether the move survives without help.

Use an independent exit test with changed evidence

The exit test should keep the decision type but change the evidence. Otherwise the team can repeat the tutor's example instead of demonstrating independent judgment.

In this simulation, two cards change: one record now includes customer contact data, and review capacity falls below expected volume.

Simulated independent output: “Request evidence before approval. The changed data card introduces customer contact data, and the reduced review capacity means the proposed human control cannot cover expected volume. The owner remains the operations manager. A read-only pilot could be reconsidered only after the data boundary is approved, volume is reduced or review capacity is restored, every output is reviewed, and the owner can pause the pilot. The workflow must not write, send, or decide on its own.”

The exit output scores 9/10 and passes. It cites the changed evidence, preserves the boundary, names the owner, connects risk to control failure, and states reversible conditions. It loses one point because this simulation uses one scorer. A real engagement should use two scorers or an adjudication log.

Make the gate executable

The brief becomes easier to reuse when its decision rules are written as configuration instead of left in a facilitator's memory. This is the complete configuration for the simulation:

decision_gate:
  owner: operations_manager
  permitted_ai_role:
    - explain_rubric
    - ask_questions
    - expose_missing_evidence
  prohibited_ai_actions:
    - choose_decision
    - write_to_system
    - send_message
    - bypass_human_review
  required_evidence_cards: [E1, E2, E3, E4, E5, E6]
  pass_score: 8
  critical_miss_veto:
    - sensitive_data_exposure
    - missing_human_owner
    - contradictory_policy
  transfer_check: 2026-09-06

Test method

Load the six practice cards and score the baseline with the five-dimension rubric. Then replace E3 with the customer-contact-data variant and E4 with the 12-per-week review-capacity variant. Remove the tutor, require a fresh go, no-go, or request-evidence decision, and apply the same rubric. Fail any approval that triggers a critical-miss veto. Record the score, decision, changed cards, owner, and reversal conditions.

Observed output from the simulation fixture

baseline: 4/10, fail
independent_exit: 9/10, pass
critical_miss_veto: false
transfer: not_measurable

These outputs are the reproducible simulation's fixture results, not a client or learner outcome. A real engagement should replace them with scored team work and retain the decision record.

Score the decision, then decide whether tutoring is done

Dimension012
BoundaryTreats the pilot as autonomousMentions a limitApplies read-only, no-send, no-write limits
EvidenceNo usable evidenceSome cardsAll material cards and unknowns
Risk and controlMisses material riskNames one sideConnects risk to control and stop condition
Decision qualityUnsupported yes/noVague conditionClear decision with reversal conditions
Ownership and overrideNo owner or overrideNames oneNames owner, triggers, and record

Pass at 8/10 or higher, but add a critical-miss veto. Any approval that ignores sensitive-data exposure, lacks a human owner, or contradicts a policy fails regardless of the total. This prevents a fluent answer from passing by accumulating points in easier categories.

The commercial implication is simple: buy enough tutoring to close the specific gap shown by the baseline, then stop when the independent gate passes. Do not buy hours as the outcome.

Illustration of a human decision owner and AI tutor separated by an explicit responsibility boundary

Check workplace transfer after the engagement

An exit test proves performance on a changed task. It does not prove that the team will use the skill at work. CDC calls delayed follow-up the best way to assess learning transfer after learners have had time to apply what they learned. (CDC delayed evaluation guidance)

For a real engagement, schedule the check before the final tutoring session. On 2026-09-06, ask the decision owner to use the same brief on a new decision. Inspect the decision record, evidence citations, override rationale, and whether the team stayed inside the boundary. Score the new case with the same rubric.

This simulation has no transfer result because no real team participated. That limitation belongs in the deliverable. A dated procedure is useful; a made-up follow-up outcome is not.

When the team should not decide alone

“The team must make the decision alone” should mean without the tutor, not without governance. A high-risk decision may still require a qualified reviewer, privacy or security review, procurement approval, or an executive go/no-go. NIST's role and oversight guidance supports keeping those responsibilities explicit.

The exception is especially important when the decision can affect people, expose sensitive data, commit money, or trigger an irreversible action. In those cases, tutoring can test whether the team recognizes the need to escalate. Passing the exit test does not grant permission to remove the human veto.

If the buyer cannot name the owner, boundary, evidence standard, independent test, and delayed check, the engagement is not ready. Fix the brief first. Then decide whether tutoring is the right purchase. Marius Manolachi offers AI consulting and AI tutoring for people who need to build capability on their own work, so this brief is also a useful acceptance test to bring into that conversation.

Questions people ask next

Should the tutor give the team the correct decision?

No. The tutor can teach the rubric, ask questions, and challenge omissions, but the team must produce the decision during the independent exit test.

What if the team passes the exit test but the decision is high risk?

Keep the qualified human decision owner and any required review or veto. A learning pass proves capability on the test, not permission to remove governance.