Field note · implementation

Why AI Fails When Operators Inherit a Workflow Without Its Assumptions

A six-case replay shows how an inherited SOP creates unsafe decisions, and how a compact handover record changes the next action.

11 minute read
  • AI implementation
  • workflow reliability
Illustration of an inherited workflow splitting into unsafe and safe operator paths

When I teach product managers to move from writing specs to building and shipping, the difficult part is rarely the first successful run. It’s the moment another person has to own the work without the original operator in the room. The procedure survives. The reasoning behind it doesn’t. That is the capability gap I teach around AI work.

Illustration of a paired workflow replay with an inherited SOP and a compact handover record

The replay result: assumptions changed the next action

In my six-case replay, the inherited SOP produced one correct first action out of six. Five proposals were unsafe or outside the operator’s permissions. The same cases with a compact handover record produced six correct first actions, zero unsafe proposals, and no clarification questions.

ConditionCorrect first actionsClarification questionsUnsafe or unauthorized proposalsMean simulated effort to accepted ownership
Inherited SOP only1/645185 seconds
SOP plus compact handover record6/60030 seconds

The time column needs a careful label. It is not human wall-clock time. I used a declared replay-effort model: 30 seconds for a case scan, 45 seconds per clarification, 60 seconds per unsafe proposal, and 90 seconds to recover after an incorrect first action. Under that model, the handover condition saved 155 simulated seconds per six-case replay.

This is not a model benchmark. No LLM or external agent endpoint was called. It is a transparent operator simulation that keeps the case data, permissions, and tools constant while changing the information carried into the handover. The result is bounded, but it gives you a reproducible way to ask a better first question: what did the operator not know?

Why an SOP is not a workflow handover

An SOP tells someone what usually happens. A handover must also tell them what is already true, which evidence is authoritative, what they may change, and when they must stop.

The distinction matters in software too. OpenAI’s Agents SDK describes a handoff as a delegation between agents, represented as a tool call. It supports structured handoff input and filtering of the history passed to the next agent. That solves the transfer mechanism. It does not, by itself, decide which source wins, who can approve, or what recovery looks like. OpenAI’s handoff documentation is useful precisely because it makes the transfer boundary explicit.

The hidden-assumptions research makes the safety issue clearer. Undocumented assumptions can be critical to correct and safe operation, or constrain where a system applies. That is the same shape as an inherited workflow: the next operator sees the visible steps but not the conditions that made those steps valid. The paper by Harel and colleagues treats discovering those assumptions as a methodical engineering problem, not as a matter of asking people to “use common sense.”

So the diagnosis is not “AI needs a better prompt” until you have checked the transfer record. A prompt can repeat a step. It cannot grant authority, choose the source of truth, or restore an exception path that nobody documented.

The sanitized fixture and its hidden assumptions

The fixture is a synthetic purchase-exception review queue. The operator reads a request, checks a policy and ledger, then routes a recommendation. No customer, vendor, or private record is involved.

The current SOP shown to the inherited operator was:

Check the request, compare the amount with the policy threshold, approve compliant requests, and escalate exceptions.

The available tools were read_queue, read_policy, read_ledger, add_internal_note, and route_to_queue. The operator could read the queue, policy, and ledger, and could write an internal note or route a case. The operator could not approve, move a request to fulfillment, edit the ledger, or send an external email.

The omitted assumptions were:

  • The ledger amount is authoritative when it conflicts with the request form.
  • The current dated policy governs, rather than the threshold printed in an old SOP.
  • The operator prepares and routes a decision. Finance Ops approves it.
  • A duplicate case stops repeated fulfillment.
  • Missing verification is held and routed internally, not sent to an external requester.

Those are not cosmetic details. Each one changes the valid next action.

The six cases that reproduced the failure

The cases were designed to isolate different assumption types while keeping the workflow small enough to rerun by hand.

CaseInput stateAssumption under testCorrect next action
C1EUR 1,200, verified vendor, clean recordOperator routes a recommendation; Finance Ops approvesAdd note and route to finance approval
C2EUR 12,000, verified vendor, clean recordAbove-threshold work enters the approval queueAdd note and route to finance approval
C3Form says EUR 4,900; ledger says EUR 5,400Ledger wins a source conflictHold, note discrepancy, route to evidence review
C4Dated policy P-2026-07 sets EUR 5,000; printed SOP says EUR 10,000; request is EUR 5,500The current policy version governsHold, cite P-2026-07, route to finance approval
C5Queue record duplicates approved case Q-099Duplicate records stop repeated fulfillmentStop, link the duplicate, close without fulfillment
C6Vendor verification is missingMissing evidence follows an internal recovery routeHold, note missing verification, route to evidence review

The inherited operator instruction allowed one clarification but required a provisional next action without waiting. When a hidden assumption mattered, the simulator asked about it and then continued with the closest interpretation of the submitted form or visible SOP. That is intentionally uncomfortable. Real handovers often create the same pressure: keep the queue moving, even though the decision rule is incomplete.

The handover record that repaired the replay

The repair was not a longer SOP. It was a compact decision artifact with fields that answer the operator’s next questions.

FieldRecord used in the replay
ObjectiveDecide the safe next state for the queued case
Accepted state and evidenceRequest, policy version, ledger status, vendor verification, and duplicate status are checked and referenced
AssumptionsLedger wins conflicts; current dated policy governs; operator routes only; duplicates stop; missing evidence is held internally
Decision rightsFinance Ops approves; the operator prepares, annotates, and routes
Allowed toolsRead queue, policy, ledger; add internal note; route to a named queue
Forbidden actionsApprove, move to fulfillment, edit ledger, or send an external email
Acceptance criteriaThe first action matches the case rule, stays within permission, and names the destination or stopping state
RecoveryConflicting values go to evidence review; missing verification goes to evidence review; duplicate goes to stop and link
Destination and stopping ruleRoute to finance approval or evidence review; stop when authority, evidence, or state is insufficient

This shape matches the useful part of current handoff guidance. Saltwater Interactive describes a real handoff as transferring completion state, evidence, the next decision, the owner of the clock, and the return path, not merely sending a notification. Its workflow handoff guide is a good comparison input, but the replay above is the evidence for this fixture.

StackOS makes a similar distinction between a handoff message and an execution packet. Its practical fields include accepted state and evidence, bounded context, authority and tools, acceptance criteria, recovery, destination, and a stopping rule. That field list informed the record here. The result does not prove that StackOS’s list is universally required. It shows what happened when these fields were made explicit in this bounded case set.

Raw results: what changed case by case

The raw replay keeps the failure visible instead of collapsing it into a pass rate.

CaseInherited first actionQuestionsUnsafeEffortHandover first actionQuestionsUnsafeEffort
C1approve_and_move01180route_to_finance_approval0030
C2route_to_finance_approval0030route_to_finance_approval0030
C3approve_using_request_amount11225hold_and_route_evidence_discrepancy0030
C4approve_using_printed_sop_threshold11225hold_and_route_current_policy0030
C5approve_and_move11225stop_and_link_duplicate0030
C6send_external_email_for_verification11225hold_and_route_missing_evidence0030
Total / mean1/6 correct451,110 / 1856/6 correct00180 / 30

The inherited SOP failed in five different ways, but the failures were not five different model problems:

  • C1 exposed decision authority.
  • C3 exposed a source-of-truth assumption.
  • C4 exposed policy versioning.
  • C5 exposed state and duplicate handling.
  • C6 exposed a recovery path and an external-action boundary.

The same classification is a useful repair queue. Fix the missing decision record first. Only call it a data, policy, tool, or model failure if the complete handover still fails.

How to tell an assumption failure from a model failure

Use a two-pass replay before changing the model.

  1. Freeze the case set, input records, tools, permissions, policy version, and success definition.
  2. Run the inherited-SOP condition and record the first action, questions, unsafe proposals, and acceptance time.
  3. Add only the missing assumptions, evidence, authority, acceptance criteria, and recovery steps.
  4. Run the same cases again.
  5. Classify the remaining failures. If the result changes, the handover was part of the failure. If it does not, inspect data, policy, tool behavior, or model behavior.

NIST’s AI Risk Management Framework puts governance across the AI lifecycle and says documentation can improve transparency, human review, and accountability. It also calls for documented roles and responsibilities. The NIST AI RMF Core supports the control logic here: a person cannot safely own a decision when the workflow does not tell them what they are responsible for or what evidence governs it.

This is also where a handover test complements an evaluation. An output score can tell you that a recommendation looks plausible. It cannot tell you whether the person receiving it was allowed to act, whether they used the authoritative record, or whether the workflow had a safe place to return an exception.

For a broader implementation sequence, use the AI workflow implementation guide. For a general packet checklist, see what an AI workflow handoff packet should contain. If you need a release gate after the packet is written, use how to score an AI workflow handoff before release.

What to repair before tuning prompts

Repair the smallest missing decision record that changes the next action.

If the operator...Repair this firstDo not do this first
Chooses the wrong record when values conflictName the authoritative source and preserve the conflicting valueAdd more examples of the same prompt
Uses an old thresholdCarry policy version and effective date with the caseIncrease model temperature or add a longer system prompt
Proposes an action outside its roleWrite allowed, forbidden, and approval-owner fieldsGrant broader permissions to make the flow “work”
Repeats a completed actionDefine duplicate identity and stop behaviorAdd retries
Contacts someone when evidence is missingDefine the internal recovery queue and external-contact authorityLet the operator improvise a message

When the record is complete and the failure remains, the next investigation is more specific. Check whether the input is wrong, the policy is contradictory, the tool returns stale state, the permission check is broken, or the model still selects the wrong action. The handover test reduces the number of plausible causes before you spend time tuning.

The practical stopping rule

Do not accept an inherited workflow because the new operator can repeat its happy path. Accept it when the operator can choose the correct next action on a changed case, stay inside decision rights, and recover when evidence or authority is missing.

The replay passes that bar only for the compact handover condition, and only for these six synthetic cases. It does not prove production readiness. It gives you a small artifact to run before you make a larger claim.

If you are handing over an AI workflow now, copy the fixture shape, write down the assumptions you normally explain aloud, and add one case for every place those assumptions change the next action. Then run the same cases twice. That is a better handover conversation than “the demo worked.”

For readers who want help turning a real workflow into a safe practice artifact, Marius Manolachi’s AI consulting and tutoring work is the next step. The page should leave you able to run the first replay yourself.

Questions people ask next

How can I tell whether the handover or the model is failing?

Replay the same cases with the source-of-truth rule, policy version, decision rights, acceptance criteria, and recovery path made explicit. If the result changes, repair the handover first. If it does not, inspect data, policy, tools, or model behavior.

What should an AI workflow handover record contain?

Record the current objective, accepted state and evidence, assumptions, policies and versions, allowed and forbidden actions, decision rights, output and acceptance criteria, exception routes, recovery steps, destination, and stopping rule.