Field note · architecture

How to Map Decision Boundaries Before Designing an AI Workflow

A worked decision-boundary map and fail-closed validator for deciding what an AI workflow may observe, propose, route, or execute.

9 minute read
  • AI architecture
  • AI workflows
  • Human oversight
Illustration of an AI workflow decision map connecting a trigger, authority boundary, approval checkpoint, recovery path, and terminal state

The architecture discussion usually starts too late. A team has already chosen an agent, a workflow engine, or a vendor before it has agreed who may make each decision.

I start with the decision boundary. In my work teaching product managers to move from writing specs to building and shipping products, the difficult question is often not “Which model should we use?” It is “What proves that this step is done, and who is allowed to say so?” That teaching context is part of Marius Manolachi's AI learning work, not a production benchmark.

Here is the artifact I would put in front of the team first: a completed map for a small internal workflow, followed by a validator that rejects four common silent gaps.

The result: a decision-boundary map that fails closed

Use a decision-boundary map before selecting workflow components. It should trace one workflow from trigger to terminal state and make every authority, effect, and stop condition visible. This applies NIST's context-and-governance emphasis to a smaller pre-build artifact, rather than pretending the table is a NIST checklist (NIST AI RMF Core, NIST Appendix C).

The representative workflow is deliberately narrow: an internal product-feedback submission becomes a reviewed Draft issue in an internal backlog. It is non-regulated. The AI cannot publish work, change priority autonomously, change permissions, or contact an external party.

The base map passed validation. Removing an owner, evidence, approval scope, or terminal state caused rejection every time.

Illustration of a completed decision-boundary table with owner, AI permission, evidence, approval, escalation, recovery, and terminal state columns

The worked map from trigger to terminal state

The map below is the reusable artifact. The owner is an accountable human role, not the model and not the workflow engine. “AI permission” describes the maximum authority for that row.

Step and decisionDecision ownerAI permissionExact side effectEvidence shown to reviewerApproval expiry or scopeEscalation pathRollback or recoveryAudit record
intake: accept a submissionProduct operations leadObserveCreate an append-only run recordAuthenticated submitter, input reference, received timeNoneMissing requester or malformed payload -> escalatedReplay by run_id after payload repairintake.accepted with actor and input hash
classify: choose category and draft priorityProduct operations leadProposeStore proposal in run state, with no backlog writeRequest text, source IDs, confidence, ambiguity flagsProposal only, no executionAmbiguous category or no source -> needs_clarificationReplace proposal and preserve its prior versionclassification.proposed with model and prompt versions
route: select a queueProduct operations leadRouteSelect one allowlisted backlog queueCategory, queue-policy version, matched queuerun_id; routing onlyNo unique queue -> escalatedReroute before issue creationroute.selected
draft_issue: create a reviewed draftProduct operations leadAct with approvalCreate one issue in Draft status in the internal backlogSource reference, title, body, labels, duplicate candidates, priority rationaleExact payload digest, project, and run_id; 30 minutes; single useMissing, expired, or changed approval -> escalatedCancel or delete only the created draft by issue_idissue.create.attempt, approved, or rejected
apply_label: add a low-impact labelProduct operations leadAct autonomouslyApply one allowlisted non-priority label to the created draftAllowlist version, issue_id, labelrun_id and issue_id; no human approvalLabel not allowlisted or issue changed -> escalatedRemove the label by issue_idlabel.added or label.removed
notify: tell the internal ownerProduct operations leadAct autonomouslySend one internal notification containing the issue linkCreated issue_id, canonical link, notification keyrun_id and issue_id; idempotent keyDelivery failure -> failedRetry with the same notification keynotification.sent or notification.failed
finalize: decide whether the run is completeProduct operations leadObserveSet exactly one terminal run stateStep statuses, issue_id if created, escalation reason if notNoneIncomplete path -> escalatedReplay from the last non-terminal steprun.terminal

The normal path is:

received -> proposed -> routed -> awaiting approval -> draft_created -> labeled -> notified -> completed

The terminal exception states are needs_clarification, escalated, rejected, and failed. draft_created is an intermediate state, not a claim that the whole workflow finished.

The exact values are example judgments. A real team must replace “Product operations lead,” the 30-minute expiry, the allowlisted labels, and the recovery operations with controls it can actually enforce.

How to choose the AI permission for each decision

Classify the decision before you classify the architecture. The question is not whether the workflow is an “agent.” The question is what effect this one row may cause.

PermissionUse it whenDo not use it when
ObserveThe AI may read bounded context or record a run factReading the source itself crosses an unresolved data boundary
ProposeA person or deterministic policy must decide whether the suggestion is acceptableThe proposal could be mistaken for an executed action
RouteThe AI may choose from a fixed set of destinationsThe route changes authority, policy, or a consequential record
Act with approvalThe side effect is consequential but can be made exact, reviewed, and recoveredThe reviewer cannot see the actual payload or approval cannot expire
Act autonomouslyThe effect is low consequence, allowlisted, idempotent, and reversibleIt affects money, permissions, external recipients, sensitive data, or an unclear target

This separation follows the direction of the primary guidance. NIST says human roles and responsibilities should be clearly defined and differentiated, and that human-AI configurations range from fully manual to fully autonomous (NIST Appendix C). Google Cloud recommends a human checkpoint for subjective judgment or final approval of critical actions, while noting that predictable, structured work may not need an agentic solution at all (Google Cloud's design-pattern guide).

For the approval row, show the exact normalized action, not a transcript. Microsoft models this as a typed request and response: the workflow pauses, emits an external request, and resumes after a response. Its checkpoint guidance also preserves pending requests when work is restored (Microsoft Agent Framework human-in-the-loop).

The validator that catches silent gaps

The validator does not judge whether the workflow is wise. It checks whether the map is complete enough to route safely. Run it before discussing an agent framework or granting a write tool.

function validate(map) {
  const errors = [];
  const required = ["owner", "evidence", "approvalScope", "sideEffect", "escalation", "recovery", "audit"];

  if (!map.terminalStates?.length) errors.push("terminal handling missing");

  for (const row of map.rows) {
    for (const field of required) {
      if (!row[field]) errors.push(`${row.id}: ${field} missing`);
    }
    if (row.mode === "act_with_approval" && !row.approvalScope) {
      errors.push(`${row.id}: approval scope required`);
    }
    if (row.mode === "act_with_approval" && !row.escalation.includes("approval")) {
      errors.push(`${row.id}: approval escalation required`);
    }
  }
  return errors;
}

const mutations = [
  ["missing owner", m => delete m.rows.find(r => r.id === "draft_issue").owner],
  ["missing evidence", m => delete m.rows.find(r => r.id === "classify").evidence],
  ["missing approval scope", m => delete m.rows.find(r => r.id === "draft_issue").approvalScope],
  ["missing terminal state", m => { m.terminalStates = []; }]
];

const mapFromTheTable = {
  terminalStates: ["completed", "needs_clarification", "escalated", "rejected", "failed"],
  rows: [
    ["intake", "observe"], ["classify", "propose"], ["route", "route"],
    ["draft_issue", "act_with_approval"], ["apply_label", "act_autonomously"],
    ["notify", "act_autonomously"], ["finalize", "observe"]
  ].map(([id, mode]) => ({
    id, mode,
    owner: "Product operations lead",
    evidence: "recorded evidence",
    approvalScope: mode === "act_with_approval" ? "exact payload, 30 min, single use" : "bounded to this run",
    sideEffect: "explicit side effect",
    escalation: mode === "act_with_approval" ? "missing approval -> escalated" : "invalid condition -> escalated",
    recovery: "named recovery action",
    audit: "named audit event"
  }))
};

for (const [name, mutate] of mutations) {
  const copy = structuredClone(mapFromTheTable);
  mutate(copy);
  console.log(name, validate(copy));
}

I ran the completed table in memory with this validator on Node v20.11.0. The base map returned []. The four mutation results were:

MutationObserved result
Remove draft_issue.ownerRejected: draft_issue: owner missing
Remove classify.evidenceRejected: classify: evidence missing
Remove draft_issue.approvalScopeRejected: missing approval scope and approval scope required
Remove all terminal statesRejected: terminal handling missing

Illustration of four incomplete AI workflow decision maps being rejected or routed to escalation by a fail-closed validator

The code uses mapFromTheTable as the serialized version of the completed table. That name is intentional: the table is the artifact, and the validator is the check you can run against your own JSON export. The result is not a production test and does not show that a real backlog integration is safe.

What the four failures reveal

Each missing field represents a different way for architecture to hide an unresolved decision.

  1. No owner means no authority. A model can propose a route, but it cannot invent the person accountable for accepting the consequences.
  2. No evidence means no review. A reviewer can click approve without knowing which source, target, or policy produced the proposal.
  3. No approval scope means approval can drift. A broad approval may be reused for a changed payload, a different issue, or a later retry.
  4. No terminal state means the workflow has no definition of done. A success message can replace a completed record, an explicit escalation, or a failed outcome.

Google's multi-agent guidance puts human oversight, carefully defined autonomy, and observability together. It specifically recommends the ability to monitor, override, and pause business-critical flows (Google Cloud's multi-agent AI system guidance). The validator turns those broad controls into a pre-build question: can a reviewer identify who owns this decision, what they can see, what the workflow may do next, and how the run ends?

When this map is not enough

This map is a starting boundary, not a compliance review or a security design. Stop and widen the review when the workflow involves regulated decisions, sensitive personal data, financial transfers, access changes, destructive deletion, external commitments, multiple tenants, or an irreversible effect. Those cases may need specialist policy owners, stronger authentication, separation of duties, retention rules, independent testing, or a human-only path.

Also stop if the recovery operation is only “try again.” A retry is not a rollback. For an external write, name the created object, expected version, idempotency key, and compensating action. If the team cannot name those things, keep the AI in observe or propose mode.

The next step is small: take one real workflow, copy the table, and ask the people who own the work to fill every blank. If a field stays blank, that is an architecture finding. It is not an invitation for the model to guess.

For the broader architecture context, connect this worksheet to the AI workflow architecture guide. If the team has already reached runtime approval design, use the human-in-the-loop approval guide and the AI agent state-machine guide next. The worksheet itself is enough to begin without choosing a vendor or hiring an implementation team.

Questions people ask next

Should every AI workflow decision require human approval?

No. Classify each decision by consequence, reversibility, authority, and ambiguity. Let the AI observe, propose, or route when the effect is bounded. Require approval for a consequential side effect, and keep the decision human-owned when the authority, evidence, or recovery path is unclear.

What should happen when nobody owns a workflow decision?

Do not infer an owner from a job title or let the model choose one. Route the run to escalation or keep the step human-only until an accountable owner accepts the decision, its evidence, its scope, and its recovery path.