Field note · opportunity

How to Choose a Safe First AI Slice for Approval Work

Use a pre-prototype scorecard to choose a safe, reviewable AI slice inside approval work without handing over the decision.

10 minute read
  • AI strategy
  • AI use cases
  • Evaluation
Illustration of a team choosing a safe first AI slice for approval work

Approval work often looks like an easy AI win. A request arrives, someone checks evidence, a decision follows, and a message or system update completes the loop. The first useful AI slice is rarely the whole loop.

The expensive part is usually hidden in the middle. The reviewer gathers missing context, corrects fields, explains exceptions, and carries the consequence of a wrong approval. Choose the smallest slice that helps this person without taking the decision away.

Illustration of a team testing approval work with a pre-prototype decision card

Choose the first slice with a veto-before-score test

My decision rule is simple: approval work earns a prototype only after it passes six vetoes and scores at least 8 out of 10 on the surviving dimensions. The first prototype should stay read-only, draft-only, or recommendation-only, with positive value after human review is included.

This is the sourceable artifact on this page. It is my decision tool, not a standard published by NIST, Microsoft, the SBA, Anthropic, or OpenAI.

Decision stateResultNext move
Prototype shadowAll six vetoes pass, the score is 8-10, and review-adjusted value is positive.Compare a proposed output with the current approval work without taking an external action.
Hold and collect evidenceAll vetoes pass, but the score is 5-7 or the net value is still unknown.Collect representative cases, a clearer rubric, or a baseline.
RedesignThe workflow may matter, but a veto fails because the action is too broad, evidence is scattered, or review duplicates the task.Narrow the candidate to extraction, classification, drafting, or recommendation.
Stop or keep human-onlyThe score is 0-4, or the first action cannot be contained.Do not prototype the proposed AI action. Find a safer adjacent job or leave approval human-owned.

NIST's AI Risk Management Framework says context, business value, risk tolerance, human oversight, and system limits should inform an initial go/no-go decision. It also says risk work is continuous, so this screen is a starting decision, not a permanent approval. NIST's AI RMF Core supports the ingredients. The vetoes and thresholds are my synthesis for this specific reader job.

Run the vetoes before you score the opportunity

Do not let theoretical time savings compensate for a missing owner or an unsafe action. A failed veto changes the work you need to do before scoring.

VetoPass whenIf it fails
Named ownerOne person owns the outcome and the escalation path.Find the owner or remove the candidate from the active queue.
Checkable outcomeA reviewer can say what is correct, incomplete, or unsafe.Define a rubric, comparison, or explicit reference result.
Available evidenceThe request contains the policy, records, or source material needed to justify a recommendation.Improve intake or narrow the task to the evidence that exists.
Explicit boundaryThe first AI action is read, extract, classify, draft, or recommend.Remove sending, approving, paying, publishing, or record-changing actions from the first slice.
Reviewer capacityThe intended reviewer can verify the result without repeating the whole job.Measure the review path or redesign the output to show evidence and uncertainty.
Reversal pathA wrong output can be rejected, corrected, or rolled back before a consequential action.Keep the action human-only or add a real control before any prototype.

Microsoft's use-case guidance asks teams to define the problem, business objective, success measurement, and accountability before comparing viability. That is why “the team will review it” is not enough. A decision needs a person, a test, and a consequence boundary. Microsoft's business-envisioning guidance is written for ISVs, but these fields are useful in a small company too.

When I taught product managers who moved from writing specifications to building and shipping products, the recurring lesson was that “done” had to be explicit. The same lesson applies here. If nobody can say what a correct approval recommendation looks like, the workflow is not ready for an AI prototype. This is a bounded teaching observation from Marius Manolachi's F-pms entity fact, not a measured study.

Score the work, not the AI idea

After the vetoes pass, score five dimensions from 0 to 2. The score describes the work you can test, not the capability of a future model.

Dimension0 points1 point2 points
RecurrenceOne-off or too rare to observeOccasionalRecurring enough to compare and improve
Evidence qualityMissing or inaccessiblePartial or scatteredSource packet is available and attributable
Decision clarityPersonal, unstable, or unexplainedPatterns exist but exceptions dominatePolicy, examples, or records make the decision reconstructable
Review efficiencyReviewer must redo the workMeaningful correction remainsReviewer can verify quickly against evidence
Consequence containmentWrong action is hard to reverseRecoverable with material costFirst action has no external side effect or is easy to undo

The maximum is 10. Prototype the shadow or draft slice at 8-10, hold at 5-7, and stop or redesign at 0-4. The threshold is deliberately visible so a team can disagree with it and change it before it starts building. Hidden thresholds produce arguments after the prototype has already consumed the budget.

The most important row is review efficiency. Approval work is not valuable because a model can produce a plausible recommendation. It is valuable only if the accountable reviewer can reach a sound decision with less effort, better evidence, better consistency, or a clearly stated non-time benefit.

Test the approval workflow without building the prototype

Use recent work to test the job before you test the model. A paper exercise is enough to expose many bad opportunities.

  1. Write the result. Replace “AI approval assistant” with a result such as “prepare a source-linked recommendation for the operations owner.” State what remains human-owned.
  2. Collect varied cases. Use recent routine, exception, rejected, and escalated approvals. If the sample contains only easy approvals, mark the evidence incomplete. Anthropic's evaluation guidance makes the same general point about defining inputs and success criteria and drawing early test cases from real failures. Its evaluation guidance is about agent evaluations, not this scorecard, so use it for test discipline rather than as a threshold authority.
  3. Reconstruct the current path. For each case, record preparation and coordination time, evidence consulted, decision, correction, escalation, and final action. Don't estimate the total from memory if the workflow has a log or queue.
  4. Describe the proposed output without calling a model. Write the fields, citations, recommendation, uncertainty, and reviewer action that a prototype would have to produce. This is a contract test, not a prompt test.
  5. Check each proposed output against evidence. A field match, policy reference, required attachment, or exact label can be a deterministic check. OpenAI's eval documentation uses human-labeled references and a string-check criterion as one example of this style of test. OpenAI's eval guidance does not say every approval task has a string answer. It shows why the check should match the work.
  6. Calculate review-adjusted value. Subtract AI preparation, human review, correction, and escalation time from the current preparation and coordination time. If the result is not positive, record another benefit hypothesis instead of forcing a time-saving claim.
  7. Make the decision. Apply the vetoes, add the five scores, record the missing evidence, and choose prototype shadow, hold, redesign, or stop.

The worksheet can be copied as a single row per workflow:

Candidate resultOwnerEvidence packetHuman actionReversible first sliceCurrent minutesProposed minutesCorrection and escalationNet valueDecision
Source-linked recommendation for an approval ownerNamed personPolicy and request records availableApprove, reject, or escalateDraft recommendation onlyRecordRecordRecordcurrent - proposed - correction - escalationApply threshold

The SBA advises small businesses to start small and test whether an AI tool adds value. It also recommends another person review AI products in some small-business use contexts. The SBA's AI guidance is general advice, not evidence that your workflow will benefit. Here, “start small” means testing the approval job and reviewer burden before building a connected system.

Keep the approval decision human in the first prototype

The safest first prototype prepares the decision. It does not own the decision.

That usually means one of four boundaries:

  • extract fields from the request and show the source location;
  • classify the request into a human-review queue;
  • draft a recommendation with evidence and uncertainty;
  • prepare the message or record change for a person to approve.

Do not treat a human button labelled “approve” as meaningful oversight if the reviewer cannot inspect the evidence, has no time to challenge the recommendation, or is effectively forced to accept a default. The control must be usable, not decorative.

Keep the work human-owned when the action affects a person's legal position, access to essential services, safety, employment, health, credit, payment release, or another consequence your organization cannot readily reverse. This is a risk boundary, not a claim that every such workflow is legally classified the same way. NIST's guidance is the right reason to pause: context, impact, risk tolerance, and human oversight must be documented before proceeding.

The exception to my normal prototype recommendation is a high-consequence approval. In that case, get the relevant domain, legal, compliance, or security review before using real data or moving beyond a tightly controlled, non-actioning test.

Read a failed score as a design signal

A failed score is useful when it tells you what to change.

FailureWhat it meansBetter next step
No ownerThe business problem is not assigned to a decision path.Name one accountable person or stop.
No checkThe team wants a plausible answer, not a testable result.Define pass, edit, reject, and escalate conditions.
Scattered evidenceThe model would be asked to compensate for an intake problem.Fix the evidence packet or narrow the scope.
Review equals original workThe AI output adds another screen, not capacity.Change the output or test a different part of the workflow.
No reversalThe first action carries too much consequence.Keep the action human-owned and prototype a proposal only.
Score below 8The opportunity is plausible but not ready to earn build effort.Collect evidence, improve the workflow, or choose another candidate.

This is where approval work often stops being an AI opportunity. The request looked repetitive, but the approval depended on missing context, personal judgment, or a downstream action nobody could undo. That is not a model failure. It is a mismatch between the proposed AI job and the real work.

If the workflow survives, the next question is how to test a safe feature with representative data. Use the parent guide, How to Prioritize AI Use Cases in a Small Business, to compare it with other candidates. Then use How to Test an AI Feature Before Production Data when the bounded prototype has a defined output and test set.

Make the next action small enough to learn from

Do not build an “AI approval system” because the scorecard is green. Build the smallest shadow, draft, or recommendation slice that can falsify the value hypothesis. Keep the reviewer, evidence packet, decision check, and reversal path visible in the test.

If you want help turning the worksheet into a capability your team can run on its own work, Marius Manolachi's AI consulting and tutoring page explains the kind of support he offers. The decision tool remains useful without that next step.

Questions people ask next

How many approval cases should I test before prototyping?

There is no universal sample size for an opportunity screen. Use a bounded set of recent cases that includes routine, exception, rejected, and escalated work. If one category is missing, mark the result incomplete instead of treating a clean routine sample as proof.

Should AI make the final approval decision?

Not in the first prototype. Keep approval ownership with the person who is accountable, and let AI extract evidence, classify a request, draft a recommendation, or prepare a response for review. Expand the action boundary only after separate evidence supports it.

What if approval work saves quality but not time?

Keep quality, traceability, or consistency as the explicit hypothesis. If review-adjusted time is zero or negative, do not call the result a time-saving opportunity. Prototype only when the non-time benefit has an owner, a check, and a safe success measure.