Field note · opportunity
What Is the Smallest Evidence Set for an AI Opportunity Decision?
Use five linked records to decide whether an AI opportunity deserves a reversible test, more discovery, or a stop before you fund a build.

When I taught product managers to move from writing specifications to building and shipping the product, the missing piece was often not the model. It was the definition of done.
An AI opportunity has the same problem. A demo can make the idea feel real while leaving the work, the decision, and the risk undefined.
The sourceable finding: for a low-to-moderate-risk AI opportunity decision, the smallest defensible packet is five linked records: a real workflow slice, a decision boundary, representative case types, a reversible test, and an accountable owner with a stop condition. Missing one means prepare, not build.

The smallest set is five linked records
Use five records for the first go, prepare, or stop decision. They are enough to make the opportunity concrete without pretending that a short discovery packet proves production readiness.
| Record | What it must answer | If it is missing |
|---|---|---|
| Workflow slice | What triggers the work, who does it, what input do they use, what action do they take, and what output is produced? | The opportunity is a slogan. |
| Decision boundary | What may AI suggest or prepare, what stays human-owned, and which errors are unacceptable? | “Automate the process” hides the dangerous part. |
| Case set | What does a normal case look like, what exception is known, and when should the system abstain? | A polished demo stands in for the real work. |
| Reversible test | What safe trial will run, how will success be checked, and what is the time or scope limit? | Interest is mistaken for evidence. |
| Owner and stop rule | Who can approve the next step, review the output, escalate a problem, or stop the test? | Nobody owns the decision when the evidence turns negative. |
This packet is my practical translation of a common pattern in primary guidance. NIST's AI Risk Management Framework says context should cover intended purpose, users, impacts, business value, risk tolerance, and system limits, and that the Map function can inform an initial go or no-go decision. It separately calls for documented test sets, metrics, tools, deployment-like conditions, and limits on generalization. (NIST AI RMF Core, NIST AI RMF Playbook)
The packet is smaller than a full risk or procurement review because it answers a narrower question: is this opportunity defined well enough to deserve a bounded next test?
The failure starts when a demo is treated as evidence
The fastest way to diagnose a weak opportunity is to ask five questions and refuse to fill the gaps with optimism.
Consider this worked scenario, which is deliberately generic and is not a client example:
“Use AI to summarize inbound requests and route them faster. The demo looks good.”
That sentence contains a possible benefit. It does not identify the real workflow slice, the source of truth, the routing decision, the ambiguous cases, the reviewer, or the point at which the experiment stops.
The failure trace looks like this:
| Check | What the card says | Diagnosis |
|---|---|---|
| Workflow slice | “Inbound requests” | Too broad. The trigger, current operator action, and output are unknown. |
| Decision boundary | “Route them faster” | Unsafe scope. It is unclear whether AI recommends a queue or changes one. |
| Case set | “The demo looks good” | No exception or abstention case has been inspected. |
| Reversible test | None | No safe way to separate a useful suggestion from a production write. |
| Owner and stop rule | None | No one has authority to narrow or stop the work. |
The correct decision is prepare. It is not a low score. It is a missing-evidence diagnosis.
That distinction matters. A score can make an undefined idea look like a weak candidate in a ranked list. A missing record tells you what to learn next.
Repair the packet before comparing AI tools
Repair the candidate in this order. Each step adds one record and removes one kind of ambiguity.
-
Write one workflow slice. Name the trigger, actor, input, current action, source of truth, and output. Keep it narrow enough that an operator can describe one pass without saying “it depends” at every step.
-
Draw the decision boundary. State what AI can draft, classify, extract, or recommend. State what remains human-owned. Write one unacceptable error in plain language. This is where “assist” becomes a real scope.
-
Choose three case types, not an arbitrary sample size. Include a normal case, a known exception, and an ambiguous or no-go case. Add more cases when the workflow has more meaningful variation or higher impact. The point is to expose the edge of the decision, not to manufacture a universal number.
-
Define a reversible test. Prefer read-only, shadow, draft, or approval-gated work for the first run. Set the success check before looking at the result. The check can combine usefulness, correction effort, review time, and unacceptable-error checks, depending on the workflow.
-
Name the owner and the stop rule. The owner is accountable for the next decision, not necessarily the person who builds the system. The stop rule should name an observation, such as an unacceptable action, an unmanageable review queue, missing source evidence, or a case outside the agreed boundary.
After repair, the worked scenario might read:
“For inbound requests in queue X, AI will suggest a category and draft a short summary from the request text. A queue owner makes the final assignment. The first test uses normal, incomplete, and ambiguous requests in shadow mode. The test stops if the output cannot be reviewed safely or if the source of truth is unclear.”
The failure-clinic record is now complete:
- Reproduction: Copy the under-specified card exactly: “Use AI to summarize inbound requests and route them faster. The demo looks good.”
- Trace: Check the five records in order. The workflow slice, decision boundary, case set, reversible test, and owner are respectively broad, unsafe, absent, absent, and absent.
- Diagnosis: The card returns prepare because it cannot state what the system may do, how an exception is handled, or who can stop the test.
- Repair: Add one concrete record for each missing decision, using shadow mode, explicit case types, human assignment, and an owner with a stop rule.
- Verification: Run the five-record check again. The repaired card returns run a bounded test because every record has a concrete answer. This verifies packet completeness only. It does not verify model quality, ROI, legal compliance, or production readiness.
This is now a decision-ready opportunity packet. It supports a bounded test. It does not support autonomous routing, a production launch, or a claim that the idea will pay back.
The case set is evidence only when it can fail
A case set is not a collection of easy examples. It is a small map of where the opportunity could be useful, uncertain, or unacceptable.
NIST recommends documenting test sets, metrics, and the tools used for evaluation so that measurement can be repeatable. It also says performance should be demonstrated under conditions similar to deployment and that limitations beyond those conditions should be documented. (NIST AI RMF Playbook)
For the initial opportunity decision, keep the case set proportional:
- Use ordinary cases to test whether the proposed help is relevant at all.
- Use a known exception to test whether the workflow becomes more expensive when reality appears.
- Use an ambiguous or no-go case to test whether the system can pause instead of forcing an answer.
Do not turn these three case types into a universal evaluation size. A hiring decision, medical workflow, financial approval, or system that changes records needs domain-specific review, stronger controls, and more evidence. The smallest packet is a starting gate, not permission to automate a consequential decision.
A reversible test needs a decision before it needs a prototype
The first test should make it cheap to learn and cheap to stop. That usually means a human reviews the output before an external action, or the system runs in shadow mode without changing the source of truth.
The UK government's AI assurance guidance defines assurance as measuring, evaluating, and communicating evidence about a system's capabilities, limitations, risks, and mitigations. It also says qualitative and quantitative techniques should be combined according to the context, and that there is no single assurance technique that works everywhere. (Introduction to AI assurance)
So the pre-set success check should contain two parts:
- Useful work: Did the output help the operator complete the slice with acceptable correction effort?
- Safe operation: Did the test avoid unacceptable actions, expose uncertainty, and stay inside the agreed boundary?
If you cannot state how both parts will be checked, the opportunity is not ready for a prototype decision. The missing work is measurement design, not model selection.
Ownership is part of the evidence
An owner is not a project-management accessory. The owner proves that the organization can act on what the test reveals.
The OECD AI Principles call for human agency and oversight, traceability, accountability, and ongoing risk management. They also say systems should be able to be overridden, repaired, or decommissioned safely when they create undue harm or undesired behavior. (OECD AI Principles)
For the first decision, record four ownership details:
- who can approve the bounded test;
- who reviews the output and knows the work well enough to spot a bad result;
- who receives an escalation;
- what observation makes the owner stop, narrow, or redesign the test.
ISO/IEC 42001 describes an AI management system as an organizational way to manage AI risks and opportunities through policies and continual improvement. That is a useful reminder that the opportunity does not belong only to the model builder. (ISO/IEC 42001:2023)
When five records are not enough
The five-record packet is not appropriate as the only evidence when any of these conditions apply:
- the system affects health, safety, fundamental rights, employment, credit, education access, or another high-impact outcome;
- the workflow uses sensitive personal data and the lawful purpose, retention, access, or security controls are not established;
- the proposed action is difficult to reverse, expensive to correct, or likely to affect someone who cannot meaningfully contest it;
- the owner, source of truth, or review capacity is unclear;
- the system is expected to act outside the cases you can inspect.
In those cases, treat the packet as an intake screen. Bring in the relevant legal, security, privacy, safety, compliance, and domain experts before running the test. NIST's framework ties scope, risk tolerance, human oversight, measurement, and risk treatment to context. The UK guidance likewise presents assurance as a combination of techniques selected for the context, not a universal checklist. (NIST AI RMF Core, Introduction to AI assurance)
What this evidence set does not prove
It does not prove return on investment, demand, model quality at scale, legal compliance, production reliability, or that an AI system is better than an ordinary software change.
It proves something narrower and useful: the team can describe the work, bound the decision, expose a failure case, run a safe next test, and name the person who will act on the result.
If you need to rank several candidates after that, use How to Prioritize AI Use Cases in a Small Business. If the candidate is already chosen and you need to check whether the process is ready for a pilot, continue with How to Tell If a Business Process Is Ready for AI Automation. The wider AI opportunity pillar is the parent collection for this decision work.
If your team can name the five records but keeps getting stuck on the next experiment, Marius Manolachi's AI consulting and tutoring work is a next step for making existing people capable of building on their own work.
The smallest useful decision is not “AI or no AI.” It is “what evidence is missing, and what is the safest next test that can produce it?” That is the decision this packet is built to support.
Continue with a related field note
- How to Measure Reversible Tasks With a One-Week Timebox
- How to Decide Evidence Quality When Source Data Conflicts
- How to Choose Evidence Quality With Incomplete Records
- Which AI Architecture Decision Should a Nontechnical Buyer Know First?
- What Evidence Justifies Observing AI in Quarterly Access Reviews?
- What Should a Professional Learn to Own AI Vendor-Dispute Evidence?