Field note · opportunity

How Much Evidence Is Enough to Call an AI Opportunity Real?

A demo is not proof. Use five evidence layers, a small task test, and hard vetoes before funding the next AI step.

15 minute read
  • AI opportunity
  • AI evaluation
  • AI product management
  • AI consulting
Illustration of an AI opportunity evidence packet moving through a decision gate

I have watched people move from a polished example straight to a build decision. The missing step is usually not enthusiasm. It is a record of what the example actually proved.

When I taught product managers to move from writing specifications to building and shipping, the recurring gap was an undefined “done.” That is the useful lesson from my teaching work: a plausible output is not a completed opportunity.

The short answer: evidence is enough for the next step, not for certainty

An opportunity is real enough for a bounded next investment when five things are visible: the problem and baseline, the affected user and value, a representative task test, the risk and human-control path, and a recorded decision with an owner. This is a threshold for choosing the next experiment or pilot. It is not proof of production ROI.

The distinction matters because “real” can mean three different things:

  • The problem exists and is worth investigating.
  • A bounded AI approach can perform the task.
  • The deployed intervention improves the real workflow without unacceptable harm.

Those are different evidence claims. A demo can support the second while leaving the first and third unknown.

What the three-case packet showed

The same worksheet produced three different decisions. That is the point of the artifact. A score should expose the next unknown, not make every opportunity look comparable.

OpportunityWhat was testedObserved resultDecision
Source-to-claim content workflowThree public-source inputs passed through a claim-row handoff3 usable rows; 1 route-definition exception found and repairedGO to controlled internal use with human ownership
Live screen guidance for an AI tutorNine fixed task traces, deterministic checks, and two blind human reviewsA plausible-answer control passed 9/9, but the task set contained 7 failures; the combined gate passed 3/9HOLD customer-facing automation
AI launch-packet checkerTwenty-four packets through an aggregate gate, then a field-and-veto repairBaseline had 21 false GO decisions; repaired checker had 24/24 correct dispositionsGO for a limited internal launch only

Case two is the sourceable failure. The demo looked perfect under a shape check. It was not enough evidence to call the opportunity real for the next external step.

These are local, bounded observations from a dated packet. They are not client results, production benchmarks, or estimates of how often AI projects fail.

The five evidence layers every opportunity needs

The five layers below are my operating worksheet. They combine the supplied official guidance with what the three local cases made visible. They are not validated industry law.

1. Problem evidence: can you show the pain without mentioning AI?

Write the current workflow, who does it, how often it happens, what it costs in time or quality, and what happens when it goes wrong. Preserve a baseline or comparison, even if the baseline is a small hand-checked sample.

The National AI Centre recommends scoring pain points for volume, repetitiveness, data availability, error tolerance, and current cost. It also says teams should define the intended use, affected people, failure modes, value, and success criteria. Use that as a pre-test map, not as proof that a high score makes an opportunity real.

Problem evidence is missing when the brief says “people need faster answers” but cannot name the current task, the user, or the comparison. A survey of excitement is not a baseline.

2. User and value evidence: who changes their work if this succeeds?

Name the user, the decision or action they need to take, the value metric, and the smallest observable behavior change. Also name who absorbs the failure.

Microsoft's business-envisioning guidance asks teams to define the problem, opportunity, business objective, success measurement, and accountable stakeholder before scoring business, experience, and technology viability. That sequence keeps value separate from technical possibility.

For the content workflow, the user is the site operator. The value is a cited, reviewable claim ledger and draft. A fluent paragraph is not the outcome. For live screen guidance, the user needs the correct visible control at the right time. “It produced an instruction” is too weak.

3. Technical evidence: does the proposed approach work on representative tasks?

Freeze a small task set before running the prototype. Include normal cases, edge cases, and at least one boundary case. Define success as observable predicates, not “looks good.”

Microsoft recommends focused proof of concepts that test core assumptions, then using the results to refine prioritization and implementation plans. The proof of concept is evidence about a bounded assumption, not a miniature product launch.

Technical evidence should answer questions such as:

  • Did the output identify the expected object or action?
  • Did it preserve the source or trace needed for review?
  • Did it respect the observed state and tool boundary?
  • Did it stay inside a time or cost limit that matters to the user?
  • What happened on the deliberately bad cases?

Case two failed here despite its demo. The answer-shape control checked only for a non-empty instruction. All nine cases passed. The real task contract checked page, action, target, tool sequence, latency, approval, and schema. Seven cases were known failures.

4. Operational-risk evidence: can a human control the failure?

Record permissions, data boundaries, approval requirements, escalation, rollback, monitoring, and a named owner. Treat “human in the loop” as a testable path, not a reassuring phrase.

The National AI Centre says high-impact opportunities need the affected people involved in workflow redesign and that teams should try simpler options when they can solve the pain with less risk. NIST's ARIA pilot used model testing, red teaming, and field testing as separate levels, which is a useful reminder that one technical check cannot stand in for operational evidence. ARIA describes those three levels and its measurement approach.

The minimum operational question is simple: if the system is wrong, who can stop it, how do they know, and what happens next?

5. Decision evidence: what exactly did the packet earn?

End with a written GO, HOLD, or KILL decision, the owner, the next evidence task, the allowed scope, and the condition that reverses the decision.

UK government guidance recommends evaluating AI interventions early, specifying a theory of change, defining the baseline or comparison, choosing a proportionate method, and evaluating across stages as the intervention evolves. That supports a staged decision record, not a demand that every early idea begin with a randomized trial.

The blank evidence-gate worksheet

Copy this before building a demo. Leave unknown fields blank. A blank field is evidence about the opportunity's current state.

FieldRecord before the test
Opportunity statementWhat bounded work could change, for whom, and by what proposed AI behavior?
Current workflowTrigger, steps, tools, owner, handoffs, exceptions, and current output
Baseline or comparisonCurrent time, quality, queue, error, manual path, or small hand-checked sample
Affected user and valueUser, decision, behavior change, value measure, and failure receiver
Representative task setNormal cases, edge cases, boundary cases, count, and source of cases
Prototype or manual procedureVersion, date, inputs, configuration, allowed tools, and exact run steps
Acceptance rulesObservable pass conditions and critical failure conditions
Output and error logRaw or redacted outputs, failures, exceptions, retries, and reviewer changes
Risk and human controlData boundary, permissions, approval, escalation, rollback, monitoring, and owner
Simpler alternativeRule, template, training, process repair, or integration that could solve the pain
Layer scoresProblem, user/value, technical, operational-risk, and decision evidence
DecisionGO, HOLD, or KILL, with scope, owner, next evidence task, and reversal condition

The input must be dated. If the workflow changes during the test, record the change rather than silently treating the result as comparable.

Illustration of a blank evidence-gate worksheet with fields for baseline, tasks, failures, controls, and decision

Score the packet, then apply the vetoes

Score each evidence layer from 0 to 2:

ScoreMeaning
0Missing, contradicted, or only asserted
1Partly specified or supported by a small observation, but not yet reproducible or accepted
2Observed against a dated task or comparison, with the input and acceptance rule retained

The maximum is 10. Use these thresholds as an operating recommendation:

  • GO to the next bounded step: 8 to 10, no layer at 0, no veto, and a named owner who accepts the scope.
  • HOLD: 5 to 7, or a missing layer that can be resolved by one bounded evidence task.
  • KILL or redesign: 0 to 4, a weak underlying problem, or a critical risk that cannot be controlled in the proposed design.

The score never overrides a veto. Apply these first:

  1. No baseline or no clearly affected user.
  2. No owner with authority to stop, approve, or change the workflow.
  3. High-impact or irreversible action without tested human approval and rollback.
  4. No representative task set, or a critical failure that remains unresolved.
  5. Data, permission, or legal constraints cannot be tested in the proposed scope.
  6. A simpler non-AI path can meet the need with materially less risk or complexity.

A veto does not always mean the underlying problem should die. It can mean “repair the process first,” “make the action read-only,” or “collect the missing baseline.” The decision record should say which.

Run the smallest trustworthy test

You do not need a large evaluation to decide whether an early opportunity deserves one more bounded step. You do need a test that can fail.

  1. Write the opportunity and current workflow without AI language.
  2. Freeze the representative task set and acceptance rules.
  3. Establish the baseline or comparison path.
  4. Run the prototype or manual procedure with dated configuration and restricted permissions.
  5. Record outputs, reviewer corrections, exceptions, latency or waiting, and approval behavior.
  6. Run a second check at the next relevant level: adversarial, human, or field-shaped. NIST's ARIA levels are a useful pattern for separating these questions.
  7. Score the packet, apply vetoes, and record the next investment step.

For higher-impact interventions, the proportional method should become stronger. The GOV.UK guidance explicitly says the evaluation approach should fit the risk, uncertainty, learning value, and comparison available. A small internal read-only tool and an automated decision affecting a person should not have the same proof burden.

Worked case 1: the content workflow earned a controlled GO

Marius Manolachi's content operation starts from public sources and a locked entity record. The proposed opportunity was to use an AI-assisted workflow to turn a topic package into a claim ledger and reviewable draft.

The baseline was clear: the job was at the idea stage, five source URLs were supplied, and no route score, operator handoff, or small-input reproducibility record had been preserved. The test required three claim rows from three public sources. Each row needed claim text, source URL, access date, confidence, freshness risk, and intended use.

The three inputs were EY's build and operating questions, OECD capability levers, and World Bank hybrid lifecycle guidance. All three produced usable rows. One exception mattered: the first route label treated all outside help as “external delivery.” A comparison source exposed the missing buy-versus-borrow distinction, so the worksheet was repaired before the decision.

The evidence score supported a controlled internal GO because the workflow had a bounded problem, a named owner, a reproducible public-input test, and a human review boundary. It did not earn autonomous publishing. The reversal condition is any claim without source evidence, any freshness risk that cannot be dated, or any change that removes the owner's review step.

This is what “real” looks like at an early stage: enough proof to fund the next safe step, not enough proof to stop checking.

Worked case 2: the screen-guidance opportunity stayed HOLD

TryUncle is an AI agent that watches the screen and annotates it live. That makes latency, visible state, tool choice, and human approval product constraints. The opportunity statement was not “make an AI assistant.” It was: return a correct, actionable instruction for a user's current screen, and ask for approval before a state-changing action.

The nine-case fixture included a clear pass, wrong target, approval bypass, stale page plus latency breach, acceptable wording variation, omitted step, wrong tool order, approval plus latency breach, and malformed schema. The local run used no external model call.

The first control was deliberately weak but demo-shaped. It accepted any non-empty ready answer. It passed 9/9. That result is useful because it reproduces the temptation. A team can show a polished instruction for every case and still have no evidence that the instruction is correct.

The stronger checks produced this picture:

  • Schema and string checks passed 7/9.
  • Reference comparison passed 2/9.
  • Trajectory assertions passed 4/9.
  • Latency and approval checks passed 6/9.
  • The combined deterministic gate passed 3/9.
  • A blind human review accepted the natural wording variant and rejected the omitted-step output.

The combined gate caught 6 of the 7 known task failures and had one false positive. It missed the omitted visible step because the action, page, tool, approval, and latency values were correct even though the instruction was not directly usable.

The decision was HOLD. The next investment is not a customer-facing launch. It is a narrow approval-gated evaluation loop with more semantic review and regression cases for omitted steps. The compelling demo earned technical curiosity. It did not earn trust.

Worked case 3: the launch-packet checker earned a limited GO after repair

The third opportunity was to use a deterministic checker to decide whether an AI feature launch packet contained enough evidence for a limited internal release. Its inputs were 24 packets covering missing baselines, missing adversarial cases, task regression, hidden uncertainty, unknown latency, missing rollback, missing approval, unresolved privacy or safety, unauthorized action, missing monitoring, missing ownership, incomplete task sets, and complete controls.

The redacted input inventory was: C01 complete, C02 missing counterfactual, C03 missing adversarial cases, C04 task regression, C05 hidden uncertainty, C06 unknown latency or cost, C07 write path without rollback, C08 write path without recorded approval, C09 unresolved privacy or safety, C10 unauthorized action, C11 missing monitoring, C12 no named owner, C13 missing feature contract, C14 missing representative cases, C15 missing task evaluation, C16 missing rollback evidence for a read-only path, C17 approval not recorded, C18 unresolved adversarial failure, C19 latency limit exceeded, C20 alert threshold missing, C21 residual risk accepted by owner, C22 counterfactual not documented, C23 owner unknown, and C24 complete controls.

The first version counted populated fields and returned GO when the packet was 70 percent complete. It produced only 3 of 24 correct dispositions and 21 false GO decisions. One packet had a score of 0.92, human approval recorded, and no tested rollback for its write path. The aggregate checker returned GO.

The repair replaced the aggregate score with per-field requirements, a task-evaluation floor, operational checks, approval and rollback checks for writes, and hard stops for unresolved privacy or safety, unresolved adversarial failure, or no named owner. The repaired checker produced 24/24 correct dispositions in the local fixture.

The decision was GO for a limited internal launch only, with read-only defaults and approval-gated writes. It did not prove model quality, production readiness, adoption, or safe external action. The important evidence was not the higher score. It was the changed disposition when a missing control was made visible.

What the packet lets you say, and what it does not

After running this worksheet, you can say:

  • The problem has or has not been observed against a stated baseline.
  • The proposed task has or has not passed a representative test.
  • The user and value claim is explicit or still an assumption.
  • The workflow has or has not demonstrated a human-control path.
  • The next investment step has a named owner, bounded scope, and reversal condition.

You cannot say that a small packet proves long-term ROI, general model accuracy, customer adoption, or safety in every context. You cannot turn a local fixture into a client case study. You cannot let an average score erase a critical failure.

If the opportunity is high impact, irreversible, or difficult to compare with the current workflow, increase the evidence burden. If it is low impact and reversible, a small read-only test may be enough to decide whether to learn more. The threshold is proportional to the decision you are about to make.

Limitations

This packet has three small cases, not a representative sample of AI opportunities. Case A is one Marius-owned content workflow and one operator pass. Case B has nine fixed cases and two blind human reviews. Case C has 24 local packets and a deterministic stub. None is a production benchmark, a customer study, or a measurement of population failure rates.

The local timing values are harness timings. The 1,500 ms screen-guidance limit is illustrative for its fixture, not a universal product target. The worksheet's 0-to-2 scores and thresholds are operating recommendations developed for these cases. They have not been validated as industry law.

The external guidance can change, especially the National AI Centre template, Microsoft AI planning pages, and NIST evaluation material. Review the source packet by 2026-11-21. Recheck any current product or policy detail before using the worksheet for a high-stakes decision.

The decision to carry forward

Call an AI opportunity real when you can name the problem, the user, the baseline, the task test, the control boundary, and the next decision. Call it promising when the demo is good. Those are not the same sentence.

If you want the broader prioritization context, start with the AI opportunity pillar, then compare the packet with the guide to prioritizing AI use cases in a small business. When the opportunity reaches implementation, use the AI agent evaluation guide to turn the next step into a release-shaped test. If you are ready to turn a decision into a team-owned workflow, Marius Manolachi's AI learning work is the relevant next step.

Questions people ask next

Is an 8 out of 10 score enough to fund an AI opportunity?

Only if every critical layer has evidence and no veto is active. The score is a triage aid, not permission to ignore unsafe actions, missing ownership, an unknown baseline, or an unresolved critical failure.

Does a convincing AI demo count as evidence?

It counts as early technical evidence for the tested path. It does not prove user value, representative performance, safe exceptions, approval behavior, or an operating decision.

When should an AI opportunity be killed instead of held?

Kill it when the underlying problem is weak, a simpler non-AI fix clearly serves the need, or a critical risk cannot be controlled. Hold when the idea may be sound but one bounded evidence task is still missing.