Field note · evaluation

Why Does AI Fail When Gaps Are Unknown?

A 36-case blind-spot fixture shows how a passing authored suite misses contradictory evidence, stale state, and failed tools.

13 minute read
  • AI evaluation
  • AI reliability
Illustration of an AI evaluation suite revealing an uncovered failure gap

In this fixture, the first failure was not a bad answer. It was a condition nobody put in the test set.

That distinction matters. If the team has only tested complete evidence, matching records, available tools, and one stable state, a clean pass says very little about what happens when one of those assumptions disappears.

The parent guide on evaluation coverage for AI features explains how to think about coverage before launch. This page focuses on the next problem: finding the gaps your team has not named yet.

The result: a passing authored suite still left 16 failures

In my pinned local fixture, the authored suite passed every repeated trial. Controlled variants then exposed 16 distinct failures that the authored cases did not enumerate.

SetCasesTrialsPassesFailsDistinct failuresWhat changed
Authored baseline12363600Nominal outcomes and obvious missing fields
Generated missing context5151231A material field was absent but not described with the baseline's missing-field wording
Generated contradictory evidence5150155Two records or policy signals disagreed
Generated state changes5150155The relevant state changed after earlier evidence was recorded
Generated tool failures5150155A current lookup or extraction step failed while stale evidence remained
Generated out-of-scope requests4121200The request mixed a supported check with legal, policy, contract, or incident work

The fixture ran three times with the same pinned configuration. It is a local workflow test, not an LLM benchmark and not a production prevalence estimate. Its useful finding is narrower: a nominal pass did not reveal the failure conditions until the test data changed shape.

Illustration of a blind-spot fixture branching authored cases into controlled variants

Why do unknown gaps survive an evaluation suite?

An evaluation suite measures performance over its case distribution. It cannot directly score a condition that the distribution never represents. If every case contains current evidence and a single consistent state, the suite has tested the workflow under those assumptions, not outside them.

NIST makes a related distinction in its report on deployed AI monitoring: pre-deployment evaluations are mostly run in controlled environments, while post-deployment monitoring is needed to see unforeseen outputs caused by changing inputs and other deployment conditions. That does not make pre-deployment evaluation useless. It tells you what it cannot establish. NIST's monitoring report treats the two activities as different evidence sources.

The gap is usually created by one of four hidden assumptions:

  • the needed context is present and named in a familiar way;
  • every source agrees with every other source;
  • the state at intake is still the state at decision time;
  • every tool call returns usable, current evidence.

Out-of-scope requests add a fifth boundary. A workflow can be good at eligibility triage and still fail when a user asks it to add legal advice or redesign policy in the same turn.

The important question is not, “Did the model pass the test?” It is, “Which conditions did the test make impossible?” Those conditions are the blind spots worth probing next.

What did the fixture actually test?

The workflow was intentionally small so the failure could be traced. It accepted a request and evidence packet, then returned one action: answer, ask_clarifying, escalate, or out_of_scope.

The pinned configuration was:

FieldValue
Fixtureblind-spot-fixture-0.1.0
WorkflowEvidence triage for approval and eligibility questions
Model under testfixture-heuristic-v1, deterministic local text classifier
Temperature0
Top-p1
Trials3 repeated runs, trial IDs 1, 2, and 3
InputRequest plus evidence text
Output contractanswer, ask_clarifying, escalate, or out_of_scope
Run date2026-08-24

The workflow's decision order was also recorded: out-of-scope marker, explicit missing-field marker, conflict marker, approval keyword, then clarification. That detail matters because the trace shows why a case passed or failed. A result without a trace only tells you that something went wrong. It does not tell you what the system noticed.

The prompt and input template were saved verbatim:

You are an evidence triage assistant. Choose exactly one action: answer, ask_clarifying, escalate, or out_of_scope. Use only the supplied request and evidence. Ask when a material field is absent. Escalate when sources disagree or current state cannot be verified. Never claim a tool completed when its result is unavailable. Refuse requests outside the supported approval and eligibility workflow.

Request: {request}
Evidence:
{evidence}
Return JSON with action and rationale.

The full authored and generated test data is versioned in the fixture specification. Each generated case changes one observable condition while keeping the workflow request stable: G01-G05 test missing context, G06-G10 contradictory evidence, G11-G15 state changes, G16-G20 tool failures, and G21-G24 mixed-scope boundaries.

NIST's work on evaluation probes makes the same operational point from another angle. Its probes are designed as adversarial verifiers, and the results are accumulated in a structured audit trail that maps claims to supporting evidence. The NIST probe project is broader than this fixture, but the design lesson transfers: preserve the evidence and the verdict together.

What did the raw pass and fail table show?

The first trial is the compact raw result. Trials 2 and 3 reproduced the same action for every case.

CaseGap axisExpectedActualResult
G01Missing contextask_clarifyingask_clarifyingpass
G02Missing contextask_clarifyingask_clarifyingpass
G03Missing contextask_clarifyingask_clarifyingpass
G04Missing contextask_clarifyingask_clarifyingpass
G05Missing contextask_clarifyinganswerfail
G06Contradictory evidenceescalateanswerfail
G07Contradictory evidenceescalateanswerfail
G08Contradictory evidenceescalateanswerfail
G09Contradictory evidenceescalateanswerfail
G10Contradictory evidenceescalateanswerfail
G11State changeescalateanswerfail
G12State changeescalateanswerfail
G13State changeescalateanswerfail
G14State changeescalateanswerfail
G15State changeescalateanswerfail
G16Tool failureask_clarifyinganswerfail
G17Tool failureask_clarifyinganswerfail
G18Tool failureask_clarifyinganswerfail
G19Tool failureask_clarifyinganswerfail
G20Tool failureask_clarifyinganswerfail
G21Out of scopeout_of_scopeout_of_scopepass
G22Out of scopeout_of_scopeout_of_scopepass
G23Out of scopeout_of_scopeout_of_scopepass
G24Out of scopeout_of_scopeout_of_scopepass

The pattern is more useful than the total. The workflow caught the out-of-scope variants because they used explicit boundary language. It missed most variants where the evidence was incomplete, stale, contradictory, or unavailable but still contained a plausible answer signal.

This is why a gap map should record conditions, not only categories. “Tool failure” is too broad. “Current lookup returned no response while a stale record still said active” tells the engineer what to assert and what to replay.

What do the representative failure traces reveal?

Three traces show the mechanism clearly.

G06: contradictory evidence

Expected: escalate
Actual: answer
Observable condition: CRM status active, billing status cancelled, policy requires agreement
Features noticed: approval keyword=true, conflict marker=false
Decision path: approval keyword -> answer

The fixture did not fail because it lacked a word like “contradiction.” It failed because the workflow had no explicit comparison step for independently named sources. The evidence looked answerable to the classifier even though the policy said disagreement should stop the decision.

G11: state changed after intake

Expected: escalate
Actual: answer
Observable condition: active at 09:00, suspended at 09:05, request at 09:06
Features noticed: approval keyword=true, conflict marker=false
Decision path: approval keyword -> answer

The current state was present in the evidence, but the workflow had no temporal rule. It treated an earlier valid state as sufficient evidence for a later decision.

G16: current tool lookup failed

Expected: ask_clarifying
Actual: answer
Observable condition: CRM lookup returned no response, cached record says active, current status required
Features noticed: approval keyword=true, tool marker=true, conflict marker=false
Decision path: approval keyword -> answer

The trace contains the tool failure. The decision path ignores it. That is the gap. A monitor or evaluator that only checks the final answer can miss the fact that the workflow used stale evidence after a current lookup failed.

The broader research points in the same direction. SLEIGHT-Bench constructs targeted variants around blind spots such as system state, omission, object reuse, and contradictory evidence, then tests whether monitors detect them. Its authors also note that synthetic variants improve scenario diversity while risking realism. The SLEIGHT-Bench method and limits are a useful model for probing, not a claim that this small fixture has the same scope.

Which discovered gaps become regression tests?

Not every generated case deserves a permanent slot. Use the observed failure plus a testable condition to decide.

GapFixture resultRegression decisionMinimum assertion
Missing approval identity or acceptanceG05 failedInclude G05 and one paraphraseThe required owner is named and has accepted responsibility
Conflicting source recordsG06-G10 all failedInclude G06 and G08, retain the rest as probesDisagreement between named authorities produces escalate
State changed after intakeG11-G15 all failedInclude G11 and G12, retain the rest as probesDecision uses current state or stops when freshness is unknown
Current tool lookup failedG16-G20 all failedInclude G16 and G19, retain the rest as probesUnavailable current evidence cannot be replaced by cached evidence
Mixed supported and out-of-scope requestG21-G24 passedKeep G21 as a boundary smoke test; generate more variantsThe workflow refuses the unsupported portion rather than answering it

The rule is simple:

  1. Include the case when the expected action differs from the baseline output.
  2. Require an observable condition that can be asserted in a trace, state record, or tool result.
  3. Prioritize cases where the wrong action could authorize, mislead, or create an irreversible side effect.
  4. Keep paraphrases and low-risk variants as discovery probes until they reveal a distinct failure boundary.

This decision makes the regression set smaller than the discovery set without throwing away the map. The permanent tests protect known boundaries. The probes keep looking for new ones.

How should you run a blind-spot fixture on your own feature?

Use one representative workflow and keep its action contract narrow. A small, inspectable fixture is more useful than a large synthetic suite whose expected outcomes nobody can explain.

  1. Freeze the contract. Write the permitted actions and the observable expected outcome for each one. Include abstention or escalation if the feature needs them.
  2. Author the baseline. Create cases for complete evidence, clear outcomes, obvious missing fields, and the normal tool path. Save the exact inputs and expected actions.
  3. Choose gap axes. Start with missing context, contradictory sources, state changes, tool failures, and out-of-scope requests. Add an axis only when it changes a real decision boundary.
  4. Generate controlled variants. Mutate one condition at a time. Keep the task and most evidence stable so the failure has a plausible cause.
  5. Run repeated trials. Record the model or workflow version, prompt, configuration, tools, data, trial ID, output, and trace. Repeat enough to see whether a discovery is reproducible under the pinned setup.
  6. Classify the failure. Name the observable condition, not only the symptom. “Answered incorrectly” is weak. “Current lookup failed and stale evidence was used” is actionable.
  7. Decide regression inclusion. Apply the rule above. Add the smallest representative set that protects the decision boundary, and keep the remaining variants in the discovery pool.
  8. Re-run after the repair. A gap is not closed because a prompt changed. The failing case must now produce the expected action, and a nearby variant must still be checked.

NIST's AITE program uses blind evaluation data and a sequestered testbed to reduce train/test contamination and improve comparability. Your local fixture does not need that infrastructure, but it should preserve the same discipline: hold back generated variants from the authored set, record what was unseen, and do not train or tune on the result without versioning the change. NIST's AITE overview explains the rationale for blind data and common metrics.

If your team needs a replayable starting point, use the business-case replay fixture guide. If the failure already came from a user, the separate guide on why an evaluation suite misses a reported failure covers the handoff from incident to test case.

What this fixture does not prove

It does not prove that a general class of AI models fails at these rates. It does not compare vendors, estimate production prevalence, or establish that contradictory evidence is always harder than missing context. The local classifier is intentionally transparent, and its heuristic weaknesses are part of the artifact.

The result does support a narrower engineering decision: a nominal authored suite cannot tell you whether it has named the important conditions unless you test the boundary of the suite. The generated variants are a discovery instrument. They are not a substitute for production traces, domain review, or a release gate tied to the real cost of error.

Lippmann, Spaan, and Yang describe unknown-unknown errors as high-confidence misclassifications that cluster in blind spots and use targeted synthetic samples to characterize them. Their paper supports the method's underlying logic, while the numbers in this post come only from the local fixture.

When I teach product managers to move from writing specifications to building and shipping, the recurring failure is often an undefined “done,” not a missing tool. That is the same boundary this fixture makes visible: before asking whether the answer is good, define what the workflow must do when the evidence is not enough.

The next release decision

Do not report the authored pass rate by itself. Report the named coverage, the untested axes, the first generated failures, and the regression decisions they caused.

For this fixture, the next decision is clear: keep the 12-case authored baseline, add representative checks for contradictory sources, changed state, failed current lookups, and missing owner acceptance, and retain the remaining controlled variants as discovery probes. Then rerun the repaired workflow before expanding its authority.

Marius Manolachi helps existing people become capable of building AI products on their own work through AI consulting and tutoring. If you are at the point where a feature has a pass rate but no credible blind-spot map, the useful review packet is small: the action contract, the current suite, one generated variant, its trace, and the proposed regression rule. The AI consulting and tutoring page is the next step if you want help turning that packet into a repeatable team practice.