Field note · implementation

Why Does AI Fail When Scores Route Poorly?

A reproducible routing fixture shows why one confidence threshold can auto-accept wrong cases, reject correct ones, and overload review.

14 minute read
  • confidence routing
  • AI evaluation
Illustration of an AI confidence score splitting into auto-accept, human-review, and reject routes

The dangerous routing bug is easy to miss. The score looks precise, the threshold looks reasonable, and the queue still receives the wrong work.

I built the small fixture below to isolate that bug. It is not a production-model benchmark. It is a deterministic test of the policy that turns a score into auto_accept, human_review, or reject.

The observed routing failure

In this fixture, the baseline >=0.80 threshold sent 3 of 8 high-score calibration cases to the wrong route. On six held-out cases, it auto-accepted four cases, three of them wrong. Raising the global threshold to >=0.90 removed those false auto-accepts, but a failure-mode-specific review rule was still needed to stop one correct low-score case from being rejected.

Held-out policyAuto-acceptHuman reviewRejectFalse auto-acceptsFalse reject escalations
Baseline, high score >=0.8041131
Calibrated global, high score >=0.9014101
Calibrated plus failure-mode review15000

That is the sourceable result of this page. It supports a diagnosis, not a universal threshold. Your first job is to find out which part of your policy is failing.

Why a high score can still send the wrong case

A confidence score normally answers a narrower question than the route asks. It may rank outputs from more to less likely to be correct. The queue needs a decision about what to do next, under a particular error cost and with a particular intervention available.

Selective prediction makes that distinction explicit: a system can abstain instead of accepting every prediction. The useful object is the trade-off between risk among accepted cases and coverage, not a single accuracy number. The ACL study on selective prediction treats abstention and confidence estimation as first-class parts of the classifier, while the NeurIPS selective-classification paper defines coverage and selective risk as the paired quantities to inspect.

The route can fail in four different ways:

Failure locationWhat you seeWhat to test
Score meaningWrong and right cases share a score bandCalibration bins by slice and task
ThresholdThe score ranks cases acceptably, but route volume or error is wrongRisk-coverage curve and threshold sensitivity
SliceEasy cases look safe while borderline, new, or contradictory cases do notCalibration and route metrics by slice
InterventionReview or retrieval cannot resolve the failureReplay the route with the actual repair step

When I teach product managers to move from writing specs to building and shipping, the failure is usually an undefined “done,” not the model. Routing has the same dependency: define what correct means, what the score predicts, and what each queue can actually repair before choosing a cutoff.

How to reproduce the failure with a labelled fixture

Use a small fixture that forces the score to meet the cases you care about. Do not sample only easy examples. The point is to make the queue see borderline, new-format, and failure-mode cases before production does.

The fixture below pins fixture-oracle-v1, a deterministic score provider. It is deliberately not a named LLM and makes no API call. correct=1 means the case label says the proposed output is correct. The calibration rows choose the repair. The held-out rows check it.

For generated language, “correct” must belong to the task rubric, not to the score itself. The NeurIPS 2024 work on selective generation uses a textual-entailment relation and human-annotated data to define the correctness target before controlling selection. Read the paper when your workflow has no simple binary answer.

IDSplitSliceCaseScoreCorrect
A1calibrationeasyExplicit evidence and clean match.941
A2calibrationeasyComplete input and known format.911
A3calibrationeasySingle rule with supporting evidence.881
A4calibrationeasyFamiliar request with one clear action.821
B1calibrationborderlinePersuasive exception hides a missing approval field.840
B2calibrationborderlineBorderline request with enough evidence.731
B3calibrationborderlineBorderline request with a subtle policy conflict.690
B4calibrationborderlineCorrect exception near the route boundary.801
O1calibrationout-of-distributionNew form layout with an unseen region code.860
O2calibrationout-of-distributionNew layout but evidence is sufficient.661
O3calibrationout-of-distributionUnseen layout with a missing field.580
O4calibrationout-of-distributionNovel layout with a correct answer scored too low.421
F1calibrationfailure-modeFluent answer contradicts retrieved evidence.890
F2calibrationfailure-modeCorrect answer with contradictory context checked.781
F3calibrationfailure-modeFluent answer repeats an unsupported premise.570
F4calibrationfailure-modeCorrect answer is cautious but scored low.311
H1held-outeasyKnown format with explicit evidence.901
H2held-outborderlineConfident wording misses the approval condition.810
H3held-outborderlineBorderline request with complete evidence.641
H4held-outout-of-distributionNew layout receives a confident wrong mapping.830
H5held-outfailure-modeFluent output ignores a contradictory source.870
H6held-outfailure-modeCautious low score accompanies a correct answer.521

Run each case three times with a fixed score offset of -0.02, 0.00, and +0.02. This is not a claim about live-model randomness. It checks whether a small score movement changes the route. B4 moves from review at .78 to auto-accept at .80, then stays auto-accepted at .82.

Failure trace from score to route

The trace matters because a bad queue can hide two different errors: a wrong case accepted too early, and a correct case rejected too aggressively. Follow each case from its pinned score to the route, then apply the smallest policy change that addresses that error.

Case traceScoreBaseline routeLabelRepaired routeWhat the trace shows
B1, persuasive exception hides approval field.84auto-accept0human review at .90 policyA high score created a false auto-accept; the global threshold moved it to review.
H6, cautious answer is correct.52reject1human review with failure-mode overrideThe score hid a correct case; a higher global threshold could not fix it, but the slice route preserved review.

This is a policy trace, not a claim that every .84 or .52 score behaves the same way elsewhere. In your fixture, record the input slice, score definition, threshold comparison, assigned route, correctness label, and the intervention's result. If one of those fields is missing, you can describe the symptom but not reproduce the failure.

Illustration of a confidence score flowing through separate auto-accept, human-review, reject, and failure-mode routes

What the score bands actually showed

The baseline score bands were not calibrated in this fixture:

Score bandCasesCorrectEmpirical correctness
0.80-1.008562.5%
0.55-0.796350.0%
Below 0.5522100.0%

The high band contains B1, O1, and F1, all wrong. The low band contains O4 and F4, both correct. A score that behaves this way cannot be trusted as a global safety boundary, even though it may still be useful as one feature in a more specific policy.

The slice view shows where the global band hides the problem:

SliceCasesAuto-accept correct / wrongReview correct / wrongReject correct / wrong
Easy44 / 00 / 00 / 0
Borderline41 / 11 / 10 / 0
Out-of-distribution40 / 11 / 11 / 0
Failure-mode40 / 11 / 11 / 0

The baseline routing counts make the operational cost visible:

SplitAuto-acceptHuman reviewRejectFalse auto-acceptsCorrect cases rejected
Calibration86232
Held-out41131

For held-out cases, baseline auto-accept coverage was 4/6, or 66.7%. Risk among auto-accepted cases was 3/4, or 75.0%. That is the number a review queue hides when it reports only overall accuracy.

Selective-classification work uses this same risk and coverage framing. A recent LLM abstention preprint also makes the practical point that a threshold should be selected on held-out labelled cases against an explicit risk target, rather than treated as meaningful by itself. Its guarantee is conditional on its assumptions. That caveat matters here.

Diagnose whether the score or the policy is broken

Use this order. It keeps a threshold change from masking a score problem.

  1. Check the score definition. Write down what a score of .80 is supposed to mean. Correctness probability, preference, evidence sufficiency, and output quality are different targets.
  2. Bin labels by score. Compare empirical correctness with the reported band. Do this overall and for each important slice. A global bin can look acceptable while a failure slice is unsafe.
  3. Draw risk against coverage. Sort by score, accept the top portion, and record error among accepted cases. If risk does not fall as coverage shrinks, the score is not ranking safety well enough for selective routing.
  4. Replay the intervention. A review route is not a repair until a reviewer can see the missing evidence and change the decision. A retrieval or re-check route needs its own held-out result.
  5. Repeat near boundaries. If small score movement flips the queue, treat the boundary as unstable until the route has a margin or a second check.

The diagnosis is usually clear after this table:

ObservationLikely breakFirst repair
Correctness is mixed inside every score bandScore is not calibrated for the taskRecalibrate with held-out labels or change the score target
Score ranking is useful but accepted error is too highThreshold is too permissiveMove the threshold and measure coverage loss
Only OOD or contradiction cases failSlice or failure mode is hiddenAdd a slice flag and a separate intervention
Review catches errors but the queue growsIntervention capacity is the bottleneckMeasure review load and narrow auto-accept

What repair worked on the held-out cases

The first repair raised only the global auto-accept threshold from .80 to .90. That produced one correct auto-accept, four reviews, one reject, and no false auto-accepts. It did not solve H6, a correct failure-mode case with a score of .52.

The second repair added one policy rule: every failure-mode case goes to human review, regardless of score. This is not “trust humans” as a universal answer. It is a claim about the intervention in this fixture: contradictory evidence needs inspection, and the score did not identify that need reliably.

Held-out policyAuto-accept correct / wrongReview correct / wrongReject correct / wrongCost units
Baseline1 / 31 / 01 / 034
Calibrated global1 / 01 / 31 / 07
Calibrated plus failure-mode review1 / 02 / 30 / 05

The cost weights are one unit for a review, ten for a false auto-accept, and three for a correct case sent to reject. They are sensitivity weights, not a forecast of money saved. Change them to your own reversibility, harm, and reviewer-time assumptions.

The Nature Machine Intelligence study on competing overconfidence and underconfidence in LLMs is a useful reason to test the failure-mode route separately: its abstract reports both persistence with an initial answer and disproportionate updating after contradictory information. That does not prove the same behavior in your workflow. It does explain why “add more feedback” is not a complete routing policy.

Verification evidence from held-out cases

A repair passes only when the policy is rerun on cases that did not choose the threshold and the route-level errors are counted again. The verification evidence for this fixture is the row-level result below, not the improved appearance of one score band.

Held-out caseLabelBaseline routeCalibrated global routeCalibrated plus failure-mode route
H1, known format1auto-acceptauto-acceptauto-accept
H2, missed approval condition0auto-accepthuman reviewhuman review
H3, complete borderline evidence1human reviewhuman reviewhuman review
H4, confident new-layout mapping0auto-accepthuman reviewhuman review
H5, contradictory source ignored0auto-accepthuman reviewhuman review
H6, cautious answer is correct1rejectrejecthuman review

The global repair therefore changes held-out results from four auto-accepts with three wrong to one correct auto-accept with no wrong auto-accepts. It still rejects H6, so it is incomplete for this fixture. The slice override moves H6 to review, leaving zero false auto-accepts and zero correct cases rejected, while reducing auto-accept coverage to 1 of 6. That is the verification decision: accept the policy only if the lower coverage and added review work fit the team's stated error and capacity limits.

The scoring script and what it does not prove

This compact script is the exact route calculation used for the fixture. The full raw input is the table above.

CASES = [
    ("A1", "easy", .94, 1), ("A2", "easy", .91, 1),
    ("A3", "easy", .88, 1), ("A4", "easy", .82, 1),
    ("B1", "borderline", .84, 0), ("B2", "borderline", .73, 1),
    ("B3", "borderline", .69, 0), ("B4", "borderline", .80, 1),
    ("O1", "ood", .86, 0), ("O2", "ood", .66, 1),
    ("O3", "ood", .58, 0), ("O4", "ood", .42, 1),
    ("F1", "failure-mode", .89, 0), ("F2", "failure-mode", .78, 1),
    ("F3", "failure-mode", .57, 0), ("F4", "failure-mode", .31, 1),
    ("H1", "easy", .90, 1), ("H2", "borderline", .81, 0),
    ("H3", "borderline", .64, 1), ("H4", "ood", .83, 0),
    ("H5", "failure-mode", .87, 0), ("H6", "failure-mode", .52, 1),
]

def route(score, calibrated=False, slice_name=None):
    if calibrated and slice_name == "failure-mode":
        return "human_review"
    high = .90 if calibrated else .80
    if score >= high: return "auto_accept"
    if score >= .55: return "human_review"
    return "reject"

for case_id, slice_name, base_score, correct in CASES:
    for offset in (-.02, 0.00, .02):
        score = round(max(0, min(1, base_score + offset)), 2)
        print(case_id, score, route(score), route(score, True, slice_name), correct)

This proves that the stated policy produces the stated route counts for this fixture. It does not prove a named model is calibrated, that six held-out cases are enough for a release decision, or that the illustrative cost weights match your business.

What to publish before changing production routing

Put these artifacts in the change review:

  • raw cases with labels, slices, and held-out membership;
  • the exact model or score provider, prompt, configuration, date, and repeat protocol;
  • score bins with empirical correctness, plus risk-coverage at the proposed coverage levels;
  • route confusion counts for auto-accept, review, and reject;
  • review-time assumptions and separate weights for false auto-accepts and false escalations;
  • at least one failure example per route that matters;
  • the repaired policy and a held-out result;
  • limits, including slices not represented in the fixture.

For the broader implementation sequence, start with How to implement confidence routing. If the problem is specifically score calibration, use How to calibrate AI confidence scores for a workflow. If the workflow should stop rather than guess, compare the route with When should an AI workflow stop instead of guessing?.

Marius Manolachi teaches teams to build AI products on their own work. The useful handoff is not a magical threshold. It is a rerunnable fixture that lets the team explain why a case entered each queue, what the intervention can fix, and what evidence would justify changing the policy.

Questions people ask next

How do I know whether the score or the threshold is broken?

Bin labelled cases by score and slice. If correctness is mixed or inverted inside the same band, the score is not calibrated for that slice. If the bands are reliable but the route still exceeds your error or review budget, the threshold or intervention policy is the problem.

Should every AI workflow use the same confidence threshold?

Only when the task, population, score meaning, error costs, and intervention are comparable and the threshold has been checked on held-out labels. Otherwise calibrate or route by slice.

Does adding human review fix poor AI routing?

Only if reviewers can resolve the failure and the queue has capacity. Review can reduce unsafe auto-accepts, but it cannot repair a score that hides a failure mode or a queue that nobody can process.