Field note · implementation

Why Does AI Fail When Review Becomes a Bottleneck?

A reproducible queueing simulation shows when faster AI drafts create review delay, rework, and escaped errors, plus a routing decision table.

11 minute read
  • AI workflows
  • AI implementation
  • AI evaluation
Illustration of an AI review queue filling faster than human reviewers can clear it

AI can make the first draft faster and the workflow slower. That is the failure people see after a good demo: the model keeps producing work, while the same small group must inspect, correct, explain, and sometimes redo it.

When I taught product managers to move from writing specifications to building and shipping products, the failure was often not the model. It was that nobody could say what done meant. (Marius Manolachi's AI teaching work) In a review workflow, “done” also needs a capacity definition.

This is a qualitative teaching observation, not a measured failure rate. The measured result below comes from the simulator.

Here is the result that matters from the model run in this article:

High-load pathMean queue delayHuman utilizationRework min / 100 tasksEscaped-error rate
Manual6370.376 min231.7%19.9474.03%
Always-review797.814 min116.4%8.9171.83%
Risk-routed0.474 min39.0%15.2420.92%
Screened with LLM judge0.237 min27.9%19.3321.43%

These are model-based results, not a client outcome. They show why faster generation does not answer the capacity question. If the human queue is the scarce resource, the workflow fails when arrival load plus correction work exceeds the attention available to reviewers.

The bottleneck is human attention, not draft generation

AI fails at this boundary when it reduces draft time without reducing the total human work needed to release a correct result. A queue forms when the rate of work arriving for people is higher than the rate at which people can review and rework it.

The queueing paper Queue & AI: When Faster Tasks Slow Down the Workflow describes this as a gap between mean task speed and system-level delay. Its mechanism fits a common implementation mistake: the team measures how quickly AI produces a draft, but not how much scarce human attention the draft consumes after it arrives.

The useful first check is simple:

base human utilization = arrival rate x mean human service time / reviewer count

That is only a first approximation. Add review routing, error probability, rework time, and service-time variance. A workflow with a 3-minute review can still be slower than expected when errors add 5-minute corrections and service times are uneven.

The exception is a high-cost or irreversible error. If a wrong release, payment, access decision, or customer commitment is expensive to reverse, keep the human boundary even when it limits throughput. Capacity is a reason to throttle or batch the work, not a reason to silently remove necessary control.

What the capacity stress test actually ran

The experiment uses a small discrete-event simulator so you can change the inputs instead of trusting a fixed benchmark. It models two reviewers, a 10,000-minute arrival window, 20 replications per cell, and lognormal human service times with a coefficient of variation of 0.8.

InputConfiguration
Reviewers2
Arrival rates0.20, 0.50, 0.75 tasks/minute
Manual completion time6.0 min
AI draft time0.8 min
Human review time3.0 min
Rework time5.0 min
AI draft error probability12%
Manual error probability4%
Review catch rate85%
Service-time coefficient of variation0.8

The four paths are:

  1. Manual: every task receives 6 minutes of human work. A 4% error probability creates rework.
  2. Always-review: every AI draft receives human review. Review catches 85% of draft errors.
  3. Risk-routed: 30% of tasks are high-risk and reviewed; 70% bypass review with a lower modeled error probability.
  4. Screened with an LLM judge: a non-human judge screens every draft. Flagged items reach a reviewer. The judge has 80% sensitivity and a 10% false-positive rate.

The simulator records mean queue delay, human utilization, rework minutes per 100 tasks, escaped first-pass errors, and mean end-to-end cycle time. The raw output is stored in raw-results.csv; the summarized output is summary-results.csv. Reproduce it with:

python3 simulator.py \
  --config config-v1.json \
  --raw-out raw-results.csv \
  --summary-out summary-results.csv

The human service-time draw uses a lognormal distribution:

sigma^2 = ln(1 + CV^2)
mu = ln(mean) - 0.5 x sigma^2

Queue delay is the mean time from a human job becoming ready to a reviewer starting it. Human utilization is busy human minutes divided by the two-reviewer capacity during the arrival window. Rework load is rework minutes per 100 original tasks. Escaped-error rate counts first-pass errors that get past the modeled review or bypass it.

The raw stdout from the run was:

version=review-bottleneck-sim-v1
seed=20260824 horizon_minutes=10000 replications=20 reviewers=2 service_time_cv=0.8
load,policy,queue_delay_min,human_utilization,rework_min_per_100,escaped_error_rate,cycle_time_min
low,manual,3.000,0.617,20.437,0.0400,9.301
low,always-review,0.262,0.310,9.008,0.0175,4.165
low,risk-routed,0.026,0.105,15.316,0.0088,1.330
low,screened-with-llm-judge,0.015,0.074,18.231,0.0143,0.975
medium,manual,2703.290,1.557,20.367,0.0401,2818.070
medium,always-review,3.659,0.773,8.919,0.0183,7.627
medium,risk-routed,0.212,0.263,15.145,0.0089,1.379
medium,screened-with-llm-judge,0.101,0.187,19.538,0.0141,1.007
high,manual,6370.376,2.317,19.947,0.0403,6633.249
high,always-review,797.814,1.164,8.917,0.0183,816.255
high,risk-routed,0.474,0.390,15.242,0.0092,1.453
high,screened-with-llm-judge,0.237,0.279,19.332,0.0143,1.040

Illustration of a queueing simulation comparing AI review policies by arrival load, human utilization, queue delay, and rework

Faster drafts help only while the human queue has room

The full summary shows a regime change, not a universal winner:

LoadPolicyQueue delay (min)Human utilizationRework min / 100 tasksEscaped-error rateCycle time (min)
LowManual3.00061.7%20.4374.00%9.301
LowAlways-review0.26231.0%9.0081.75%4.165
LowRisk-routed0.02610.5%15.3160.88%1.330
LowScreened judge0.0157.4%18.2311.43%0.975
MediumManual2703.290155.7%20.3674.01%2818.070
MediumAlways-review3.65977.3%8.9191.83%7.627
MediumRisk-routed0.21226.3%15.1450.89%1.379
MediumScreened judge0.10118.7%19.5381.41%1.007
HighManual6370.376231.7%19.9474.03%6633.249
HighAlways-review797.814116.4%8.9171.83%816.255
HighRisk-routed0.47439.0%15.2420.92%1.453
HighScreened judge0.23727.9%19.3321.43%1.040

At low load, always-review buys a lower modeled rework load without using much human capacity. At medium load, it still works in this configuration, but utilization has reached 77.3% and the queue is no longer negligible. At high load, it crosses capacity. The queue delay becomes the result.

Risk routing keeps the queue short in all three regimes because it sends only 30% of tasks to review. It also produces fewer escaped errors than always-review in this model because the reviewed path is concentrated on higher-risk work. That is an assumption about routing quality, not proof that a confidence score will work in your process.

Screening with an LLM judge creates the shortest queue in this run. It also creates more rework than risk routing because false accepts let some errors through and false rejects send correct work into unnecessary review. The second queueing paper, When to Screen, When to Bypass, describes this trade-off directly: screening can amplify human capacity when reviewers are scarce, but can create a rework trap when workers are already stretched.

Use this decision table before you add more review

The table turns observed conditions into an operating choice. The thresholds are starting points for a pilot, not laws. Rerun the model with your measured service and error data.

Observed conditionDefault actionWhat must be trueVeto
Human utilization below 70%, queue delay stable, rework lowGate with human reviewReview catches the errors that matter and the reviewer can finish within the target SLADo not gate every low-risk item if it starves high-risk work
Utilization 70% to 90%, queue rising, risk labels have evidenceRisk-routeHigh-risk recall is measured and low-risk errors are cheap to reverseDo not trust an uncalibrated confidence score
Utilization above 90% for repeated intervals, low-risk work is reversibleAutomate or screen the low-risk pathEscaped-error monitoring, rollback, and a human escalation path existDo not auto-release irreversible or high-impact decisions
Arrival rate is bursty, review is cheaper in groups, backlog is predictableBatch reviewThe user can tolerate delay and the batch has a priority and expiry ruleDo not batch time-critical or safety-sensitive work
Utilization above 100%, rework is high, or error cost is highKeep the decision manual and throttle demandThe team accepts lower throughput until capacity or routing evidence improvesDo not call the workflow automated while people still absorb unpriced corrections

The decision rule is: choose the least human work that keeps escaped-error risk inside the policy limit, then verify that the resulting utilization stays below capacity under the highest normal load. If no path meets both conditions, the right implementation is admission control, more staffed capacity, or no automation yet.

For a broader implementation pattern, start with the queue-backed AI workflow guide. If reviewers work asynchronously, compare the capacity assumptions with the time-zone implementation guide. The assigned parent guide, How to implement AI with review capacity limits, is the place to connect this stress test to a larger rollout plan.

Monitor the workflow after the pilot

Passing a pre-launch evaluation does not prove that a reviewable workflow will remain healthy under changing demand. NIST's AI 800-4 monitoring report distinguishes operational monitoring from human-factors monitoring and notes that controlled pre-deployment evaluations cannot capture all real-world dynamics.

For this failure mode, monitor at least these fields:

FieldWhy it mattersTrigger to investigate
Arrival rateShows whether the queue input changedSustained increase above the tested regime
Human utilizationShows whether review capacity is being consumedAbove the team's agreed ceiling
Queue delay and backlog ageShows whether users are waitingRising trend or SLA breach
Review timeShows whether task mix or policy changedDrift from the simulator input
Rework minutesPrices correction work that draft metrics omitIncrease after model or prompt changes
Escaped errors and overridesTests whether routing protects the right casesAny high-severity escape or rising override rate
Reviewer ownership and incident statusMakes control actionableNo named owner, stale review, or unresolved incident

The NIST AI RMF Playbook's Govern guidance calls for monitoring frequency, defined roles, periodic review, incident response, and appeal or override processes. For a small team, that can be a short runbook: one owner checks the queue and rework dashboard each day, a second person reviews escaped errors weekly, and a named decision-maker can pause the automated path.

Run the test on your own workflow

Use the artifact before launch or before increasing arrival volume.

  1. Measure arrivals for a representative window. Keep bursts visible instead of replacing them with one average if the workflow is seasonal or deadline-driven.
  2. Measure human service time separately for first review, correction, escalation, and final approval.
  3. Estimate error probability and escaped-error probability from a labeled sample. If you cannot separate them, model a range rather than one confident value.
  4. Run manual, always-review, risk-routed, and screened paths with the same arrival stream and reviewer capacity.
  5. Choose a policy only after checking queue delay, utilization, rework, and escaped errors together. Record the policy, thresholds, owner, rollback trigger, and next review date.

Do not replace the simulator with a faster demo. The question is not whether AI can produce a draft. The question is whether the whole system can finish correct work with the people you actually have.

What this simulation cannot tell you

This is a bounded model, not a benchmark of AI systems. It assumes Poisson arrivals, independent errors, homogeneous reviewers, fixed risk-routing accuracy, no priority classes, no reviewer learning, no batching delay, and unlimited draft and judge capacity. It does not model financial loss, customer harm, compliance requirements, or the cost of hiring.

The most important unknown is routing quality. The risk-routed path wins on queue relief because its high-risk fraction and error probabilities are supplied assumptions. A bad router can turn that advantage into silent escapes. The screened path exposes the opposite trade-off: it relieves the queue more aggressively, but false accepts and false rejects add correction work.

Treat the values in the table as a reproducible starting point. Replace the inputs with your own traces, rerun the command, and keep the raw CSV beside the decision record. That is how you find out whether AI is helping the workflow or only moving the bottleneck into a less visible human queue.

Questions people ask next

Should every AI output go through human review?

No. Review every output only when the risk and available capacity justify it. If review utilization approaches capacity, validate a risk-routing or screening policy and keep a human gate for high-cost or irreversible outcomes.

Is an LLM judge a replacement for a human reviewer?

No. An LLM judge changes which items reach the human queue. False accepts can increase escaped errors, while false rejects create rework. Measure both queue relief and correction load before trusting the screen.

What should we measure after adding AI review?

Track arrival rate, human utilization, queue delay, review time, rework minutes, escaped errors, overrides, and incidents. Compare them with a pre-launch baseline and assign a named owner for review and rollback.