Field note · implementation
Why Can an AI Patch Pass Tests but Fail Its Contract?
AI patches fail review when intent, scope, and acceptance evidence stay implicit. A five-case paired diff shows what makes changes inspectable.

An AI-assisted change can pass a demo and still fail the first serious review. The code runs, the screenshot looks right, and nobody can answer a basic question: what exactly was this patch allowed to change?
When I taught product managers to move from writing specs to building and shipping, the recurring problem was not a measured model failure. It was a qualitative teaching observation: people often lacked a shared definition of done. That same gap makes AI-generated code hard to inspect.

The failure starts when the reviewer must reconstruct intent
AI changes become hard to inspect when the patch makes the reviewer infer the purpose, allowed scope, and proof of correctness from code alone. A working result does not tell you whether the change satisfies the requirement, preserves existing behavior, or quietly adds work nobody requested.
Google’s code-review guidance treats design, functionality, complexity, tests, naming, comments, style, and documentation as separate review concerns, not one pass/fail feeling. It also says a reviewer should understand every assigned line and the surrounding context. (Google Engineering Practices)
That gives a useful diagnosis:
| Inspectability question | What the reviewer needs to see | Typical opaque-patch failure |
|---|---|---|
| Why was this change made? | An intent note tied to the requirement | The reviewer reverse-engineers purpose from implementation |
| Where may it change code? | A touched-file boundary | Refactors and “helpful” modules expand the review surface |
| What must be true? | Acceptance checks and boundary cases | A happy-path test stands in for the contract |
| What must not change? | Explicit regression checks | Existing behavior moves without being named |
| How do I decide? | A review order and decision rubric | The reviewer scans, guesses, and gives a vague approval |
The problem is not that humans cannot read code. The problem is that the patch has hidden the map a human needs to read it quickly.
What the paired five-case experiment found
In the fixture, both AI-generated conditions passed their visible tests. Only the constrained sequence passed all five independent seeded checks.
| Condition | Visible tests | Seeded cases passed | Reviewer reconstruction accuracy | Missed seeded failures | Decision |
|---|---|---|---|---|---|
| Unconstrained one-pass patch | 4/4 | 3/5 | 4/5 | 1 | Request changes |
| Constrained intent-labeled sequence | 6/6 | 5/5 | 5/5 | 0 | Approve |
The reviewer accuracy and timing are from one instrumented Codex pass, not a human study. The final harness run took 37.693 ms for the unconstrained condition and 37.990 ms for the constrained condition. Those values measure file reads and oracle execution, not how long an engineer would spend thinking.
The trace is case-level rather than a single green or red label. The independent oracle records C1 through C5: the unconstrained patch fails C1, completed-task exclusion, and C4, scope and existing-order preservation; the constrained sequence passes all five. To reproduce and verify the repair, rerun the Python 3 test and oracle commands recorded in the archived method. This verifies the fixture, not a general failure rate.
The sourceable result is narrower than “AI makes code worse”: a visible green test run did not distinguish the two conditions, while the independent oracle did. The constrained change made the missing review information explicit enough for the recorded reviewer to reconstruct all five cases.
Why did the opaque patch pass tests while failing the feature contract?
It proved only the happy path and changed unrelated behavior in the same diff.
The feature request added Board.list_due(today) to a small task board. The required contract was simple:
- Return open tasks with a due date.
- Include today and exclude tomorrow.
- Sort by due date, then task ID.
- Preserve the existing priority order of
list_open(). - Do not mutate stored tasks.
The unconstrained patch did three things that made the review harder:
- It filtered for a due date but forgot
status == "open", so a completed task appeared in the due queue. - It changed
list_open()from priority order to due-date order, which was outside the request. - It added
src/reporting.py, a future-dashboard helper that the feature did not need.
Its tests still passed because they covered one open task due today and changed the existing ordering test to match the new behavior. That is a review failure hiding inside a test success. Google’s guidance makes the relevant distinction: tests should be useful and should fail when the code is broken, and a reviewer should consider the change in context rather than trust a nearby green assertion. (What to Look For in a Code Review)
The 2023 HumanEval study in the evidence bundle is a useful boundary on the larger claim. It measured generated-code quality across tools and found different correctness rates, but it did not test this patch-shape intervention or human inspectability. (Yetiştiren et al., 2023)
Which artifacts made the constrained change easier to review?
The constrained condition added evidence before adding code: an intent note, a touched-file boundary, named acceptance checks, and tests for each case.
The intent note said the new method was a read-only due-task query and that the priority ordering of list_open() must remain unchanged. The allowed paths were src/board.py, tests/test_board.py, and docs/due-task-queue.md. That boundary turned scope from an opinion into a check.
The tests then mirrored the contract instead of merely exercising the implementation:
| Acceptance check | Test evidence |
|---|---|
| Completed tasks excluded | Add an overdue completed task and assert it is absent |
| Today included, tomorrow excluded | Use adjacent dates around the review date |
| Stable ordering | Add two same-day tasks in reverse insertion order |
| Existing behavior preserved | Keep the original priority-order test unchanged |
| No mutation | Compare the stored task before and after the read |
This sequence also follows Google’s review navigation advice: start with the description and the main part of the change, then examine the rest in an appropriate order. Google explicitly recommends asking for smaller changes when a review is too large to understand. (Navigating a CL in Review)
GitHub’s review model gives the team a concrete decision vocabulary: comment, approve, or request changes. Use it. A review that ends with “looks fine” but never records which contract was checked is hard to reproduce and hard to learn from. (GitHub pull request reviews)
What should you require before approving an AI code change?
Require a small review packet whenever the reviewer cannot explain the change without reconstructing its intent from the diff.
Use this five-question gate:
- Intent: Can the author state the behavior being added in one or two sentences?
- Boundary: Can the author list the files and behaviors that are out of scope?
- Acceptance: Does every requirement have a visible check, including at least one boundary or regression case?
- Sequence: Can the reviewer inspect the change in small logical steps rather than one mixed refactor?
- Decision: Does the recorded review say approve, comment, or request changes, and name the unresolved risk?
If any answer is no, split or reshape the change before asking for approval. This is a decision rule from the fixture, not a universal threshold. It is useful because it catches the exact failure the experiment exposed: the patch can be executable and still leave the reviewer without enough evidence to judge it.
The parent implementation guide, How to Implement Reviewable AI Code Changes, covers the broader workflow. For the final pre-merge pass, use How to Review AI-Generated Code Before Merging. This article adds the narrower green-test failure diagnosis and the paired artifact.
What this experiment does not prove
It does not show that constrained prompts always produce better code. It does not estimate human review speed. It does not compare models, temperatures, repositories, or teams. It uses one small fixture, one feature, one run date, one model record, and one non-independent reviewer pass.
The broader literature also argues for caution. The 2026 study of LLM requirement-conformance review evaluated paired correct and buggy implementations across five models and reported false acceptance, false rejection, and explanation-alignment problems. That supports testing model judgements against executable evidence. It does not prove that the five-case result here will repeat at another scale. (Jin and Chen, 2026)
The practical conclusion is smaller and more useful: if an AI change is hard to inspect, first make intent, scope, acceptance checks, and regression boundaries visible. Then review the code. If your team is trying to make that capability routine, Marius Manolachi’s AI consulting and tutoring work is built around helping existing people build AI products on their own work.