Why AI coding agents fail on real pull requests: every failure from 12 attempts
32 AI coding agent misses from our own runs, sorted into 8 classes: wrong answers, gates, caps, lost context and more. Counts, not rates.
TL;DR
- We found 32 misses in our own runs and sorted them into 8 classes. A miss is a run that did not verify, resolve or pass. They come from 12 real pull-request attempts (11 misses), 33 SWE-bench Verified attempts (8) and 152 hard-set calls (13).
- On real pull-request tasks, the stop records do not prove the root cause. The 11 misses show 4 gates, 3 caps, 1 lost context and 3 stops that name only "4 consecutive failures". 5 of the 11 passed the offline gates and 2 were empty. One uvicorn patch also failed worker-written cross-tests.
- On strict answer keys, wrong answers lead. 12 of the 32 are wrong answers, and 5 more are right answers in the wrong format. Haiku 4.5 gave 8 of the 12 and all 5.
- These are counts, not rates, and we chose the classes.
Why AI coding agents fail: 8 classes, 32 misses
Wrong-answer labels lead and gate labels come next. One model supplies most wrong answers and every format miss. These selected runs mix different tasks and units. The counts do not estimate how often each cause happens.
| Class | Real PRs (of 11) | SWE-bench (of 8) | Hard set (of 13) | Total (of 32) |
|---|---|---|---|---|
| Wrong answer | 0 | 4 | 8 | 12 |
| Gates and checks | 4 | 2 | 0 | 6 |
| Format miss | 0 | 0 | 5 | 5 |
| Budget or time cap | 3 | 0 | 0 | 3 |
| Stop cause not named | 3 | 0 | 0 | 3 |
| Environment and tooling | 0 | 1 | 0 | 1 |
| Lost context | 1 | 0 | 0 | 1 |
| Missing fact | 0 | 1 | 0 | 1 |
Each miss gets one class from the recorded failure or stop. A class is our reading, not proof of the root cause. We note other findings without counting them again. Below are all 12 real pull-request attempts: three tasks, four builds, one attempt each.
| Task | Build | Ended | Class |
|---|---|---|---|
| fastify/session | Baseline | Escalated: 4 failures in a row | Not named |
| fastify/session | Fix wave 1 | $5 cap | Cap |
| fastify/session | f0ac3a8a | Escalated: 6 plans turned back | Gate |
| fastify/session | 236c0d3f | Delivered | Verified |
| h3js/h3 | Baseline | Escalated: 4 failures in a row | Not named |
| h3js/h3 | Fix wave 1 | Escalated: 4 failures in a row | Not named |
| h3js/h3 | f0ac3a8a | Delivered | Gate |
| h3js/h3 | 236c0d3f | Escalated: loop guard | Lost context |
| Kludex/uvicorn | Baseline | $5 cap | Cap |
| Kludex/uvicorn | Fix wave 1 | 20 minute cap | Cap |
| Kludex/uvicorn | f0ac3a8a | Delivered | Gate |
| Kludex/uvicorn | 236c0d3f | Delivered | Gate |
Wrong answers and format misses: 12 and 5
Four SWE-bench instances ended unresolved with a non-empty patch and no recorded platform cause: django-10554, django-14034, django-15022 and sympy-20428. We class them as wrong fixes because they failed official grading. That result does not identify why they failed. In 3 of the 4, no public panel run solved the instance either (a calculation). Agent resolved 25 of 33 (76%, 95% interval 59% to 87%).
On the hard set, Haiku 4.5 gave 8 wrong answers and 5 right answers in the wrong format, such as inside a code fence. Haiku passed 11 of 24 strictly (46%, 95% interval 28% to 65%). Four other configurations passed 24 of 24 each (95% interval 86% to 100%). Two passed 16 of 16 each (81% to 100%). The set hits a ceiling for those six configurations. Their intervals overlap, so pass rate cannot rank them.
Gates and checks: 6
A gate is a platform check that refuses work or judges a result. We put six misses in this class. Some stopped work; others left a delivered patch unverified.
- fastify/session, f0ac3a8a: the plan gates turned back 6 plans. The patch was empty.
- h3, f0ac3a8a: the patch passed the gates. The account-route label did not match, so it counts as unverified.
- uvicorn, f0ac3a8a: the acceptance check and the ws17 cross-tests did not run.
- uvicorn, 236c0d3f: the full suite passed, but two worker-written cross-tests failed under websockets 17.0. The base tests timed out instead of failing, which also blocked verification. A parser missed test ids with spaces; the command exit code still showed the cross-test failure.
- sympy-18189 (SWE-bench): a plan gate read the test convention too literally and refused an existing test file.
- astropy-7336 (SWE-bench): the worker edited a test file that the official patch renames, so the official tests never ran.
For fastify, the record does not say if the gates were right.
Budget or time caps: 3
fastify/session hit the $5 cap at $4.85, and uvicorn hit it at $4.88 (list-price calculations). A second uvicorn run hit the 20 minute cap. Together they cost $14.36 (a calculation), and none reached verified delivery. The first run (2026-10-02) also stopped at $4.89 (a list-price calculation), outside the 12.
Without caps, uvicorn cost $4.74 and $4.57 (list-price calculations). The capped runs stopped before completion, so their spend does not show the cost of finishing. The builds also differ; these figures cannot show what removing the cap caused.
Stop cause not named: 3
fastify/session and h3 at baseline, and h3 at fix wave 1, ended with "4 consecutive failures". Each patch had passed the gates. The stop note names no failing step, so we guess no class.
Environment and tooling: 1
SWE-bench django-14631: the sandbox lacked lint tools. The pipeline held delivery, and the patch was empty. Both capped uvicorn runs also show step-fail:install. We class those as caps.
Two groups sit outside the 32. Six SWE-bench instances could not import compiled extensions in campaign 1, so a declared rule replaced them. And 30 earlier Codex CLI attempts stopped before any model call, because the CLI showed no signed-in account. The hard-set study reports them as not scored.
Lost context: 1
h3 on build 236c0d3f lost a formatting failure across a context fold and ended with an empty patch. The stop note shows the loop guard: the worker kept repeating one check. The same task passed the gates on the build before. One attempt cannot separate a regression from chance.
Missing or stale facts: 1, plus 15 memory sessions
SWE-bench sympy-22456: the worker asked a requirements question. Nobody answered, by declared rule. The patch was empty.
The agent memory study tests a related missing-fact problem. Without the late-fee rate on file, 0 of 15 sessions passed the late-fee checks (95% interval 0% to 20%). Sonnet 5.5 asked in 7 of 9 (45% to 94%) and guessed in 2 of 9 (6% to 55%). Sonnet flagged the missing or guessed rate in 9 of 9 final messages (95% interval 70% to 100%).
Haiku 4.5 guessed without flagging it in 6 of 6 (61% to 100%). With the rate on file, Sonnet passed the late-fee checks in 15 of 15 (80% to 100%). Haiku passed them in 10 of 10 (72% to 100%). These figures measure hidden-test results, not whether the whole task passed.
The README also held a stale test command. Sonnet ran that broken command in 12 of 15 no-memory sessions (95% interval 55% to 93%). It ran it in 0 of 15 with the curated file (0% to 20%). The 32 exclude memory sessions, which use another unit. More: Does CLAUDE.md help?
What changed across builds
Verified delivery went 0, 0, 0, then 1 of 3. Guardrail refusals went 26, 28, 25, then 19.
In the two capped slices, none of the 6 attempts reached "Delivered": 3 hit a cap and 3 escalated. In the two uncapped slices, 4 of 6 reached it and 1 verified. The 5 misses were 4 gates and 1 lost context.
A refusal is not always a failure, because some catch real problems. Each build changed the platform, and the last two removed the caps. No pair of slices is a matched comparison, so we do not credit the builds.
What to fix first (our reading, not tested)
- Name the failing step in every stop record. 3 of 11 real-PR misses carry only a count.
- Test each gate on a known-good patch and on an empty one. We labelled 6 of 32 misses as gates, versus 3 as caps. This is a count comparison, not a cause estimate.
- Carry the latest failure across a context fold. One miss shows why.
- Write down facts the repository cannot show. No no-rate session passed the late-fee checks. With the rate on file, 25 of 25 passed them (pooled calculation; 95% interval 87% to 100%).
How we measured
Classes come from the why codes, stop labels and stop notes in each study's run data. The calibration extract supplies its stop notes. The SWE-bench attempts and platform findings supply its grading results and holds. The hard-set receipts separate wrong answers from format misses. The Sonnet memory receipts and Haiku memory receipts supply the memory counts.
Class totals are calculations over these records, not population rates, so they carry no interval. Other intervals are 95% Wilson. Memory intervals pool repeated sessions on the same tasks and describe this sample; they are not independent-task estimates. Costs are notional list-price calculations for subscription calls, not invoices. "Delivered" means the internal in-review state, not a merged upstream pull request.
Caveats
- One attempt per cell finds defects. It does not measure a rate.
- The 12 attempts cover 3 tasks, not 12.
- These are systems, not models. The pipeline is part of every real-PR miss.
- The class for each miss is our judgment. A failed test, stop label or gate does not by itself prove a root cause.
What to read next
- Coding calibration study
- SWE-bench Verified study
- Why we count every failed attempt
- Harness vs model: where agent gains come from
See how Agent works on a pull request in your own repository: Try Agent.