• Coding Agents
  • Failures
  • Calibration
  • SWE-bench

Why AI coding agents fail on real pull requests: every failure from 12 attempts

32 AI coding agent misses from our own runs, sorted into 8 classes: wrong answers, gates, caps, lost context and more. Counts, not rates.

TL;DR

  • We found 32 misses in our own runs and sorted them into 8 classes. A miss is a run that did not verify, resolve or pass. They come from 12 real pull-request attempts (11 misses), 33 SWE-bench Verified attempts (8) and 152 hard-set calls (13).
  • On real pull-request tasks, the stop records do not prove the root cause. The 11 misses show 4 gates, 3 caps, 1 lost context and 3 stops that name only "4 consecutive failures". 5 of the 11 passed the offline gates and 2 were empty. One uvicorn patch also failed worker-written cross-tests.
  • On strict answer keys, wrong answers lead. 12 of the 32 are wrong answers, and 5 more are right answers in the wrong format. Haiku 4.5 gave 8 of the 12 and all 5.
  • These are counts, not rates, and we chose the classes.

Why AI coding agents fail: 8 classes, 32 misses

Wrong-answer labels lead and gate labels come next. One model supplies most wrong answers and every format miss. These selected runs mix different tasks and units. The counts do not estimate how often each cause happens.

ClassReal PRs (of 11)SWE-bench (of 8)Hard set (of 13)Total (of 32)
Wrong answer04812
Gates and checks4206
Format miss0055
Budget or time cap3003
Stop cause not named3003
Environment and tooling0101
Lost context1001
Missing fact0101

Each miss gets one class from the recorded failure or stop. A class is our reading, not proof of the root cause. We note other findings without counting them again. Below are all 12 real pull-request attempts: three tasks, four builds, one attempt each.

  • Functional pass (offline gates)
  • Verified delivery
  • Pull request opened
Baseline (capped)
Fix wave 1 (capped)
Uncapped, build f0ac3a8a
Uncapped, build 236c0d3f

One square per item of n; 3 strips per row, one per class.

4 rows, 3 series: Functional pass (offline gates), Verified delivery, Pull request opened. Functional pass (offline gates): highest Baseline (capped) 2 (n 3). Lowest Uncapped, build 236c0d3f 1 (n 3). Verified delivery: highest Uncapped, build 236c0d3f 1 (n 3). Lowest Uncapped, build f0ac3a8a 0 (n 3).

Notesn = 3 per row

Tasks per slice: functional pass, verified delivery, pull request opened

One attempt per task per slice. The two capped slices stopped at 20 minutes or $5; the uncapped slices had no ceiling. Each slice is a different platform build, so a change is not a matched improvement.

Source: Coding calibration: fastify/session, h3, uvicorn

TaskBuildEndedClass
fastify/sessionBaselineEscalated: 4 failures in a rowNot named
fastify/sessionFix wave 1$5 capCap
fastify/sessionf0ac3a8aEscalated: 6 plans turned backGate
fastify/session236c0d3fDeliveredVerified
h3js/h3BaselineEscalated: 4 failures in a rowNot named
h3js/h3Fix wave 1Escalated: 4 failures in a rowNot named
h3js/h3f0ac3a8aDeliveredGate
h3js/h3236c0d3fEscalated: loop guardLost context
Kludex/uvicornBaseline$5 capCap
Kludex/uvicornFix wave 120 minute capCap
Kludex/uvicornf0ac3a8aDeliveredGate
Kludex/uvicorn236c0d3fDeliveredGate

Wrong answers and format misses: 12 and 5

Four SWE-bench instances ended unresolved with a non-empty patch and no recorded platform cause: django-10554, django-14034, django-15022 and sympy-20428. We class them as wrong fixes because they failed official grading. That result does not identify why they failed. In 3 of the 4, no public panel run solved the instance either (a calculation). Agent resolved 25 of 33 (76%, 95% interval 59% to 87%).

On the hard set, Haiku 4.5 gave 8 wrong answers and 5 right answers in the wrong format, such as inside a code fence. Haiku passed 11 of 24 strictly (46%, 95% interval 28% to 65%). Four other configurations passed 24 of 24 each (95% interval 86% to 100%). Two passed 16 of 16 each (81% to 100%). The set hits a ceiling for those six configurations. Their intervals overlap, so pass rate cannot rank them.

  • Strict pass
  • Format miss (correct answer, wrong format)
  • Wrong answer
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

One square per call; counts at the right are exact and in legend order.

7 rows, 3 series: Strict pass, Format miss (correct answer, wrong format), Wrong answer. Strict pass: highest Claude Sonnet 5.5 · Claude Code 24 (n 24). Lowest Claude Haiku 4.5 · Claude Code 11 (n 24). Format miss (correct answer, wrong format): highest Claude Haiku 4.5 · Claude Code 5 (n 24). Lowest GPT-6.1 Sol (high) · Codex CLI 0 (n 16).

Notesn 16–24 per row

Counts per configuration: strict passes, format misses and wrong answers

A format miss is a reply that the strict validator rejected (for example, wrapped in a code fence or in prose although the prompt said not to) whose extracted answer passes the same validator. It is not a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Gates and checks: 6

A gate is a platform check that refuses work or judges a result. We put six misses in this class. Some stopped work; others left a delivered patch unverified.

  • fastify/session, f0ac3a8a: the plan gates turned back 6 plans. The patch was empty.
  • h3, f0ac3a8a: the patch passed the gates. The account-route label did not match, so it counts as unverified.
  • uvicorn, f0ac3a8a: the acceptance check and the ws17 cross-tests did not run.
  • uvicorn, 236c0d3f: the full suite passed, but two worker-written cross-tests failed under websockets 17.0. The base tests timed out instead of failing, which also blocked verification. A parser missed test ids with spaces; the command exit code still showed the cross-test failure.
  • sympy-18189 (SWE-bench): a plan gate read the test convention too literally and refused an existing test file.
  • astropy-7336 (SWE-bench): the worker edited a test file that the official patch renames, so the official tests never ran.

For fastify, the record does not say if the gates were right.

Budget or time caps: 3

fastify/session hit the $5 cap at $4.85, and uvicorn hit it at $4.88 (list-price calculations). A second uvicorn run hit the 20 minute cap. Together they cost $14.36 (a calculation), and none reached verified delivery. The first run (2026-10-02) also stopped at $4.89 (a list-price calculation), outside the 12.

Baseline (capped)

fastify/session
h3js/h3
Kludex/uvicorn

Fix wave 1 (capped)

fastify/session
h3js/h3
Kludex/uvicorn

Uncapped, build f0ac3a8a

fastify/session
h3js/h3
Kludex/uvicorn

Uncapped, build 236c0d3f

fastify/session
h3js/h3
Kludex/uvicorn

One panel per series, all on the same axis.

3 rows, 4 series: Baseline (capped), Fix wave 1 (capped), Uncapped, build f0ac3a8a, Uncapped, build 236c0d3f. Baseline (capped): highest Kludex/uvicorn $4.88. Lowest fastify/session $3.43. Fix wave 1 (capped): highest fastify/session $4.85. Lowest h3js/h3 $3.29.

Notes

List-price estimates of subscription calls. Failed and capped attempts count.

Source: Coding calibration: fastify/session, h3, uvicorn

Without caps, uvicorn cost $4.74 and $4.57 (list-price calculations). The capped runs stopped before completion, so their spend does not show the cost of finishing. The builds also differ; these figures cannot show what removing the cap caused.

Stop cause not named: 3

fastify/session and h3 at baseline, and h3 at fix wave 1, ended with "4 consecutive failures". Each patch had passed the gates. The stop note names no failing step, so we guess no class.

Environment and tooling: 1

SWE-bench django-14631: the sandbox lacked lint tools. The pipeline held delivery, and the patch was empty. Both capped uvicorn runs also show step-fail:install. We class those as caps.

Two groups sit outside the 32. Six SWE-bench instances could not import compiled extensions in campaign 1, so a declared rule replaced them. And 30 earlier Codex CLI attempts stopped before any model call, because the CLI showed no signed-in account. The hard-set study reports them as not scored.

Lost context: 1

h3 on build 236c0d3f lost a formatting failure across a context fold and ended with an empty patch. The stop note shows the loop guard: the worker kept repeating one check. The same task passed the gates on the build before. One attempt cannot separate a regression from chance.

Missing or stale facts: 1, plus 15 memory sessions

SWE-bench sympy-22456: the worker asked a requirements question. Nobody answered, by declared rule. The patch was empty.

The agent memory study tests a related missing-fact problem. Without the late-fee rate on file, 0 of 15 sessions passed the late-fee checks (95% interval 0% to 20%). Sonnet 5.5 asked in 7 of 9 (45% to 94%) and guessed in 2 of 9 (6% to 55%). Sonnet flagged the missing or guessed rate in 9 of 9 final messages (95% interval 70% to 100%).

Haiku 4.5 guessed without flagging it in 6 of 6 (61% to 100%). With the rate on file, Sonnet passed the late-fee checks in 15 of 15 (80% to 100%). Haiku passed them in 10 of 10 (72% to 100%). These figures measure hidden-test results, not whether the whole task passed.

The README also held a stale test command. Sonnet ran that broken command in 12 of 15 no-memory sessions (95% interval 55% to 93%). It ran it in 0 of 15 with the curated file (0% to 20%). The 32 exclude memory sessions, which use another unit. More: Does CLAUDE.md help?

What changed across builds

Verified delivery went 0, 0, 0, then 1 of 3. Guardrail refusals went 26, 28, 25, then 19.

Baseline (capped)
Fix wave 1 (capped)
Uncapped, build f0ac3a8a
Uncapped, build 236c0d3f

Hover or focus a bar for its ratio to Uncapped, build 236c0d3f (the highlighted row): a ratio of the two values shown, not a measurement.

4 rows. Highest Fix wave 1 (capped) 28. Lowest Uncapped, build 236c0d3f 19.

Notes

Tool calls the platform turned back (plan schema, delivery contract, outgoing text, loops)

A refusal is a guardrail working, not always a failure: some catch real problems, some were platform defects that the next build fixed.

Source: Coding calibration: fastify/session, h3, uvicorn

In the two capped slices, none of the 6 attempts reached "Delivered": 3 hit a cap and 3 escalated. In the two uncapped slices, 4 of 6 reached it and 1 verified. The 5 misses were 4 gates and 1 lost context.

A refusal is not always a failure, because some catch real problems. Each build changed the platform, and the last two removed the caps. No pair of slices is a matched comparison, so we do not credit the builds.

What to fix first (our reading, not tested)

  1. Name the failing step in every stop record. 3 of 11 real-PR misses carry only a count.
  2. Test each gate on a known-good patch and on an empty one. We labelled 6 of 32 misses as gates, versus 3 as caps. This is a count comparison, not a cause estimate.
  3. Carry the latest failure across a context fold. One miss shows why.
  4. Write down facts the repository cannot show. No no-rate session passed the late-fee checks. With the rate on file, 25 of 25 passed them (pooled calculation; 95% interval 87% to 100%).

How we measured

Classes come from the why codes, stop labels and stop notes in each study's run data. The calibration extract supplies its stop notes. The SWE-bench attempts and platform findings supply its grading results and holds. The hard-set receipts separate wrong answers from format misses. The Sonnet memory receipts and Haiku memory receipts supply the memory counts.

Class totals are calculations over these records, not population rates, so they carry no interval. Other intervals are 95% Wilson. Memory intervals pool repeated sessions on the same tasks and describe this sample; they are not independent-task estimates. Costs are notional list-price calculations for subscription calls, not invoices. "Delivered" means the internal in-review state, not a merged upstream pull request.

Caveats

  • One attempt per cell finds defects. It does not measure a rate.
  • The 12 attempts cover 3 tasks, not 12.
  • These are systems, not models. The pipeline is part of every real-PR miss.
  • The class for each miss is our judgment. A failed test, stop label or gate does not by itself prove a root cause.

See how Agent works on a pull request in your own repository: Try Agent.

The data behind this post

  • Calibration
  • Coding Agents

Coding calibration: what broke on three real pull requests

One attempt per task, four platform builds, failures kept: how an AI worker did on real fastify/session, h3 and uvicorn issues, and what broke.

1 of 3Verified deliveries, latest build · n = 3

4 chartsUpdated October 5, 2026

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.