Agent on SWE-bench Verified vs 11 public models
How does Agent, a full worker pipeline on one model, do on SWE-bench Verified next to public single-model runs on the very same instances?
Published · 6 charts · Download the data or a carousel
76%
The answer
Agent resolved 25 of 33 attempted instances (75.8%, 95% interval 59% to 87%). On the same instances the 11 public mini-SWE-agent v2 runs resolved between 21 and 28 (panel mean 74.1%). Every interval overlaps, so this sample cannot rank Agent above or below any panel model. Agent spent a notional $2.81 and 49 model calls per attempt, with a median of 9.6 minutes; it is slower and more expensive per instance than a bare bash agent because it onboards, plans, verifies and reviews. It resolved 1 of 4 instances that no panel model solved.
Live story
Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.
Agent Benchmarks, October 2026: 28 studies in one film
The headline of each of our 28 open benchmark studies, with sample sizes and intervals. Calculations labelled; failures counted.
Transcript
- October 2026 roundup · 28 studies. Agent Benchmarks: every headline. Recorded runs, 95% intervals where they exist, every failure counted. Calculations labelled.
- SWE-bench Verified: Agent resolved 25 of 33. The public panel averaged 74.1%; the intervals overlap, so no rank. Agent resolved (one attempt each): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Model cost per resolved instance (calculation, notional): $3.71 (n = 25). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
- SWE-bench, interim (3 of 8 pairs graded): Opus 5.5 as Agent’s brain resolved 2 of 3, Sonnet 5.5 1 of 3. Exact McNemar p = 1.0: no difference yet. Opus cost 2.6× as much, a calculation. Opus 5.5 in Agent: resolved (interim): 67% (2/3) (n = 3, 95% CI 21–94%). Sonnet 5.5 in Agent: same instances: 33% (1/3) (n = 3, 95% CI 6–79%). List-price cost, Opus vs Sonnet (calculation): 2.6× (n = 3). Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
- Blind critics preferred the AI pull request to the merged human one on 9 of 12 tasks at the latest attempt, 6 of 12 at the first. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
- Five short tasks, nine setups: 127 of 130 calls passed, so speed and tokens separate them. Fastest median: Fable 5.1 at 1.9 s. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Fastest median total time (Fable 5.1): 1.9 s (n = 15). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.
- Eight hard tasks: Sonnet 5.5, Opus 5.5, Opus 5.5 high, GPT-6.1 Sol medium, Fable 5.1 and GPT-6.1 Sol high passed every call (24/24 or 16/16). Haiku 4.5 passed 11/24. Chart: Eight hard tasks · strict pass rate · 95% intervals (n = 16–24 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
- Coding agents on hidden-test tasks: 36 of 36 sessions passed, so time separates them. Claude Code with Sonnet 5.5: median 23.1 s; Codex CLI took 4.9× as long, with extra standing instructions. Sessions that passed every hidden check: 100% (36/36) (n = 36, 95% CI 90–100%). Median time, Sonnet 5.5 in Claude Code: 23.1 s (n = 12). Median time, Codex CLI vs Claude Code with Sonnet: 4.9× (n = 12). Caveat: Codex CLI ran with the tester’s global AGENTS.md: --ignore-user-config does not switch that file off. 12 of 12 Codex sessions read the tester’s notes, 10 wrote a WORKLOG.md and 9 reported a commit attempt. Its time, tool calls and lines changed include that work. Claude Code ran with setting sources off, and no Claude session did any of it.
- Effort ladder: 11 model and effort settings passed 176 of 176 calls strictly (16/16 each, 81%–100%). No effort level wins any of 172 rows. Calls that passed strictly, low to high effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference). Effort comparison rows where one level wins: 0 (of 172 rows · 16 pairs). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Caching, a list-price calculation: 50% less on Sonnet 5.5, 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every time. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- Agent memory: Sonnet 5.5 followed team-only rules 40% (6/15) of the time without memory and 100% (15/15) with an 11-line file. Without the late-fee rate it asked 7/9 times. Team knowledge, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). Asked for the missing rate: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
- System One arena: on 1,085 checkable decisions, Jev 1.13 answered 76.8% right; the best open model, Clef 27B, 69.9%. Jev 1.13: 76.8% (805/1048) (n = 1048, 95% CI 74–79%). Clef 27B: 69.9% (733/1048) (n = 1048, 95% CI 67–73%). Checkable decisions: 1,085. Caveat: Local models ran quantized on one Mac; the vendors measured full-precision weights on data-centre GPUs. Latency on other hardware will differ, and Jev latency includes the network.
- Routing: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82 exact, and the intervals overlap. List-price cost is where they differ. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
- Routing overhead: an in-process policy decides in 1.42 µs; Sonnet 5.5 as a router takes 2.60 s through the CLI, about 1.8 million times longer. Chart: Time per routing decision · log scale · median to p95 (n = 82–20000 each). Caveat: The policy is timed in process and the LLM routers through a CLI: this compares the two ways of routing as deployed, not two models on equal footing. A direct API call would skip the CLI time (shown separately).
- Provider prices, reported by OpenRouter: open-weight models vary up to 12.6x across providers. Closed models: one price, plus a 5.5% credit fee. Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Gateway price equals the first-party price: 10 of 10. Endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
- Recorded SWE-bench tokens, repriced: $87.23 on Sonnet 5.5, $143.83 on Opus 5.5, $343.33 with no prompt cache. At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Sonnet 5.5 without the prompt cache: $343.33 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
- What a CLI adds: for a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens. Codex CLI vs OpenAI API, one-line answer: 3.5x slower (n = 30). Input tokens the Codex CLI sends for one line: 19,551 (n = 15). Scheduler repair, median: Claude Code vs Codex CLI: 15.0 s vs 61.2 s (n = 3, models differ too). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
- Three real pull requests: the latest build verified 1 of 3, at $11.06 notional. Guardrail refusals went 26 → 19. Verified deliveries, latest build: 1 of 3 (n = 3). Notional cost, latest build, 3 tasks: $11.06 (n = 3). Guardrail refusals, first vs latest slice: 26 → 19 (n = 3). Caveat: One attempt per cell: these are defect-finding runs, not rates.
- Agent-loop strict passes: Haiku 54% (13/24); Sonnet 100% (16/16). Attempts excluded for outside reads: 2 of 56. Claude Haiku 4.5 strict pass rate, agent loop: 54% (13/24) (n = 24, 95% CI 35–72%). Claude Sonnet 5.5 strict pass rate, agent loop: 100% (16/16) (n = 16, 95% CI 81–100%). Agent-loop attempts left out for reading outside the work folder: 2 of 56 (n = 56). Caveat: The Claude single-call cells ran in another batch on 2026-10-06 (03:23 to 04:02 UTC), with the same CLI version, tasks and validators; provider load can differ by hour.
- Haiku exact routing: thinking off 87% (71/82), on 89% (73/82). Paired McNemar p = 0.754; day and account differ. Claude Haiku 4.5 (thinking off): exact routing decisions: 87% (71/82) (n = 82, 95% CI 78–92%). Claude Haiku 4.5 (thinking on): exact routing decisions: 89% (73/82) (n = 82, 95% CI 80–94%). Paired exact test, thinking off vs on (exact McNemar p): 0.754 (n = 82). Caveat: The thinking-on and Sonnet arms are the recorded routing run of 2026-10-05, reused. The thinking-off arm ran on a different day and on a different Claude subscription account, so this is a confounded comparison. Day, account, CLI version and input-token differences can affect the results. The data cannot isolate the effect of thinking.
- Format misses: instructions 35% (17/48), JSON schema 0% (0/48). The schema still gave wrong values: 13% (6/48). Format-miss rate with instructions only, all models: 35% (17/48) (n = 48, 95% CI 23–50%). Format-miss rate with a JSON schema, all models: 0% (0/48) (n = 48, 95% CI 0–7%). Wrong-values rate with a JSON schema, all models: 13% (6/48) (n = 48, 95% CI 6–25%). Caveat: Small samples: Haiku 24 calls per mode, Sonnet 12 calls per mode and GPT-6.1 Sol 12 calls per mode. A 12/12 result has a 95% interval of 76% to 100%. This is not a minimum detectable difference.
- Later Claude sessions with substantial cache reuse: new folders 0 of 2 (95% interval 0% to 66%); a fixed folder 2 of 2 (95% interval 34% to 100%). This analysis is exploratory. Later sessions with at least 50% of turn-1 input cached, A: new folder each time: 0 of 2 (95% interval 0% to 66%) (n = 2, 95% CI 0–66%). Later sessions with at least 50% of turn-1 input cached, B: fixed folder: 2 of 2 (95% interval 34% to 100%) (n = 2, 95% CI 34–100%). Caveat: The surviving protocol file was created after all counted calls. Its claimed 00:32 UTC declaration is not supported by its file birth time. Amendment 1 and 2 state 00:36 and 00:37 UTC, but separate pre-edit copies do not verify those times. The current summary was regenerated at 07:41 UTC. Treat the analysis as exploratory.
- Cache break-even, a calculation: a new prefix with a one-hour write needs 2 reuses (the 3rd request). Recorded session payback: 2 turns (6 of 6 sessions). Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation): 2 reuses (the 3rd request). Turn at which a recorded Claude Code session’s total input cost with the cache first fell below its cost with no cache (calculation on recorded tokens): 2 turns (6 of 6 sessions) (n = 6). Calculation, not a run. Caveat: The retained Claude protocol file was created at 14:43:55 UTC, after the first counted session at 14:35:39 UTC on 2026-10-06. Its declaration says 14:25 UTC, but file times do not verify that claim. Treat this as a retrospective protocol record.
- Exact unseen routing decisions: Jev 82% (46/56), Haiku 79% (44/56), Sonnet 88% (49/56). The intervals overlap; no rank. Jev 1.13 (TypeSafe): exact on unseen decisions: 82% (46/56) (n = 56, 95% CI 70–90%). Claude Haiku 4.5 · Claude Code: exact on unseen decisions: 79% (44/56) (n = 56, 95% CI 66–87%). Claude Sonnet 5.5 (low) · Claude Code: exact on unseen decisions: 88% (49/56) (n = 56, 95% CI 76–94%). Caveat: The case author and two of the three routers are Claude models, so a same-family label bias is possible. A second labeller (GPT-6.1 Sol) labelled every case blind; the secondary scoring keeps only the keys where its label is inside our acceptable set.
- Reasoning cost at high vs low effort, a calculation from recorded calls: Sonnet 2.2x ($0.0043 at low, $0.0094 at high); Opus 3.6x ($0.0050 at low, $0.0180 at high). Reasoning cost per call, high ÷ low effort, Claude Sonnet 5.5 · Claude Code (calculation, means): 2.2x ($0.0043 at low, $0.0094 at high) (n = 16). Reasoning cost per call, high ÷ low effort, Claude Opus 5.5 · Claude Code (calculation, means): 3.6x ($0.0050 at low, $0.0180 at high) (n = 16). Calculation, not a run. Caveat: The hard-set protocol file was created after its first counted call. Both effort-ladder protocol files were created after their batches ended. The short-set file predates its first call, but its top-up amendment timing is unverified. Batch receipts preserve protocol text, but we cannot verify all rules were written before inference. Treat these as exploratory calculations, not preregistered tests.
- Speed anatomy, a calculation: same-text token count ratio 1.8x. Extra time to first text with the larger prompt: Haiku +0.9 s; Sonnet +1.6 s. The same 4,327-character reply: most tokens ÷ fewest tokens across models (calculation): 1.8x (n = 23). Haiku: extra time to first text at 64k vs 1k (calculation): +0.9 s (n = 6). Sonnet: extra time to first text at 64k vs 1k (calculation): +1.6 s (n = 6). Calculation, not a run. Caveat: 4 calls per model in part A and 3 per size in part B. Medians of so few calls move with one slow call, and the ranges are not confidence intervals. The 95% Wilson intervals on the lookup rates are wide.
- Cost per correct answer, a calculation: Sonnet every time $0.0132; Haiku with one retry then Sonnet $0.0566. No policy ran. Cost per correct answer, Sonnet 5.5 every time (calculation): $0.0132 (n = 24). Cost per correct answer, Haiku with one retry then Sonnet (calculation): $0.0566 (n = 48). Calculation, not a run. Caveat: Advance registration is not verified. The hard-set protocol file birth time is 2026-10-06 04:03:37 UTC; its first counted call started at 03:23:59 UTC. The ladder protocol file birth time is 14:35:11 UTC; its first Claude call started at 14:21:50 UTC. Both files claim advance declaration, but the available file times do not support that claim. Copying could explain the times; we cannot establish it.
- Strict passes on four selected harder tasks: Sol 69% (11/16); Opus 42% (5/12); Sonnet 38% (6/16). Tasks were selected with a Sonnet pilot. GPT-6.1 Sol (medium) (Codex CLI): strict pass rate on the harder tasks: 69% (11/16) (n = 16, 95% CI 44–86%). Opus 5.5 (Claude Code): strict pass rate on the harder tasks: 42% (5/12) (n = 12, 95% CI 19–68%). Sonnet 5.5 (Claude Code): strict pass rate on the harder tasks: 38% (6/16) (n = 16, 95% CI 18–61%). Caveat: Selection effect: the study picked tasks that Sonnet did not pass twice in the pilot. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls (calculation: 25% and 38%; the intervals overlap). Selection can produce this pattern, but the run does not establish its cause.
- Latency budget, a calculation: steps within an assumed 300 ms budget 3 of 14 (median: 3 of 14); within 1,500 ms 4 of 14 (median: 8 of 14), at p95 or the observed maximum. No voice turn ran. Steps that fit a 300 ms budget at the slow end (calculation): 3 of 14 (median: 3 of 14) (n = 14). Steps that fit a 1,500 ms budget at the slow end (calculation): 4 of 14 (median: 8 of 14) (n = 14). Calculation, not a run. Caveat: Timing samples omit 0 failed or untimed matched explorer calls and 0 failed or untimed Claude head-to-head calls. A failure is not a fast successful step. No pass rate is estimated here.
- Routing cost at an assumed million decisions a day, a calculation: Jev $33.70; Sonnet $7,324. No load test ran. Daily cost at 1 million decisions, Jev (calculation): $33.70 (n = 82). Daily cost at 1 million decisions, Sonnet 5.5 router (calculation): $7,324 (n = 82). Calculation, not a run. Caveat: The routing and live Jev protocols predate the first counted calls by file birth time. The overhead protocol predates its microbenchmark output. Claude ran 82 calls per router, below its 120-call cap. Jev ran 246 counted calls, below its 300-call cap. One Haiku pilot and one Jev probe are excluded. All counted calls completed without call errors. The policy microbenchmark used 5,000 warm-up iterations and 64 synthetic contexts; the minimum and individual timing samples are unavailable, so we cannot rebuild its quantiles or full range. No separate validator-control receipt or Sonnet pilot is retained.
- 28 studies, one place. Intervals, sources and every failure kept.
SWE-bench Verified: Agent vs 11 public model runs
Agent resolved 25 of 33 SWE-bench Verified instances; every 95% interval overlaps the public panel, so the data does not rank it.
Transcript
- SWE-bench Verified · same 33 instances. Agent vs 11 public model runs. One attempt each. No retries. 95% intervals on individual run rates. The panel mean has no interval.
- Agent resolved 25 of 33. The public panel averaged 74.1% on the same instances. Agent resolved (full pipeline, Sonnet 5.5): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Agent model cost per attempt (calculation, notional): $2.81 (n = 33). Caveat: Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
- Panel runs resolved 21 to 28 of 33. Every interval overlaps, so this sample cannot rank Agent. Chart: Resolved rate on the same 33 SWE-bench Verified instances (n = 33 each). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
- Hardest band: on the 4 instances no panel model solved, Agent solved 1. Too few to call a rate. Chart: Resolved rate by difficulty band (n = 4–14 each). Caveat: Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
- Open benchmarks: intervals, sources and every failure kept.
Key numbers
72% (18/25)
Agent resolved, campaign 1 sample of 25
95% CI 52%–86% · n = 25
76% (19/25)
Agent resolved, original seed draw of 25 (no replacements)
95% CI 57%–89% · n = 25
74.1%
Public panel mean on the same 33 instances
n = 33
$2.81
Agent model cost per attempt (notional)
n = 33
$3.71
Agent model cost per resolved instance (notional)
n = 25
9.6min
Median worker time per attempt
n = 33
49
Model calls per attempt
n = 33
One real attempt as a receipt: outcome, minutes, model calls and notional cost
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
Every interval overlaps every other: this chart does not order these rows.
| Item | Resolved rate | 95% interval | n |
|---|---|---|---|
| GPT 5.2 (high) | 85% | 69%–93% | 33 |
| Gemini 3 Flash (high) | 82% | 66%–91% | 33 |
| GLM 5 (high) | 79% | 62%–89% | 33 |
| Agent (Sonnet 5.5, full pipeline) | 76% | 59%–87% | 33 |
| Claude 4.5 Sonnet (high) | 76% | 59%–87% | 33 |
| Claude 4.5 Haiku (high) | 76% | 59%–87% | 33 |
| Claude 4.5 Opus (high) | 73% | 56%–85% | 33 |
| DeepSeek V3.2 (high) | 73% | 56%–85% | 33 |
| MiniMax M2.5 (high) | 70% | 53%–83% | 33 |
| Claude 4.6 Opus | 70% | 53%–83% | 33 |
| Kimi K2.5 (high) | 70% | 53%–83% | 33 |
| GPT 5 mini | 64% | 47%–78% | 33 |
12 rows. Highest GPT 5.2 (high) 85% (95% interval 69%–93%, n 33). Lowest GPT 5 mini 64% (95% interval 47%–78%, n 33). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 33 per row
Agent vs 11 public mini-SWE-agent v2 runs, one attempt each
Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.
Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs
- Agent
- Public panel mean (square)
Gap labels, Public panel mean vs Agent: Public panel mean is x percentage points higher (+) or lower (−) than Agent, calculated from the two values shown; lines are the 95% Wilson interval.
| Difficulty band | Agent | Public panel mean | 95% interval | n |
|---|---|---|---|---|
| No panel model solved it | 25% | 0% | Agent: 4.6%–70% | 4 |
| Under half solved it | 75% | 27% | Agent: 30%–95% | 4 |
| Half or more solved it | 82% | 85% | Agent: 52%–95% | 11 |
| Every panel model solved it | 86% | 100% | Agent: 60%–96% | 14 |
4 difficulty bands, 2 series: Agent, Public panel mean. Agent: highest Every panel model solved it 86% (95% interval 60%–96%, n 14). Lowest No panel model solved it 25% (95% interval 4.6%–70%, n 4). All intervals overlap. Public panel mean: highest Every panel model solved it 100% (n 14). Lowest No panel model solved it 0% (n 4). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 4–14 per row
Band = how many of the 11 public panel models solved the instance
Agent whiskers are 95% Wilson intervals. The "no panel model solved it" band has 4 instances; Agent resolved 1 (matplotlib__matplotlib-21568).
Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, SWE-bench campaign rules and sample design
- Public panel (mini-SWE-agent v2)
- Agent
| Point | Series | Mean model cost per instance (USD) | Resolved rate | n |
|---|---|---|---|---|
| Claude 4.5 Opus (high) | Public panel (mini-SWE-agent v2) | $0.86 | 73% | 33 |
| Gemini 3 Flash (high) | Public panel (mini-SWE-agent v2) | $0.36 | 82% | 33 |
| MiniMax M2.5 (high) | Public panel (mini-SWE-agent v2) | $0.075 | 70% | 33 |
| Claude 4.6 Opus | Public panel (mini-SWE-agent v2) | $0.61 | 70% | 33 |
| GLM 5 (high) | Public panel (mini-SWE-agent v2) | $0.53 | 79% | 33 |
| GPT 5.2 (high) | Public panel (mini-SWE-agent v2) | $0.53 | 85% | 33 |
| Claude 4.5 Sonnet (high) | Public panel (mini-SWE-agent v2) | $0.69 | 76% | 33 |
| Kimi K2.5 (high) | Public panel (mini-SWE-agent v2) | $0.18 | 70% | 33 |
| DeepSeek V3.2 (high) | Public panel (mini-SWE-agent v2) | $0.46 | 73% | 33 |
| Claude 4.5 Haiku (high) | Public panel (mini-SWE-agent v2) | $0.36 | 76% | 33 |
| GPT 5 mini | Public panel (mini-SWE-agent v2) | $0.051 | 64% | 33 |
| Agent | Agent | $2.81 | 76% | 33 |
12 points: Resolved rate against Mean model cost per instance (USD). Mean model cost per instance (USD) runs from $0.051 to $2.81; Resolved rate from 64% to 85%. Highlighted: Agent.
Notesn = 33 per point
Same 33 instances. Panel = API list price; Agent = notional subscription estimate
Agent's cost includes repository onboarding, planning, verification and review; it is a list-price estimate for subscription calls, not an invoice. Panel costs are published API costs.
Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, Anthropic list prices (Claude models)
Model calls per instance
Mean over the same 33 instances
Hover or focus a bar for its ratio to Agent (the highlighted row): a ratio of the two values shown, not a measurement.
| Item | Mean calls | n |
|---|---|---|
| DeepSeek V3.2 (high) | 88.2 | 33 |
| GLM 5 (high) | 77.5 | 33 |
| Claude 4.5 Haiku (high) | 68.5 | 33 |
| MiniMax M2.5 (high) | 58.4 | 33 |
| Kimi K2.5 (high) | 56.7 | 33 |
| Gemini 3 Flash (high) | 54.2 | 33 |
| Claude 4.5 Sonnet (high) | 51 | 33 |
| Agent | 49.5 | 33 |
| Claude 4.5 Opus (high) | 35.9 | 33 |
| GPT 5.2 (high) | 35.6 | 33 |
| Claude 4.6 Opus | 28.9 | 33 |
| GPT 5 mini | 20.8 | 33 |
12 rows. Highest DeepSeek V3.2 (high) 88.2 (n 33). Lowest GPT 5 mini 20.8 (n 33).
Notesn = 33 per row
A panel call is one bash-agent step. An Agent call is one model request of any stage (research, plan, act, verify, review).
Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs
Parts sorted by value, largest first
- Act (edit and run)
- Research
- Verify
- Other
- Review
- Context compaction
- Onboarding notes
Shares are calculated from the values shown; rounding can make the sum of the parts differ from the stated total by a cent.
| Item | Cost |
|---|---|
| Act (edit and run) | $44.05 |
| Research | $20.56 |
| Verify | $8.72 |
| Other | $6.24 |
| Review | $5.58 |
| Context compaction | $5.41 |
| Onboarding notes | $2.09 |
7 rows. Highest Act (edit and run) $44.05. Lowest Onboarding notes $2.09.
Notes
Share of notional model cost by stage, all 33 attempts
Total $92.64 over 33 attempts.
Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)
- Agent
- Public panel mean, same instances
| Item | Agent | Public panel mean, same instances | 95% interval | n |
|---|---|---|---|---|
| Campaign 1: 25-instance sample | 72% | 71% | Agent: 52%–86% | 25 |
| Campaign 2: 8 compiled-extension instances | 88% | 84% | Agent: 53%–98% | 8 |
| Original seed draw of 25 | 76% | 71% | Agent: 57%–89% | 25 |
| All 33 attempted | 76% | 74% | Agent: 59%–87% | 33 |
4 rows, 2 series: Agent, Public panel mean, same instances. Agent: highest Campaign 2: 8 compiled-extension instances 88% (95% interval 53%–98%, n 8). Lowest Campaign 1: 25-instance sample 72% (95% interval 52%–86%, n 25). All intervals overlap. Public panel mean, same instances: highest Campaign 2: 8 compiled-extension instances 84% (n 8). Lowest Original seed draw of 25 71% (n 25). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 8–33 per row
Agent resolved rate and 95% Wilson interval per declared view
The original draw and "all 33" mix two platform builds.
Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, SWE-bench campaign rules and sample design
Tables
Every attempt
Sorted by how many panel systems solved each instance, most first
Each tile is one row; from top: Panel solved, then Agent outcome.
Shade is the share k of n in each cell.
| Instance | Panel solved | Agent outcome | Minutes | Calls | Cost (notional) | Campaign |
|---|---|---|---|---|---|---|
| astropy__astropy-13579 | 11/11 | resolved | 54.1 min | 65 | $4.00 | 2 |
| astropy__astropy-14096 | 10/11 | resolved | 10.4 min | 64 | $3.35 | 2 |
| astropy__astropy-7336 | 11/11 | unresolved | 23 min | 48 | $2.03 | 2 |
| django__django-10554 | 0/11 | unresolved | 4.7 min | 34 | $1.52 | 1 |
| django__django-11333 | 11/11 | resolved | 8 min | 46 | $2.51 | 1 |
| django__django-11885 | 3/11 | resolved | 7.3 min | 38 | $2.27 | 1 |
| django__django-12193 | 10/11 | resolved | 5.7 min | 43 | $2.06 | 1 |
| django__django-12741 | 11/11 | resolved | 10.5 min | 54 | $2.98 | 1 |
| django__django-13158 | 10/11 | resolved | 7 min | 43 | $2.15 | 1 |
| django__django-13925 | 8/11 | resolved | 6.8 min | 43 | $2.29 | 1 |
| django__django-14034 | 0/11 | unresolved | 11.5 min | 54 | $3.03 | 1 |
| django__django-14373 | 11/11 | resolved | 5.7 min | 38 | $1.74 | 1 |
| django__django-14559 | 11/11 | resolved | 7.2 min | 47 | $2.71 | 1 |
| django__django-14631 | 9/11 | empty patch | 38.2 min | 47 | $2.80 | 1 |
| django__django-14672 | 11/11 | resolved | 9.1 min | 58 | $2.99 | 1 |
| django__django-15022 | 4/11 | unresolved | 15 min | 62 | $4.10 | 1 |
| django__django-15280 | 7/11 | resolved | 11 min | 59 | $3.67 | 1 |
| django__django-15572 | 10/11 | resolved | 5.3 min | 42 | $1.81 | 1 |
| django__django-15732 | 3/11 | resolved | 10.9 min | 59 | $2.90 | 1 |
| django__django-16255 | 10/11 | resolved | 7.6 min | 47 | $2.55 | 1 |
| django__django-16333 | 11/11 | resolved | 13 min | 37 | $1.61 | 1 |
| django__django-17029 | 11/11 | resolved | 7.8 min | 42 | $2.28 | 1 |
| matplotlib__matplotlib-21568 | 0/11 | resolved | 9.4 min | 53 | $3.02 | 2 |
| matplotlib__matplotlib-24149 | 10/11 | resolved | 41.1 min | 64 | $3.64 | 2 |
| matplotlib__matplotlib-24970 | 11/11 | resolved | 33 min | 66 | $3.90 | 2 |
| pydata__xarray-6721 | 11/11 | resolved | 16.2 min | 57 | $4.09 | 1 |
| scikit-learn__scikit-learn-14894 | 11/11 | resolved | 35.2 min | 50 | $3.23 | 2 |
| scikit-learn__scikit-learn-25973 | 10/11 | resolved | 8.6 min | 44 | $2.43 | 2 |
| sphinx-doc__sphinx-10435 | 2/11 | resolved | 16.8 min | 67 | $4.83 | 1 |
| sphinx-doc__sphinx-9698 | 11/11 | resolved | 9.6 min | 49 | $2.69 | 1 |
| sympy__sympy-18189 | 11/11 | empty patch | 1.6 min | 13 | $0.89 | 1 |
| sympy__sympy-20428 | 0/11 | unresolved | 21 min | 71 | $4.81 | 1 |
| sympy__sympy-22456 | 9/11 | empty patch | 6.4 min | 28 | $1.75 | 1 |
Counts from the table, not an intervaln = 11 per cell
33 rows by 2 columns. 14 of 33 k/n cells are full (n of n) and 4 are zero. Agent outcome: 25 of 33 resolved.
An empty patch counts as unresolved here; the Table view gives each outcome.
Public panel on the same instances
| System | Resolved of 33 | Rate | Full Verified board (500) | Mean cost per instance, these 33 |
|---|---|---|---|---|
| GPT 5.2 (high) | 28 | 85% | 73% | $0.53 |
| Gemini 3 Flash (high) | 27 | 82% | 76% | $0.36 |
| GLM 5 (high) | 26 | 79% | 73% | $0.53 |
| Agent (Sonnet 5.5, full pipeline) | 25 | 76% | — | $2.81 |
| Claude 4.5 Sonnet (high) | 25 | 76% | 71% | $0.69 |
| Claude 4.5 Haiku (high) | 25 | 76% | 67% | $0.36 |
| Claude 4.5 Opus (high) | 24 | 73% | 77% | $0.86 |
| DeepSeek V3.2 (high) | 24 | 73% | 70% | $0.46 |
| MiniMax M2.5 (high) | 23 | 70% | 76% | $0.075 |
| Claude 4.6 Opus | 23 | 70% | 76% | $0.61 |
| Kimi K2.5 (high) | 23 | 70% | 71% | $0.18 |
| GPT 5 mini | 21 | 64% | 56% | $0.051 |
Method
- Sample: 25 of the 500 Verified instances, stratified by public difficulty (seed 20261004). Difficulty is how many of the 11 public mini-SWE-agent v2 runs solved the instance.
- Campaign 1 ran 25 instances on platform build f0ac3a8a. Six compiled-extension instances could not import in the worker checkout, so the declared rule replaced them in the same band.
- Campaign 2 ran those compiled instances (plus two replacement candidates) on build 236c0d3f after a sandbox fix.
- One attempt per instance. No retries, no operator answers: an escalation is graded on what was delivered.
- Grading uses the official SWE-bench harness and instance images (under amd64 emulation). The gold patch resolved on the host for every instance.
- Agent runs its full pipeline (onboarding, research, plan, act, verify, review) on claude-sonnet-5-5 through a subscription CLI. Costs are list-price estimates of the recorded tokens.
Caveats
- n = 33: intervals are wide. This is a defect-finding run, not a ranking.
- Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
- Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
- Campaign 1 ended up 19 django, 3 sympy, 2 sphinx and 1 xarray after replacements. Campaign 2 covers the compiled repositories.
- Verified issues are public (2015 to 2023) and likely in every model’s training data; contamination is uncontrolled for all systems.
- Difficulty bands come from the panel’s own results, so a system outside the panel tends to look better than the panel on hard bands and worse on easy ones (regression to the mean). Read the band chart with that selection effect in mind.
- Three campaign-1 empty patches were platform holds before delivery (missing lint tools, a too-literal plan gate, an unanswered question), not wrong fixes. They count as failures here.
Sources
Agent on SWE-bench Verified, campaign 1 (25 instances)
Stratified sample of 25 Verified instances (seed 20261004), one attempt each, official grading harness. Fixed model claude-sonnet-5-5, platform build f0ac3a8a.
Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)
The 6 compiled-extension instances that campaign 1 could not run, plus 2 replacement candidates. One attempt each, platform build 236c0d3f.
SWE-bench Verified leaderboard, mini-SWE-agent v2 runs
Public per-instance results of 11 models under mini-SWE-agent 2.0.0 (bash only, one attempt). Costs are API list prices as published.
SWE-bench campaign rules and sample design
Rules declared before the first run: escalations are graded as delivered, gold must resolve on the host, blocked instances are replaced in the same difficulty band, no second attempts. The excluded and replaced instances are listed in the exclusions extract.
Anthropic list prices (Claude models)
Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “Agent on SWE-bench Verified vs 11 public models”, updated October 5, 2026, https://agent.sasid.ai/benchmarks/swe-bench-verified.
Explainers that cite this study
Read the methods and terms in the context of these recorded results.
- Benchmark saturation: when every model scores 100%
- Cost per correct answer: the LLM price that counts failures
- How many runs does an LLM eval need? Sample size, with real intervals
- LLM as judge: how blind review works, and where it fails
- SWE-bench Verified, explained
- What is an agent harness? The code around the model, measured
More comparisons based on this study (58)
These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.
- Agent vs Claude Haiku 4.5
- Agent vs Claude Opus 4.5
- Agent vs Claude Opus 4.6
- Agent vs Claude Sonnet 4.5
- Agent vs DeepSeek V3.2
- Agent vs Gemini 3 Flash
- Agent vs GLM 5
- Agent vs GPT 5.2
- Agent vs GPT 5 mini
- Agent vs Kimi K2.5
- Agent vs MiniMax M2.5
- Claude Haiku 4.5 vs Kimi K2.5
- Claude Haiku 4.5 vs MiniMax M2.5
- Claude Opus 4.5 vs Claude Opus 4.6
- Claude Opus 4.5 vs DeepSeek V3.2
- Claude Opus 4.5 vs GPT 5 mini
- Claude Opus 4.5 vs Kimi K2.5
- Claude Opus 4.5 vs MiniMax M2.5
- Claude Opus 4.6 vs DeepSeek V3.2
- Claude Opus 4.6 vs GPT 5 mini
- Claude Opus 4.6 vs Kimi K2.5
- Claude Opus 4.6 vs MiniMax M2.5
- Claude Sonnet 4.5 vs Claude Opus 4.5
- Claude Sonnet 4.5 vs Claude Opus 4.6
- Claude Sonnet 4.5 vs DeepSeek V3.2
- Claude Sonnet 4.5 vs GPT 5 mini
- Claude Sonnet 4.5 vs Kimi K2.5
- Claude Sonnet 4.5 vs MiniMax M2.5
- DeepSeek V3.2 vs GPT 5 mini
- DeepSeek V3.2 vs Kimi K2.5
- DeepSeek V3.2 vs MiniMax M2.5
- Gemini 3 Flash vs Claude Opus 4.5
- Gemini 3 Flash vs Claude Opus 4.6
- Gemini 3 Flash vs Claude Sonnet 4.5
- Gemini 3 Flash vs DeepSeek V3.2
- Gemini 3 Flash vs GLM 5
- Gemini 3 Flash vs GPT 5 mini
- Gemini 3 Flash vs Kimi K2.5
- Gemini 3 Flash vs MiniMax M2.5
- GLM 5 vs Claude Opus 4.5
- GLM 5 vs Claude Opus 4.6
- GLM 5 vs Claude Sonnet 4.5
- GLM 5 vs DeepSeek V3.2
- GLM 5 vs GPT 5 mini
- GLM 5 vs Kimi K2.5
- GLM 5 vs MiniMax M2.5
- GPT 5.2 vs Claude Opus 4.5
- GPT 5.2 vs Claude Opus 4.6
- GPT 5.2 vs Claude Sonnet 4.5
- GPT 5.2 vs DeepSeek V3.2
- GPT 5.2 vs Gemini 3 Flash
- GPT 5.2 vs GLM 5
- GPT 5.2 vs GPT 5 mini
- GPT 5.2 vs Kimi K2.5
- GPT 5.2 vs MiniMax M2.5
- Kimi K2.5 vs GPT 5 mini
- MiniMax M2.5 vs GPT 5 mini
- MiniMax M2.5 vs Kimi K2.5
More write-ups that cite this study (1)
Models and comparisons in this study
Write-ups on this study
AI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
AI coding cost per developer: a formula built on recorded work
AI coding cost per developer: a monthly budget formula using recorded tokens and list-price calculations. Replace our task volume and pass rate with yours.
Claude Code cost per task: a price ladder from one decision to one agent run
$0.087 per Claude Code coding session at API list prices, $0.0036 per short call, $2.81 per full agent attempt. Calculations on recorded tokens, not bills.
Devin's $0.60 per task and our $3.71 per resolved task are different numbers
Devin reports $0.60 per task. Our list-price calculation gives $3.71 per resolved SWE-bench task. Compare the units with a table and buyer checklist.
How long does an AI coding agent take per task? Minutes, calls and where the time goes
Agent took a median 9.6 minutes per SWE-bench attempt (n = 33, range 1.6 to 54.1). Compare single calls, repairs and full tasks with limits.
How many runs do you need to compare two AI models? A sample-size table
Calculation: 100 runs can separate an observed 8-point gap at a perfect score. See Wilson intervals, paired tests and limits for LLM eval sample sizes.
Is Claude Haiku cheaper than Sonnet? Cost per correct answer, with retries
Haiku 4.5 lists at half the price of Sonnet 5.5, yet calculated cost per correct hard answer was 4.7x higher. Retry maths, two exceptions and every source.
Most AI model comparisons are ties: 44 of 1,196 rows show a clear gap
44 of 1,196 AI model comparison rows show a gap under our overlap rules. Most gaps are timing rows. No effort row separates quality.
What is the best AI model for coding? Our data says four models tie
4 models tied at the top of our hard coding set: Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol. Only Haiku 4.5 separated. A tier list built from intervals.
Why AI coding agents fail on real pull requests: every failure from 12 attempts
32 AI coding agent misses from our own runs, sorted into 8 classes: wrong answers, gates, caps, lost context and more. Counts, not rates.
Your AI agent says it is done. Is it? Claimed vs verified in our runs
6 of 6 Haiku 4.5 sessions invented a late-fee rate and reported done. Sonnet 5.5 flagged the gap in 9 of 9. Claimed vs verified, with small n.
AI coding benchmarks roundup, October 2026: sixteen studies, every number in one place
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
An AI model leaderboard without a composite score: why, and how to read ours
Our leaderboard lists 46 models, CLIs, routers and providers with their best-supported facts, each with n and an interval. No single score. Here is why.
How much does prompt caching actually save? Measured in Claude Code and Codex
Claude Code read 97% of later-turn input from the cache. At list price that halved a 5-turn session and cut an agent bill about 3.9x. Turn 1 costs more.
How to estimate your AI coding bill from real token mixes
Estimate AI coding costs from recorded token mixes: an agent task at $2.64, a hard call at $0.014, a routing decision at $0.005. Calculations, limits stated.
Opus 5.5 vs Sonnet 5.5 on real GitHub issues: an early look, limits first
Interim: 3 of 8 paired SWE-bench Verified issues graded. Opus 5.5 resolved 2, Sonnet 5.5 1; McNemar p = 1.0. Opus cost 2.6x, a calculation. Limits first.
Harness vs model: where do AI coding agent gains really come from?
Is it the model or the harness? Our data shows the harness clearly moves speed, tokens and cost. Whether it moves accuracy, our samples cannot yet say.
SWE-bench Verified: an agent pipeline vs 11 public models on the same 33 tasks
Agent resolved 25 of 33 SWE-bench Verified instances (76%). Eleven public models solved 21 to 28 of the same ones. Why that is a tie, not a win.
What one resolved SWE-bench task really costs an AI coding agent
A full agent pipeline spent a notional $3.71 per resolved SWE-bench instance. Where the money went, what caching saved, and the public panel range.
Why we count every failed attempt: our rules for honest AI benchmarks
How we benchmark AI agents: declare the protocol first, keep every failure, show intervals, publish our own confounds and label interim results.
More studies
All benchmarksOpus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)
Interim: 3 of 8 paired SWE-bench Verified instances graded. Opus 5.5 resolved 2, Sonnet 5.5 1 (McNemar p = 1.0). Cost 2.6×, a list-price calculation.
Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
96 calls. Strict passes, schema vs instructions: Haiku 18/24 vs 0/24, Sonnet 12/12 vs 12/12, GPT-6.1 Sol 12/12 vs 12/12. With intervals.
Does a new Claude Code session reuse the prompt cache of an earlier one?
30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.