• SWE-bench
  • Coding Agents
  • Leaderboard
  • Cost
  • Claude Sonnet

Agent on SWE-bench Verified vs 11 public models

How does Agent, a full worker pipeline on one model, do on SWE-bench Verified next to public single-model runs on the very same instances?

Published · 6 charts · Download the data or a carousel

76%

95% CI 59%–87% · n = 33

25/33 · Agent resolved, all 33 attempted instances

Both campaigns, one attempt each, failures and empty patches included.

The answer

Agent resolved 25 of 33 attempted instances (75.8%, 95% interval 59% to 87%). On the same instances the 11 public mini-SWE-agent v2 runs resolved between 21 and 28 (panel mean 74.1%). Every interval overlaps, so this sample cannot rank Agent above or below any panel model. Agent spent a notional $2.81 and 49 model calls per attempt, with a median of 9.6 minutes; it is slower and more expensive per instance than a bare bash agent because it onboards, plans, verifies and reviews. It resolved 1 of 4 instances that no panel model solved.

Live story

Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.

Live story · 211 sAgent Benchmarks, October 2026: 28 studies in one film

Agent Benchmarks, October 2026: 28 studies in one film

The headline of each of our 28 open benchmark studies, with sample sizes and intervals. Calculations labelled; failures counted.

Transcript
  1. October 2026 roundup · 28 studies. Agent Benchmarks: every headline. Recorded runs, 95% intervals where they exist, every failure counted. Calculations labelled.
  2. SWE-bench Verified: Agent resolved 25 of 33. The public panel averaged 74.1%; the intervals overlap, so no rank. Agent resolved (one attempt each): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Model cost per resolved instance (calculation, notional): $3.71 (n = 25). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
  3. SWE-bench, interim (3 of 8 pairs graded): Opus 5.5 as Agent’s brain resolved 2 of 3, Sonnet 5.5 1 of 3. Exact McNemar p = 1.0: no difference yet. Opus cost 2.6× as much, a calculation. Opus 5.5 in Agent: resolved (interim): 67% (2/3) (n = 3, 95% CI 21–94%). Sonnet 5.5 in Agent: same instances: 33% (1/3) (n = 3, 95% CI 6–79%). List-price cost, Opus vs Sonnet (calculation): 2.6× (n = 3). Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
  4. Blind critics preferred the AI pull request to the merged human one on 9 of 12 tasks at the latest attempt, 6 of 12 at the first. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
  5. Five short tasks, nine setups: 127 of 130 calls passed, so speed and tokens separate them. Fastest median: Fable 5.1 at 1.9 s. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Fastest median total time (Fable 5.1): 1.9 s (n = 15). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.
  6. Eight hard tasks: Sonnet 5.5, Opus 5.5, Opus 5.5 high, GPT-6.1 Sol medium, Fable 5.1 and GPT-6.1 Sol high passed every call (24/24 or 16/16). Haiku 4.5 passed 11/24. Chart: Eight hard tasks · strict pass rate · 95% intervals (n = 16–24 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
  7. Coding agents on hidden-test tasks: 36 of 36 sessions passed, so time separates them. Claude Code with Sonnet 5.5: median 23.1 s; Codex CLI took 4.9× as long, with extra standing instructions. Sessions that passed every hidden check: 100% (36/36) (n = 36, 95% CI 90–100%). Median time, Sonnet 5.5 in Claude Code: 23.1 s (n = 12). Median time, Codex CLI vs Claude Code with Sonnet: 4.9× (n = 12). Caveat: Codex CLI ran with the tester’s global AGENTS.md: --ignore-user-config does not switch that file off. 12 of 12 Codex sessions read the tester’s notes, 10 wrote a WORKLOG.md and 9 reported a commit attempt. Its time, tool calls and lines changed include that work. Claude Code ran with setting sources off, and no Claude session did any of it.
  8. Effort ladder: 11 model and effort settings passed 176 of 176 calls strictly (16/16 each, 81%–100%). No effort level wins any of 172 rows. Calls that passed strictly, low to high effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference). Effort comparison rows where one level wins: 0 (of 172 rows · 16 pairs). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  9. Caching, a list-price calculation: 50% less on Sonnet 5.5, 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every time. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  10. Agent memory: Sonnet 5.5 followed team-only rules 40% (6/15) of the time without memory and 100% (15/15) with an 11-line file. Without the late-fee rate it asked 7/9 times. Team knowledge, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). Asked for the missing rate: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
  11. System One arena: on 1,085 checkable decisions, Jev 1.13 answered 76.8% right; the best open model, Clef 27B, 69.9%. Jev 1.13: 76.8% (805/1048) (n = 1048, 95% CI 74–79%). Clef 27B: 69.9% (733/1048) (n = 1048, 95% CI 67–73%). Checkable decisions: 1,085. Caveat: Local models ran quantized on one Mac; the vendors measured full-precision weights on data-centre GPUs. Latency on other hardware will differ, and Jev latency includes the network.
  12. Routing: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82 exact, and the intervals overlap. List-price cost is where they differ. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
  13. Routing overhead: an in-process policy decides in 1.42 µs; Sonnet 5.5 as a router takes 2.60 s through the CLI, about 1.8 million times longer. Chart: Time per routing decision · log scale · median to p95 (n = 82–20000 each). Caveat: The policy is timed in process and the LLM routers through a CLI: this compares the two ways of routing as deployed, not two models on equal footing. A direct API call would skip the CLI time (shown separately).
  14. Provider prices, reported by OpenRouter: open-weight models vary up to 12.6x across providers. Closed models: one price, plus a 5.5% credit fee. Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Gateway price equals the first-party price: 10 of 10. Endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
  15. Recorded SWE-bench tokens, repriced: $87.23 on Sonnet 5.5, $143.83 on Opus 5.5, $343.33 with no prompt cache. At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Sonnet 5.5 without the prompt cache: $343.33 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  16. What a CLI adds: for a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens. Codex CLI vs OpenAI API, one-line answer: 3.5x slower (n = 30). Input tokens the Codex CLI sends for one line: 19,551 (n = 15). Scheduler repair, median: Claude Code vs Codex CLI: 15.0 s vs 61.2 s (n = 3, models differ too). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
  17. Three real pull requests: the latest build verified 1 of 3, at $11.06 notional. Guardrail refusals went 26 → 19. Verified deliveries, latest build: 1 of 3 (n = 3). Notional cost, latest build, 3 tasks: $11.06 (n = 3). Guardrail refusals, first vs latest slice: 26 → 19 (n = 3). Caveat: One attempt per cell: these are defect-finding runs, not rates.
  18. Agent-loop strict passes: Haiku 54% (13/24); Sonnet 100% (16/16). Attempts excluded for outside reads: 2 of 56. Claude Haiku 4.5 strict pass rate, agent loop: 54% (13/24) (n = 24, 95% CI 35–72%). Claude Sonnet 5.5 strict pass rate, agent loop: 100% (16/16) (n = 16, 95% CI 81–100%). Agent-loop attempts left out for reading outside the work folder: 2 of 56 (n = 56). Caveat: The Claude single-call cells ran in another batch on 2026-10-06 (03:23 to 04:02 UTC), with the same CLI version, tasks and validators; provider load can differ by hour.
  19. Haiku exact routing: thinking off 87% (71/82), on 89% (73/82). Paired McNemar p = 0.754; day and account differ. Claude Haiku 4.5 (thinking off): exact routing decisions: 87% (71/82) (n = 82, 95% CI 78–92%). Claude Haiku 4.5 (thinking on): exact routing decisions: 89% (73/82) (n = 82, 95% CI 80–94%). Paired exact test, thinking off vs on (exact McNemar p): 0.754 (n = 82). Caveat: The thinking-on and Sonnet arms are the recorded routing run of 2026-10-05, reused. The thinking-off arm ran on a different day and on a different Claude subscription account, so this is a confounded comparison. Day, account, CLI version and input-token differences can affect the results. The data cannot isolate the effect of thinking.
  20. Format misses: instructions 35% (17/48), JSON schema 0% (0/48). The schema still gave wrong values: 13% (6/48). Format-miss rate with instructions only, all models: 35% (17/48) (n = 48, 95% CI 23–50%). Format-miss rate with a JSON schema, all models: 0% (0/48) (n = 48, 95% CI 0–7%). Wrong-values rate with a JSON schema, all models: 13% (6/48) (n = 48, 95% CI 6–25%). Caveat: Small samples: Haiku 24 calls per mode, Sonnet 12 calls per mode and GPT-6.1 Sol 12 calls per mode. A 12/12 result has a 95% interval of 76% to 100%. This is not a minimum detectable difference.
  21. Later Claude sessions with substantial cache reuse: new folders 0 of 2 (95% interval 0% to 66%); a fixed folder 2 of 2 (95% interval 34% to 100%). This analysis is exploratory. Later sessions with at least 50% of turn-1 input cached, A: new folder each time: 0 of 2 (95% interval 0% to 66%) (n = 2, 95% CI 0–66%). Later sessions with at least 50% of turn-1 input cached, B: fixed folder: 2 of 2 (95% interval 34% to 100%) (n = 2, 95% CI 34–100%). Caveat: The surviving protocol file was created after all counted calls. Its claimed 00:32 UTC declaration is not supported by its file birth time. Amendment 1 and 2 state 00:36 and 00:37 UTC, but separate pre-edit copies do not verify those times. The current summary was regenerated at 07:41 UTC. Treat the analysis as exploratory.
  22. Cache break-even, a calculation: a new prefix with a one-hour write needs 2 reuses (the 3rd request). Recorded session payback: 2 turns (6 of 6 sessions). Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation): 2 reuses (the 3rd request). Turn at which a recorded Claude Code session’s total input cost with the cache first fell below its cost with no cache (calculation on recorded tokens): 2 turns (6 of 6 sessions) (n = 6). Calculation, not a run. Caveat: The retained Claude protocol file was created at 14:43:55 UTC, after the first counted session at 14:35:39 UTC on 2026-10-06. Its declaration says 14:25 UTC, but file times do not verify that claim. Treat this as a retrospective protocol record.
  23. Exact unseen routing decisions: Jev 82% (46/56), Haiku 79% (44/56), Sonnet 88% (49/56). The intervals overlap; no rank. Jev 1.13 (TypeSafe): exact on unseen decisions: 82% (46/56) (n = 56, 95% CI 70–90%). Claude Haiku 4.5 · Claude Code: exact on unseen decisions: 79% (44/56) (n = 56, 95% CI 66–87%). Claude Sonnet 5.5 (low) · Claude Code: exact on unseen decisions: 88% (49/56) (n = 56, 95% CI 76–94%). Caveat: The case author and two of the three routers are Claude models, so a same-family label bias is possible. A second labeller (GPT-6.1 Sol) labelled every case blind; the secondary scoring keeps only the keys where its label is inside our acceptable set.
  24. Reasoning cost at high vs low effort, a calculation from recorded calls: Sonnet 2.2x ($0.0043 at low, $0.0094 at high); Opus 3.6x ($0.0050 at low, $0.0180 at high). Reasoning cost per call, high ÷ low effort, Claude Sonnet 5.5 · Claude Code (calculation, means): 2.2x ($0.0043 at low, $0.0094 at high) (n = 16). Reasoning cost per call, high ÷ low effort, Claude Opus 5.5 · Claude Code (calculation, means): 3.6x ($0.0050 at low, $0.0180 at high) (n = 16). Calculation, not a run. Caveat: The hard-set protocol file was created after its first counted call. Both effort-ladder protocol files were created after their batches ended. The short-set file predates its first call, but its top-up amendment timing is unverified. Batch receipts preserve protocol text, but we cannot verify all rules were written before inference. Treat these as exploratory calculations, not preregistered tests.
  25. Speed anatomy, a calculation: same-text token count ratio 1.8x. Extra time to first text with the larger prompt: Haiku +0.9 s; Sonnet +1.6 s. The same 4,327-character reply: most tokens ÷ fewest tokens across models (calculation): 1.8x (n = 23). Haiku: extra time to first text at 64k vs 1k (calculation): +0.9 s (n = 6). Sonnet: extra time to first text at 64k vs 1k (calculation): +1.6 s (n = 6). Calculation, not a run. Caveat: 4 calls per model in part A and 3 per size in part B. Medians of so few calls move with one slow call, and the ranges are not confidence intervals. The 95% Wilson intervals on the lookup rates are wide.
  26. Cost per correct answer, a calculation: Sonnet every time $0.0132; Haiku with one retry then Sonnet $0.0566. No policy ran. Cost per correct answer, Sonnet 5.5 every time (calculation): $0.0132 (n = 24). Cost per correct answer, Haiku with one retry then Sonnet (calculation): $0.0566 (n = 48). Calculation, not a run. Caveat: Advance registration is not verified. The hard-set protocol file birth time is 2026-10-06 04:03:37 UTC; its first counted call started at 03:23:59 UTC. The ladder protocol file birth time is 14:35:11 UTC; its first Claude call started at 14:21:50 UTC. Both files claim advance declaration, but the available file times do not support that claim. Copying could explain the times; we cannot establish it.
  27. Strict passes on four selected harder tasks: Sol 69% (11/16); Opus 42% (5/12); Sonnet 38% (6/16). Tasks were selected with a Sonnet pilot. GPT-6.1 Sol (medium) (Codex CLI): strict pass rate on the harder tasks: 69% (11/16) (n = 16, 95% CI 44–86%). Opus 5.5 (Claude Code): strict pass rate on the harder tasks: 42% (5/12) (n = 12, 95% CI 19–68%). Sonnet 5.5 (Claude Code): strict pass rate on the harder tasks: 38% (6/16) (n = 16, 95% CI 18–61%). Caveat: Selection effect: the study picked tasks that Sonnet did not pass twice in the pilot. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls (calculation: 25% and 38%; the intervals overlap). Selection can produce this pattern, but the run does not establish its cause.
  28. Latency budget, a calculation: steps within an assumed 300 ms budget 3 of 14 (median: 3 of 14); within 1,500 ms 4 of 14 (median: 8 of 14), at p95 or the observed maximum. No voice turn ran. Steps that fit a 300 ms budget at the slow end (calculation): 3 of 14 (median: 3 of 14) (n = 14). Steps that fit a 1,500 ms budget at the slow end (calculation): 4 of 14 (median: 8 of 14) (n = 14). Calculation, not a run. Caveat: Timing samples omit 0 failed or untimed matched explorer calls and 0 failed or untimed Claude head-to-head calls. A failure is not a fast successful step. No pass rate is estimated here.
  29. Routing cost at an assumed million decisions a day, a calculation: Jev $33.70; Sonnet $7,324. No load test ran. Daily cost at 1 million decisions, Jev (calculation): $33.70 (n = 82). Daily cost at 1 million decisions, Sonnet 5.5 router (calculation): $7,324 (n = 82). Calculation, not a run. Caveat: The routing and live Jev protocols predate the first counted calls by file birth time. The overhead protocol predates its microbenchmark output. Claude ran 82 calls per router, below its 120-call cap. Jev ran 246 counted calls, below its 300-call cap. One Haiku pilot and one Jev probe are excluded. All counted calls completed without call errors. The policy microbenchmark used 5,000 warm-up iterations and 64 synthetic contexts; the minimum and individual timing samples are unavailable, so we cannot rebuild its quantiles or full range. No separate validator-control receipt or Sonnet pilot is retained.
  30. 28 studies, one place. Intervals, sources and every failure kept.
Live story · 28 sSWE-bench Verified: Agent vs 11 public model runs

SWE-bench Verified: Agent vs 11 public model runs

Agent resolved 25 of 33 SWE-bench Verified instances; every 95% interval overlaps the public panel, so the data does not rank it.

Transcript
  1. SWE-bench Verified · same 33 instances. Agent vs 11 public model runs. One attempt each. No retries. 95% intervals on individual run rates. The panel mean has no interval.
  2. Agent resolved 25 of 33. The public panel averaged 74.1% on the same instances. Agent resolved (full pipeline, Sonnet 5.5): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Agent model cost per attempt (calculation, notional): $2.81 (n = 33). Caveat: Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
  3. Panel runs resolved 21 to 28 of 33. Every interval overlaps, so this sample cannot rank Agent. Chart: Resolved rate on the same 33 SWE-bench Verified instances (n = 33 each). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
  4. Hardest band: on the 4 instances no panel model solved, Agent solved 1. Too few to call a rate. Chart: Resolved rate by difficulty band (n = 4–14 each). Caveat: Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
  5. Open benchmarks: intervals, sources and every failure kept.

Key numbers

72% (18/25)

Agent resolved, campaign 1 sample of 25

95% CI 52%–86% · n = 25

76% (19/25)

Agent resolved, original seed draw of 25 (no replacements)

95% CI 57%–89% · n = 25

74.1%

Public panel mean on the same 33 instances

n = 33

$2.81

Agent model cost per attempt (notional)

n = 33

$3.71

Agent model cost per resolved instance (notional)

n = 25

9.6min

Median worker time per attempt

n = 33

49

Model calls per attempt

n = 33

One real attempt as a receipt: outcome, minutes, model calls and notional cost

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

GPT 5.2 (high)
Gemini 3 Flash (high)
GLM 5 (high)
Agent (Sonnet 5.5, full pipeline)
Claude 4.5 Sonnet (high)
Claude 4.5 Haiku (high)
Claude 4.5 Opus (high)
DeepSeek V3.2 (high)
MiniMax M2.5 (high)
Claude 4.6 Opus
Kimi K2.5 (high)
GPT 5 mini

Every interval overlaps every other: this chart does not order these rows.

12 rows. Highest GPT 5.2 (high) 85% (95% interval 69%–93%, n 33). Lowest GPT 5 mini 64% (95% interval 47%–78%, n 33). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 33 per row

Agent vs 11 public mini-SWE-agent v2 runs, one attempt each

Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

Share card (PNG)
  • Agent
  • Public panel mean (square)
In chart order.
No panel model solved it
Under half solved it
Half or more solved it
Every panel model solved it

Gap labels, Public panel mean vs Agent: Public panel mean is x percentage points higher (+) or lower (−) than Agent, calculated from the two values shown; lines are the 95% Wilson interval.

4 difficulty bands, 2 series: Agent, Public panel mean. Agent: highest Every panel model solved it 86% (95% interval 60%–96%, n 14). Lowest No panel model solved it 25% (95% interval 4.6%–70%, n 4). All intervals overlap. Public panel mean: highest Every panel model solved it 100% (n 14). Lowest No panel model solved it 0% (n 4). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 4–14 per row

Band = how many of the 11 public panel models solved the instance

Agent whiskers are 95% Wilson intervals. The "no panel model solved it" band has 4 instances; Agent resolved 1 (matplotlib__matplotlib-21568).

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, SWE-bench campaign rules and sample design

Share card (PNG)
  • Public panel (mini-SWE-agent v2)
  • Agent
Better: upper left

12 points: Resolved rate against Mean model cost per instance (USD). Mean model cost per instance (USD) runs from $0.051 to $2.81; Resolved rate from 64% to 85%. Highlighted: Agent.

Notesn = 33 per point

Same 33 instances. Panel = API list price; Agent = notional subscription estimate

Agent's cost includes repository onboarding, planning, verification and review; it is a list-price estimate for subscription calls, not an invoice. Panel costs are published API costs.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, Anthropic list prices (Claude models)

Share card (PNG)

Model calls per instance

Mean over the same 33 instances

DeepSeek V3.2 (high)
GLM 5 (high)
Claude 4.5 Haiku (high)
MiniMax M2.5 (high)
Kimi K2.5 (high)
Gemini 3 Flash (high)
Claude 4.5 Sonnet (high)
Agent
Claude 4.5 Opus (high)
GPT 5.2 (high)
Claude 4.6 Opus
GPT 5 mini

Hover or focus a bar for its ratio to Agent (the highlighted row): a ratio of the two values shown, not a measurement.

12 rows. Highest DeepSeek V3.2 (high) 88.2 (n 33). Lowest GPT 5 mini 20.8 (n 33).

Notesn = 33 per row

A panel call is one bash-agent step. An Agent call is one model request of any stage (research, plan, act, verify, review).

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

Share card (PNG)
$92.64Total over 33 attempts (from the note)

Parts sorted by value, largest first

  1. Act (edit and run)
  2. Research
  3. Verify
  4. Other
  5. Review
  6. Context compaction
  7. Onboarding notes

Shares are calculated from the values shown; rounding can make the sum of the parts differ from the stated total by a cent.

7 rows. Highest Act (edit and run) $44.05. Lowest Onboarding notes $2.09.

Notes

Share of notional model cost by stage, all 33 attempts

Total $92.64 over 33 attempts.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

Share card (PNG)
  • Agent
  • Public panel mean, same instances
Campaign 1: 25-instance sample
Campaign 2: 8 compiled-extension instances
Original seed draw of 25
All 33 attempted

4 rows, 2 series: Agent, Public panel mean, same instances. Agent: highest Campaign 2: 8 compiled-extension instances 88% (95% interval 53%–98%, n 8). Lowest Campaign 1: 25-instance sample 72% (95% interval 52%–86%, n 25). All intervals overlap. Public panel mean, same instances: highest Campaign 2: 8 compiled-extension instances 84% (n 8). Lowest Original seed draw of 25 71% (n 25). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 8–33 per row

Agent resolved rate and 95% Wilson interval per declared view

The original draw and "all 33" mix two platform builds.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, SWE-bench campaign rules and sample design

Share card (PNG)

Tables

Every attempt

Sorted by how many panel systems solved each instance, most first

Each tile is one row; from top: Panel solved, then Agent outcome.

11/1111/1111/1111/1111/1111/1111/1111/1111/1111/1111/1111/1111/1111/1110/1110/1110/1110/1110/1110/1110/119/119/118/117/114/113/113/112/110/110/110/110/11

Shade is the share k of n in each cell.

Counts from the table, not an intervaln = 11 per cell

33 rows by 2 columns. 14 of 33 k/n cells are full (n of n) and 4 are zero. Agent outcome: 25 of 33 resolved.

An empty patch counts as unresolved here; the Table view gives each outcome.

Public panel on the same instances

SystemResolved of 33RateFull Verified board (500)Mean cost per instance, these 33
GPT 5.2 (high)2885%73%$0.53
Gemini 3 Flash (high)2782%76%$0.36
GLM 5 (high)2679%73%$0.53
Agent (Sonnet 5.5, full pipeline)2576%—$2.81
Claude 4.5 Sonnet (high)2576%71%$0.69
Claude 4.5 Haiku (high)2576%67%$0.36
Claude 4.5 Opus (high)2473%77%$0.86
DeepSeek V3.2 (high)2473%70%$0.46
MiniMax M2.5 (high)2370%76%$0.075
Claude 4.6 Opus2370%76%$0.61
Kimi K2.5 (high)2370%71%$0.18
GPT 5 mini2164%56%$0.051

Method

  1. Sample: 25 of the 500 Verified instances, stratified by public difficulty (seed 20261004). Difficulty is how many of the 11 public mini-SWE-agent v2 runs solved the instance.
  2. Campaign 1 ran 25 instances on platform build f0ac3a8a. Six compiled-extension instances could not import in the worker checkout, so the declared rule replaced them in the same band.
  3. Campaign 2 ran those compiled instances (plus two replacement candidates) on build 236c0d3f after a sandbox fix.
  4. One attempt per instance. No retries, no operator answers: an escalation is graded on what was delivered.
  5. Grading uses the official SWE-bench harness and instance images (under amd64 emulation). The gold patch resolved on the host for every instance.
  6. Agent runs its full pipeline (onboarding, research, plan, act, verify, review) on claude-sonnet-5-5 through a subscription CLI. Costs are list-price estimates of the recorded tokens.

Caveats

  • n = 33: intervals are wide. This is a defect-finding run, not a ranking.
  • Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
  • Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
  • Campaign 1 ended up 19 django, 3 sympy, 2 sphinx and 1 xarray after replacements. Campaign 2 covers the compiled repositories.
  • Verified issues are public (2015 to 2023) and likely in every model’s training data; contamination is uncontrolled for all systems.
  • Difficulty bands come from the panel’s own results, so a system outside the panel tends to look better than the panel on hard bands and worse on easy ones (regression to the mean). Read the band chart with that selection effect in mind.
  • Three campaign-1 empty patches were platform holds before delivery (missing lint tools, a too-literal plan gate, an unanswered question), not wrong fixes. They count as failures here.

Sources

  • Agent on SWE-bench Verified, campaign 1 (25 instances)

    Our recorded runs ·

    Stratified sample of 25 Verified instances (seed 20261004), one attempt each, official grading harness. Fixed model claude-sonnet-5-5, platform build f0ac3a8a.

    Raw data: swebench/attempts.json, swebench/exclusions.json

  • Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

    Our recorded runs ·

    The 6 compiled-extension instances that campaign 1 could not run, plus 2 replacement candidates. One attempt each, platform build 236c0d3f.

    Raw data: swebench/attempts.json

  • SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

    Public leaderboard ·

    Public per-instance results of 11 models under mini-SWE-agent 2.0.0 (bash only, one attempt). Costs are API list prices as published.

    Raw data: swebench/panel.json

  • SWE-bench campaign rules and sample design

    Protocol ·

    Rules declared before the first run: escalations are graded as delivered, gold must resolve on the host, blocked instances are replaced in the same difficulty band, no second attempts. The excluded and replaced instances are listed in the exclusions extract.

    Raw data: swebench/exclusions.json

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Agent on SWE-bench Verified vs 11 public models”, updated October 5, 2026, https://agent.sasid.ai/benchmarks/swe-bench-verified.

Explainers that cite this study

Read the methods and terms in the context of these recorded results.

More comparisons based on this study (58)

These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.

More write-ups that cite this study (1)

Models and comparisons in this study

More studies

All benchmarks
  • Prompt Caching
  • Cache Reuse

Does a new Claude Code session reuse the prompt cache of an earlier one?

30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.

0of 2 (95% interval 0% to 66%) · Later sessions with at least 50% of turn-1 input cached, A: new folder each time · n = 2

4 chartsUpdated October 7, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.