• SWE-bench
  • Coding Agents
  • Leaderboard
  • Benchmarks

SWE-bench Verified: an agent pipeline vs 11 public models on the same 33 tasks

Agent resolved 25 of 33 SWE-bench Verified instances (76%). Eleven public models solved 21 to 28 of the same ones. Why that is a tie, not a win.

TL;DR

  • Agent resolved 25 of 33 SWE-bench Verified instances: 75.8%, with a 95% interval of 59% to 87%.
  • On the very same 33 instances, 11 public mini-SWE-agent v2 runs resolved 21 to 28. The panel mean is 74.1%.
  • Every interval overlaps. This sample cannot rank Agent above or below any panel model, and we do not claim it does.
  • Agent is slower and costs more per instance: a median 9.6 minutes, about 49 model calls and a notional $2.81 per attempt.
  • It resolved 1 of 4 instances that no panel model solved. It also missed 2 of 14 that every panel model solved.

Full data, every attempt and every interval: /benchmarks/swe-bench-verified.

Why compare on the same instances?

Most leaderboard comparisons put a number from 500 instances next to a number from some other subset. That mixes difficulty with skill. We wanted a fair look, so we did something simpler. We took the public per-instance results of 11 models under mini-SWE-agent 2.0.0 and scored them on exactly the instances we ran.

So every row in the chart below answers one question: of these 33 issues, how many did each system fix?

GPT 5.2 (high)
Gemini 3 Flash (high)
GLM 5 (high)
Agent (Sonnet 5.5, full pipeline)
Claude 4.5 Sonnet (high)
Claude 4.5 Haiku (high)
Claude 4.5 Opus (high)
DeepSeek V3.2 (high)
MiniMax M2.5 (high)
Claude 4.6 Opus
Kimi K2.5 (high)
GPT 5 mini

Every interval overlaps every other: this chart does not order these rows.

12 rows. Highest GPT 5.2 (high) 85% (95% interval 69%–93%, n 33). Lowest GPT 5 mini 64% (95% interval 47%–78%, n 33). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 33 per row

Agent vs 11 public mini-SWE-agent v2 runs, one attempt each

Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

Live story · 28 sSWE-bench Verified: Agent vs 11 public model runs

SWE-bench Verified: Agent vs 11 public model runs

Agent resolved 25 of 33 SWE-bench Verified instances; every 95% interval overlaps the public panel, so the data does not rank it.

Transcript
  1. SWE-bench Verified · same 33 instances. Agent vs 11 public model runs. One attempt each. No retries. 95% intervals on individual run rates. The panel mean has no interval.
  2. Agent resolved 25 of 33. The public panel averaged 74.1% on the same instances. Agent resolved (full pipeline, Sonnet 5.5): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Agent model cost per attempt (calculation, notional): $2.81 (n = 33). Caveat: Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
  3. Panel runs resolved 21 to 28 of 33. Every interval overlaps, so this sample cannot rank Agent. Chart: Resolved rate on the same 33 SWE-bench Verified instances (n = 33 each). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
  4. Hardest band: on the 4 instances no panel model solved, Agent solved 1. Too few to call a rate. Chart: Resolved rate by difficulty band (n = 4–14 each). Caveat: Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
  5. Open benchmarks: intervals, sources and every failure kept.

GPT 5.2 (high) leads the panel with 28 of 33. Gemini 3 Flash (high) has 27 and GLM 5 (high) has 26. Agent, Claude 4.5 Sonnet (high) and Claude 4.5 Haiku (high) each have 25. GPT 5 mini sits at the bottom with 21.

Now look at the whiskers. With n = 33, one extra solved instance moves a rate by three points. The interval for 25 of 33 runs from 59% to 87%. The interval for 28 of 33 runs from 69% to 93%. They overlap by a wide margin. If someone tells you a 3-instance gap on a sample this size is a ranking, ask for the interval.

What Agent actually is

The panel is a set of models in one bash-only harness. Each model gets a shell and the issue, and it works until it submits. That is a clean way to compare models.

Agent is not a model. It is a full worker pipeline on one model, claude-sonnet-5-5. For each instance it:

  1. Onboards the repository and writes notes.
  2. Researches the issue.
  3. Writes a plan.
  4. Edits and runs code.
  5. Verifies the change.
  6. Reviews its own work before delivery.

That is why we call this a systems comparison, not a model comparison. It also means the result says nothing about Sonnet 5.5 as a bare model. For that question, see harness vs model: where the gains come from.

Where the hard instances are

The most useful view is by difficulty. We define difficulty by the public data itself: how many of the 11 panel models solved the instance.

  • Agent
  • Public panel mean (square)
In chart order.
No panel model solved it
Under half solved it
Half or more solved it
Every panel model solved it

Gap labels, Public panel mean vs Agent: Public panel mean is x percentage points higher (+) or lower (−) than Agent, calculated from the two values shown; lines are the 95% Wilson interval.

4 difficulty bands, 2 series: Agent, Public panel mean. Agent: highest Every panel model solved it 86% (95% interval 60%–96%, n 14). Lowest No panel model solved it 25% (95% interval 4.6%–70%, n 4). All intervals overlap. Public panel mean: highest Every panel model solved it 100% (n 14). Lowest No panel model solved it 0% (n 4). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 4–14 per row

Band = how many of the 11 public panel models solved the instance

Agent whiskers are 95% Wilson intervals. The "no panel model solved it" band has 4 instances; Agent resolved 1 (matplotlib__matplotlib-21568).

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, SWE-bench campaign rules and sample design

The bands tell a more interesting story than the headline:

  • No panel model solved it (4 instances). Agent resolved 1: matplotlib__matplotlib-21568. The panel mean here is 0 by definition.
  • Under half solved it (4 instances). Agent resolved 3 of 4. The panel mean is 27.3%.
  • Half or more solved it (11 instances). Agent resolved 9 of 11. The panel mean is 85.1%.
  • Every panel model solved it (14 instances). Agent resolved 12 of 14. The panel mean is 100% by definition.

Agent looks a bit better than the panel on hard instances and a bit worse on easy ones. With four instances per hard band, that is a pattern to test, not a finding. The bands also come from the panel's own results, so any system outside the panel tends to look better than the panel on the hard bands and worse on the easy ones (regression to the mean). Part of this pattern is that selection effect. The two easy misses are instructive, though. One (astropy__astropy-7336) was a plain wrong fix. The other (sympy__sympy-18189) was an empty patch after 1.6 minutes, which was a platform hold, not a coding failure. We count it anyway.

Every way to slice the run

We ran two campaigns. Campaign 1 drew a stratified sample of 25 instances with a fixed seed. Six of them had compiled extensions that would not import in the worker checkout. The rule we declared before the first run said: replace a blocked instance with another from the same difficulty band. Campaign 2 then ran those compiled instances, plus two replacement candidates, on a newer platform build after a sandbox fix.

That gives several honest ways to report the result. We show all of them.

  • Agent
  • Public panel mean, same instances
Campaign 1: 25-instance sample
Campaign 2: 8 compiled-extension instances
Original seed draw of 25
All 33 attempted

4 rows, 2 series: Agent, Public panel mean, same instances. Agent: highest Campaign 2: 8 compiled-extension instances 88% (95% interval 53%–98%, n 8). Lowest Campaign 1: 25-instance sample 72% (95% interval 52%–86%, n 25). All intervals overlap. Public panel mean, same instances: highest Campaign 2: 8 compiled-extension instances 84% (n 8). Lowest Original seed draw of 25 71% (n 25). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 8–33 per row

Agent resolved rate and 95% Wilson interval per declared view

The original draw and "all 33" mix two platform builds.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, SWE-bench campaign rules and sample design

  • Campaign 1, 25-instance sample: 18 of 25 (72%), panel mean 70.9%.
  • Campaign 2, 8 compiled-extension instances: 7 of 8 (87.5%), panel mean 84.1%.
  • Original seed draw of 25, no replacements: 19 of 25 (76%), panel mean 70.9%.
  • All 33 attempted: 25 of 33 (75.8%), panel mean 74.1%.

In every view, Agent sits within a few points of the panel mean. No view lets us claim more. The original draw and the 33-instance figure mix two platform builds, and we say so on the chart.

Calls, time and money

The pipeline is busier per instance than a typical panel run, but not by as much as you might expect.

Model calls per instance

Mean over the same 33 instances

DeepSeek V3.2 (high)
GLM 5 (high)
Claude 4.5 Haiku (high)
MiniMax M2.5 (high)
Kimi K2.5 (high)
Gemini 3 Flash (high)
Claude 4.5 Sonnet (high)
Agent
Claude 4.5 Opus (high)
GPT 5.2 (high)
Claude 4.6 Opus
GPT 5 mini

Hover or focus a bar for its ratio to Agent (the highlighted row): a ratio of the two values shown, not a measurement.

12 rows. Highest DeepSeek V3.2 (high) 88.2 (n 33). Lowest GPT 5 mini 20.8 (n 33).

Notesn = 33 per row

A panel call is one bash-agent step. An Agent call is one model request of any stage (research, plan, act, verify, review).

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

Agent made a mean of 49.5 model calls per instance. The panel ranges from 20.8 (GPT 5 mini) to 88.2 (DeepSeek V3.2 high). Claude 4.5 Sonnet (high) made 51. A panel call is one bash step. An Agent call is one request in any stage. So the call counts are similar, but an Agent call carries more context: notes, plans and verification output.

The cost difference is larger. Agent's notional cost was $2.81 per attempt and $3.71 per resolved instance. The panel's published costs per instance range from $0.051 to $0.861. We break that down in what one resolved SWE-bench task really costs.

Time per attempt had a median of 9.6 minutes, with a range of 1.6 to 54.1 minutes. The long tail came from the compiled-extension repositories (astropy, matplotlib, scikit-learn), where builds and tests are slow.

How we measured

  • Sample. 25 of the 500 Verified instances, stratified by public difficulty, seed 20261004. Campaign 2 added the compiled-extension instances and two replacements, for 33 attempted in total.
  • Attempts. One attempt per instance. No retries. No operator answers. An escalation is graded on what was delivered at that point.
  • Grading. The official SWE-bench harness and instance images, under amd64 emulation. Before we counted anything, we checked that the gold patch resolved on our host for every instance.
  • Model. Agent ran its full pipeline on claude-sonnet-5-5 through a subscription CLI. Costs are list-price estimates of the recorded tokens.
  • Panel. Public per-instance results of 11 models under mini-SWE-agent 2.0.0, one attempt each, as published on swebench.com. Costs are the published API list costs.

All of this is in the study's method section, with sources, at /benchmarks/swe-bench-verified.

Caveats

  • n = 33. The intervals are wide. We ran this to find defects in our pipeline, not to claim a rank.
  • Systems, not models. The panel compares models in one harness. Agent is a full pipeline on one model. Different Claude versions are involved (Sonnet 5.5 in Agent, Claude 4.5 and 4.6 models in the panel).
  • Costs are not like for like. Panel costs are published API costs. Agent's costs are notional estimates of subscription calls, not invoices.
  • Repository mix. After replacements, campaign 1 was 19 django, 3 sympy, 2 sphinx and 1 xarray instances. Campaign 2 covers the compiled repositories.
  • Contamination. Verified issues are public, from 2015 to 2023. They are likely in every model's training data. Nobody controls for that here, including us.
  • Three empty patches in campaign 1 were platform holds before delivery: missing lint tools, a plan gate that was too literal, and an unanswered question. They were not wrong fixes. They count as failures anyway, because that is what a user would have seen. We explain that rule in why we count every failed attempt.

Try it on your own repository

Leaderboards are a starting point. The real test is your codebase, your tickets and your review bar. Agent onboards a repository, plans, codes, verifies and reviews before it hands over a pull request, and it shows you every step and every receipt. Try Agent on a real issue and judge the result yourself.

The data behind this post

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.