Explainer · SWE-bench Verified

SWE-bench Verified, explained

Definition

SWE-bench Verified is a benchmark of 500 real GitHub issues from popular Python repositories, each checked by people to have a clear problem statement and a fair test. A coding agent gets the repository and the issue text, writes a patch, and the patch counts as resolved only when the project's hidden tests that the real fix made pass now pass, and the tests that passed before still pass. The score is the share of instances resolved.

Agent team · · 4 min read · Every number is from the public studies

How it works

How one SWE-bench Verified instance is scored

A real GitHub issue, a repository at the commit before the fix, and tests the agent never sees.

Measured
  1. The issue

    The agent gets the issue text and the repository at the commit before the human fix.

    • No test files from the fix
    • Same instance list for every model
  2. The change

    The agent reads, edits and runs code until it decides it is done.

    • Every attempt counted, failures too
    • Cost and time recorded per attempt
  3. Hidden tests

    The tests from the real fix run against the change.

    • FAIL_TO_PASS: the bug is fixed
    • PASS_TO_PASS: nothing else broke
  4. Resolved?

    76% (25/33)

    Resolved only when both test sets pass.

    • 95% Wilson interval on the rate
    • n shown with every rate

Four stages: the issue, the agent's change, the hidden tests, and the score. Agent resolved 76% (25/33) of the instances it attempted.

Source: SWE-bench Verified study

How an instance is scored

Each instance comes from a merged pull request that fixed an issue. The benchmark keeps the repository at the commit before the fix and two groups of tests:

  • Fail-to-pass tests: they failed before the real fix and passed after it. The agent's patch must make them pass.
  • Pass-to-pass tests: they passed before and after. The patch must not break them.

The agent never sees these tests. It sees the issue text and the code, like a developer who picks up the ticket. A patch that compiles, looks right and passes the agent's own checks still scores zero if a hidden test fails.

"Verified" is the subset that annotators confirmed as solvable from the issue text, with tests that do not demand one exact implementation. It removes many instances of the original SWE-bench that were unfair or ambiguous.

What a resolved rate tells you

A resolved rate is a proportion, so its precision depends on the number of instances. A Wilson 95% interval around 25 of 33 runs from 59% to 87%: almost 30 points wide. The same rate on all 500 instances would give an interval about four times narrower, because the width shrinks with the square root of n.

We ran Agent, our full worker pipeline on Claude Sonnet 5.5, on 33 Verified instances and compared it with 11 public single-model runs (mini-SWE-agent v2) on the very same instances:

GPT 5.2 (high)
Gemini 3 Flash (high)
GLM 5 (high)
Agent (Sonnet 5.5, full pipeline)
Claude 4.5 Sonnet (high)
Claude 4.5 Haiku (high)
Claude 4.5 Opus (high)
DeepSeek V3.2 (high)
MiniMax M2.5 (high)
Claude 4.6 Opus
Kimi K2.5 (high)
GPT 5 mini

Every interval overlaps every other: this chart does not order these rows.

12 rows. Highest GPT 5.2 (high) 85% (95% interval 69%–93%, n 33). Lowest GPT 5 mini 64% (95% interval 47%–78%, n 33). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 33 per row

Agent vs 11 public mini-SWE-agent v2 runs, one attempt each

Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

  • Agent resolved 25 of 33 (76%, 95% interval 59% to 87%).
  • The 11 public runs resolved between 21 and 28 of the same 33 (panel mean 74.1%).
  • Every interval overlaps every other, so this sample cannot rank Agent above or below any panel model.

That last point is the main lesson for reading any SWE-bench table: two scores a few points apart on a small sample are a tie. See Wilson confidence intervals for AI benchmarks.

Difficulty is uneven

Instances differ a lot in difficulty. One way to see it is to count how many of the 11 public models solved each instance:

  • Agent
  • Public panel mean (square)
In chart order.
No panel model solved it
Under half solved it
Half or more solved it
Every panel model solved it

Gap labels, Public panel mean vs Agent: Public panel mean is x percentage points higher (+) or lower (−) than Agent, calculated from the two values shown; lines are the 95% Wilson interval.

4 difficulty bands, 2 series: Agent, Public panel mean. Agent: highest Every panel model solved it 86% (95% interval 60%–96%, n 14). Lowest No panel model solved it 25% (95% interval 4.6%–70%, n 4). All intervals overlap. Public panel mean: highest Every panel model solved it 100% (n 14). Lowest No panel model solved it 0% (n 4). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 4–14 per row

Band = how many of the 11 public panel models solved the instance

Agent whiskers are 95% Wilson intervals. The "no panel model solved it" band has 4 instances; Agent resolved 1 (matplotlib__matplotlib-21568).

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, SWE-bench campaign rules and sample design

  • On the 14 instances every panel model solved, Agent resolved 86% (60% to 96%).
  • On the 4 instances no panel model solved, Agent resolved 1 (25%, 5% to 70%).

A score therefore depends on which instances are in the sample. Comparing two systems is only fair on the same instances, which is why the chart above uses the same 33 for all twelve rows.

Cost and time are part of the result

A resolved rate says nothing about what it took. Agent spent a notional $2.81 and 49 model calls per attempt, with a median of 9.6 minutes, because it onboards, plans, verifies and reviews before it delivers. A bare agent loop is cheaper and faster per instance. The cost thought experiment prices the same recorded tokens at other models' list prices.

How to read a SWE-bench Verified claim

  • Check n. Was it all 500 instances or a sample? A sample needs an interval.
  • Check the harness. A model score depends on the agent loop around it: tools, retries, time and cost limits. A model in one harness is not the same as the model in another.
  • Check the instance set. Two scores on different subsets do not compare.
  • Check attempts. One attempt per instance, or the best of several? Best-of-k inflates the rate.
  • Check for contamination. The issues and fixes are public, so a model may have seen them in training.

Frequently asked questions

What is a good SWE-bench Verified score?

There is no fixed threshold, because the score depends on the harness, the attempt policy and the instance set. On our 33-instance sample, 11 public model runs resolved between 21 and 28 instances; differences of that size on 33 instances fall inside overlapping 95% intervals.

What is the difference between SWE-bench and SWE-bench Verified?

SWE-bench is the original, larger set of GitHub issues. SWE-bench Verified is a 500-instance subset that people checked for a clear issue text and fair tests, so a failure is more likely to be the agent's fault than the benchmark's.

Why does Agent cost more per instance than a bare model run?

Agent runs a full pipeline: it onboards the repository, plans, implements, verifies and reviews. That took 49 model calls and a notional $2.81 per attempt on our sample. A bare agent loop makes fewer calls, so it is cheaper per instance; the resolved rates on the same instances overlap.

Can a patch pass its own tests and still fail SWE-bench?

Yes. The scoring uses hidden tests from the real fix. A patch can pass the agent's own checks and still fail a hidden fail-to-pass test, or break a pass-to-pass test, and then it is not resolved.

Watch the data

Live story · 28 sSWE-bench Verified: Agent vs 11 public model runs

SWE-bench Verified: Agent vs 11 public model runs

Agent resolved 25 of 33 SWE-bench Verified instances; every 95% interval overlaps the public panel, so the data does not rank it.

Transcript
  1. SWE-bench Verified · same 33 instances. Agent vs 11 public model runs. One attempt each. No retries. 95% intervals on individual run rates. The panel mean has no interval.
  2. Agent resolved 25 of 33. The public panel averaged 74.1% on the same instances. Agent resolved (full pipeline, Sonnet 5.5): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Agent model cost per attempt (calculation, notional): $2.81 (n = 33). Caveat: Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
  3. Panel runs resolved 21 to 28 of 33. Every interval overlaps, so this sample cannot rank Agent. Chart: Resolved rate on the same 33 SWE-bench Verified instances (n = 33 each). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
  4. Hardest band: on the 4 instances no panel model solved, Agent solved 1. Too few to call a rate. Chart: Resolved rate by difficulty band (n = 4–14 each). Caveat: Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
  5. Open benchmarks: intervals, sources and every failure kept.

The data behind this explainer

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.