Explainer · SWE-bench Verified
SWE-bench Verified, explained
Definition
SWE-bench Verified is a benchmark of 500 real GitHub issues from popular Python repositories, each checked by people to have a clear problem statement and a fair test. A coding agent gets the repository and the issue text, writes a patch, and the patch counts as resolved only when the project's hidden tests that the real fix made pass now pass, and the tests that passed before still pass. The score is the share of instances resolved.
Agent team · · 4 min read · Every number is from the public studies
How it works
How one SWE-bench Verified instance is scored
A real GitHub issue, a repository at the commit before the fix, and tests the agent never sees.
The issue
The agent gets the issue text and the repository at the commit before the human fix.
- No test files from the fix
- Same instance list for every model
The change
The agent reads, edits and runs code until it decides it is done.
- Every attempt counted, failures too
- Cost and time recorded per attempt
Hidden tests
The tests from the real fix run against the change.
- FAIL_TO_PASS: the bug is fixed
- PASS_TO_PASS: nothing else broke
76% (25/33)
Resolved only when both test sets pass.
- 95% Wilson interval on the rate
- n shown with every rate
| Stage | What it does | Checks |
|---|---|---|
| 1. The issue | The agent gets the issue text and the repository at the commit before the human fix. | No test files from the fix; Same instance list for every model |
| 2. The change | The agent reads, edits and runs code until it decides it is done. | Every attempt counted, failures too; Cost and time recorded per attempt |
| 3. Hidden tests | The tests from the real fix run against the change. | FAIL_TO_PASS: the bug is fixed; PASS_TO_PASS: nothing else broke |
| 4. Resolved? | Resolved only when both test sets pass. (76% (25/33)) | 95% Wilson interval on the rate; n shown with every rate |
Four stages: the issue, the agent's change, the hidden tests, and the score. Agent resolved 76% (25/33) of the instances it attempted.
Source: SWE-bench Verified study
How an instance is scored
Each instance comes from a merged pull request that fixed an issue. The benchmark keeps the repository at the commit before the fix and two groups of tests:
- Fail-to-pass tests: they failed before the real fix and passed after it. The agent's patch must make them pass.
- Pass-to-pass tests: they passed before and after. The patch must not break them.
The agent never sees these tests. It sees the issue text and the code, like a developer who picks up the ticket. A patch that compiles, looks right and passes the agent's own checks still scores zero if a hidden test fails.
"Verified" is the subset that annotators confirmed as solvable from the issue text, with tests that do not demand one exact implementation. It removes many instances of the original SWE-bench that were unfair or ambiguous.
What a resolved rate tells you
A resolved rate is a proportion, so its precision depends on the number of instances. A Wilson 95% interval around 25 of 33 runs from 59% to 87%: almost 30 points wide. The same rate on all 500 instances would give an interval about four times narrower, because the width shrinks with the square root of n.
We ran Agent, our full worker pipeline on Claude Sonnet 5.5, on 33 Verified instances and compared it with 11 public single-model runs (mini-SWE-agent v2) on the very same instances:
- Agent resolved 25 of 33 (76%, 95% interval 59% to 87%).
- The 11 public runs resolved between 21 and 28 of the same 33 (panel mean 74.1%).
- Every interval overlaps every other, so this sample cannot rank Agent above or below any panel model.
That last point is the main lesson for reading any SWE-bench table: two scores a few points apart on a small sample are a tie. See Wilson confidence intervals for AI benchmarks.
Difficulty is uneven
Instances differ a lot in difficulty. One way to see it is to count how many of the 11 public models solved each instance:
- On the 14 instances every panel model solved, Agent resolved 86% (60% to 96%).
- On the 4 instances no panel model solved, Agent resolved 1 (25%, 5% to 70%).
A score therefore depends on which instances are in the sample. Comparing two systems is only fair on the same instances, which is why the chart above uses the same 33 for all twelve rows.
Cost and time are part of the result
A resolved rate says nothing about what it took. Agent spent a notional $2.81 and 49 model calls per attempt, with a median of 9.6 minutes, because it onboards, plans, verifies and reviews before it delivers. A bare agent loop is cheaper and faster per instance. The cost thought experiment prices the same recorded tokens at other models' list prices.
How to read a SWE-bench Verified claim
- Check n. Was it all 500 instances or a sample? A sample needs an interval.
- Check the harness. A model score depends on the agent loop around it: tools, retries, time and cost limits. A model in one harness is not the same as the model in another.
- Check the instance set. Two scores on different subsets do not compare.
- Check attempts. One attempt per instance, or the best of several? Best-of-k inflates the rate.
- Check for contamination. The issues and fixes are public, so a model may have seen them in training.
Frequently asked questions
What is a good SWE-bench Verified score?
There is no fixed threshold, because the score depends on the harness, the attempt policy and the instance set. On our 33-instance sample, 11 public model runs resolved between 21 and 28 instances; differences of that size on 33 instances fall inside overlapping 95% intervals.
What is the difference between SWE-bench and SWE-bench Verified?
SWE-bench is the original, larger set of GitHub issues. SWE-bench Verified is a 500-instance subset that people checked for a clear issue text and fair tests, so a failure is more likely to be the agent's fault than the benchmark's.
Why does Agent cost more per instance than a bare model run?
Agent runs a full pipeline: it onboards the repository, plans, implements, verifies and reviews. That took 49 model calls and a notional $2.81 per attempt on our sample. A bare agent loop makes fewer calls, so it is cheaper per instance; the resolved rates on the same instances overlap.
Can a patch pass its own tests and still fail SWE-bench?
Yes. The scoring uses hidden tests from the real fix. A patch can pass the agent's own checks and still fail a hidden fail-to-pass test, or break a pass-to-pass test, and then it is not resolved.
Watch the data
SWE-bench Verified: Agent vs 11 public model runs
Agent resolved 25 of 33 SWE-bench Verified instances; every 95% interval overlaps the public panel, so the data does not rank it.
Transcript
- SWE-bench Verified · same 33 instances. Agent vs 11 public model runs. One attempt each. No retries. 95% intervals on individual run rates. The panel mean has no interval.
- Agent resolved 25 of 33. The public panel averaged 74.1% on the same instances. Agent resolved (full pipeline, Sonnet 5.5): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Agent model cost per attempt (calculation, notional): $2.81 (n = 33). Caveat: Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
- Panel runs resolved 21 to 28 of 33. Every interval overlaps, so this sample cannot rank Agent. Chart: Resolved rate on the same 33 SWE-bench Verified instances (n = 33 each). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
- Hardest band: on the 4 instances no panel model solved, Agent solved 1. Too few to call a rate. Chart: Resolved rate by difficulty band (n = 4–14 each). Caveat: Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
- Open benchmarks: intervals, sources and every failure kept.