Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)
Does Agent resolve more SWE-bench Verified instances with Claude Opus 5.5 as its brain than with Claude Sonnet 5.5, and at what cost and time?
Published · 5 charts · Download the data or a carousel
67%
95% CI 21%–94% · n = 3
2/3 · Claude Opus 5.5 in Agent: resolved, interim (3 of 8 pairs graded)
33%
95% CI 6.2%–79% · n = 3
1/3 · Claude Sonnet 5.5 in Agent: resolved on the same 3 instances
p = 1.0 · Exact McNemar p on the graded pairs · n = 3
1 pair where only Opus resolved, 0 pairs where only Sonnet did. The test reads only these; p = 1.0 is no evidence of a difference.
The answer
Interim, not a final result: 3 of 8 declared pairs are graded. With Claude Opus 5.5 as its brain, Agent resolved 2 of 3 (95% Wilson 21–94%); with Claude Sonnet 5.5 it resolved 1 of 3 (6–79%) on the same instances. The one discordant instance (django__django-10554) went to Opus, and the exact McNemar p is 1.0, so these pairs show no difference. On the same instances Opus cost 2.6× as much, a list-price calculation ($22.76 vs $8.64), and used 1.9× the worker minutes. The arms ran on different platform builds, and the other 5 pairs remain not started in this dataset; the recorded usage-reset date was 2026-10-09.
Key numbers
3 of 8
Declared pairs graded so far
n = 8
2.6×
Opus vs Sonnet list-price cost on the same instances (calculation)
n = 3
1.9×
Opus vs Sonnet worker minutes on the same instances (ratio of totals)
n = 3
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
| Item | Resolved | 95% interval | n |
|---|---|---|---|
| Claude Opus 5.5 (Agent, new build) | 67% | 21%–94% | 3 |
| Claude Sonnet 5.5 (Agent, older builds) | 33% | 6.2%–79% | 3 |
2 rows. Highest Claude Opus 5.5 (Agent, new build) 67% (95% interval 21%–94%, n 3). Lowest Claude Sonnet 5.5 (Agent, older builds) 33% (95% interval 6.2%–79%, n 3). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 3 per row
One attempt per arm per instance, official grader · 95% Wilson intervals
Interim: 3 of 8 declared pairs are graded; 5 were never started; no resumed attempts are included here. With n = 3 the intervals span most of the axis, so this chart supports no ranking. The instances were chosen to hold Sonnet misses and resolves in equal numbers, so neither rate estimates SWE-bench Verified as a whole.
Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)
| Item | List-price cost per attempt | n |
|---|---|---|
| Claude Opus 5.5 (Agent, new build) | $7.59 | 3 |
| Claude Sonnet 5.5 (Agent, older builds) | $2.88 | 3 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 (Agent, new build) $7.59 (n 3). Lowest Claude Sonnet 5.5 (Agent, older builds) $2.88 (n 3).
Notesn = 3 per row
Mean over the same 3 instances; notional cost from the platform price table
Calculation, not an invoice: the platform price table at each run commit applied to the recorded tokens of subscription runs (Opus 5.5: $4 input, $20 output, $0.20 cache read per million tokens). Totals $22.76 vs $8.64: 2.6×, a ratio of two calculations.
Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Repricing calculation, Anthropic list prices (Claude models)
- Claude Sonnet 5.5 (Agent, older builds)
- Claude Opus 5.5 (Agent, new build) (square)
Gap labels, Claude Opus 5.5 (Agent, new build) vs Claude Sonnet 5.5 (Agent, older builds): Claude Opus 5.5 (Agent, new build) is x% higher (+) or lower (−) than Claude Sonnet 5.5 (Agent, older builds), calculated from the two values shown (the change counted from Claude Sonnet 5.5 (Agent, older builds)’s value).
| Instance | Claude Sonnet 5.5 (Agent, older builds) | Claude Opus 5.5 (Agent, new build) | n |
|---|---|---|---|
| django-10554 | $1.52 | $9.99 | 1 |
| matplotlib-21568 | $3.02 | $4.90 | 1 |
| django-15022 | $4.10 | $7.87 | 1 |
List-price calculation, not a run. 3 instances, 2 series: Claude Sonnet 5.5 (Agent, older builds), Claude Opus 5.5 (Agent, new build). Claude Sonnet 5.5 (Agent, older builds): highest django-15022 $4.10 (n 1). Lowest django-10554 $1.52 (n 1). Claude Opus 5.5 (Agent, new build): highest django-10554 $9.99 (n 1). Lowest matplotlib-21568 $4.90 (n 1).
Notesn = 1 per row
One attempt per arm; notional cost from the platform price table
Calculation, not an invoice. Outcomes: django-10554 Opus resolved, Sonnet unresolved; matplotlib-21568 Opus resolved, Sonnet resolved; django-15022 Opus unresolved, Sonnet unresolved.
Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Repricing calculation, Anthropic list prices (Claude models)
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Worker minutes per attempt | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 (Agent, new build) | 20.3 min | 9.8 min–25.3 min | 3 |
| Claude Sonnet 5.5 (Agent, older builds) | 9.4 min | 4.7 min–15 min | 3 |
2 rows. Slowest Claude Opus 5.5 (Agent, new build) 20.3 min (range 9.8 min–25.3 min, n 3). Fastest Claude Sonnet 5.5 (Agent, older builds) 9.4 min (range 4.7 min–15 min, n 3). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Median minutes; whiskers = fastest and slowest of 3 attempts (not an interval)
Worker minutes from the run reports. Opus attempts ran one at a time on the newer build; the Sonnet attempts ran in earlier campaigns. 3 attempts per arm is too few to call a difference; a range is not a confidence interval.
Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)
- Claude Sonnet 5.5 (Agent, older builds)
- Claude Opus 5.5 (Agent, new build) (square)
Gap labels, Claude Opus 5.5 (Agent, new build) vs Claude Sonnet 5.5 (Agent, older builds): Claude Opus 5.5 (Agent, new build) is x% higher (+) or lower (−) than Claude Sonnet 5.5 (Agent, older builds), calculated from the two values shown (the change counted from Claude Sonnet 5.5 (Agent, older builds)’s value).
| Pipeline stage | Claude Sonnet 5.5 (Agent, older builds) | Claude Opus 5.5 (Agent, new build) | n |
|---|---|---|---|
| Research | $2.34 | $6.28 | 3 |
| Plan | $1.16 | $1.11 | 3 |
| Implement | $1.32 | $4.94 | 3 |
| Validate | $1.42 | $2.97 | 3 |
| Review | $0.51 | $2.01 | 3 |
| Deliver | $1.56 | $4.34 | 3 |
List-price calculation, not a run. 6 pipeline stages, 2 series: Claude Sonnet 5.5 (Agent, older builds), Claude Opus 5.5 (Agent, new build). Claude Sonnet 5.5 (Agent, older builds): highest Research $2.34 (n 3). Lowest Review $0.51 (n 3). Claude Opus 5.5 (Agent, new build): highest Research $6.28 (n 3). Lowest Plan $1.11 (n 3).
Notesn = 3 per row
List-price cost per stage, summed over the same 3 instances
Calculation, not an invoice: stage costs from the run reports (run-mode stages). Onboarding and calls outside a stage are not in these bars, so the stages sum to less than the attempt totals.
Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Repricing calculation, Anthropic list prices (Claude models)
Tables
Same 3 instances, two brains
Every case once: both right, only one right, both wrong
Only the 1 discordant decisions count: 1 vs 0. Exact McNemar p = 1: no evidence of a difference.
- Both right
- 1 33%Agree and right: says nothing about which is better.
- Only Claude Opus 5.5 right
- 1 33%Discordant: Claude Opus 5.5 right where Claude Sonnet 5.5 was wrong.
- Only Claude Sonnet 5.5 right
- 0 0%Discordant: Claude Sonnet 5.5 right where Claude Opus 5.5 was wrong.
- Both wrong
- 1 33%Agree and wrong: says nothing about which is better.
One square per paired decision (n = 3). Squares start in one grid and sort into the four outcomes; their order inside a quadrant carries no meaning.
| Pair | Both resolved | Only Opus resolved | Only Sonnet resolved | Neither resolved | Exact McNemar p |
|---|---|---|---|---|---|
| Claude Opus 5.5 vs Claude Sonnet 5.5 | 1 | 1 | 0 | 1 | 1 |
Counts of cases from the table; the exact McNemar p reads only the discordant casesn = 3 cases per pair
Claude Opus 5.5 vs Claude Sonnet 5.5: 1 vs 0 discordant cases of 3, exact McNemar p = 1.
Graded instances: resolved or not, per brain
Instance by column: each cell is k of n from the table
- Claude Opus 5.5 (Agent, new build)
- Claude Sonnet 5.5 (Agent, older builds)
Filled means yes, empty means no. Totals and other columns are in the Table view.
| Instance | Difficulty band | Claude Opus 5.5 (Agent, new build) | Claude Sonnet 5.5 (Agent, older builds) |
|---|---|---|---|
| django__django-10554 | No panel model solved it | resolved | unresolved |
| matplotlib__matplotlib-21568 | No panel model solved it | resolved | resolved |
| django__django-15022 | Under half solved it | unresolved | unresolved |
Counts from the table
3 rows by 2 columns. Claude Opus 5.5 (Agent, new build): 2 of 3 resolved.
All 8 declared instances, in run order (an escalation is graded as delivered)
| Instance | Difficulty band | Panel solved (of 11) | Opus 5.5 | Sonnet 5.5 | Opus cost (calc.) | Sonnet cost (calc.) | Opus min | Sonnet min |
|---|---|---|---|---|---|---|---|---|
| django__django-10554 | No panel model solved it | 0 | resolved | unresolved (escalated) | $9.99 | $1.52 | 25.3 min | 4.7 min |
| matplotlib__matplotlib-21568 | No panel model solved it | 0 | resolved (escalated) | resolved | $4.90 | $3.02 | 9.8 min | 9.4 min |
| django__django-15022 | Under half solved it | 4 | unresolved (escalated) | unresolved | $7.87 | $4.10 | 20.3 min | 15 min |
| django__django-11885 | Under half solved it | 3 | not started | resolved | — | $2.27 | — | 7.3 min |
| django__django-14631 | Half or more solved it | 9 | not started | empty patch (escalated) | — | $2.80 | — | 38.2 min |
| django__django-12193 | Half or more solved it | 10 | not started | resolved | — | $2.06 | — | 5.7 min |
| sympy__sympy-18189 | Every panel model solved it | 11 | not started | empty patch (escalated) | — | $0.89 | — | 1.6 min |
| django__django-11333 | Every panel model solved it | 11 | not started | resolved | — | $2.51 | — | 8 min |
Method
- A paired probe declared before any Opus run: 8 SWE-bench Verified instances from the 33 that Agent attempted with Sonnet, 2 per difficulty band (1 Sonnet miss and 1 Sonnet resolve each), picked by a fixed rule.
- Opus arm: Agent with every model call on Claude Opus 5.5 (routing off), balanced mode, real onboarding, cold start, no cost or time cap, one attempt per instance, on platform build 4f6f4027. Sonnet arm: the earlier campaign attempt on the same instance with the same flags and task spec (builds f0ac3a8a and 236c0d3f).
- Grading: Official SWE-bench harness 5.0.2 with the official instance images, amd64 emulation; the gold patch must resolve on this host first. An escalation is graded as delivered; nobody answers it.
- A usage gate starts an instance only when the subscription has 3% or more left in both its weekly and 5-hour windows. The run ledger resumes after a halt and never repeats an attempt.
- Notional cost: the platform price table at each run commit applied to the recorded tokens. The exact McNemar test reads only the discordant pairs.
Caveats
- Interim: 3 of 8 declared pairs. n = 3 supports no ranking: the 95% Wilson intervals overlap almost completely.
- Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
- 2 of the Sonnet misses in the declared set (django__django-14631, sympy__sympy-18189) were empty patches from platform holds, not wrong fixes. They are not graded for Opus yet.
- The instances were chosen to hold Sonnet misses and resolves in equal numbers, so neither rate estimates SWE-bench Verified as a whole.
- Costs are list-price calculations on subscription runs, not invoices. Grading ran under amd64 emulation; contamination is not controlled.
Sources
Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim)
8 Verified instances declared before the first run, 2 per difficulty band. One Opus attempt each on platform build 4f6f4027, paired with the earlier Sonnet attempt on the same instance; official grader. Interim: a usage gate stopped the campaign after 3 instances; the other 5 resume after the reset on 2026-10-09.
Agent on SWE-bench Verified, campaign 1 (25 instances)
Stratified sample of 25 Verified instances (seed 20261004), one attempt each, official grading harness. Fixed model claude-sonnet-5-5, platform build f0ac3a8a.
Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)
The 6 compiled-extension instances that campaign 1 could not run, plus 2 replacement candidates. One attempt each, platform build 236c0d3f.
Repricing calculation
Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.
Anthropic list prices (Claude models)
Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/swe-bench-opus-vs-sonnet.
Models and comparisons in this study
Write-ups on this study
Most AI model comparisons are ties: 44 of 1,196 rows show a clear gap
44 of 1,196 AI model comparison rows show a gap under our overlap rules. Most gaps are timing rows. No effort row separates quality.
AI coding benchmarks roundup, October 2026: sixteen studies, every number in one place
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
Claude Sonnet 5.5 vs Opus 5.5: when is Opus worth the price?
Sonnet 5.5 and Opus 5.5 tied on every quality test we ran, easy, hard and agentic. Opus cost 1.6x to 2.6x per unit of work. Where the gap comes from.
Opus 5.5 vs Sonnet 5.5 on real GitHub issues: an early look, limits first
Interim: 3 of 8 paired SWE-bench Verified issues graded. Opus 5.5 resolved 2, Sonnet 5.5 1; McNemar p = 1.0. Opus cost 2.6x, a calculation. Limits first.
Why we count every failed attempt: our rules for honest AI benchmarks
How we benchmark AI agents: declare the protocol first, keep every failure, show intervals, publish our own confounds and label interim results.
More studies
All benchmarksGPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
176 calls on 8 hard tasks at low, medium, high and default effort: Sonnet, Opus and GPT-6.1 Sol. Pass rate with 95% intervals, time, tokens, cost per pass.