• SWE-bench
  • Head to head
  • Claude Opus
  • Claude Sonnet

Opus 5.5 vs Sonnet 5.5 on real GitHub issues: an early look, limits first

Interim: 3 of 8 paired SWE-bench Verified issues graded. Opus 5.5 resolved 2, Sonnet 5.5 1; McNemar p = 1.0. Opus cost 2.6x, a calculation. Limits first.

Read this first. This is an interim result, and it is small:

  • 3 of 8 declared pairs are graded. The other 5 have not started: a usage gate stopped the run. They resume after the usage reset on 2026-10-09.
  • n = 3 supports no ranking. The two 95% intervals overlap almost completely.
  • The two arms ran on different platform builds. A gap can come from the platform, not the model.
  • We chose the instances to hold Sonnet misses and resolves in equal numbers. Neither rate estimates SWE-bench Verified as a whole.
  • Costs are list-price calculations, not invoices.

TL;DR

  • With Claude Opus 5.5 as its brain, our agent resolved 2 of 3 real GitHub issues (95% interval 21% to 94%). With Claude Sonnet 5.5, it resolved 1 of 3 of the same issues (6% to 79%).
  • One pair differs: django__django-10554, which Opus resolved and Sonnet did not. The exact McNemar test gives p = 1.0. These pairs show no difference.
  • Even all 8 pairs cannot prove one (our calculation). At most 3 pairs can break the same way. A 3-to-0 split gives an exact p of 0.25. This probe is for cost and behaviour; for quality it can only show a direction.
  • Opus cost 2.6x as much on the same issues: $22.76 against $8.64 (a calculation). It used 1.9x the worker minutes: 55.3 against 29.1.
  • It matches our other data. In every study where both models ran, no quality row separates them. Opus cost 1.6x to 2.6x as much per pass or per attempt (calculations).

Every pair, cost and declared instance: /benchmarks/swe-bench-opus-vs-sonnet.

Live story · 101 sClaude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say

Claude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say

66 comparison rows from 10 studies: 0 rows favour Sonnet 5.5, 0 favour Opus 5.5, 66 are ties or unclear. Cost rows are calculations.

Transcript
  1. Comparison · 66 rows · 10 studies. Sonnet 5.5 vs Opus 5.5. A winner only where the 95% intervals or run ranges do not overlap.
  2. 66 comparison rows from 10 studies. None separates them: intervals or run ranges overlap, or none was recorded. Rows where Sonnet 5.5 is ahead: 0 (of 66). Rows where Opus 5.5 is ahead: 0 (of 66). Ties or unclear: 66 (16 ties · 50 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
  3. SWE-bench pairs, interim: pass rate 33% vs 67%, tie: 95% intervals overlap. None of the 3 rows separates them. Table: SWE-bench pairs, interim · Agent · n = 3 per side. Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
  4. Five short tasks: pass rate 80% vs 100%, tie: 95% intervals overlap. None of the 8 rows separates them. Table: Five short tasks · Claude Code · n = 15 per side. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
  5. Eight hard tasks: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 6 rows separates them. Table: Eight hard tasks · Claude Code · n = 24 per side. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  6. Coding agents, hidden tests: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Coding agents, hidden tests · Claude Code · n = 12 per side. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
  7. Effort ladder, default effort: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Effort ladder, default effort · Claude Code · n = 16 per side. Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  8. Caching sessions: pass rate not measured. None of the 4 rows separates them. Table: Caching sessions · Claude Code · n = 15, 3, 12 per side. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  9. Prompt cache break-even: after how many reuses does a cached prefix cost less?: pass rate not measured. None of the 9 rows separates them. Table: Prompt cache break-even: after how many reuses does a cached prefix cost less? · no cache · n = per side. Caveat: The source has 30 attempted Claude turns, 0 failed turns and 0 turns without token usage. Failed turns with usage remain in cost totals. Missing usage cannot be priced. No quality rate or cache-caused speed effect is claimed.
  10. How much of an AI bill is thinking? Reasoning tokens by model and effort: pass rate not measured. None of the 8 rows separates them. Table: How much of an AI bill is thinking? Reasoning tokens by model and effort · Claude Code · n = 24, 16, 15 per side. Caveat: The calls used one shared Mac and network. Other work and provider load were not controlled. Latency includes those effects.
  11. Where the seconds go: first text, output speed and prompt size for 6 LLMs: pass rate 100% vs 56%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: Where the seconds go: first text, output speed and prompt size for 6 LLMs · Claude Code · n = 4, 3, 9 per side. Caveat: First text includes CLI start-up and reasoning time. It is not the API’s time to first token; a direct API call would skip the CLI start-up (the CLI reported ready after a median 0.33 s to 0.84 s per cell).
  12. GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks: pass rate 38% vs 42%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks · Claude Code · n = per side. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.
  13. No winner where the data shows none. Every row and its reason online.

The question, and why we asked it this way

Our SWE-bench Verified run used Agent, our full coding pipeline, with Claude Sonnet 5.5 as its brain. It resolved 25 of 33 instances. The obvious next question: would Opus 5.5 do better on the same real issues, and at what cost?

A second 33-instance run on Opus costs a lot of subscription time. So we designed a smaller, paired probe:

  • 8 instances from the 33 that Sonnet attempted.
  • 2 per difficulty band, with 1 Sonnet miss and 1 Sonnet resolve in each band.
  • Picked by a fixed rule, declared before any Opus run.

Pairing is the key. Both brains face the same issue, so the comparison does not depend on how hard the issues are. The test reads only the pairs where the two results differ.

Why we publish an interim result

A usage gate starts an instance only when the subscription has 3% or more left in both of its usage windows. It stopped the campaign after 3 instances.

We could wait for all 8. We publish now, with the limits up front, because a hidden interim result is how cherry-picking starts. If a team publishes only when the result looks good, the readers never see the other runs. So the rules for this study are:

  • The title, the tags and the answer say "interim".
  • Only graded pairs count. The study text follows the graded count.
  • The 5 instances that have not started stay in the declared table, marked "not started". Every started attempt counts.

The three graded pairs

Claude Opus 5.5 (Agent, new build)
Claude Sonnet 5.5 (Agent, older builds)

2 rows. Highest Claude Opus 5.5 (Agent, new build) 67% (95% interval 21%–94%, n 3). Lowest Claude Sonnet 5.5 (Agent, older builds) 33% (95% interval 6.2%–79%, n 3). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 3 per row

One attempt per arm per instance, official grader · 95% Wilson intervals

Interim: 3 of 8 declared pairs are graded; 5 were never started; no resumed attempts are included here. With n = 3 the intervals span most of the axis, so this chart supports no ranking. The instances were chosen to hold Sonnet misses and resolves in equal numbers, so neither rate estimates SWE-bench Verified as a whole.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

  • django__django-10554. No public panel model of 11 solved it. Opus resolved it; Sonnet did not (it escalated). Opus cost $9.99 and took 25.3 min; Sonnet cost $1.52 and took 4.7 min.
  • matplotlib__matplotlib-21568. No public panel model solved this one either. Both brains resolved it. The Opus attempt escalated; an escalation is graded on what it delivered, and nobody answers it. Opus cost $4.90 in 9.8 min; Sonnet $3.02 in 9.4 min.
  • django__django-15022. 4 of the 11 panel models solved it. Neither brain resolved it. Opus cost $7.87 in 20.3 min; Sonnet $4.10 in 15.0 min.

Opus resolved 2 of 3 (67%) and Sonnet 1 of 3 (33%). The intervals, 21% to 94% and 6% to 79%, cover most of the axis. The full grid is on the study page: same 3 instances, two brains.

What the paired test says

The exact McNemar test ignores the pairs where both brains agree: one where both resolved, one where neither did. It reads only the pairs that differ. Here that is 1 pair for Opus and 0 for Sonnet. One pair is not evidence, and the test gives p = 1.0.

Now look ahead to all 8 pairs (our calculation). Sonnet missed 4 of the 8 declared instances. Opus can break a pair its way only on those, and it already missed one. So at most 3 pairs can favour Opus. The same holds for Sonnet: at most 3. The most one-sided result, 3 to 0, gives an exact two-sided p of 0.25. So even a clean sweep would not reach the usual 0.05 line.

We knew this when we declared the probe. Its job is to measure what Opus costs and how it behaves on real issues, on a budget. A quality claim needs more pairs.

Cost: the clearer signal

Calculation
  • Claude Sonnet 5.5 (Agent, older builds)
  • Claude Opus 5.5 (Agent, new build) (square)
In chart order.
django-10554
matplotlib-21568
django-15022

Gap labels, Claude Opus 5.5 (Agent, new build) vs Claude Sonnet 5.5 (Agent, older builds): Claude Opus 5.5 (Agent, new build) is x% higher (+) or lower (−) than Claude Sonnet 5.5 (Agent, older builds), calculated from the two values shown (the change counted from Claude Sonnet 5.5 (Agent, older builds)’s value).

List-price calculation, not a run. 3 instances, 2 series: Claude Sonnet 5.5 (Agent, older builds), Claude Opus 5.5 (Agent, new build). Claude Sonnet 5.5 (Agent, older builds): highest django-15022 $4.10 (n 1). Lowest django-10554 $1.52 (n 1). Claude Opus 5.5 (Agent, new build): highest django-10554 $9.99 (n 1). Lowest matplotlib-21568 $4.90 (n 1).

Notesn = 1 per row

One attempt per arm; notional cost from the platform price table

Calculation, not an invoice. Outcomes: django-10554 Opus resolved, Sonnet unresolved; matplotlib-21568 Opus resolved, Sonnet resolved; django-15022 Opus unresolved, Sonnet unresolved.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Repricing calculation, Anthropic list prices (Claude models)

On the same 3 issues, the list-price cost was $22.76 for Opus and $8.64 for Sonnet: 2.6x. Per attempt, that is a mean of $7.59 against $2.88. These are calculations: the platform price table at each run's commit, applied to the recorded tokens of subscription runs.

Per issue, the gap moved a lot (our calculation):

  • matplotlib-21568, both resolved: 1.6x.
  • django-15022, neither resolved: 1.9x.
  • django-10554, only Opus resolved: 6.6x. Sonnet escalated after 4.7 minutes. Opus kept going for 25.3 minutes and found the fix.

Opus 5.5 lists at $4 input, $20 output and $0.20 cache read per million tokens: twice Sonnet's input and output price, and the same cache-read price. A cost ratio above 2x means the Opus attempts also used more tokens. Part of that can be the newer build, so we do not credit it all to the model.

Calculation
Claude Opus 5.5 (Agent, new build)
Claude Sonnet 5.5 (Agent, older builds)

List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 (Agent, new build) $7.59 (n 3). Lowest Claude Sonnet 5.5 (Agent, older builds) $2.88 (n 3).

Notesn = 3 per row

Mean over the same 3 instances; notional cost from the platform price table

Calculation, not an invoice: the platform price table at each run commit applied to the recorded tokens of subscription runs (Opus 5.5: $4 input, $20 output, $0.20 cache read per million tokens). Totals $22.76 vs $8.64: 2.6×, a ratio of two calculations.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Repricing calculation, Anthropic list prices (Claude models)

Where the money went

Calculation
  • Claude Sonnet 5.5 (Agent, older builds)
  • Claude Opus 5.5 (Agent, new build) (square)
Sorted by gap, largest first.
Review
Implement
Deliver
Research
Validate
Plan

Gap labels, Claude Opus 5.5 (Agent, new build) vs Claude Sonnet 5.5 (Agent, older builds): Claude Opus 5.5 (Agent, new build) is x% higher (+) or lower (−) than Claude Sonnet 5.5 (Agent, older builds), calculated from the two values shown (the change counted from Claude Sonnet 5.5 (Agent, older builds)’s value).

List-price calculation, not a run. 6 pipeline stages, 2 series: Claude Sonnet 5.5 (Agent, older builds), Claude Opus 5.5 (Agent, new build). Claude Sonnet 5.5 (Agent, older builds): highest Research $2.34 (n 3). Lowest Review $0.51 (n 3). Claude Opus 5.5 (Agent, new build): highest Research $6.28 (n 3). Lowest Plan $1.11 (n 3).

Notesn = 3 per row

List-price cost per stage, summed over the same 3 instances

Calculation, not an invoice: stage costs from the run reports (run-mode stages). Onboarding and calls outside a stage are not in these bars, so the stages sum to less than the attempt totals.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Repricing calculation, Anthropic list prices (Claude models)

Summed over the 3 issues, Opus spent more than Sonnet in every pipeline stage but one:

  • Research: $6.28 against $2.34.
  • Implement: $4.94 against $1.32.
  • Deliver: $4.34 against $1.56.
  • Validate: $2.97 against $1.42.
  • Review: $2.01 against $0.51.
  • Plan: $1.11 against $1.16. The one stage where Opus spent less.

In dollars, the largest extra spend was in research (+$3.94) and implementation (+$3.62). In percent, review grew the most (+294%), from a small base. These differences are our arithmetic on the chart values. These bars leave out onboarding and calls outside a stage, so they sum to less than the attempt totals.

Time

Entrance: medians race at 870× real time
Claude Opus 5.5 (Agent, new build)
Claude Sonnet 5.5 (Agent, older builds)

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest Claude Opus 5.5 (Agent, new build) 20.3 min (range 9.8 min–25.3 min, n 3). Fastest Claude Sonnet 5.5 (Agent, older builds) 9.4 min (range 4.7 min–15 min, n 3). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Median minutes; whiskers = fastest and slowest of 3 attempts (not an interval)

Worker minutes from the run reports. Opus attempts ran one at a time on the newer build; the Sonnet attempts ran in earlier campaigns. 3 attempts per arm is too few to call a difference; a range is not a confidence interval.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

Median worker time per attempt: 20.3 min for Opus, 9.4 min for Sonnet. The fastest-to-slowest ranges overlap (9.8 to 25.3 min, and 4.7 to 15.0 min). Three attempts per arm is too few to call a difference. The Opus attempts ran one at a time on the newer build; the Sonnet attempts ran in earlier campaigns.

How this fits our other Sonnet vs Opus data

Other studies ran both models on the same work:

So far, the pattern is steady: no quality gap that our samples can show, and a cost gap of 1.6x to 2.6x (calculations; 1.6x on five easy tasks, 2.6x here and on the coding-agent tasks). More on the trade-off: Claude Sonnet vs Opus: when is Opus worth the price?

What happens next

  • The 5 remaining pairs resume after the usage reset on 2026-10-09: django__django-11885, django__django-14631, django__django-12193, sympy__sympy-18189 and django__django-11333. The study page updates when they are graded.
  • Two of them were platform holds, not wrong fixes. Sonnet's attempts on django__django-14631 and sympy__sympy-18189 were empty patches from platform holds. If Opus resolves them on the newer build, the build can be the cause, not the model.
  • What would make this a quality test: more pairs, and both arms on the same build. A same-build Sonnet arm would remove the build confound.

How we measured

  • Paired probe, declared before any Opus run: 8 SWE-bench Verified instances from the 33 that Agent attempted with Sonnet, 2 per difficulty band (1 Sonnet miss and 1 Sonnet resolve each), picked by a fixed rule.
  • Opus arm: Agent with every model call on Claude Opus 5.5 (routing off), balanced mode, real onboarding, cold start, no cost or time cap, one attempt per instance.
  • Sonnet arm: the earlier campaign attempt on the same instance, with the same flags and task spec.
  • Grading: the official SWE-bench harness 5.0.2 with the official instance images, under amd64 emulation. The gold patch must resolve on this host first. An escalation is graded on what it delivered.
  • Usage gate: an instance starts only with 3% or more left in both the weekly and the 5-hour window. The run ledger resumes after a halt and never repeats an attempt.
  • Statistics: 95% Wilson intervals; the exact McNemar test on the discordant pairs. Notional cost: the platform price table at each run's commit × the recorded tokens.

Caveats

  • Interim: 3 of 8 declared pairs. n = 3 supports no ranking.
  • Different builds: the Opus arm ran on a newer platform build than the two builds behind the Sonnet attempts. Platform changes can move results by themselves.
  • Chosen instances: equal numbers of Sonnet misses and resolves, so neither rate estimates SWE-bench Verified as a whole.
  • Costs are calculations on subscription runs, not invoices.
  • Grading ran under amd64 emulation; contamination is not controlled.

Pay for the bigger model only where it helps

Agent records which model did each step, what it cost and whether the result passed validation. Try Agent and see where a bigger model changes the outcome on your own issues.

The data behind this post

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.