• SWE-bench
  • Claude Opus
  • Claude Sonnet
  • Agent Harness
  • Interim

Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)

Does Agent resolve more SWE-bench Verified instances with Claude Opus 5.5 as its brain than with Claude Sonnet 5.5, and at what cost and time?

Published · 5 charts · Download the data or a carousel

67%

95% CI 21%–94% · n = 3

2/3 · Claude Opus 5.5 in Agent: resolved, interim (3 of 8 pairs graded)

33%

95% CI 6.2%–79% · n = 3

1/3 · Claude Sonnet 5.5 in Agent: resolved on the same 3 instances

p = 1.0 · Exact McNemar p on the graded pairs · n = 3

1 pair where only Opus resolved, 0 pairs where only Sonnet did. The test reads only these; p = 1.0 is no evidence of a difference.

The answer

Interim, not a final result: 3 of 8 declared pairs are graded. With Claude Opus 5.5 as its brain, Agent resolved 2 of 3 (95% Wilson 21–94%); with Claude Sonnet 5.5 it resolved 1 of 3 (6–79%) on the same instances. The one discordant instance (django__django-10554) went to Opus, and the exact McNemar p is 1.0, so these pairs show no difference. On the same instances Opus cost 2.6× as much, a list-price calculation ($22.76 vs $8.64), and used 1.9× the worker minutes. The arms ran on different platform builds, and the other 5 pairs remain not started in this dataset; the recorded usage-reset date was 2026-10-09.

Key numbers

3 of 8

Declared pairs graded so far

n = 8

2.6×

Opus vs Sonnet list-price cost on the same instances (calculation)

n = 3

1.9×

Opus vs Sonnet worker minutes on the same instances (ratio of totals)

n = 3

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Claude Opus 5.5 (Agent, new build)
Claude Sonnet 5.5 (Agent, older builds)

2 rows. Highest Claude Opus 5.5 (Agent, new build) 67% (95% interval 21%–94%, n 3). Lowest Claude Sonnet 5.5 (Agent, older builds) 33% (95% interval 6.2%–79%, n 3). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 3 per row

One attempt per arm per instance, official grader · 95% Wilson intervals

Interim: 3 of 8 declared pairs are graded; 5 were never started; no resumed attempts are included here. With n = 3 the intervals span most of the axis, so this chart supports no ranking. The instances were chosen to hold Sonnet misses and resolves in equal numbers, so neither rate estimates SWE-bench Verified as a whole.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

Share card (PNG)
Calculation
Claude Opus 5.5 (Agent, new build)
Claude Sonnet 5.5 (Agent, older builds)

List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 (Agent, new build) $7.59 (n 3). Lowest Claude Sonnet 5.5 (Agent, older builds) $2.88 (n 3).

Notesn = 3 per row

Mean over the same 3 instances; notional cost from the platform price table

Calculation, not an invoice: the platform price table at each run commit applied to the recorded tokens of subscription runs (Opus 5.5: $4 input, $20 output, $0.20 cache read per million tokens). Totals $22.76 vs $8.64: 2.6×, a ratio of two calculations.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Repricing calculation, Anthropic list prices (Claude models)

Share card (PNG)
Calculation
  • Claude Sonnet 5.5 (Agent, older builds)
  • Claude Opus 5.5 (Agent, new build) (square)
In chart order.
django-10554
matplotlib-21568
django-15022

Gap labels, Claude Opus 5.5 (Agent, new build) vs Claude Sonnet 5.5 (Agent, older builds): Claude Opus 5.5 (Agent, new build) is x% higher (+) or lower (−) than Claude Sonnet 5.5 (Agent, older builds), calculated from the two values shown (the change counted from Claude Sonnet 5.5 (Agent, older builds)’s value).

List-price calculation, not a run. 3 instances, 2 series: Claude Sonnet 5.5 (Agent, older builds), Claude Opus 5.5 (Agent, new build). Claude Sonnet 5.5 (Agent, older builds): highest django-15022 $4.10 (n 1). Lowest django-10554 $1.52 (n 1). Claude Opus 5.5 (Agent, new build): highest django-10554 $9.99 (n 1). Lowest matplotlib-21568 $4.90 (n 1).

Notesn = 1 per row

One attempt per arm; notional cost from the platform price table

Calculation, not an invoice. Outcomes: django-10554 Opus resolved, Sonnet unresolved; matplotlib-21568 Opus resolved, Sonnet resolved; django-15022 Opus unresolved, Sonnet unresolved.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Repricing calculation, Anthropic list prices (Claude models)

Share card (PNG)
Entrance: medians race at 870× real time
Claude Opus 5.5 (Agent, new build)
Claude Sonnet 5.5 (Agent, older builds)

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest Claude Opus 5.5 (Agent, new build) 20.3 min (range 9.8 min–25.3 min, n 3). Fastest Claude Sonnet 5.5 (Agent, older builds) 9.4 min (range 4.7 min–15 min, n 3). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Median minutes; whiskers = fastest and slowest of 3 attempts (not an interval)

Worker minutes from the run reports. Opus attempts ran one at a time on the newer build; the Sonnet attempts ran in earlier campaigns. 3 attempts per arm is too few to call a difference; a range is not a confidence interval.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

Share card (PNG)
Calculation
  • Claude Sonnet 5.5 (Agent, older builds)
  • Claude Opus 5.5 (Agent, new build) (square)
Sorted by gap, largest first.
Review
Implement
Deliver
Research
Validate
Plan

Gap labels, Claude Opus 5.5 (Agent, new build) vs Claude Sonnet 5.5 (Agent, older builds): Claude Opus 5.5 (Agent, new build) is x% higher (+) or lower (−) than Claude Sonnet 5.5 (Agent, older builds), calculated from the two values shown (the change counted from Claude Sonnet 5.5 (Agent, older builds)’s value).

List-price calculation, not a run. 6 pipeline stages, 2 series: Claude Sonnet 5.5 (Agent, older builds), Claude Opus 5.5 (Agent, new build). Claude Sonnet 5.5 (Agent, older builds): highest Research $2.34 (n 3). Lowest Review $0.51 (n 3). Claude Opus 5.5 (Agent, new build): highest Research $6.28 (n 3). Lowest Plan $1.11 (n 3).

Notesn = 3 per row

List-price cost per stage, summed over the same 3 instances

Calculation, not an invoice: stage costs from the run reports (run-mode stages). Onboarding and calls outside a stage are not in these bars, so the stages sum to less than the attempt totals.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Repricing calculation, Anthropic list prices (Claude models)

Share card (PNG)

Tables

Same 3 instances, two brains

Every case once: both right, only one right, both wrong

Only the 1 discordant decisions count: 1 vs 0. Exact McNemar p = 1: no evidence of a difference.

1Both right1Only Claude Opus 5.5 right0Only Claude Sonnet 5.5 right1Both wrong
Both right
1 33%Agree and right: says nothing about which is better.
Only Claude Opus 5.5 right
1 33%Discordant: Claude Opus 5.5 right where Claude Sonnet 5.5 was wrong.
Only Claude Sonnet 5.5 right
0 0%Discordant: Claude Sonnet 5.5 right where Claude Opus 5.5 was wrong.
Both wrong
1 33%Agree and wrong: says nothing about which is better.

One square per paired decision (n = 3). Squares start in one grid and sort into the four outcomes; their order inside a quadrant carries no meaning.

Counts of cases from the table; the exact McNemar p reads only the discordant casesn = 3 cases per pair

Claude Opus 5.5 vs Claude Sonnet 5.5: 1 vs 0 discordant cases of 3, exact McNemar p = 1.

Graded instances: resolved or not, per brain

Instance by column: each cell is k of n from the table

InstanceClaude Opus 5.5 (Agent, new build)1Claude Sonnet 5.5 (Agent, older builds)2
django__django-10554✓–
matplotlib__matplotlib-21568✓✓
django__django-15022––
  1. Claude Opus 5.5 (Agent, new build)
  2. Claude Sonnet 5.5 (Agent, older builds)

Filled means yes, empty means no. Totals and other columns are in the Table view.

Counts from the table

3 rows by 2 columns. Claude Opus 5.5 (Agent, new build): 2 of 3 resolved.

All 8 declared instances, in run order (an escalation is graded as delivered)

InstanceDifficulty bandPanel solved (of 11)Opus 5.5Sonnet 5.5Opus cost (calc.)Sonnet cost (calc.)Opus minSonnet min
django__django-10554No panel model solved it0resolvedunresolved (escalated)$9.99$1.5225.3 min4.7 min
matplotlib__matplotlib-21568No panel model solved it0resolved (escalated)resolved$4.90$3.029.8 min9.4 min
django__django-15022Under half solved it4unresolved (escalated)unresolved$7.87$4.1020.3 min15 min
django__django-11885Under half solved it3not startedresolved—$2.27—7.3 min
django__django-14631Half or more solved it9not startedempty patch (escalated)—$2.80—38.2 min
django__django-12193Half or more solved it10not startedresolved—$2.06—5.7 min
sympy__sympy-18189Every panel model solved it11not startedempty patch (escalated)—$0.89—1.6 min
django__django-11333Every panel model solved it11not startedresolved—$2.51—8 min

Method

  1. A paired probe declared before any Opus run: 8 SWE-bench Verified instances from the 33 that Agent attempted with Sonnet, 2 per difficulty band (1 Sonnet miss and 1 Sonnet resolve each), picked by a fixed rule.
  2. Opus arm: Agent with every model call on Claude Opus 5.5 (routing off), balanced mode, real onboarding, cold start, no cost or time cap, one attempt per instance, on platform build 4f6f4027. Sonnet arm: the earlier campaign attempt on the same instance with the same flags and task spec (builds f0ac3a8a and 236c0d3f).
  3. Grading: Official SWE-bench harness 5.0.2 with the official instance images, amd64 emulation; the gold patch must resolve on this host first. An escalation is graded as delivered; nobody answers it.
  4. A usage gate starts an instance only when the subscription has 3% or more left in both its weekly and 5-hour windows. The run ledger resumes after a halt and never repeats an attempt.
  5. Notional cost: the platform price table at each run commit applied to the recorded tokens. The exact McNemar test reads only the discordant pairs.

Caveats

  • Interim: 3 of 8 declared pairs. n = 3 supports no ranking: the 95% Wilson intervals overlap almost completely.
  • Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
  • 2 of the Sonnet misses in the declared set (django__django-14631, sympy__sympy-18189) were empty patches from platform holds, not wrong fixes. They are not graded for Opus yet.
  • The instances were chosen to hold Sonnet misses and resolves in equal numbers, so neither rate estimates SWE-bench Verified as a whole.
  • Costs are list-price calculations on subscription runs, not invoices. Grading ran under amd64 emulation; contamination is not controlled.

Sources

  • Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim)

    Our recorded runs ·

    8 Verified instances declared before the first run, 2 per difficulty band. One Opus attempt each on platform build 4f6f4027, paired with the earlier Sonnet attempt on the same instance; official grader. Interim: a usage gate stopped the campaign after 3 instances; the other 5 resume after the reset on 2026-10-09.

    Raw data: swebench-opus/instances.json, swebench/attempts.json

  • Agent on SWE-bench Verified, campaign 1 (25 instances)

    Our recorded runs ·

    Stratified sample of 25 Verified instances (seed 20261004), one attempt each, official grading harness. Fixed model claude-sonnet-5-5, platform build f0ac3a8a.

    Raw data: swebench/attempts.json, swebench/exclusions.json

  • Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

    Our recorded runs ·

    The 6 compiled-extension instances that campaign 1 could not run, plus 2 replacement candidates. One attempt each, platform build 236c0d3f.

    Raw data: swebench/attempts.json

  • Repricing calculation

    Calculation ·

    Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/swe-bench-opus-vs-sonnet.

Models and comparisons in this study

More studies

All benchmarks

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.