• Head to head
  • Claude Sonnet
  • Claude Opus
  • LLM pricing

Claude Sonnet 5.5 vs Opus 5.5: when is Opus worth the price?

Sonnet 5.5 and Opus 5.5 tied on every quality test we ran, easy, hard and agentic. Opus cost 1.6x to 2.6x per unit of work. Where the gap comes from.

TL;DR

  • In every study where both ran, we found no quality gap between Claude Sonnet 5.5 and Claude Opus 5.5. Both passed 24 of 24 hard tasks. On five easy tasks, Opus passed 15/15 and Sonnet 12/15, and those intervals overlap. New: in Claude Code on hidden-test repository tasks, both passed 12 of 12; as our agent's brain on 3 graded SWE-bench pairs (interim), Opus resolved 2 and Sonnet 1 (exact McNemar p = 1.0).
  • Sonnet was faster in our runs: a median 7.7 s against 9.2 s per hard call, and 2.31 s against 2.75 s per easy call. The ranges overlap, so this is not a tested difference.
  • Opus cost more for the same work (list-price calculations): 2.0x per hard pass, 1.6x per easy pass, 2.6x per coding-agent pass and 2.6x on the same SWE-bench issues. Repricing Sonnet's own agent tokens at Opus prices gives only 1.6x; when Opus really ran, it used more tokens.
  • Opus 5.5 lists at twice Sonnet 5.5's input and output price, but its cache-read price is the same. That is why long agent loops cost less than double on Opus.
  • When is Opus worth it? On our data, we cannot show a case. The answer is: when your own tasks show a pass-rate gap that pays for a 1.6x to 2.6x price.

The side-by-side rows: /compare/claude-sonnet-5-5-vs-claude-opus-5-5.

Live story · 101 sClaude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say

Claude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say

66 comparison rows from 10 studies: 0 rows favour Sonnet 5.5, 0 favour Opus 5.5, 66 are ties or unclear. Cost rows are calculations.

Transcript
  1. Comparison · 66 rows · 10 studies. Sonnet 5.5 vs Opus 5.5. A winner only where the 95% intervals or run ranges do not overlap.
  2. 66 comparison rows from 10 studies. None separates them: intervals or run ranges overlap, or none was recorded. Rows where Sonnet 5.5 is ahead: 0 (of 66). Rows where Opus 5.5 is ahead: 0 (of 66). Ties or unclear: 66 (16 ties · 50 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
  3. SWE-bench pairs, interim: pass rate 33% vs 67%, tie: 95% intervals overlap. None of the 3 rows separates them. Table: SWE-bench pairs, interim · Agent · n = 3 per side. Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
  4. Five short tasks: pass rate 80% vs 100%, tie: 95% intervals overlap. None of the 8 rows separates them. Table: Five short tasks · Claude Code · n = 15 per side. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
  5. Eight hard tasks: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 6 rows separates them. Table: Eight hard tasks · Claude Code · n = 24 per side. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  6. Coding agents, hidden tests: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Coding agents, hidden tests · Claude Code · n = 12 per side. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
  7. Effort ladder, default effort: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Effort ladder, default effort · Claude Code · n = 16 per side. Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  8. Caching sessions: pass rate not measured. None of the 4 rows separates them. Table: Caching sessions · Claude Code · n = 15, 3, 12 per side. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  9. Prompt cache break-even: after how many reuses does a cached prefix cost less?: pass rate not measured. None of the 9 rows separates them. Table: Prompt cache break-even: after how many reuses does a cached prefix cost less? · no cache · n = per side. Caveat: The source has 30 attempted Claude turns, 0 failed turns and 0 turns without token usage. Failed turns with usage remain in cost totals. Missing usage cannot be priced. No quality rate or cache-caused speed effect is claimed.
  10. How much of an AI bill is thinking? Reasoning tokens by model and effort: pass rate not measured. None of the 8 rows separates them. Table: How much of an AI bill is thinking? Reasoning tokens by model and effort · Claude Code · n = 24, 16, 15 per side. Caveat: The calls used one shared Mac and network. Other work and provider load were not controlled. Latency includes those effects.
  11. Where the seconds go: first text, output speed and prompt size for 6 LLMs: pass rate 100% vs 56%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: Where the seconds go: first text, output speed and prompt size for 6 LLMs · Claude Code · n = 4, 3, 9 per side. Caveat: First text includes CLI start-up and reasoning time. It is not the API’s time to first token; a direct API call would skip the CLI start-up (the CLI reported ready after a median 0.33 s to 0.84 s per cell).
  12. GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks: pass rate 38% vs 42%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks · Claude Code · n = per side. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.
  13. No winner where the data shows none. Every row and its reason online.

The question

Sonnet is Anthropic's mid-tier model. Opus is the larger one, and it costs more. Most teams that use Claude for coding ask the same question: is the bigger model worth the price for my work?

We have five sources that touch it:

  1. A hard head-to-head: 8 hard tasks, strict validators, 24 calls per configuration.
  2. An easy head-to-head: 5 short tasks, 15 calls per configuration.
  3. A repricing: the tokens our agent recorded on SWE-bench Verified, priced at each model's list price. A calculation, not a run.
  4. Coding agents on hidden tests: Claude Code with each model on 6 small repository tasks, 12 sessions each, graded by tests the agent could not see.
  5. A paired SWE-bench probe (interim): our agent with Opus as its brain on SWE-bench Verified instances that it had attempted with Sonnet. 3 of 8 declared pairs are graded so far.

None of them is a perfect answer. Together they give a clear picture of the price side and an honest blank on the quality side.

Quality: a tie at the ceiling

On the hard set, Sonnet 5.5, Opus 5.5 and Opus 5.5 at high effort each passed 24 of 24 strictly. A perfect 24/24 has a 95% interval of 86% to 100%, so all three sit in the same band.

On the easy set, Opus passed 15 of 15 at low, default and high effort. Sonnet passed 12 of 15 (80%, 95% interval 55% to 93%). All three Sonnet misses were on one arithmetic task, and each one ended on the right number with extra working lines. The exact-text validator rejects that by design. The intervals overlap, so the comparison page marks the row a tie.

Read that carefully. A tie at the ceiling does not mean the two models are equal. It means our tasks are not hard enough to find the difference. Opus may well win on longer, more open work. We have not measured that, and we will not claim it either way.

What we can say: on tasks of the kind we tested, paying for Opus bought no extra passes.

The two newer studies agree. In Claude Code, on 6 repository tasks with hidden tests, both models passed 12 of 12 sessions (76% to 100%): another ceiling. As our agent's brain on real SWE-bench issues, Opus resolved 2 of 3 and Sonnet 1 of 3 on the same instances. Only one pair differs, the exact McNemar p is 1.0, and the arms ran on different platform builds. See the early look.

Speed: Sonnet a little faster, not proven

Median total time per call:

  • Hard set: Sonnet 7.7 s, Opus 9.2 s, Opus high 11.0 s.
  • Easy set: Sonnet 2.31 s, Opus 2.75 s, Opus low 2.83 s, Opus high 2.71 s.

The single-call ranges overlap on both sets. On the hard set, Sonnet ran from 2.26 s to 34.79 s and Opus from 4.24 s to 27.21 s. So the medians describe these runs; they do not rank the models.

Effort did little for Opus. On the easy set, low, default and high landed within 0.12 s of each other. On the hard set, high effort added about 2 s of median time and no passes, because default effort already passed everything.

Price: where "2x" comes from and where it does not

Here are the list prices our calculations use, per million tokens:

ModelInputCache readOutput
Claude Sonnet 5.5$2$0.20$10
Claude Opus 5.5$4$0.20$20

One-hour cache writes cost twice the input price for both: $4 on Sonnet and $8 on Opus. So Opus is exactly 2x on everything except cache reads, which cost the same.

That one column decides how much more Opus costs. A short call has few cache reads, so it pays close to 2x. A long agent loop re-reads its context on every turn, so most of its tokens cost the same on both models.

Short calls: close to 2x

Calculation
Claude Fable 5.1
Claude Sonnet 5.5
Claude Opus 5.5 (high)
Claude Opus 5.5
Claude Opus 5.5 (low)
Claude Haiku 4.5
GPT-6.1 Sol (high)
GPT-6.1 Sol (medium)
GPT-6.1 Sol (low)

Every interval overlaps every other: this chart does not order these rows.

List-price calculation, not a run. 9 rows. Highest GPT-6.1 Sol (high) · Codex CLI $0.01 (range $0.0066–$0.028, n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.0036 (range $0.0034–$0.01, n 15). All run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 10–15 per row

Reported tokens × list price; the calls ran on subscriptions

Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.

Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

On the easy set, the median call cost $0.0036 on Sonnet and $0.0069 on Opus, about 1.9x (our calculation from the chart). Cache reads were about two thirds of the input tokens on these calls (1,401 of about 2,086 for both), but input is cheap; the output and the other input carry most of the price.

Per passing answer, Sonnet's three format misses count against it. It still came out lowest: $0.0062 against $0.0101 for Opus, a ratio of 1.6x.

Hard calls: about 2x

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

On the hard set, every call in both configurations passed, so the cost per pass equals the cost per call:

  • Sonnet 5.5: $0.0143
  • Opus 5.5: $0.0282, 2.0x Sonnet
  • Opus 5.5 (high): $0.0334, 2.3x Sonnet

Opus wrote slightly fewer output tokens (a median 945 against 1,050), but each one costs twice as much. High effort wrote 1,052.

Long agent loops: about 1.6x

Calculation
Largest value is 310x the smallest; Log shows the small bars.
Claude Fable 5.1
Claude Opus 5
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol
Claude Haiku 4.5
Gemini 3.x Flash
Jev 1.13 (router)

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 8 rows. Highest Claude Fable 5.1 $12.85. Lowest Jev 1.13 (router) $0.042.

Notes

Cost per resolved SWE-bench instance if 162.9M input and 1.8M output tokens had been billed at each model's list price

Calculation, not a run: tokens recorded by Agent on claude-sonnet-5-5 (33 attempts, 25 resolved) times list prices effective 2026-09-21. Another model would use a different number of tokens and resolve a different set. Jev is a routing model and cannot do this work; its bar is a price floor only.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models), Google Gemini list prices, OpenAI list prices, Jev 1.13 list price

Our agent's SWE-bench run used 162.9M input tokens, and 94.0% of them were cache reads. At list price, those exact tokens cost $87.23 on Sonnet and $143.83 on Opus: 1.6x. Per resolved instance, that is $3.49 against $5.75. Only the Sonnet figure matches the model that produced the tokens; Opus would have taken a different path.

You can check the 1.6x by hand. At Sonnet prices, cache reads were $30.62 of the $87.23. Opus doubles everything else and keeps cache reads the same: 2 × ($87.23 − $30.62) + $30.62 = $143.84. The one-cent difference is rounding.

Real agent runs: about 2.6x

The repricing keeps Sonnet's tokens and changes only the price. When Opus really does the work, it takes its own path.

Calculation
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Opus 5.5 · Claude Code $0.22 (n 12). Lowest Claude Sonnet 5.5 · Claude Code $0.085 (n 12).

Notesn = 12 per row

Reported tokens of all 12 sessions × list price, divided by the passes

Calculation, not a bill: both CLIs ran on flat subscriptions. Claude cache writes are priced at 2× input, as in the other studies (Claude Code’s own estimate gives the same totals); Codex cached input at its cache-read price. Codex input includes its own system prompt and, here, the tester’s AGENTS.md.

Sources: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

  • Coding agents on hidden tests (Claude Code): $0.085 per pass on Sonnet and $0.22 on Opus, 2.6x. Opus wrote 1.9x the output tokens by median (1.7x by mean) and read more context per session.
  • SWE-bench Verified, interim (our agent): on the same 3 issues, $8.64 on Sonnet and $22.76 on Opus, 2.6x. On the one issue where only Opus succeeded, it cost 6.6x as much (our calculation).

Both values are above 2x. With the cache-read price the same on both models, a ratio above 2x is possible only if Opus used more tokens. So the 1.6x of the repricing is not a forecast: it holds the tokens fixed, and in both real runs Opus used more. The SWE-bench probe has only 3 pairs and a build difference, so treat its 2.6x as a direction.

Does routing to Opus pay?

A common plan is to send only the "strong" steps to Opus and the rest to cheaper models. We priced that as a calculation on 2,362 recorded calls from 50 benchmark runs.

Calculation
all Fable 5.1
all Opus 5.5
policy (Opus strong, Haiku ancillary)
all Sonnet 5.5
split (Sonnet main line, Haiku ancillary)
all Haiku 4.5

Hover or focus a bar for its ratio to all Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 6 rows. Highest all Fable 5.1 $370. Lowest all Haiku 4.5 $54.27.

Notes

50 benchmark runs, 2,362 model calls, repriced

Calculation, not a run: every recorded call ran on Sonnet 5.5 with routing off. Same tokens on every model; a different model or mix would take a different path.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Anthropic list prices (Claude models)

  • All Sonnet 5.5: $108.54
  • Policy (Opus on strong stages, Haiku on side jobs): $161.62, 1.49x all-Sonnet
  • All Opus 5.5: $170.92

The strong stages hold most of the spend, so the policy costs nearly as much as all-Opus. Routing to Opus is not a small premium on a few calls. It is most of the Opus premium. For the full breakdown, see What if every call ran on Opus?.

So when is Opus worth it?

Use a simple test. Opus pays off when it raises your pass rate enough to cover its price. If Opus costs r times Sonnet per call, Opus is cheaper per success only when its pass rate is more than r times Sonnet's.

With our measured ratios:

  • On short calls (about 1.9x), Opus must pass about 1.9 times as often. If Sonnet passes 50% of a task type, Opus must pass about 95%.
  • In long agent loops priced on Sonnet's own tokens (about 1.6x), the bar is about 1.6 times. If Sonnet already passes 80% or more, no pass rate can repay a 1.6x price on cost alone.
  • In the agent runs where Opus really did the work (about 2.6x), the bar is 2.6 times. If Sonnet already passes 39% or more, no pass rate can repay that price on cost alone (our calculation).

That leaves three good reasons to pay for Opus:

  1. Your own tasks show a gap. Measure it. Ours did not.
  2. A failure costs much more than a call. If a wrong answer costs an hour of review, a few extra passes can pay for a lot of tokens.
  3. Your work is long and cache-heavy, so the price gap can fall toward 1.6x. Measure it: in our real agent runs, Opus used more tokens and cost 2.6x.

And one bad reason: "bigger is safer". On every task we ran, Sonnet passed at least as often within the intervals, ran a little faster and cost less.

How we measured

  • Hard and easy sets: Claude Code, one turn, tools off, fresh folder, deterministic validators. Hard: 3 repetitions × 8 tasks. Easy: 3 repetitions × 5 tasks.
  • Costs: reported tokens × list price per call, cache reads and writes priced separately. Calculations; the calls ran on a subscription.
  • Repricing: the agent's recorded SWE-bench tokens at each list price. A calculation, not a run.
  • Coding agents: Claude Code 2.1.286 headless, default effort, normal file and shell tools, 6 tasks × 2 repetitions per model, hidden tests graded after each session.
  • SWE-bench probe: 8 instances declared before any Opus run, one attempt per arm, official harness; interim at 3 graded pairs. Notional cost from the platform price table.
  • Ratios in this post are our arithmetic on the published values.

Caveats

  • Ceiling effects. Both models scored 24/24 on the hard set, so it cannot separate them.
  • Short tasks. The hard and easy sets are single-turn; the coding-agent tasks are small repositories. Long, multi-step work may show a quality gap that these tasks do not. The SWE-bench probe has 3 graded pairs so far.
  • Small samples. 15 to 24 calls per configuration. Time ranges overlap.
  • Repricing is not a run. Opus would make different calls and use different tokens.
  • List prices change. The cache-read price matters most for agent loops; check it before you decide.

Pay for the model only where it helps

Agent records the model, the tokens and the result of every step, so you can see where a bigger model changes the outcome. Try Agent and compare on your own work.

The data behind this post

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.