35 measured metrics · 31 calculated · 10 studies

Claude Sonnet 5.5vsClaude Opus 5.5

No row separates them: 16 ties, 50 unclear.

The verdict

Claude Sonnet 5.5 and Claude Opus 5.5 share 35 measured metrics and 31 list-price calculations from 10 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 16 ties and 50 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Output tokens per call (Output tokens), 1.7x (Claude Sonnet 5.5 larger).

Watch it build

A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.

Live story · 101 sClaude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say

Claude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say

66 comparison rows from 10 studies: 0 rows favour Sonnet 5.5, 0 favour Opus 5.5, 66 are ties or unclear. Cost rows are calculations.

Transcript
  1. Comparison · 66 rows · 10 studies. Sonnet 5.5 vs Opus 5.5. A winner only where the 95% intervals or run ranges do not overlap.
  2. 66 comparison rows from 10 studies. None separates them: intervals or run ranges overlap, or none was recorded. Rows where Sonnet 5.5 is ahead: 0 (of 66). Rows where Opus 5.5 is ahead: 0 (of 66). Ties or unclear: 66 (16 ties · 50 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
  3. SWE-bench pairs, interim: pass rate 33% vs 67%, tie: 95% intervals overlap. None of the 3 rows separates them. Table: SWE-bench pairs, interim · Agent · n = 3 per side. Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
  4. Five short tasks: pass rate 80% vs 100%, tie: 95% intervals overlap. None of the 8 rows separates them. Table: Five short tasks · Claude Code · n = 15 per side. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
  5. Eight hard tasks: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 6 rows separates them. Table: Eight hard tasks · Claude Code · n = 24 per side. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  6. Coding agents, hidden tests: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Coding agents, hidden tests · Claude Code · n = 12 per side. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
  7. Effort ladder, default effort: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Effort ladder, default effort · Claude Code · n = 16 per side. Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  8. Caching sessions: pass rate not measured. None of the 4 rows separates them. Table: Caching sessions · Claude Code · n = 15, 3, 12 per side. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  9. Prompt cache break-even: after how many reuses does a cached prefix cost less?: pass rate not measured. None of the 9 rows separates them. Table: Prompt cache break-even: after how many reuses does a cached prefix cost less? · no cache · n = per side. Caveat: The source has 30 attempted Claude turns, 0 failed turns and 0 turns without token usage. Failed turns with usage remain in cost totals. Missing usage cannot be priced. No quality rate or cache-caused speed effect is claimed.
  10. How much of an AI bill is thinking? Reasoning tokens by model and effort: pass rate not measured. None of the 8 rows separates them. Table: How much of an AI bill is thinking? Reasoning tokens by model and effort · Claude Code · n = 24, 16, 15 per side. Caveat: The calls used one shared Mac and network. Other work and provider load were not controlled. Latency includes those effects.
  11. Where the seconds go: first text, output speed and prompt size for 6 LLMs: pass rate 100% vs 56%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: Where the seconds go: first text, output speed and prompt size for 6 LLMs · Claude Code · n = 4, 3, 9 per side. Caveat: First text includes CLI start-up and reasoning time. It is not the API’s time to first token; a direct API call would skip the CLI start-up (the CLI reported ready after a median 0.33 s to 0.84 s per cell).
  12. GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks: pass rate 38% vs 42%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks · Claude Code · n = per side. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.
  13. No winner where the data shows none. Every row and its reason online.

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Claude Sonnet 5.5
  • Claude Opus 5.5
  • 95% interval
  • fastest–slowest run (not an interval)
  • where the two overlap
  • hollow: list-price calculation

These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.

Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)

1 tie · 2 unclear
  • Resolved on the same 3 SWE-bench Verified instances (interim): Claude Sonnet 5.5 33% (1/3) (n 3, 95% interval 6.2%–79%); Claude Opus 5.5 67% (2/3) (n 3, 95% interval 21%–94%). Tie.

    Calculation: at these rates, about 31 runs per side would separate them.

  • List-price cost per attempt (calculation), calculation: Claude Sonnet 5.5 $2.88 (n 3); Claude Opus 5.5 $7.59 (n 3). Unclear.
  • Worker time per attempt: Claude Sonnet 5.5 9.4 min (n 3, run range 4.7 min–15 min); Claude Opus 5.5 20.3 min (n 3, run range 9.8 min–25.3 min). Unclear.

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

2 ties · 6 unclear
  • Pass rate on five validated tasks: Claude Sonnet 5.5 80% (12/15) (n 15, 95% interval 55%–93%); Claude Opus 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.

    Calculation: at these rates, about 33 runs per side would separate them.

  • Total time per call: Claude Sonnet 5.5 2.31 s (n 15, run range 2.2 s–7.7 s); Claude Opus 5.5 2.75 s (n 15, run range 2.5 s–8.9 s). Unclear.
  • Time to first useful output: Claude Sonnet 5.5 1.56 s (n 15, run range 1 s–6.4 s); Claude Opus 5.5 1.92 s (n 15, run range 1.6 s–7.2 s). Unclear.
  • Input tokens per call: what the CLI sends (Cache read): Claude Sonnet 5.5 1,401 (n 15); Claude Opus 5.5 1,401 (n 15). Tie.
  • Input tokens per call: what the CLI sends (Other input): Claude Sonnet 5.5 685 (n 15); Claude Opus 5.5 680 (n 15). Unclear.
  • Output tokens per call (Output tokens): Claude Sonnet 5.5 107 (n 15); Claude Opus 5.5 64 (n 15). Unclear.
  • List-price cost per call (calculation), calculation: Claude Sonnet 5.5 $0.0036 (n 15, run range $0.0034–$0.01); Claude Opus 5.5 $0.0069 (n 15, run range $0.0059–$0.022). Unclear.
  • List-price cost per passing answer (calculation), calculation: Claude Sonnet 5.5 $0.0062 (n 15); Claude Opus 5.5 $0.010 (n 15). Unclear.

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

2 ties · 4 unclear
  • Pass rate on eight hard tasks (Strict pass): Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%); Claude Opus 5.5 100% (24/24) (n 24, 95% interval 86%–100%). Tie.
  • Pass rate on eight hard tasks (Lenient (format misses counted)): Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%); Claude Opus 5.5 100% (24/24) (n 24, 95% interval 86%–100%). Tie.
  • Total time per call on hard tasks (separate batches): Claude Sonnet 5.5 7.75 s (n 24, run range 2.3 s–34.8 s); Claude Opus 5.5 9.18 s (n 24, run range 4.2 s–27.2 s). Unclear.
  • Time to first useful output on hard tasks: Claude Sonnet 5.5 5.95 s (n 24, run range 0.9 s–30.6 s); Claude Opus 5.5 6.78 s (n 24, run range 2.4 s–21.8 s). Unclear.
  • Output tokens per call on hard tasks (Output tokens): Claude Sonnet 5.5 1,050 (n 24); Claude Opus 5.5 945 (n 24). Unclear.
  • List-price cost per strict pass on hard tasks (calculation), calculation: Claude Sonnet 5.5 $0.014 (n 24); Claude Opus 5.5 $0.028 (n 24). Unclear.

Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks

2 ties · 2 unclear
  • Coding sessions that passed every hidden check: Claude Sonnet 5.5 100% (12/12) (n 12, 95% interval 76%–100%); Claude Opus 5.5 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
  • Time per coding session: Claude Sonnet 5.5 23.1 s (n 12, run range 18.7 s–44.5 s); Claude Opus 5.5 56.9 s (n 12, run range 29.8 s–186 s). Unclear.
  • Tool calls per coding session: Claude Sonnet 5.5 7.5 (n 12, run range 3–14); Claude Opus 5.5 7.5 (n 12, run range 5–14). Tie.
  • List-price cost per passing coding session (calculation), calculation: Claude Sonnet 5.5 $0.085 (n 12); Claude Opus 5.5 $0.22 (n 12). Unclear.

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks

1 tie · 3 unclear
  • Strict pass rate by effort on eight hard tasks: Claude Sonnet 5.5 100% (16/16) (n 16, 95% interval 81%–100%); Claude Opus 5.5 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Total time per call by effort on hard tasks: Claude Sonnet 5.5 7.97 s (n 16, run range 2.3 s–21.6 s); Claude Opus 5.5 9.18 s (n 16, run range 4.2 s–27.2 s). Unclear.
  • Output tokens per call by effort on hard tasks (Output tokens): Claude Sonnet 5.5 1,054 (n 16); Claude Opus 5.5 945 (n 16). Unclear.
  • List-price cost per strict pass by effort (calculation), calculation: Claude Sonnet 5.5 $0.014 (n 16); Claude Opus 5.5 $0.029 (n 16). Unclear.

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

4 unclear
  • List-price cost of 5-question sessions with and without the cache (calculation) (With the cache, as recorded), calculation: Claude Sonnet 5.5 $0.14 (n 15); Claude Opus 5.5 $0.26 (n 15). Unclear.
  • List-price cost of 5-question sessions with and without the cache (calculation) (Without a cache: every input token at the input price), calculation: Claude Sonnet 5.5 $0.27 (n 15); Claude Opus 5.5 $0.54 (n 15). Unclear.
  • Time per turn: first turn vs later turns in a cached session (Turn 1 (writes the ledger to the cache)): Claude Sonnet 5.5 1.64 s (n 3, run range 1.6 s–1.8 s); Claude Opus 5.5 1.90 s (n 3, run range 1.8 s–4.4 s). Unclear.
  • Time per turn: first turn vs later turns in a cached session (Turns 2-5 (read the ledger from the cache)): Claude Sonnet 5.5 1.61 s (n 12, run range 1.4 s–5.6 s); Claude Opus 5.5 2.40 s (n 12, run range 1.6 s–12.7 s). Unclear.

Prompt cache break-even: after how many reuses does a cached prefix cost less?

9 unclear
  • Cost of a reused prefix with and without the cache, by session length (calculation): 1 turn, calculation: Claude Sonnet 5.5 $15.66; Claude Opus 5.5 $31.31. Unclear.
  • Cost of a reused prefix with and without the cache, by session length (calculation): 2 turns, calculation: Claude Sonnet 5.5 $31.32; Claude Opus 5.5 $62.62. Unclear.
  • Cost of a reused prefix with and without the cache, by session length (calculation): 3 turns, calculation: Claude Sonnet 5.5 $46.99; Claude Opus 5.5 $93.94. Unclear.
  • Cost of a reused prefix with and without the cache, by session length (calculation): 5 turns, calculation: Claude Sonnet 5.5 $78.31; Claude Opus 5.5 $156.56. Unclear.
  • Cost of a reused prefix with and without the cache, by session length (calculation): 10 turns, calculation: Claude Sonnet 5.5 $156.62; Claude Opus 5.5 $313.12. Unclear.
  • Cost of a reused prefix with and without the cache, by session length (calculation): 20 turns, calculation: Claude Sonnet 5.5 $313.24; Claude Opus 5.5 $626.24. Unclear.
  • One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (No cache (the same either way)), calculation: Claude Sonnet 5.5 $156.62; Claude Opus 5.5 $313.12. Unclear.
  • One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (One 10-turn session, 1-hour cache), calculation: Claude Sonnet 5.5 $45.42; Claude Opus 5.5 $76.71. Unclear.
  • One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix)), calculation: Claude Sonnet 5.5 $313.24; Claude Opus 5.5 $626.24. Unclear.

How much of an AI bill is thinking? Reasoning tokens by model and effort

1 tie · 7 unclear
  • Reasoning share of output tokens per call on hard tasks (calculation), calculation: Claude Sonnet 5.5 54.5% (n 24, run range 0%–96%); Claude Opus 5.5 54.8% (n 24, run range 30%–96%). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Claude Sonnet 5.5 $0.0067 (n 24); Claude Opus 5.5 $0.013 (n 24). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Claude Sonnet 5.5 $0.0037 (n 24); Claude Opus 5.5 $0.0080 (n 24). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Claude Sonnet 5.5 $0.0040 (n 24); Claude Opus 5.5 $0.0077 (n 24). Unclear.
  • Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass), calculation: Claude Sonnet 5.5 $0.0063 (n 16); Claude Opus 5.5 $0.013 (n 16). Unclear.
  • Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass), calculation: Claude Sonnet 5.5 $0.014 (n 16); Claude Opus 5.5 $0.029 (n 16). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Claude Sonnet 5.5 54.5% (n 24, run range 0%–96%); Claude Opus 5.5 54.8% (n 24, run range 30%–96%). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Claude Sonnet 5.5 0% (n 15, run range 0%–73%); Claude Opus 5.5 0% (n 15, run range 0%–93%). Tie.

Where the seconds go: first text, output speed and prompt size for 6 LLMs

1 tie · 9 unclear
  • Time to first text: a 250-line answer, six models: Claude Sonnet 5.5 1.96 s (n 4, run range 0.9 s–4.1 s); Claude Opus 5.5 1.97 s (n 4, run range 1.7 s–2.4 s). Unclear.
  • Output speed after the first text: visible tokens per second (calculation), calculation: Claude Sonnet 5.5 232 (n 4, run range 230–233); Claude Opus 5.5 156 (n 4, run range 155–156). Unclear.
  • Output speed in characters per second after the first text (calculation), calculation: Claude Sonnet 5.5 517 (n 4, run range 513–519); Claude Opus 5.5 347 (n 4, run range 345–349). Unclear.
  • Time to first text as the prompt grows: 1k, calculation: Claude Sonnet 5.5 1.45 s (n 3, run range 1.2 s–1.7 s); Claude Opus 5.5 1.51 s (n 3, run range 1.5 s–2 s). Unclear.
  • Time to first text as the prompt grows: 16k, calculation: Claude Sonnet 5.5 1.78 s (n 3, run range 1.6 s–2.1 s); Claude Opus 5.5 1.74 s (n 3, run range 1.7 s–3 s). Unclear.
  • Time to first text as the prompt grows: 64k, calculation: Claude Sonnet 5.5 3.07 s (n 3, run range 1.4 s–3.6 s); Claude Opus 5.5 1.79 s (n 3, run range 1.7 s–3.7 s). Unclear.
  • Total time per call by prompt size (1k prompt): Claude Sonnet 5.5 1.78 s (n 3, run range 1.6 s–2.1 s); Claude Opus 5.5 1.83 s (n 3, run range 1.8 s–2.4 s). Unclear.
  • Total time per call by prompt size (16k prompt): Claude Sonnet 5.5 2.10 s (n 3, run range 2 s–2.5 s); Claude Opus 5.5 2.36 s (n 3, run range 2.1 s–3.4 s). Unclear.
  • Total time per call by prompt size (64k prompt): Claude Sonnet 5.5 3.44 s (n 3, run range 1.7 s–4.4 s); Claude Opus 5.5 2.35 s (n 3, run range 2.3 s–4.3 s). Unclear.
  • Exact lookup answers at the 1k, 16k and 64k prompt-size targets: Claude Sonnet 5.5 100% (9/9) (n 9, 95% interval 70%–100%); Claude Opus 5.5 56% (5/9) (n 9, 95% interval 27%–81%). Tie.

    Calculation: at these rates, about 13 runs per side would separate them.

GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

6 ties · 4 unclear
  • Pass rate on 4 harder tasks (Strict pass): Claude Sonnet 5.5 38% (6/16) (n 16, 95% interval 18%–61%); Claude Opus 5.5 42% (5/12) (n 12, 95% interval 19%–68%). Tie.

    Calculation: at these rates, about 2,094 runs per side would separate them.

  • Pass rate on 4 harder tasks (Lenient (format misses counted)): Claude Sonnet 5.5 38% (6/16) (n 16, 95% interval 18%–61%); Claude Opus 5.5 50% (6/12) (n 12, 95% interval 25%–75%). Tie.

    Calculation: at these rates, about 233 runs per side would separate them.

  • Calls that tried a tool although tools were off: Claude Sonnet 5.5 31% (5/16) (n 16, 95% interval 14%–56%); Claude Opus 5.5 42% (5/12) (n 12, 95% interval 19%–68%). Unclear.
  • Strict pass rate by task: 10x10 nonogram: Claude Sonnet 5.5 100% (4/4) (n 4, 95% interval 51%–100%); Claude Opus 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.
  • Strict pass rate by task: Sudoku, 22 givens: Claude Sonnet 5.5 0% (0/4) (n 4, 95% interval 0%–49%); Claude Opus 5.5 0% (0/3) (n 3, 95% interval 0%–56%). Tie.
  • Strict pass rate by task: 6x6 Skyscrapers: Claude Sonnet 5.5 0% (0/4) (n 4, 95% interval 0%–49%); Claude Opus 5.5 33% (1/3) (n 3, 95% interval 6.2%–79%). Tie.

    Calculation: at these rates, about 20 runs per side would separate them.

  • Strict pass rate by task: Seeded shuffle output: Claude Sonnet 5.5 50% (2/4) (n 4, 95% interval 15%–85%); Claude Opus 5.5 33% (1/3) (n 3, 95% interval 6.2%–79%). Tie.

    Calculation: at these rates, about 127 runs per side would separate them.

  • Total time per call on harder tasks: Claude Sonnet 5.5 70.4 s (n 12, run range 4.3 s–210 s); Claude Opus 5.5 80.3 s (n 9, run range 3.8 s–280 s). Unclear.
  • Output tokens per call on harder tasks (Output tokens): Claude Sonnet 5.5 9,287 (n 12, run range 407–27,921); Claude Opus 5.5 8,420 (n 9, run range 279–40,044). Unclear.
  • List-price cost per strict pass on harder tasks (calculation), calculation: Claude Sonnet 5.5 $0.24 (n 16); Claude Opus 5.5 $0.59 (n 12). Unclear.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

66 rows from 10 studies. No row separates them: 16 ties, 50 unclear.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Claude Sonnet 5.5

  • List-price cost per attempt (calculation): $2.88 vs $7.59. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call (calculation): $0.0036 vs $0.0069. A list-price calculation, not a measured difference. Calculation
  • List-price cost per passing answer (calculation): $0.0062 vs $0.010. A list-price calculation, not a measured difference. Calculation
  • List-price cost per strict pass on hard tasks (calculation): $0.014 vs $0.028. A list-price calculation, not a measured difference. Calculation
  • List-price cost per passing coding session (calculation): $0.085 vs $0.22. A list-price calculation, not a measured difference. Calculation
  • List-price cost per strict pass by effort (calculation): $0.014 vs $0.029. A list-price calculation, not a measured difference. Calculation
  • List-price cost of 5-question sessions with and without the cache (calculation) (With the cache, as recorded): $0.14 vs $0.26. A list-price calculation, not a measured difference. Calculation
  • List-price cost of 5-question sessions with and without the cache (calculation) (Without a cache: every input token at the input price): $0.27 vs $0.54. A list-price calculation, not a measured difference. Calculation
  • Cost of a reused prefix with and without the cache, by session length (calculation): 1 turn: $15.66 vs $31.31. A list-price calculation, not a measured difference. Calculation
  • Cost of a reused prefix with and without the cache, by session length (calculation): 2 turns: $31.32 vs $62.62. A list-price calculation, not a measured difference. Calculation
  • Cost of a reused prefix with and without the cache, by session length (calculation): 3 turns: $46.99 vs $93.94. A list-price calculation, not a measured difference. Calculation
  • Cost of a reused prefix with and without the cache, by session length (calculation): 5 turns: $78.31 vs $156.56. A list-price calculation, not a measured difference. Calculation
  • Cost of a reused prefix with and without the cache, by session length (calculation): 10 turns: $156.62 vs $313.12. A list-price calculation, not a measured difference. Calculation
  • Cost of a reused prefix with and without the cache, by session length (calculation): 20 turns: $313.24 vs $626.24. A list-price calculation, not a measured difference. Calculation
  • One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (No cache (the same either way)): $156.62 vs $313.12. A list-price calculation, not a measured difference. Calculation
  • One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (One 10-turn session, 1-hour cache): $45.42 vs $76.71. A list-price calculation, not a measured difference. Calculation
  • One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix)): $313.24 vs $626.24. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0067 vs $0.013. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)): $0.0037 vs $0.0080. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)): $0.0040 vs $0.0077. A list-price calculation, not a measured difference. Calculation
  • Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass): $0.0063 vs $0.013. A list-price calculation, not a measured difference. Calculation
  • Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass): $0.014 vs $0.029. A list-price calculation, not a measured difference. Calculation
  • List-price cost per strict pass on harder tasks (calculation): $0.24 vs $0.59. A list-price calculation, not a measured difference. Calculation

When to pick Claude Opus 5.5

  • Time to first text as the prompt grows: 64k: 1.79 s vs 3.07 s. A list-price calculation, not a measured difference. Calculation

Side by side

The study charts, showing only these two. Open a study for every configuration.

Claude Opus 5.5 (Agent, new build)
Claude Sonnet 5.5 (Agent, older builds)

2 rows. Highest Claude Opus 5.5 (Agent, new build) 67% (95% interval 21%–94%, n 3). Lowest Claude Sonnet 5.5 (Agent, older builds) 33% (95% interval 6.2%–79%, n 3). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 3 per row

One attempt per arm per instance, official grader · 95% Wilson intervals

Interim: 3 of 8 declared pairs are graded; 5 were never started; no resumed attempts are included here. With n = 3 the intervals span most of the axis, so this chart supports no ranking. The instances were chosen to hold Sonnet misses and resolves in equal numbers, so neither rate estimates SWE-bench Verified as a whole.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

Calculation
Claude Opus 5.5 (Agent, new build)
Claude Sonnet 5.5 (Agent, older builds)

List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 (Agent, new build) $7.59 (n 3). Lowest Claude Sonnet 5.5 (Agent, older builds) $2.88 (n 3).

Notesn = 3 per row

Mean over the same 3 instances; notional cost from the platform price table

Calculation, not an invoice: the platform price table at each run commit applied to the recorded tokens of subscription runs (Opus 5.5: $4 input, $20 output, $0.20 cache read per million tokens). Totals $22.76 vs $8.64: 2.6×, a ratio of two calculations.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Repricing calculation, Anthropic list prices (Claude models)

Entrance: medians race at 870× real time
Claude Opus 5.5 (Agent, new build)
Claude Sonnet 5.5 (Agent, older builds)

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest Claude Opus 5.5 (Agent, new build) 20.3 min (range 9.8 min–25.3 min, n 3). Fastest Claude Sonnet 5.5 (Agent, older builds) 9.4 min (range 4.7 min–15 min, n 3). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Median minutes; whiskers = fastest and slowest of 3 attempts (not an interval)

Worker minutes from the run reports. Opus attempts ran one at a time on the newer build; the Sonnet attempts ran in earlier campaigns. 3 attempts per arm is too few to call a difference; a range is not a confidence interval.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code

2 rows. Highest Claude Opus 5.5 (high) · Claude Code 100% (95% interval 80%–100%, n 15). Lowest Claude Sonnet 5.5 · Claude Code 80% (95% interval 55%–93%, n 15). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 15 per row

Every call counts; failures and timeouts are non-passes

Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.

Source: Provider head-to-head: Claude Code models vs Codex efforts

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Claude Sonnet 5.5 or Claude Opus 5.5?
Claude Sonnet 5.5 and Claude Opus 5.5 share 35 measured metrics and 31 list-price calculations from 10 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 16 ties and 50 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).
How were Claude Sonnet 5.5 and Claude Opus 5.5 measured?
They share 35 measured metrics and 31 list-price calculations from 10 public studies: Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim); Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks; Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks; Prompt caching and run-to-run consistency in Claude Code and Codex CLI; Prompt cache break-even: after how many reuses does a cached prefix cost less?; How much of an AI bill is thinking? Reasoning tokens by model and effort; Where the seconds go: first text, output speed and prompt size for 6 LLMs; GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Every row names its configuration, its sample size and its interval or range.
How do Claude Sonnet 5.5 and Claude Opus 5.5 compare on resolved on the same 3 SWE-bench Verified instances (interim)?
Claude Sonnet 5.5: 33% (1/3) (Agent · older builds · SWE-bench Verified, interim paired probe; n = 3; 95% interval 6.2% to 79%). Claude Opus 5.5: 67% (2/3) (Agent · new build · SWE-bench Verified, interim paired probe; n = 3; 95% interval 21% to 94%). The 95% intervals overlap (Claude Sonnet 5.5 6% to 79%; Claude Opus 5.5 21% to 94%), so this sample cannot separate them.
How do Claude Sonnet 5.5 and Claude Opus 5.5 compare on worker time per attempt?
Claude Sonnet 5.5: 9.4 min (Agent · older builds · SWE-bench Verified, interim paired probe; n = 3; run range 4.7 min to 15 min). Claude Opus 5.5: 20.3 min (Agent · new build · SWE-bench Verified, interim paired probe; n = 3; run range 9.8 min to 25.3 min). The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 4.7 min to 15.0 min; Claude Opus 5.5 9.8 min to 25.3 min); the medians alone do not show a reliable difference. A range is not a confidence interval.
How do Claude Sonnet 5.5 and Claude Opus 5.5 compare on pass rate on five validated tasks?
Claude Sonnet 5.5: 80% (12/15) (Claude Code · five short validated tasks; n = 15; 95% interval 55% to 93%). Claude Opus 5.5: 100% (15/15) (Claude Code · five short validated tasks; n = 15; 95% interval 80% to 100%). The 95% intervals overlap (Claude Sonnet 5.5 55% to 93%; Claude Opus 5.5 80% to 100%), so this sample cannot separate them.
How do Claude Sonnet 5.5 and Claude Opus 5.5 compare on total time per call?
Claude Sonnet 5.5: 2.31 s (Claude Code · five short validated tasks; n = 15; run range 2.2 s to 7.7 s). Claude Opus 5.5: 2.75 s (Claude Code · five short validated tasks; n = 15; run range 2.5 s to 8.9 s). The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.17 s to 7.73 s; Claude Opus 5.5 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
How do Claude Sonnet 5.5 and Claude Opus 5.5 compare on time to first useful output?
Claude Sonnet 5.5: 1.56 s (Claude Code · five short validated tasks; n = 15; run range 1 s to 6.4 s). Claude Opus 5.5: 1.92 s (Claude Code · five short validated tasks; n = 15; run range 1.6 s to 7.2 s). The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.99 s to 6.39 s; Claude Opus 5.5 1.56 s to 7.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.

The studies behind this page

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Includes calculations
  • Prompt Caching
  • Break Even

Prompt cache break-even: after how many reuses does a cached prefix cost less?

A calculation on list prices: reuses before a cached prompt prefix costs less, per Claude and GPT model, and cost per 1,000 sessions of 1 to 20 turns.

2reuses (the 3rd request) · Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation)

3 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.