35 measured metrics · 31 calculated · 10 studies
Claude Sonnet 5.5vsClaude Opus 5.5
No row separates them: 16 ties, 50 unclear.
The verdict
Claude Sonnet 5.5 and Claude Opus 5.5 share 35 measured metrics and 31 list-price calculations from 10 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 16 ties and 50 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | Claude Sonnet 5.5 | Claude Opus 5.5 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 80% (12/15)Claude Code · five short validated tasks | 100% (15/15)Claude Code · five short validated tasks | 15 | 95% CI: 55%–93% vs 80%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 55% to 93%; Claude Opus 5.5 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 2.31 sClaude Code · five short validated tasks | 2.75 sClaude Code · five short validated tasks | 15 | range: 2.2 s–7.7 s vs 2.5 s–8.9 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.17 s to 7.73 s; Claude Opus 5.5 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 1.56 sClaude Code · five short validated tasks | 1.92 sClaude Code · five short validated tasks | 15 | range: 1 s–6.4 s vs 1.6 s–7.2 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.99 s to 6.39 s; Claude Opus 5.5 1.56 s to 7.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 1,401Claude Code · five short validated tasks | 1,401Claude Code · five short validated tasks | 15 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 107Claude Code · five short validated tasks | 64Claude Code · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.0062Claude Code · five short validated tasks | $0.010Claude Code · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.0062 vs $0.010) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Output tokens per call (Output tokens), 1.7x (Claude Sonnet 5.5 larger).
Watch it build
A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.
Claude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say
66 comparison rows from 10 studies: 0 rows favour Sonnet 5.5, 0 favour Opus 5.5, 66 are ties or unclear. Cost rows are calculations.
Transcript
- Comparison · 66 rows · 10 studies. Sonnet 5.5 vs Opus 5.5. A winner only where the 95% intervals or run ranges do not overlap.
- 66 comparison rows from 10 studies. None separates them: intervals or run ranges overlap, or none was recorded. Rows where Sonnet 5.5 is ahead: 0 (of 66). Rows where Opus 5.5 is ahead: 0 (of 66). Ties or unclear: 66 (16 ties · 50 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
- SWE-bench pairs, interim: pass rate 33% vs 67%, tie: 95% intervals overlap. None of the 3 rows separates them. Table: SWE-bench pairs, interim · Agent · n = 3 per side. Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
- Five short tasks: pass rate 80% vs 100%, tie: 95% intervals overlap. None of the 8 rows separates them. Table: Five short tasks · Claude Code · n = 15 per side. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
- Eight hard tasks: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 6 rows separates them. Table: Eight hard tasks · Claude Code · n = 24 per side. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
- Coding agents, hidden tests: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Coding agents, hidden tests · Claude Code · n = 12 per side. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
- Effort ladder, default effort: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Effort ladder, default effort · Claude Code · n = 16 per side. Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
- Caching sessions: pass rate not measured. None of the 4 rows separates them. Table: Caching sessions · Claude Code · n = 15, 3, 12 per side. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- Prompt cache break-even: after how many reuses does a cached prefix cost less?: pass rate not measured. None of the 9 rows separates them. Table: Prompt cache break-even: after how many reuses does a cached prefix cost less? · no cache · n = per side. Caveat: The source has 30 attempted Claude turns, 0 failed turns and 0 turns without token usage. Failed turns with usage remain in cost totals. Missing usage cannot be priced. No quality rate or cache-caused speed effect is claimed.
- How much of an AI bill is thinking? Reasoning tokens by model and effort: pass rate not measured. None of the 8 rows separates them. Table: How much of an AI bill is thinking? Reasoning tokens by model and effort · Claude Code · n = 24, 16, 15 per side. Caveat: The calls used one shared Mac and network. Other work and provider load were not controlled. Latency includes those effects.
- Where the seconds go: first text, output speed and prompt size for 6 LLMs: pass rate 100% vs 56%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: Where the seconds go: first text, output speed and prompt size for 6 LLMs · Claude Code · n = 4, 3, 9 per side. Caveat: First text includes CLI start-up and reasoning time. It is not the API’s time to first token; a direct API call would skip the CLI start-up (the CLI reported ready after a median 0.33 s to 0.84 s per cell).
- GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks: pass rate 38% vs 42%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks · Claude Code · n = per side. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.
- No winner where the data shows none. Every row and its reason online.
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- Claude Sonnet 5.5
- Claude Opus 5.5
- 95% interval
- fastest–slowest run (not an interval)
- where the two overlap
- hollow: list-price calculation
These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.
Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)
- Resolved on the same 3 SWE-bench Verified instances (interim)33% (1/3)n 367% (2/3)n 3TieResolved on the same 3 SWE-bench Verified instances (interim): Claude Sonnet 5.5 33% (1/3) (n 3, 95% interval 6.2%–79%); Claude Opus 5.5 67% (2/3) (n 3, 95% interval 21%–94%). Tie.
Calculation: at these rates, about 31 runs per side would separate them.
- List-price cost per attempt (calculation)Calculation$2.88n 3$7.59n 3UnclearList-price cost per attempt (calculation), calculation: Claude Sonnet 5.5 $2.88 (n 3); Claude Opus 5.5 $7.59 (n 3). Unclear.
- Worker time per attempt9.4 minn 320.3 minn 3UnclearWorker time per attempt: Claude Sonnet 5.5 9.4 min (n 3, run range 4.7 min–15 min); Claude Opus 5.5 20.3 min (n 3, run range 9.8 min–25.3 min). Unclear.
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
- Pass rate on five validated tasks80% (12/15)n 15100% (15/15)n 15TiePass rate on five validated tasks: Claude Sonnet 5.5 80% (12/15) (n 15, 95% interval 55%–93%); Claude Opus 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
Calculation: at these rates, about 33 runs per side would separate them.
- Total time per call2.31 sn 152.75 sn 15UnclearTotal time per call: Claude Sonnet 5.5 2.31 s (n 15, run range 2.2 s–7.7 s); Claude Opus 5.5 2.75 s (n 15, run range 2.5 s–8.9 s). Unclear.
- Time to first useful output1.56 sn 151.92 sn 15UnclearTime to first useful output: Claude Sonnet 5.5 1.56 s (n 15, run range 1 s–6.4 s); Claude Opus 5.5 1.92 s (n 15, run range 1.6 s–7.2 s). Unclear.
- Input tokens per call: what the CLI sends (Cache read)1,401n 151,401n 15TieInput tokens per call: what the CLI sends (Cache read): Claude Sonnet 5.5 1,401 (n 15); Claude Opus 5.5 1,401 (n 15). Tie.
- Input tokens per call: what the CLI sends (Other input)685n 15680n 15UnclearInput tokens per call: what the CLI sends (Other input): Claude Sonnet 5.5 685 (n 15); Claude Opus 5.5 680 (n 15). Unclear.
- Output tokens per call (Output tokens)107n 1564n 15UnclearOutput tokens per call (Output tokens): Claude Sonnet 5.5 107 (n 15); Claude Opus 5.5 64 (n 15). Unclear.
- List-price cost per call (calculation)Calculation$0.0036n 15$0.0069n 15UnclearList-price cost per call (calculation), calculation: Claude Sonnet 5.5 $0.0036 (n 15, run range $0.0034–$0.01); Claude Opus 5.5 $0.0069 (n 15, run range $0.0059–$0.022). Unclear.
- List-price cost per passing answer (calculation)Calculation$0.0062n 15$0.010n 15UnclearList-price cost per passing answer (calculation), calculation: Claude Sonnet 5.5 $0.0062 (n 15); Claude Opus 5.5 $0.010 (n 15). Unclear.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
- Pass rate on eight hard tasks (Strict pass)100% (24/24)n 24100% (24/24)n 24TiePass rate on eight hard tasks (Strict pass): Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%); Claude Opus 5.5 100% (24/24) (n 24, 95% interval 86%–100%). Tie.
- Pass rate on eight hard tasks (Lenient (format misses counted))100% (24/24)n 24100% (24/24)n 24TiePass rate on eight hard tasks (Lenient (format misses counted)): Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%); Claude Opus 5.5 100% (24/24) (n 24, 95% interval 86%–100%). Tie.
- Total time per call on hard tasks (separate batches)7.75 sn 249.18 sn 24UnclearTotal time per call on hard tasks (separate batches): Claude Sonnet 5.5 7.75 s (n 24, run range 2.3 s–34.8 s); Claude Opus 5.5 9.18 s (n 24, run range 4.2 s–27.2 s). Unclear.
- Time to first useful output on hard tasks5.95 sn 246.78 sn 24UnclearTime to first useful output on hard tasks: Claude Sonnet 5.5 5.95 s (n 24, run range 0.9 s–30.6 s); Claude Opus 5.5 6.78 s (n 24, run range 2.4 s–21.8 s). Unclear.
- Output tokens per call on hard tasks (Output tokens)1,050n 24945n 24UnclearOutput tokens per call on hard tasks (Output tokens): Claude Sonnet 5.5 1,050 (n 24); Claude Opus 5.5 945 (n 24). Unclear.
- List-price cost per strict pass on hard tasks (calculation)Calculation$0.014n 24$0.028n 24UnclearList-price cost per strict pass on hard tasks (calculation), calculation: Claude Sonnet 5.5 $0.014 (n 24); Claude Opus 5.5 $0.028 (n 24). Unclear.
Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
- Coding sessions that passed every hidden check100% (12/12)n 12100% (12/12)n 12TieCoding sessions that passed every hidden check: Claude Sonnet 5.5 100% (12/12) (n 12, 95% interval 76%–100%); Claude Opus 5.5 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
- Time per coding session23.1 sn 1256.9 sn 12UnclearTime per coding session: Claude Sonnet 5.5 23.1 s (n 12, run range 18.7 s–44.5 s); Claude Opus 5.5 56.9 s (n 12, run range 29.8 s–186 s). Unclear.
- Tool calls per coding session7.5n 127.5n 12TieTool calls per coding session: Claude Sonnet 5.5 7.5 (n 12, run range 3–14); Claude Opus 5.5 7.5 (n 12, run range 5–14). Tie.
- List-price cost per passing coding session (calculation)Calculation$0.085n 12$0.22n 12UnclearList-price cost per passing coding session (calculation), calculation: Claude Sonnet 5.5 $0.085 (n 12); Claude Opus 5.5 $0.22 (n 12). Unclear.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
- Strict pass rate by effort on eight hard tasks100% (16/16)n 16100% (16/16)n 16TieStrict pass rate by effort on eight hard tasks: Claude Sonnet 5.5 100% (16/16) (n 16, 95% interval 81%–100%); Claude Opus 5.5 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
- Total time per call by effort on hard tasks7.97 sn 169.18 sn 16UnclearTotal time per call by effort on hard tasks: Claude Sonnet 5.5 7.97 s (n 16, run range 2.3 s–21.6 s); Claude Opus 5.5 9.18 s (n 16, run range 4.2 s–27.2 s). Unclear.
- Output tokens per call by effort on hard tasks (Output tokens)1,054n 16945n 16UnclearOutput tokens per call by effort on hard tasks (Output tokens): Claude Sonnet 5.5 1,054 (n 16); Claude Opus 5.5 945 (n 16). Unclear.
- List-price cost per strict pass by effort (calculation)Calculation$0.014n 16$0.029n 16UnclearList-price cost per strict pass by effort (calculation), calculation: Claude Sonnet 5.5 $0.014 (n 16); Claude Opus 5.5 $0.029 (n 16). Unclear.
Prompt caching and run-to-run consistency in Claude Code and Codex CLI
- List-price cost of 5-question sessions with and without the cache (calculation) (With the cache, as recorded)Calculation$0.14n 15$0.26n 15UnclearList-price cost of 5-question sessions with and without the cache (calculation) (With the cache, as recorded), calculation: Claude Sonnet 5.5 $0.14 (n 15); Claude Opus 5.5 $0.26 (n 15). Unclear.
- List-price cost of 5-question sessions with and without the cache (calculation) (Without a cache: every input token at the input price)Calculation$0.27n 15$0.54n 15UnclearList-price cost of 5-question sessions with and without the cache (calculation) (Without a cache: every input token at the input price), calculation: Claude Sonnet 5.5 $0.27 (n 15); Claude Opus 5.5 $0.54 (n 15). Unclear.
- Time per turn: first turn vs later turns in a cached session (Turn 1 (writes the ledger to the cache))1.64 sn 31.90 sn 3UnclearTime per turn: first turn vs later turns in a cached session (Turn 1 (writes the ledger to the cache)): Claude Sonnet 5.5 1.64 s (n 3, run range 1.6 s–1.8 s); Claude Opus 5.5 1.90 s (n 3, run range 1.8 s–4.4 s). Unclear.
- Time per turn: first turn vs later turns in a cached session (Turns 2-5 (read the ledger from the cache))1.61 sn 122.40 sn 12UnclearTime per turn: first turn vs later turns in a cached session (Turns 2-5 (read the ledger from the cache)): Claude Sonnet 5.5 1.61 s (n 12, run range 1.4 s–5.6 s); Claude Opus 5.5 2.40 s (n 12, run range 1.6 s–12.7 s). Unclear.
Prompt cache break-even: after how many reuses does a cached prefix cost less?
- Cost of a reused prefix with and without the cache, by session length (calculation): 1 turnCalculation$15.66$31.31UnclearCost of a reused prefix with and without the cache, by session length (calculation): 1 turn, calculation: Claude Sonnet 5.5 $15.66; Claude Opus 5.5 $31.31. Unclear.
- Cost of a reused prefix with and without the cache, by session length (calculation): 2 turnsCalculation$31.32$62.62UnclearCost of a reused prefix with and without the cache, by session length (calculation): 2 turns, calculation: Claude Sonnet 5.5 $31.32; Claude Opus 5.5 $62.62. Unclear.
- Cost of a reused prefix with and without the cache, by session length (calculation): 3 turnsCalculation$46.99$93.94UnclearCost of a reused prefix with and without the cache, by session length (calculation): 3 turns, calculation: Claude Sonnet 5.5 $46.99; Claude Opus 5.5 $93.94. Unclear.
- Cost of a reused prefix with and without the cache, by session length (calculation): 5 turnsCalculation$78.31$156.56UnclearCost of a reused prefix with and without the cache, by session length (calculation): 5 turns, calculation: Claude Sonnet 5.5 $78.31; Claude Opus 5.5 $156.56. Unclear.
- Cost of a reused prefix with and without the cache, by session length (calculation): 10 turnsCalculation$156.62$313.12UnclearCost of a reused prefix with and without the cache, by session length (calculation): 10 turns, calculation: Claude Sonnet 5.5 $156.62; Claude Opus 5.5 $313.12. Unclear.
- Cost of a reused prefix with and without the cache, by session length (calculation): 20 turnsCalculation$313.24$626.24UnclearCost of a reused prefix with and without the cache, by session length (calculation): 20 turns, calculation: Claude Sonnet 5.5 $313.24; Claude Opus 5.5 $626.24. Unclear.
- One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (No cache (the same either way))Calculation$156.62$313.12UnclearOne 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (No cache (the same either way)), calculation: Claude Sonnet 5.5 $156.62; Claude Opus 5.5 $313.12. Unclear.
- One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (One 10-turn session, 1-hour cache)Calculation$45.42$76.71UnclearOne 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (One 10-turn session, 1-hour cache), calculation: Claude Sonnet 5.5 $45.42; Claude Opus 5.5 $76.71. Unclear.
- One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix))Calculation$313.24$626.24UnclearOne 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix)), calculation: Claude Sonnet 5.5 $313.24; Claude Opus 5.5 $626.24. Unclear.
How much of an AI bill is thinking? Reasoning tokens by model and effort
- Reasoning share of output tokens per call on hard tasks (calculation)Calculation54.5%n 2454.8%n 24UnclearReasoning share of output tokens per call on hard tasks (calculation), calculation: Claude Sonnet 5.5 54.5% (n 24, run range 0%–96%); Claude Opus 5.5 54.8% (n 24, run range 30%–96%). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation$0.0067n 24$0.013n 24UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Claude Sonnet 5.5 $0.0067 (n 24); Claude Opus 5.5 $0.013 (n 24). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation$0.0037n 24$0.0080n 24UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Claude Sonnet 5.5 $0.0037 (n 24); Claude Opus 5.5 $0.0080 (n 24). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation$0.0040n 24$0.0077n 24UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Claude Sonnet 5.5 $0.0040 (n 24); Claude Opus 5.5 $0.0077 (n 24). Unclear.
- Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation$0.0063n 16$0.013n 16UnclearReasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass), calculation: Claude Sonnet 5.5 $0.0063 (n 16); Claude Opus 5.5 $0.013 (n 16). Unclear.
- Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation$0.014n 16$0.029n 16UnclearReasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass), calculation: Claude Sonnet 5.5 $0.014 (n 16); Claude Opus 5.5 $0.029 (n 16). Unclear.
- Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation54.5%n 2454.8%n 24UnclearReasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Claude Sonnet 5.5 54.5% (n 24, run range 0%–96%); Claude Opus 5.5 54.8% (n 24, run range 30%–96%). Unclear.
- Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation0%n 150%n 15TieReasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Claude Sonnet 5.5 0% (n 15, run range 0%–73%); Claude Opus 5.5 0% (n 15, run range 0%–93%). Tie.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
- Time to first text: a 250-line answer, six models1.96 sn 41.97 sn 4UnclearTime to first text: a 250-line answer, six models: Claude Sonnet 5.5 1.96 s (n 4, run range 0.9 s–4.1 s); Claude Opus 5.5 1.97 s (n 4, run range 1.7 s–2.4 s). Unclear.
- Output speed after the first text: visible tokens per second (calculation)Calculation232n 4156n 4UnclearOutput speed after the first text: visible tokens per second (calculation), calculation: Claude Sonnet 5.5 232 (n 4, run range 230–233); Claude Opus 5.5 156 (n 4, run range 155–156). Unclear.
- Output speed in characters per second after the first text (calculation)Calculation517n 4347n 4UnclearOutput speed in characters per second after the first text (calculation), calculation: Claude Sonnet 5.5 517 (n 4, run range 513–519); Claude Opus 5.5 347 (n 4, run range 345–349). Unclear.
- Time to first text as the prompt grows: 1kCalculation1.45 sn 31.51 sn 3UnclearTime to first text as the prompt grows: 1k, calculation: Claude Sonnet 5.5 1.45 s (n 3, run range 1.2 s–1.7 s); Claude Opus 5.5 1.51 s (n 3, run range 1.5 s–2 s). Unclear.
- Time to first text as the prompt grows: 16kCalculation1.78 sn 31.74 sn 3UnclearTime to first text as the prompt grows: 16k, calculation: Claude Sonnet 5.5 1.78 s (n 3, run range 1.6 s–2.1 s); Claude Opus 5.5 1.74 s (n 3, run range 1.7 s–3 s). Unclear.
- Time to first text as the prompt grows: 64kCalculation3.07 sn 31.79 sn 3UnclearTime to first text as the prompt grows: 64k, calculation: Claude Sonnet 5.5 3.07 s (n 3, run range 1.4 s–3.6 s); Claude Opus 5.5 1.79 s (n 3, run range 1.7 s–3.7 s). Unclear.
- Total time per call by prompt size (1k prompt)1.78 sn 31.83 sn 3UnclearTotal time per call by prompt size (1k prompt): Claude Sonnet 5.5 1.78 s (n 3, run range 1.6 s–2.1 s); Claude Opus 5.5 1.83 s (n 3, run range 1.8 s–2.4 s). Unclear.
- Total time per call by prompt size (16k prompt)2.10 sn 32.36 sn 3UnclearTotal time per call by prompt size (16k prompt): Claude Sonnet 5.5 2.10 s (n 3, run range 2 s–2.5 s); Claude Opus 5.5 2.36 s (n 3, run range 2.1 s–3.4 s). Unclear.
- Total time per call by prompt size (64k prompt)3.44 sn 32.35 sn 3UnclearTotal time per call by prompt size (64k prompt): Claude Sonnet 5.5 3.44 s (n 3, run range 1.7 s–4.4 s); Claude Opus 5.5 2.35 s (n 3, run range 2.3 s–4.3 s). Unclear.
- Exact lookup answers at the 1k, 16k and 64k prompt-size targets100% (9/9)n 956% (5/9)n 9TieExact lookup answers at the 1k, 16k and 64k prompt-size targets: Claude Sonnet 5.5 100% (9/9) (n 9, 95% interval 70%–100%); Claude Opus 5.5 56% (5/9) (n 9, 95% interval 27%–81%). Tie.
Calculation: at these rates, about 13 runs per side would separate them.
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
- Pass rate on 4 harder tasks (Strict pass)38% (6/16)n 1642% (5/12)n 12TiePass rate on 4 harder tasks (Strict pass): Claude Sonnet 5.5 38% (6/16) (n 16, 95% interval 18%–61%); Claude Opus 5.5 42% (5/12) (n 12, 95% interval 19%–68%). Tie.
Calculation: at these rates, about 2,094 runs per side would separate them.
- Pass rate on 4 harder tasks (Lenient (format misses counted))38% (6/16)n 1650% (6/12)n 12TiePass rate on 4 harder tasks (Lenient (format misses counted)): Claude Sonnet 5.5 38% (6/16) (n 16, 95% interval 18%–61%); Claude Opus 5.5 50% (6/12) (n 12, 95% interval 25%–75%). Tie.
Calculation: at these rates, about 233 runs per side would separate them.
- Calls that tried a tool although tools were off31% (5/16)n 1642% (5/12)n 12UnclearCalls that tried a tool although tools were off: Claude Sonnet 5.5 31% (5/16) (n 16, 95% interval 14%–56%); Claude Opus 5.5 42% (5/12) (n 12, 95% interval 19%–68%). Unclear.
- Strict pass rate by task: 10x10 nonogram100% (4/4)n 4100% (3/3)n 3TieStrict pass rate by task: 10x10 nonogram: Claude Sonnet 5.5 100% (4/4) (n 4, 95% interval 51%–100%); Claude Opus 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.
- Strict pass rate by task: Sudoku, 22 givens0% (0/4)n 40% (0/3)n 3TieStrict pass rate by task: Sudoku, 22 givens: Claude Sonnet 5.5 0% (0/4) (n 4, 95% interval 0%–49%); Claude Opus 5.5 0% (0/3) (n 3, 95% interval 0%–56%). Tie.
- Strict pass rate by task: 6x6 Skyscrapers0% (0/4)n 433% (1/3)n 3TieStrict pass rate by task: 6x6 Skyscrapers: Claude Sonnet 5.5 0% (0/4) (n 4, 95% interval 0%–49%); Claude Opus 5.5 33% (1/3) (n 3, 95% interval 6.2%–79%). Tie.
Calculation: at these rates, about 20 runs per side would separate them.
- Strict pass rate by task: Seeded shuffle output50% (2/4)n 433% (1/3)n 3TieStrict pass rate by task: Seeded shuffle output: Claude Sonnet 5.5 50% (2/4) (n 4, 95% interval 15%–85%); Claude Opus 5.5 33% (1/3) (n 3, 95% interval 6.2%–79%). Tie.
Calculation: at these rates, about 127 runs per side would separate them.
- Total time per call on harder tasks70.4 sn 1280.3 sn 9UnclearTotal time per call on harder tasks: Claude Sonnet 5.5 70.4 s (n 12, run range 4.3 s–210 s); Claude Opus 5.5 80.3 s (n 9, run range 3.8 s–280 s). Unclear.
- Output tokens per call on harder tasks (Output tokens)9,287n 128,420n 9UnclearOutput tokens per call on harder tasks (Output tokens): Claude Sonnet 5.5 9,287 (n 12, run range 407–27,921); Claude Opus 5.5 8,420 (n 9, run range 279–40,044). Unclear.
- List-price cost per strict pass on harder tasks (calculation)Calculation$0.24n 16$0.59n 12UnclearList-price cost per strict pass on harder tasks (calculation), calculation: Claude Sonnet 5.5 $0.24 (n 16); Claude Opus 5.5 $0.59 (n 12). Unclear.
| Metric | Claude Sonnet 5.5 | Claude Opus 5.5 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Resolved on the same 3 SWE-bench Verified instances (interim) | 33% (1/3)Agent · older builds · SWE-bench Verified, interim paired probe | 67% (2/3)Agent · new build · SWE-bench Verified, interim paired probe | 3 | 95% CI: 6.2%–79% vs 21%–94% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 6% to 79%; Claude Opus 5.5 21% to 94%), so this sample cannot separate them. | Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim) |
| List-price cost per attempt (calculation)Calculation | $2.88Agent · older builds · SWE-bench Verified, interim paired probe | $7.59Agent · new build · SWE-bench Verified, interim paired probe | 3 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($2.88 vs $7.59, 2.6x) is not tested against run-to-run variation. | Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim) |
| Worker time per attempt | 9.4 minAgent · older builds · SWE-bench Verified, interim paired probe | 20.3 minAgent · new build · SWE-bench Verified, interim paired probe | 3 | range: 4.7 min–15 min vs 9.8 min–25.3 min | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 4.7 min to 15.0 min; Claude Opus 5.5 9.8 min to 25.3 min); the medians alone do not show a reliable difference. A range is not a confidence interval. | Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim) |
| Pass rate on five validated tasks | 80% (12/15)Claude Code · five short validated tasks | 100% (15/15)Claude Code · five short validated tasks | 15 | 95% CI: 55%–93% vs 80%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 55% to 93%; Claude Opus 5.5 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 2.31 sClaude Code · five short validated tasks | 2.75 sClaude Code · five short validated tasks | 15 | range: 2.2 s–7.7 s vs 2.5 s–8.9 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.17 s to 7.73 s; Claude Opus 5.5 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 1.56 sClaude Code · five short validated tasks | 1.92 sClaude Code · five short validated tasks | 15 | range: 1 s–6.4 s vs 1.6 s–7.2 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.99 s to 6.39 s; Claude Opus 5.5 1.56 s to 7.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 1,401Claude Code · five short validated tasks | 1,401Claude Code · five short validated tasks | 15 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Other input) | 685Claude Code · five short validated tasks | 680Claude Code · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 107Claude Code · five short validated tasks | 64Claude Code · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per call (calculation)Calculation | $0.0036Claude Code · five short validated tasks | $0.0069Claude Code · five short validated tasks | 15 | range: $0.0034–$0.01 vs $0.0059–$0.022 | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 $0.0034 to $0.010; Claude Opus 5.5 $0.0059 to $0.022); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.0062Claude Code · five short validated tasks | $0.010Claude Code · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.0062 vs $0.010) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Pass rate on eight hard tasks (Strict pass) | 100% (24/24)Claude Code · eight hard validated tasks | 100% (24/24)Claude Code · eight hard validated tasks | 24 | 95% CI: 86%–100% vs 86%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; Claude Opus 5.5 86% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 100% (24/24)Claude Code · eight hard validated tasks | 100% (24/24)Claude Code · eight hard validated tasks | 24 | 95% CI: 86%–100% vs 86%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; Claude Opus 5.5 86% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Total time per call on hard tasks (separate batches) | 7.75 sClaude Code · eight hard validated tasks | 9.18 sClaude Code · eight hard validated tasks | 24 | range: 2.3 s–34.8 s vs 4.2 s–27.2 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; Claude Opus 5.5 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Time to first useful output on hard tasks | 5.95 sClaude Code · eight hard validated tasks | 6.78 sClaude Code · eight hard validated tasks | 24 | range: 0.9 s–30.6 s vs 2.4 s–21.8 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.86 s to 30.6 s; Claude Opus 5.5 2.39 s to 21.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call on hard tasks (Output tokens) | 1,050Claude Code · eight hard validated tasks | 945Claude Code · eight hard validated tasks | 24 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass on hard tasks (calculation)Calculation | $0.014Claude Code · eight hard validated tasks | $0.028Claude Code · eight hard validated tasks | 24 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.014 vs $0.028) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Coding sessions that passed every hidden check | 100% (12/12)Claude Code · six small repository tasks with hidden tests | 100% (12/12)Claude Code · six small repository tasks with hidden tests | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; Claude Opus 5.5 76% to 100%), so this sample cannot separate them. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Time per coding session | 23.1 sClaude Code · six small repository tasks with hidden tests | 56.9 sClaude Code · six small repository tasks with hidden tests | 12 | range: 18.7 s–44.5 s vs 29.8 s–186 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 18.7 s to 44.5 s; Claude Opus 5.5 29.8 s to 185.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Tool calls per coding session | 7.5Claude Code · six small repository tasks with hidden tests | 7.5Claude Code · six small repository tasks with hidden tests | 12 | range: 3–14 vs 5–14 | Tie | Same value. More or fewer is not better by itself for this metric. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| List-price cost per passing coding session (calculation)Calculation | $0.085Claude Code · six small repository tasks with hidden tests | $0.22Claude Code · six small repository tasks with hidden tests | 12 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.085 vs $0.22, 2.6x) is not tested against run-to-run variation. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Strict pass rate by effort on eight hard tasks | 100% (16/16)Claude Code · eight hard validated tasks, effort ladder | 100% (16/16)Claude Code · eight hard validated tasks, effort ladder | 16 | 95% CI: 81%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 81% to 100%; Claude Opus 5.5 81% to 100%), so this sample cannot separate them. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Total time per call by effort on hard tasks | 7.97 sClaude Code · eight hard validated tasks, effort ladder | 9.18 sClaude Code · eight hard validated tasks, effort ladder | 16 | range: 2.3 s–21.6 s vs 4.2 s–27.2 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 21.6 s; Claude Opus 5.5 4.24 s to 27.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call by effort on hard tasks (Output tokens) | 1,054Claude Code · eight hard validated tasks, effort ladder | 945Claude Code · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass by effort (calculation)Calculation | $0.014Claude Code · eight hard validated tasks, effort ladder | $0.029Claude Code · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.014 vs $0.029, 2.1x) is not tested against run-to-run variation. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| List-price cost of 5-question sessions with and without the cache (calculation) (With the cache, as recorded)Calculation | $0.14Claude Code · calculation: 5-turn cached sessions over a fixed ledger | $0.26Claude Code · calculation: 5-turn cached sessions over a fixed ledger | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.14 vs $0.26) is not tested against run-to-run variation. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| List-price cost of 5-question sessions with and without the cache (calculation) (Without a cache: every input token at the input price)Calculation | $0.27Claude Code · calculation: 5-turn cached sessions over a fixed ledger | $0.54Claude Code · calculation: 5-turn cached sessions over a fixed ledger | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.27 vs $0.54, 2.0x) is not tested against run-to-run variation. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Time per turn: first turn vs later turns in a cached session (Turn 1 (writes the ledger to the cache)) | 1.64 sClaude Code · 5-turn cached sessions over a fixed ledger | 1.90 sClaude Code · 5-turn cached sessions over a fixed ledger | 3 | range: 1.6 s–1.8 s vs 1.8 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.58 s to 1.79 s; Claude Opus 5.5 1.78 s to 4.36 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Time per turn: first turn vs later turns in a cached session (Turns 2-5 (read the ledger from the cache)) | 1.61 sClaude Code · 5-turn cached sessions over a fixed ledger | 2.40 sClaude Code · 5-turn cached sessions over a fixed ledger | 12 | range: 1.4 s–5.6 s vs 1.6 s–12.7 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.35 s to 5.63 s; Claude Opus 5.5 1.63 s to 12.7 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Cost of a reused prefix with and without the cache, by session length (calculation): 1 turnCalculation | $15.66no cache | $31.31no cache | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($15.66 vs $31.31) is not tested against run-to-run variation. | Prompt cache break-even: after how many reuses does a cached prefix cost less? |
| Cost of a reused prefix with and without the cache, by session length (calculation): 2 turnsCalculation | $31.32no cache | $62.62no cache | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($31.32 vs $62.62) is not tested against run-to-run variation. | Prompt cache break-even: after how many reuses does a cached prefix cost less? |
| Cost of a reused prefix with and without the cache, by session length (calculation): 3 turnsCalculation | $46.99no cache | $93.94no cache | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($46.99 vs $93.94) is not tested against run-to-run variation. | Prompt cache break-even: after how many reuses does a cached prefix cost less? |
| Cost of a reused prefix with and without the cache, by session length (calculation): 5 turnsCalculation | $78.31no cache | $156.56no cache | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($78.31 vs $156.56) is not tested against run-to-run variation. | Prompt cache break-even: after how many reuses does a cached prefix cost less? |
| Cost of a reused prefix with and without the cache, by session length (calculation): 10 turnsCalculation | $156.62no cache | $313.12no cache | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($156.62 vs $313.12) is not tested against run-to-run variation. | Prompt cache break-even: after how many reuses does a cached prefix cost less? |
| Cost of a reused prefix with and without the cache, by session length (calculation): 20 turnsCalculation | $313.24no cache | $626.24no cache | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($313.24 vs $626.24) is not tested against run-to-run variation. | Prompt cache break-even: after how many reuses does a cached prefix cost less? |
| One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (No cache (the same either way))Calculation | $156.62 | $313.12cache read $0.2 per M | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($156.62 vs $313.12) is not tested against run-to-run variation. | Prompt cache break-even: after how many reuses does a cached prefix cost less? |
| One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (One 10-turn session, 1-hour cache)Calculation | $45.42 | $76.71cache read $0.2 per M | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($45.42 vs $76.71) is not tested against run-to-run variation. | Prompt cache break-even: after how many reuses does a cached prefix cost less? |
| One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix))Calculation | $313.24 | $626.24cache read $0.2 per M | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($313.24 vs $626.24) is not tested against run-to-run variation. | Prompt cache break-even: after how many reuses does a cached prefix cost less? |
| Reasoning share of output tokens per call on hard tasks (calculation)Calculation | 54.5%Claude Code | 54.8%Claude Code | 24 | range: 0%–96% vs 30%–96% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation | $0.0067Claude Code | $0.013Claude Code | 24 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation | $0.0037Claude Code | $0.0080Claude Code | 24 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation | $0.0040Claude Code | $0.0077Claude Code | 24 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation | $0.0063Claude Code | $0.013Claude Code | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation | $0.014Claude Code | $0.029Claude Code | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation | 54.5%Claude Code | 54.8%Claude Code | 24 | range: 0%–96% vs 30%–96% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation | 0%Claude Code | 0%Claude Code | 15 | range: 0%–73% vs 0%–93% | Tie | Same value. More or fewer is not better by itself for this metric. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Time to first text: a 250-line answer, six models | 1.96 sClaude Code | 1.97 sClaude Code | 4 | range: 0.9 s–4.1 s vs 1.7 s–2.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.88 s to 4.09 s; Claude Opus 5.5 1.70 s to 2.35 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 232Claude Code | 156Claude Code | 4 | range: 230–233 vs 155–156 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 517Claude Code | 347Claude Code | 4 | range: 513–519 vs 345–349 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 1kCalculation | 1.45 sClaude Code | 1.51 sClaude Code | 3 | range: 1.2 s–1.7 s vs 1.5 s–2 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.23 s to 1.72 s; Claude Opus 5.5 1.46 s to 2.01 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 16kCalculation | 1.78 sClaude Code | 1.74 sClaude Code | 3 | range: 1.6 s–2.1 s vs 1.7 s–3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.64 s to 2.11 s; Claude Opus 5.5 1.70 s to 2.97 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 64kCalculation | 3.07 sClaude Code | 1.79 sClaude Code | 3 | range: 1.4 s–3.6 s vs 1.7 s–3.7 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.38 s to 3.61 s; Claude Opus 5.5 1.72 s to 3.72 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (1k prompt) | 1.78 sClaude Code | 1.83 sClaude Code | 3 | range: 1.6 s–2.1 s vs 1.8 s–2.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.57 s to 2.12 s; Claude Opus 5.5 1.82 s to 2.41 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (16k prompt) | 2.10 sClaude Code | 2.36 sClaude Code | 3 | range: 2 s–2.5 s vs 2.1 s–3.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.98 s to 2.48 s; Claude Opus 5.5 2.11 s to 3.40 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (64k prompt) | 3.44 sClaude Code | 2.35 sClaude Code | 3 | range: 1.7 s–4.4 s vs 2.3 s–4.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.74 s to 4.38 s; Claude Opus 5.5 2.26 s to 4.29 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Exact lookup answers at the 1k, 16k and 64k prompt-size targets | 100% (9/9)Claude Code | 56% (5/9)Claude Code | 9 | 95% CI: 70%–100% vs 27%–81% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 70% to 100%; Claude Opus 5.5 27% to 81%), so this sample cannot separate them. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Pass rate on 4 harder tasks (Strict pass) | 38% (6/16)Claude Code | 42% (5/12)Claude Code | 16 / 12 | 95% CI: 18%–61% vs 19%–68% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 18% to 61%; Claude Opus 5.5 19% to 68%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Pass rate on 4 harder tasks (Lenient (format misses counted)) | 38% (6/16)Claude Code | 50% (6/12)Claude Code | 16 / 12 | 95% CI: 18%–61% vs 25%–75% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 18% to 61%; Claude Opus 5.5 25% to 75%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Calls that tried a tool although tools were off | 31% (5/16)Claude Code | 42% (5/12)Claude Code | 16 / 12 | 95% CI: 14%–56% vs 19%–68% | Unclear | More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 10x10 nonogram | 100% (4/4)Claude Code | 100% (3/3)Claude Code | 4 / 3 | 95% CI: 51%–100% vs 44%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 51% to 100%; Claude Opus 5.5 44% to 100%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Sudoku, 22 givens | 0% (0/4)Claude Code | 0% (0/3)Claude Code | 4 / 3 | 95% CI: 0%–49% vs 0%–56% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 0% to 49%; Claude Opus 5.5 0% to 56%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 6x6 Skyscrapers | 0% (0/4)Claude Code | 33% (1/3)Claude Code | 4 / 3 | 95% CI: 0%–49% vs 6.2%–79% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 0% to 49%; Claude Opus 5.5 6% to 79%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Seeded shuffle output | 50% (2/4)Claude Code | 33% (1/3)Claude Code | 4 / 3 | 95% CI: 15%–85% vs 6.2%–79% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 15% to 85%; Claude Opus 5.5 6% to 79%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Total time per call on harder tasks | 70.4 sClaude Code | 80.3 sClaude Code | 12 / 9 | range: 4.3 s–210 s vs 3.8 s–280 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 4.32 s to 210.1 s; Claude Opus 5.5 3.82 s to 279.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Output tokens per call on harder tasks (Output tokens) | 9,287Claude Code | 8,420Claude Code | 12 / 9 | range: 407–27,921 vs 279–40,044 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| List-price cost per strict pass on harder tasks (calculation)Calculation | $0.24Claude Code | $0.59Claude Code | 16 / 12 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.24 vs $0.59, 2.5x) is not tested against run-to-run variation. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
66 rows from 10 studies. No row separates them: 16 ties, 50 unclear.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick Claude Sonnet 5.5
- List-price cost per attempt (calculation): $2.88 vs $7.59. A list-price calculation, not a measured difference. Calculation
- List-price cost per call (calculation): $0.0036 vs $0.0069. A list-price calculation, not a measured difference. Calculation
- List-price cost per passing answer (calculation): $0.0062 vs $0.010. A list-price calculation, not a measured difference. Calculation
- List-price cost per strict pass on hard tasks (calculation): $0.014 vs $0.028. A list-price calculation, not a measured difference. Calculation
- List-price cost per passing coding session (calculation): $0.085 vs $0.22. A list-price calculation, not a measured difference. Calculation
- List-price cost per strict pass by effort (calculation): $0.014 vs $0.029. A list-price calculation, not a measured difference. Calculation
- List-price cost of 5-question sessions with and without the cache (calculation) (With the cache, as recorded): $0.14 vs $0.26. A list-price calculation, not a measured difference. Calculation
- List-price cost of 5-question sessions with and without the cache (calculation) (Without a cache: every input token at the input price): $0.27 vs $0.54. A list-price calculation, not a measured difference. Calculation
- Cost of a reused prefix with and without the cache, by session length (calculation): 1 turn: $15.66 vs $31.31. A list-price calculation, not a measured difference. Calculation
- Cost of a reused prefix with and without the cache, by session length (calculation): 2 turns: $31.32 vs $62.62. A list-price calculation, not a measured difference. Calculation
- Cost of a reused prefix with and without the cache, by session length (calculation): 3 turns: $46.99 vs $93.94. A list-price calculation, not a measured difference. Calculation
- Cost of a reused prefix with and without the cache, by session length (calculation): 5 turns: $78.31 vs $156.56. A list-price calculation, not a measured difference. Calculation
- Cost of a reused prefix with and without the cache, by session length (calculation): 10 turns: $156.62 vs $313.12. A list-price calculation, not a measured difference. Calculation
- Cost of a reused prefix with and without the cache, by session length (calculation): 20 turns: $313.24 vs $626.24. A list-price calculation, not a measured difference. Calculation
- One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (No cache (the same either way)): $156.62 vs $313.12. A list-price calculation, not a measured difference. Calculation
- One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (One 10-turn session, 1-hour cache): $45.42 vs $76.71. A list-price calculation, not a measured difference. Calculation
- One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation) (Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix)): $313.24 vs $626.24. A list-price calculation, not a measured difference. Calculation
- List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0067 vs $0.013. A list-price calculation, not a measured difference. Calculation
- List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)): $0.0037 vs $0.0080. A list-price calculation, not a measured difference. Calculation
- List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)): $0.0040 vs $0.0077. A list-price calculation, not a measured difference. Calculation
- Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass): $0.0063 vs $0.013. A list-price calculation, not a measured difference. Calculation
- Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass): $0.014 vs $0.029. A list-price calculation, not a measured difference. Calculation
- List-price cost per strict pass on harder tasks (calculation): $0.24 vs $0.59. A list-price calculation, not a measured difference. Calculation
When to pick Claude Opus 5.5
- Time to first text as the prompt grows: 64k: 1.79 s vs 3.07 s. A list-price calculation, not a measured difference. Calculation
Side by side
The study charts, showing only these two. Open a study for every configuration.
| Item | Resolved | 95% interval | n |
|---|---|---|---|
| Claude Opus 5.5 (Agent, new build) | 67% | 21%–94% | 3 |
| Claude Sonnet 5.5 (Agent, older builds) | 33% | 6.2%–79% | 3 |
2 rows. Highest Claude Opus 5.5 (Agent, new build) 67% (95% interval 21%–94%, n 3). Lowest Claude Sonnet 5.5 (Agent, older builds) 33% (95% interval 6.2%–79%, n 3). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 3 per row
One attempt per arm per instance, official grader · 95% Wilson intervals
Interim: 3 of 8 declared pairs are graded; 5 were never started; no resumed attempts are included here. With n = 3 the intervals span most of the axis, so this chart supports no ranking. The instances were chosen to hold Sonnet misses and resolves in equal numbers, so neither rate estimates SWE-bench Verified as a whole.
Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)
| Item | List-price cost per attempt | n |
|---|---|---|
| Claude Opus 5.5 (Agent, new build) | $7.59 | 3 |
| Claude Sonnet 5.5 (Agent, older builds) | $2.88 | 3 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 (Agent, new build) $7.59 (n 3). Lowest Claude Sonnet 5.5 (Agent, older builds) $2.88 (n 3).
Notesn = 3 per row
Mean over the same 3 instances; notional cost from the platform price table
Calculation, not an invoice: the platform price table at each run commit applied to the recorded tokens of subscription runs (Opus 5.5: $4 input, $20 output, $0.20 cache read per million tokens). Totals $22.76 vs $8.64: 2.6×, a ratio of two calculations.
Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Repricing calculation, Anthropic list prices (Claude models)
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Worker minutes per attempt | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 (Agent, new build) | 20.3 min | 9.8 min–25.3 min | 3 |
| Claude Sonnet 5.5 (Agent, older builds) | 9.4 min | 4.7 min–15 min | 3 |
2 rows. Slowest Claude Opus 5.5 (Agent, new build) 20.3 min (range 9.8 min–25.3 min, n 3). Fastest Claude Sonnet 5.5 (Agent, older builds) 9.4 min (range 4.7 min–15 min, n 3). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Median minutes; whiskers = fastest and slowest of 3 attempts (not an interval)
Worker minutes from the run reports. Opus attempts ran one at a time on the newer build; the Sonnet attempts ran in earlier campaigns. 3 attempts per arm is too few to call a difference; a range is not a confidence interval.
Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)
| Item | Pass rate | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 80% | 55%–93% | 15 |
| Claude Opus 5.5 (high) · Claude Code | 100% | 80%–100% | 15 |
2 rows. Highest Claude Opus 5.5 (high) · Claude Code 100% (95% interval 80%–100%, n 15). Lowest Claude Sonnet 5.5 · Claude Code 80% (95% interval 55%–93%, n 15). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 15 per row
Every call counts; failures and timeouts are non-passes
Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.
Source: Provider head-to-head: Claude Code models vs Codex efforts
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 2.3 s | 2.2 s–7.7 s | 15 |
| Claude Opus 5.5 · Claude Code | 2.8 s | 2.5 s–8.9 s | 15 |
2 rows. Slowest Claude Opus 5.5 · Claude Code 2.8 s (range 2.5 s–8.9 s, n 15). Fastest Claude Sonnet 5.5 · Claude Code 2.3 s (range 2.2 s–7.7 s, n 15). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 15 per row
Median per configuration; whiskers = fastest and slowest call
One host, one network, one day. Whiskers are a range, not a confidence interval.
Source: Provider head-to-head: Claude Code models vs Codex efforts
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first useful output | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1.6 s | 1 s–6.4 s | 15 |
| Claude Opus 5.5 · Claude Code | 1.9 s | 1.6 s–7.2 s | 15 |
2 rows. Slowest Claude Opus 5.5 · Claude Code 1.9 s (range 1.6 s–7.2 s, n 15). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1 s–6.4 s, n 15). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 15 per row
Median per configuration; whiskers = fastest and slowest call
One host, one network, one day. Whiskers are a range, not a confidence interval.
Source: Provider head-to-head: Claude Code models vs Codex efforts
- Cache read
- Other input
Totals are the sum of the parts shown. Shares are calculated from the same values.
| Item | Cache read | Other input | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1,401 | 685 | 15 |
| Claude Opus 5.5 · Claude Code | 1,401 | 680 | 15 |
2 rows, 2 series: Cache read, Other input. Cache read: all at 1,401. Other input: highest Claude Sonnet 5.5 · Claude Code 685 (n 15). Lowest Claude Opus 5.5 · Claude Code 680 (n 15).
Notesn = 15 per row
Mean per call, split into prompt-cache reads and other input
The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.
Source: Provider head-to-head: Claude Code models vs Codex efforts
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 107 | 0 | 15 |
| Claude Opus 5.5 · Claude Code | 64 | 0 | 15 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 · Claude Code 107 (n 15). Lowest Claude Opus 5.5 · Claude Code 64 (n 15). Reasoning tokens: all at 0.
Notesn = 15 per row
Median per configuration; reasoning tokens where the CLI reports them
Vendors count reasoning differently; compare within a vendor first. A median of 0 can mean the CLI did not report reasoning for most calls.
Source: Provider head-to-head: Claude Code models vs Codex efforts
| Item | Cost per call | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.0036 | $0.0034–$0.01 | 15 |
| Claude Opus 5.5 · Claude Code | $0.0069 | $0.0059–$0.022 | 15 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 · Claude Code $0.0069 (range $0.0059–$0.022, n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.0036 (range $0.0034–$0.01, n 15). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 15 per row
Reported tokens × list price; the calls ran on subscriptions
Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.
Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
| Item | Cost per pass | n |
|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.0062 | 15 |
| Claude Opus 5.5 · Claude Code | $0.01 | 15 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 · Claude Code $0.01 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.0062 (n 15).
Notesn = 15 per row
All calls in a configuration, failures included, divided by its passes
Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.
Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Strict pass
- Lenient (format misses counted)
| Item | Strict pass | Lenient (format misses counted) | 95% interval | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 100% | 100% | Strict pass: 86%–100%; Lenient (format misses counted): 86%–100% | 24 |
| Claude Opus 5.5 · Claude Code | 100% | 100% | Strict pass: 86%–100%; Lenient (format misses counted): 86%–100% | 24 |
2 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: all at 100%. Lenient (format misses counted): all at 100%.
NotesWhiskers: 95% Wilson intervaln = 24 per row
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call on hard tasks (separate batches) | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 7.8 s | 2.3 s–34.8 s | 24 |
| Claude Opus 5.5 · Claude Code | 9.2 s | 4.2 s–27.2 s | 24 |
2 rows. Slowest Claude Opus 5.5 · Claude Code 9.2 s (range 4.2 s–27.2 s, n 24). Fastest Claude Sonnet 5.5 · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 24 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first useful output on hard tasks | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 6 s | 0.9 s–30.6 s | 24 |
| Claude Opus 5.5 · Claude Code | 6.8 s | 2.4 s–21.8 s | 24 |
2 rows. Slowest Claude Opus 5.5 · Claude Code 6.8 s (range 2.4 s–21.8 s, n 24). Fastest Claude Sonnet 5.5 · Claude Code 6 s (range 0.9 s–30.6 s, n 24). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 24 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1,050 | 585 | 24 |
| Claude Opus 5.5 · Claude Code | 945 | 529 | 24 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 · Claude Code 1,050 (n 24). Lowest Claude Opus 5.5 · Claude Code 945 (n 24). Reasoning tokens: highest Claude Sonnet 5.5 · Claude Code 585 (n 24). Lowest Claude Opus 5.5 · Claude Code 529 (n 24).
Notesn = 24 per row
Median per configuration; reasoning tokens as the CLI reports them
Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.014 | 24 |
| Claude Opus 5.5 · Claude Code | $0.028 | 24 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 · Claude Code $0.028 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).
Notesn = 24 per row
All calls in a configuration, failures and format misses included, divided by its strict passes
Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.
Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
| Item | Passed every hidden check | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 100% | 76%–100% | 12 |
| Claude Opus 5.5 · Claude Code | 100% | 76%–100% | 12 |
2 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 12 per row
A pass needs every hidden check · 12 sessions per agent · 95% Wilson intervals
6 tasks × 2 repetitions per agent. All 36 sessions passed, so the chart sits at its ceiling: this task set cannot rank the agents on quality. Before the first session, every base repository failed its hidden checks and every reference patch passed them.
Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Wall time per session | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 23.1 s | 18.7 s–44.5 s | 12 |
| Claude Opus 5.5 · Claude Code | 56.9 s | 29.8 s–186 s | 12 |
2 rows. Slowest Claude Opus 5.5 · Claude Code 56.9 s (range 29.8 s–186 s, n 12). Fastest Claude Sonnet 5.5 · Claude Code 23.1 s (range 18.7 s–44.5 s, n 12). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 12 per row
Median wall time; whiskers = fastest and slowest of 12 sessions (not an interval)
CLI process start to exit. One session at a time per lane; the Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts. Codex time includes reading the tester’s notes, writing a work log and trying to commit. A range is not a confidence interval.
Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks
| Item | Tool calls per session | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 7.5 | 3–14 | 12 |
| Claude Opus 5.5 · Claude Code | 7.5 | 5–14 | 12 |
2 rows. All at 7.5.
NotesLines: lowest–highest run (not an interval)n = 12 per row
Median; whiskers = fewest and most of 12 sessions (not an interval)
Claude Code counts its tool calls (Bash, Read, Edit, Write, Glob, Grep). Codex CLI counts shell commands and file changes; it has no separate read tool, so it reads files with shell commands. Turns are not compared: Codex reports one turn per run.
Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks
| Item | List-price cost per pass | n |
|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.085 | 12 |
| Claude Opus 5.5 · Claude Code | $0.22 | 12 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 · Claude Code $0.22 (n 12). Lowest Claude Sonnet 5.5 · Claude Code $0.085 (n 12).
Notesn = 12 per row
Reported tokens of all 12 sessions × list price, divided by the passes
Calculation, not a bill: both CLIs ran on flat subscriptions. Claude cache writes are priced at 2× input, as in the other studies (Claude Code’s own estimate gives the same totals); Codex cached input at its cache-read price. Codex input includes its own system prompt and, here, the tester’s AGENTS.md.
Sources: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
| Item | Strict pass | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 (low) · Claude Code | 100% | 81%–100% | 16 |
| Claude Opus 5.5 (low) · Claude Code | 100% | 81%–100% | 16 |
2 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 16 per row
Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.
Source: Effort ladder: the hard task set at each effort level
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call by effort on hard tasks | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 8 s | 2.3 s–21.6 s | 16 |
| Claude Opus 5.5 · Claude Code | 9.2 s | 4.2 s–27.2 s | 16 |
2 rows. Slowest Claude Opus 5.5 · Claude Code 9.2 s (range 4.2 s–27.2 s, n 16). Fastest Claude Sonnet 5.5 · Claude Code 8 s (range 2.3 s–21.6 s, n 16). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 16 per row
Median per configuration; whiskers = fastest and slowest call
Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.
Source: Effort ladder: the hard task set at each effort level
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1,054 | 668 | 16 |
| Claude Opus 5.5 · Claude Code | 945 | 538 | 16 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 · Claude Code 1,054 (n 16). Lowest Claude Opus 5.5 · Claude Code 945 (n 16). Reasoning tokens: highest Claude Sonnet 5.5 · Claude Code 668 (n 16). Lowest Claude Opus 5.5 · Claude Code 538 (n 16).
Notesn = 16 per row
Median per configuration; reasoning tokens as the CLI reports them
Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean "not reported". Their content is never captured. More tokens is not better or worse by itself.
Source: Effort ladder: the hard task set at each effort level
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.014 | 16 |
| Claude Opus 5.5 · Claude Code | $0.029 | 16 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 · Claude Code $0.029 (n 16). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 16).
Notesn = 16 per row
All calls in a configuration divided by its strict passes
Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.
Sources: Effort ladder: the hard task set at each effort level, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- With the cache, as recorded
- Without a cache: every input token at the input price (square)
Gap labels, Without a cache: every input token at the input price vs With the cache, as recorded: Without a cache: every input token at the input price is x% higher (+) or lower (−) than With the cache, as recorded, calculated from the two values shown (the change counted from With the cache, as recorded’s value).
| Item | With the cache, as recorded | Without a cache: every input token at the input price | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.14 | $0.27 | 15 |
| Claude Opus 5.5 · Claude Code | $0.26 | $0.54 | 15 |
List-price calculation, not a run. 2 rows, 2 series: With the cache, as recorded, Without a cache: every input token at the input price. With the cache, as recorded: highest Claude Opus 5.5 · Claude Code $0.26 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.14 (n 15). Without a cache: every input token at the input price: highest Claude Opus 5.5 · Claude Code $0.54 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.27 (n 15).
Notesn = 15 per row
All recorded turns per model; the same reported tokens priced two ways
Calculation, not a bill: the calls ran on a subscription. Cache reads at the cache-read price, 1-hour cache writes at 2× the input price (every write in this run was a 1-hour write). Codex CLI is not priced here: it reports no cache-write count.
Sources: Caching sessions and repeated prompts (Claude Code and Codex CLI), Cost with and without the prompt cache (calculation), Anthropic list prices (Claude models)
- Turn 1 (writes the ledger to the cache)
- Turns 2-5 (read the ledger from the cache)
| Item | Turn 1 (writes the ledger to the cache) | Turns 2-5 (read the ledger from the cache) | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1.6 s | 1.6 s | Turn 1 (writes the ledger to the cache): 1.6 s–1.8 s; Turns 2-5 (read the ledger from the cache): 1.4 s–5.6 s | 3 |
| Claude Opus 5.5 · Claude Code | 1.9 s | 2.4 s | Turn 1 (writes the ledger to the cache): 1.8 s–4.4 s; Turns 2-5 (read the ledger from the cache): 1.6 s–12.7 s | 3 |
2 rows, 2 series: Turn 1 (writes the ledger to the cache), Turns 2-5 (read the ledger from the cache). Turn 1 (writes the ledger to the cache): slowest Claude Opus 5.5 · Claude Code 1.9 s (range 1.8 s–4.4 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1.6 s–1.8 s, n 3). All run ranges overlap. Turns 2-5 (read the ledger from the cache): slowest Claude Opus 5.5 · Claude Code 2.4 s (range 1.6 s–12.7 s, n 12). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1.4 s–5.6 s, n 12). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 3–12 per row
Median; whiskers = fastest and slowest turn
Whiskers are a range (fastest and slowest turn), not a confidence interval. Turns ask different questions: the slow later turns are the counting question, which produced the most output.
Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)
- Claude Sonnet 5.5 (no cache)
- Claude Sonnet 5.5 (1-hour cache write)
- Claude Opus 5.5 (no cache)
- Claude Opus 5.5 (1-hour cache write, read $0.2 per M)
- Claude Opus 5.5 (1-hour cache write, read $0.4 per M)
| Turns in the session (requests that send the prefix) | Claude Sonnet 5.5 (no cache) | Claude Sonnet 5.5 (1-hour cache write) | Claude Opus 5.5 (no cache) | Claude Opus 5.5 (1-hour cache write, read $0.2 per M) | Claude Opus 5.5 (1-hour cache write, read $0.4 per M) |
|---|---|---|---|---|---|
| 1 turn | $15.66 | $31.32 | $31.31 | $62.62 | $62.62 |
| 2 turns | $31.32 | $32.89 | $62.62 | $64.19 | $65.76 |
| 3 turns | $46.99 | $34.46 | $93.94 | $65.76 | $68.89 |
| 5 turns | $78.31 | $37.59 | $157 | $68.89 | $75.15 |
| 10 turns | $157 | $45.42 | $313 | $76.71 | $90.80 |
| 20 turns | $313 | $61.08 | $626 | $92.37 | $122 |
List-price calculation, not a run. 6 rows, 5 series: Claude Sonnet 5.5 (no cache), Claude Sonnet 5.5 (1-hour cache write), Claude Opus 5.5 (no cache), Claude Opus 5.5 (1-hour cache write, read $0.2 per M), Claude Opus 5.5 (1-hour cache write, read $0.4 per M). Claude Sonnet 5.5 (no cache): highest 20 turns $313. Lowest 1 turn $15.66. Claude Sonnet 5.5 (1-hour cache write): highest 20 turns $61.08. Lowest 1 turn $31.32.
Notes
Prefix of 7,831 tokens (Sonnet 5.5) and 7,828 tokens (Opus 5.5), the rounded mean turn-1 input; n = 3 Sonnet sessions (range 7,831 to 7,832) and 3 Opus sessions (the same in every session); USD per 1,000 sessions
Calculation, not a bill. The chart multiplies list prices by a prefix of 7,831 tokens (Sonnet) or 7,828 tokens (Opus). It uses a 1-hour write at 2× input and treats the whole prefix as new. It leaves out output and the tokens that each turn adds. The cache costs more at 1 to 2 turns and less from 3. The second Opus line uses a $0.4 cache read, which another table of the product lists. We did not check the vendor price. The table adds the 5-minute write.
Sources: Cost with and without the prompt cache (calculation), Caching sessions and repeated prompts (Claude Code and Codex CLI), Anthropic list prices (Claude models), OpenAI list prices
No cache (the same either way)
One 10-turn session, 1-hour cache
Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix)
One panel per series, all on the same axis.
| Item | No cache (the same either way) | One 10-turn session, 1-hour cache | Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix) |
|---|---|---|---|
| Claude Sonnet 5.5 | $157 | $45.42 | $313 |
| Claude Opus 5.5 (cache read $0.2 per M) | $313 | $76.71 | $626 |
List-price calculation, not a run. 2 rows, 3 series: No cache (the same either way), One 10-turn session, 1-hour cache, Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix). No cache (the same either way): highest Claude Opus 5.5 (cache read $0.2 per M) $313. Lowest Claude Sonnet 5.5 $157. One 10-turn session, 1-hour cache: highest Claude Opus 5.5 (cache read $0.2 per M) $76.71. Lowest Claude Sonnet 5.5 $45.42.
Notes
USD per 1,000 workloads of 10 requests that send the same prefix; each recorded session started in a new temporary folder, n = 4 later sessions; 0 met the read-more-than-write proxy, 95% Wilson 0% to 49%
Calculation, not a bill: the same prefix, prices and 1-hour write as the cost-curve chart. Each recorded session ran in a new temporary working folder. On turn 1, 0 of 4 later sessions read more tokens than they wrote. This proxy cannot identify which earlier call supplied a cache entry. The split calculation assumes no shared reuse and a wholly new prefix, so it charges ten full writes. With reuse across sessions they would cost the same as one 10-turn session. The cause of the missing reuse was not tested here. The caching study protocol names the random temporary folder as one untested hypothesis. The table adds the 5-minute variants.
Sources: Cost with and without the prompt cache (calculation), Caching sessions and repeated prompts (Claude Code and Codex CLI), Anthropic list prices (Claude models), OpenAI list prices
- Reasoning (output tokens)
- Remaining output (visible-answer estimate)
- Input (prompt, cache priced)
Totals are the sum of the parts shown. Shares are calculated from the same values.
| Item | Reasoning (output tokens) | Remaining output (visible-answer estimate) | Input (prompt, cache priced) | n |
|---|---|---|---|---|
| Claude Opus 5.5 · Claude Code | $0.013 | $0.008 | $0.0077 | 24 |
| Claude Sonnet 5.5 · Claude Code | $0.0067 | $0.0037 | $0.004 | 24 |
List-price calculation, not a run. 2 rows, 3 series: Reasoning (output tokens), Remaining output (visible-answer estimate), Input (prompt, cache priced). Reasoning (output tokens): highest Claude Opus 5.5 · Claude Code $0.013 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.0067 (n 24). Remaining output (visible-answer estimate): highest Claude Opus 5.5 · Claude Code $0.008 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.0037 (n 24).
Notesn = 24 per row
Mean per call on the hard tasks; the three parts add up to the call
Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
- Reasoning cost per strict pass
- Total cost per strict pass (square)
Gap labels, Total cost per strict pass vs Reasoning cost per strict pass: Total cost per strict pass is x% higher (+) or lower (−) than Reasoning cost per strict pass, calculated from the two values shown (the change counted from Reasoning cost per strict pass’s value).
| Item | Reasoning cost per strict pass | Total cost per strict pass | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.0063 | $0.014 | 16 |
| Claude Opus 5.5 · Claude Code | $0.013 | $0.029 | 16 |
List-price calculation, not a run. 2 rows, 2 series: Reasoning cost per strict pass, Total cost per strict pass. Reasoning cost per strict pass: highest Claude Opus 5.5 · Claude Code $0.013 (n 16). Lowest Claude Sonnet 5.5 · Claude Code $0.0063 (n 16). Total cost per strict pass: highest Claude Opus 5.5 · Claude Code $0.029 (n 16). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 16).
Notesn = 16 per row
List price ÷ strict passes; every cell is 8 tasks × 2 repetitions
Calculation, not a bill: reported tokens × list price, divided by the cell's strict passes; the calls ran on flat subscriptions. The effort-ladder cells: new calls plus reference cells reused from the hard head-to-head (Claude repetitions 1-2 only). "Default" means the effort flag was not passed. The total is the same value as the effort-ladder cost-per-pass chart. Effort levels are not the same scale across vendors, and the reference cells ran in a different batch and hour.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
- Eight hard tasks
- Five short tasks (square)
Gap labels, Five short tasks vs Eight hard tasks: Five short tasks is x percentage points higher (+) or lower (−) than Eight hard tasks, calculated from the two values shown; lines are the lowest–highest run (not an interval).
| Item | Eight hard tasks | Five short tasks | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 55% | 0% | Eight hard tasks: 0%–96%; Five short tasks: 0%–73% | 24 |
| Claude Opus 5.5 · Claude Code | 55% | 0% | Eight hard tasks: 30%–96%; Five short tasks: 0%–93% | 24 |
List-price calculation, not a run. 2 rows, 2 series: Eight hard tasks, Five short tasks. Eight hard tasks: highest Claude Opus 5.5 · Claude Code 55% (range 30%–96%, n 24). Lowest Claude Sonnet 5.5 · Claude Code 55% (range 0%–96%, n 24). All run ranges overlap. Five short tasks: all at 0%.
NotesLines: lowest–highest run (not an interval)n 15–24 per row
Median call per configuration; five short tasks and eight hard tasks
Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first text | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 2 s | 0.9 s–4.1 s | 4 |
| Claude Opus 5.5 · Claude Code | 2 s | 1.7 s–2.4 s | 4 |
2 rows. Slowest Claude Opus 5.5 · Claude Code 2 s (range 1.7 s–2.4 s, n 4). Fastest Claude Sonnet 5.5 · Claude Code 2 s (range 0.9 s–4.1 s, n 4). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = fastest and slowest call
Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.
Source: LLM speed anatomy
| Item | Visible tokens per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 232 | 230–233 | 4 |
| Claude Opus 5.5 · Claude Code | 156 | 155–156 | 4 |
List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 · Claude Code 232 (range 230–233, n 4). Lowest Claude Opus 5.5 · Claude Code 156 (range 155–156, n 4). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = slowest and fastest call
Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.
Source: LLM speed anatomy
| Item | Characters per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 517 | 513–519 | 4 |
| Claude Opus 5.5 · Claude Code | 347 | 345–349 | 4 |
List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 · Claude Code 517 (range 513–519, n 4). Lowest Claude Opus 5.5 · Claude Code 347 (range 345–349, n 4). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call
Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
- Claude Haiku 4.5 · Claude Code
- Claude Sonnet 5.5 · Claude Code
- Claude Opus 5.5 · Claude Code
- GPT-6.1 Sol (low) · Codex CLI
| Prompt-size target (approximate Haiku tokens; calibration calculation) | Claude Haiku 4.5 · Claude Code | Claude Sonnet 5.5 · Claude Code | Claude Opus 5.5 · Claude Code | GPT-6.1 Sol (low) · Codex CLI | Range (lowest–highest run) | n |
|---|---|---|---|---|---|---|
| 1k | 1.9 s | 1.5 s | 1.5 s | 3.4 s | Claude Haiku 4.5 · Claude Code: 1.9 s–2 s; Claude Sonnet 5.5 · Claude Code: 1.2 s–1.7 s; Claude Opus 5.5 · Claude Code: 1.5 s–2 s; GPT-6.1 Sol (low) · Codex CLI: 3.4 s–4.8 s | 3 |
| 16k | 2.3 s | 1.8 s | 1.7 s | 4 s | Claude Haiku 4.5 · Claude Code: 2.2 s–2.5 s; Claude Sonnet 5.5 · Claude Code: 1.6 s–2.1 s; Claude Opus 5.5 · Claude Code: 1.7 s–3 s; GPT-6.1 Sol (low) · Codex CLI: 3.3 s–4.3 s | 3 |
| 64k | 2.8 s | 3.1 s | 1.8 s | 3.9 s | Claude Haiku 4.5 · Claude Code: 2.5 s–2.9 s; Claude Sonnet 5.5 · Claude Code: 1.4 s–3.6 s; Claude Opus 5.5 · Claude Code: 1.7 s–3.7 s; GPT-6.1 Sol (low) · Codex CLI: 3.4 s–4.4 s | 3 |
List-price calculation, not a run. 3 rows, 4 series: Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (low) · Codex CLI. Claude Haiku 4.5 · Claude Code: slowest 64k 2.8 s (range 2.5 s–2.9 s, n 3). Fastest 1k 1.9 s (range 1.9 s–2 s, n 3). Not all run ranges overlap. Claude Sonnet 5.5 · Claude Code: slowest 64k 3.1 s (range 1.4 s–3.6 s, n 3). Fastest 1k 1.5 s (range 1.2 s–1.7 s, n 3). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Median of 3 calls per size; whiskers = fastest and slowest call
Each call used a new ledger seed. Cache-read counts stayed within the short-prompt baseline (see the cache table). This does not identify which tokens were cached. Sizes name the text we send; each model’s reported input tokens are in the table and include the CLI’s own prefix. The size calibration subtracts estimated prefixes from probe input counts; these are calculations, not measured prefix counts for each call. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
1k prompt
16k prompt
64k prompt
One panel per series, all on the same axis; whiskers are the fastest–slowest run (not an interval).
| Item | 1k prompt | 16k prompt | 64k prompt | Range (lowest–highest run) | n |
|---|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1.8 s | 2.1 s | 3.4 s | 1k prompt: 1.6 s–2.1 s; 16k prompt: 2 s–2.5 s; 64k prompt: 1.7 s–4.4 s | 3 |
| Claude Opus 5.5 · Claude Code | 1.8 s | 2.4 s | 2.4 s | 1k prompt: 1.8 s–2.4 s; 16k prompt: 2.1 s–3.4 s; 64k prompt: 2.3 s–4.3 s | 3 |
2 rows, 3 series: 1k prompt, 16k prompt, 64k prompt. 1k prompt: slowest Claude Opus 5.5 · Claude Code 1.8 s (range 1.8 s–2.4 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 1.8 s (range 1.6 s–2.1 s, n 3). All run ranges overlap. 16k prompt: slowest Claude Opus 5.5 · Claude Code 2.4 s (range 2.1 s–3.4 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 2.1 s (range 2 s–2.5 s, n 3). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Median of 3 calls per bar; whiskers = fastest and slowest call
Whole call: CLI start-up, first text and the one-line answer. Whiskers are a range of calls, not a confidence interval. Each call used a new ledger.
Source: LLM speed anatomy
| Item | Exact answer | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 100% | 70%–100% | 9 |
| Claude Opus 5.5 · Claude Code | 56% | 27%–81% | 9 |
2 rows. Highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 70%–100%, n 9). Lowest Claude Opus 5.5 · Claude Code 56% (95% interval 27%–81%, n 9). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 9 per row
All sizes together per model; whiskers = 95% Wilson intervals
Whiskers are 95% Wilson intervals. One lookup question per call; a reply with extra words is a format miss, not a pass. With 9 calls per model, a perfect score still has a wide interval.
Source: LLM speed anatomy
- Strict pass
- Lenient (format misses counted)
| Item | Strict pass | Lenient (format misses counted) | 95% interval | n |
|---|---|---|---|---|
| Claude Opus 5.5 · Claude Code | 42% | 50% | Strict pass: 19%–68%; Lenient (format misses counted): 25%–75% | 12 |
| Claude Sonnet 5.5 · Claude Code | 38% | 38% | Strict pass: 18%–61%; Lenient (format misses counted): 18%–61% | 16 |
2 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Opus 5.5 · Claude Code 42% (95% interval 19%–68%, n 12). Lowest Claude Sonnet 5.5 · Claude Code 38% (95% interval 18%–61%, n 16). All intervals overlap. Lenient (format misses counted): highest Claude Opus 5.5 · Claude Code 50% (95% interval 25%–75%, n 12). Lowest Claude Sonnet 5.5 · Claude Code 38% (95% interval 18%–61%, n 16). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 12–16 per row
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row. Counted calls are new calls.
Source: Harder tasks head-to-head
| Item | Tool attempt | 95% interval | n |
|---|---|---|---|
| Claude Opus 5.5 · Claude Code | 42% | 19%–68% | 12 |
| Claude Sonnet 5.5 · Claude Code | 31% | 14%–56% | 16 |
2 rows. Highest Claude Opus 5.5 · Claude Code 42% (95% interval 19%–68%, n 12). Lowest Claude Sonnet 5.5 · Claude Code 31% (95% interval 14%–56%, n 16). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 12–16 per row
Share of calls whose reply held tool-call markup or whose tool call the CLI could not parse
Every prompt asked for the answer only. Both CLIs disabled tools. The Codex runner adds a developer instruction not to call tools. The Claude Code runner adds no such instruction. The validator decides the score. Tool-call markup is a separate flag; a flagged call can still pass if its whole reply meets the validator. It is a behaviour, not a quality score: more is not better. No call with a tool attempt passed. Whiskers are 95% Wilson intervals. The flag comes from the reply text, which is not published.
Source: Harder tasks head-to-head
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | GPT-6.1 Sol (medium) · Codex CLI | Claude Opus 5.5 · Claude Code | Claude Sonnet 5.5 · Claude Code | Claude Haiku 4.5 · Claude Code | 95% interval | n |
|---|---|---|---|---|---|---|
| 10x10 nonogram | 75% | 100% | 100% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 30%–95%; Claude Opus 5.5 · Claude Code: 44%–100%; Claude Sonnet 5.5 · Claude Code: 51%–100%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
| Sudoku, 22 givens | 25% | 0% | 0% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 4.6%–70%; Claude Opus 5.5 · Claude Code: 0%–56%; Claude Sonnet 5.5 · Claude Code: 0%–49%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
| 6x6 Skyscrapers | 100% | 33% | 0% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 51%–100%; Claude Opus 5.5 · Claude Code: 6.2%–79%; Claude Sonnet 5.5 · Claude Code: 0%–49%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
| Seeded shuffle output | 75% | 33% | 50% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 30%–95%; Claude Opus 5.5 · Claude Code: 6.2%–79%; Claude Sonnet 5.5 · Claude Code: 15%–85%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
4 rows, 4 series: GPT-6.1 Sol (medium) · Codex CLI, Claude Opus 5.5 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Haiku 4.5 · Claude Code. GPT-6.1 Sol (medium) · Codex CLI: highest 6x6 Skyscrapers 100% (95% interval 51%–100%, n 4). Lowest Sudoku, 22 givens 25% (95% interval 4.6%–70%, n 4). All intervals overlap. Claude Opus 5.5 · Claude Code: highest 10x10 nonogram 100% (95% interval 44%–100%, n 3). Lowest Sudoku, 22 givens 0% (95% interval 0%–56%, n 3). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 3–4 per row
One bar per configuration and task; each bar rests on only a few calls, so the 95% intervals are wide
Each bar is 3 to 4 calls, so one call moves a bar by a quarter or a third. Whiskers are 95% Wilson intervals; with this few calls they overlap almost everywhere.
Source: Harder tasks head-to-head
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 · Claude Code | 80.3 s | 3.8 s–280 s | 9 |
| Claude Sonnet 5.5 · Claude Code | 70.4 s | 4.3 s–210 s | 12 |
2 rows. Slowest Claude Opus 5.5 · Claude Code 80.3 s (range 3.8 s–280 s, n 9). Fastest Claude Sonnet 5.5 · Claude Code 70.4 s (range 4.3 s–210 s, n 12). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 9–12 per row
Median per configuration; whiskers = fastest and slowest call
Median and range over the calls that completed. Completed calls include wrong answers and format misses. Only timeouts and tool-call parse errors are excluded from this run’s timings. Both count as non-passes in the outcomes chart. One Mac, one network, one session. Arena servers shared the Mac during part of the run. Host load was not controlled, so these times cannot isolate model speed. Whiskers are a range, not a confidence interval. Times include the CLI start-up and the CLI’s own system prompt. Highlighted: configurations that passed every call.
Source: Harder tasks head-to-head
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Opus 5.5 · Claude Code | 8,420 | 8,352 | Output tokens: 279–40,044; Reasoning tokens: 21–9,897 | 9 |
| Claude Sonnet 5.5 · Claude Code | 9,287 | 6,557 | Output tokens: 407–27,921; Reasoning tokens: 63–27,902 | 12 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 · Claude Code 9,287 (range 407–27,921, n 12). Lowest Claude Opus 5.5 · Claude Code 8,420 (range 279–40,044, n 9). All run ranges overlap. Reasoning tokens: highest Claude Opus 5.5 · Claude Code 8,352 (range 21–9,897, n 9). Lowest Claude Sonnet 5.5 · Claude Code 6,557 (range 63–27,902, n 12). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 9–12 per row
Median per configuration; reasoning tokens as the CLI reports them
Medians and minimum-to-maximum token ranges cover completed calls only. Ranges are not confidence intervals. The chart omits unknown reasoning counts. Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. Claude Code used an output-token cap setting of 16,000. Some reported totals exceeded it. Codex CLI had no cap. More tokens is not better or worse by itself.
Source: Harder tasks head-to-head
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.24 | 16 |
| Claude Opus 5.5 · Claude Code | $0.59 | 12 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 · Claude Code $0.59 (n 12). Lowest Claude Sonnet 5.5 · Claude Code $0.24 (n 16).
Notesn 12–16 per row
All calls in a configuration, failures and format misses included, divided by its strict passes
Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Timeout calls report no tokens and are not priced. Unpriced calls: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16. These cells show a lower bound. Assume each unpriced call cost its cell’s median priced call. This sensitivity calculation gives GPT-6.1 Sol (medium) $0.100, Opus 5.5 $0.633 and Sonnet 5.5 $0.303. Opus 5.5 figures are provisional: its cache-read price is under re-check. Highlights mark the observed frontier of these lower-bound costs. Unknown timeout costs can change it; this is not a cost ranking. Claude Haiku 4.5 · Claude Code had no strict pass, so it has no cost per pass.
Sources: Harder tasks head-to-head, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Which is better, Claude Sonnet 5.5 or Claude Opus 5.5?
- Claude Sonnet 5.5 and Claude Opus 5.5 share 35 measured metrics and 31 list-price calculations from 10 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 16 ties and 50 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).
- How were Claude Sonnet 5.5 and Claude Opus 5.5 measured?
- They share 35 measured metrics and 31 list-price calculations from 10 public studies: Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim); Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks; Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks; Prompt caching and run-to-run consistency in Claude Code and Codex CLI; Prompt cache break-even: after how many reuses does a cached prefix cost less?; How much of an AI bill is thinking? Reasoning tokens by model and effort; Where the seconds go: first text, output speed and prompt size for 6 LLMs; GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Every row names its configuration, its sample size and its interval or range.
- How do Claude Sonnet 5.5 and Claude Opus 5.5 compare on resolved on the same 3 SWE-bench Verified instances (interim)?
- Claude Sonnet 5.5: 33% (1/3) (Agent · older builds · SWE-bench Verified, interim paired probe; n = 3; 95% interval 6.2% to 79%). Claude Opus 5.5: 67% (2/3) (Agent · new build · SWE-bench Verified, interim paired probe; n = 3; 95% interval 21% to 94%). The 95% intervals overlap (Claude Sonnet 5.5 6% to 79%; Claude Opus 5.5 21% to 94%), so this sample cannot separate them.
- How do Claude Sonnet 5.5 and Claude Opus 5.5 compare on worker time per attempt?
- Claude Sonnet 5.5: 9.4 min (Agent · older builds · SWE-bench Verified, interim paired probe; n = 3; run range 4.7 min to 15 min). Claude Opus 5.5: 20.3 min (Agent · new build · SWE-bench Verified, interim paired probe; n = 3; run range 9.8 min to 25.3 min). The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 4.7 min to 15.0 min; Claude Opus 5.5 9.8 min to 25.3 min); the medians alone do not show a reliable difference. A range is not a confidence interval.
- How do Claude Sonnet 5.5 and Claude Opus 5.5 compare on pass rate on five validated tasks?
- Claude Sonnet 5.5: 80% (12/15) (Claude Code · five short validated tasks; n = 15; 95% interval 55% to 93%). Claude Opus 5.5: 100% (15/15) (Claude Code · five short validated tasks; n = 15; 95% interval 80% to 100%). The 95% intervals overlap (Claude Sonnet 5.5 55% to 93%; Claude Opus 5.5 80% to 100%), so this sample cannot separate them.
- How do Claude Sonnet 5.5 and Claude Opus 5.5 compare on total time per call?
- Claude Sonnet 5.5: 2.31 s (Claude Code · five short validated tasks; n = 15; run range 2.2 s to 7.7 s). Claude Opus 5.5: 2.75 s (Claude Code · five short validated tasks; n = 15; run range 2.5 s to 8.9 s). The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.17 s to 7.73 s; Claude Opus 5.5 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
- How do Claude Sonnet 5.5 and Claude Opus 5.5 compare on time to first useful output?
- Claude Sonnet 5.5: 1.56 s (Claude Code · five short validated tasks; n = 15; run range 1 s to 6.4 s). Claude Opus 5.5: 1.92 s (Claude Code · five short validated tasks; n = 15; run range 1.6 s to 7.2 s). The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.99 s to 6.39 s; Claude Opus 5.5 1.56 s to 7.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
The studies behind this page
Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)
Interim: 3 of 8 paired SWE-bench Verified instances graded. Opus 5.5 resolved 2, Sonnet 5.5 1 (McNemar p = 1.0). Cost 2.6×, a list-price calculation.
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.
Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
36 graded sessions: Claude Code with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol. All passed every hidden test; time, tool calls and diffs differ.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
176 calls on 8 hard tasks at low, medium, high and default effort: Sonnet, Opus and GPT-6.1 Sol. Pass rate with 95% intervals, time, tokens, cost per pass.
Prompt caching and run-to-run consistency in Claude Code and Codex CLI
135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.
Prompt cache break-even: after how many reuses does a cached prefix cost less?
A calculation on list prices: reuses before a cached prompt prefix costs less, per Claude and GPT model, and cost per 1,000 sessions of 1 to 20 turns.
How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.