Which Claude model should you use? A task-by-task guide from our measurements
Which Claude model fits your job? Sonnet passed 24/24 hard calls (95% interval 86%–100%) at $0.01435 per pass (calculation). The top models hit a ceiling.
TL;DR
- Start with Claude Sonnet 5.5. It passed 24 of 24 calls on eight hard tasks (95% interval 86% to 100%) at $0.01435 per pass. Sonnet had the lowest recorded cost per pass in both head-to-heads (calculations).
- Opus 5.5 and Fable 5.1 also passed 24 of 24 (95% interval 86% to 100% each), at $0.02824 and $0.09331 per pass (calculations). Our tasks hit a ceiling, so we found no quality gain.
- Haiku 4.5 costs half as much per token, not per pass. It passed 11/24 hard calls (95% interval 28% to 65%). Repeated prompts gave 0/10 exact-number passes (0% to 28%) and 1/10 JSON passes (2% to 40%). Sonnet passed 10/10 on both repeated prompts (72% to 100% each).
- Start at low effort. All 11 effort cells passed 16 of 16 (95% interval 81% to 100% each). Picks are our reading of small samples.
Claude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say
66 comparison rows from 10 studies: 0 rows favour Sonnet 5.5, 0 favour Opus 5.5, 66 are ties or unclear. Cost rows are calculations.
Transcript
- Comparison · 66 rows · 10 studies. Sonnet 5.5 vs Opus 5.5. A winner only where the 95% intervals or run ranges do not overlap.
- 66 comparison rows from 10 studies. None separates them: intervals or run ranges overlap, or none was recorded. Rows where Sonnet 5.5 is ahead: 0 (of 66). Rows where Opus 5.5 is ahead: 0 (of 66). Ties or unclear: 66 (16 ties · 50 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
- SWE-bench pairs, interim: pass rate 33% vs 67%, tie: 95% intervals overlap. None of the 3 rows separates them. Table: SWE-bench pairs, interim · Agent · n = 3 per side. Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
- Five short tasks: pass rate 80% vs 100%, tie: 95% intervals overlap. None of the 8 rows separates them. Table: Five short tasks · Claude Code · n = 15 per side. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
- Eight hard tasks: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 6 rows separates them. Table: Eight hard tasks · Claude Code · n = 24 per side. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
- Coding agents, hidden tests: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Coding agents, hidden tests · Claude Code · n = 12 per side. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
- Effort ladder, default effort: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Effort ladder, default effort · Claude Code · n = 16 per side. Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
- Caching sessions: pass rate not measured. None of the 4 rows separates them. Table: Caching sessions · Claude Code · n = 15, 3, 12 per side. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- Prompt cache break-even: after how many reuses does a cached prefix cost less?: pass rate not measured. None of the 9 rows separates them. Table: Prompt cache break-even: after how many reuses does a cached prefix cost less? · no cache · n = per side. Caveat: The source has 30 attempted Claude turns, 0 failed turns and 0 turns without token usage. Failed turns with usage remain in cost totals. Missing usage cannot be priced. No quality rate or cache-caused speed effect is claimed.
- How much of an AI bill is thinking? Reasoning tokens by model and effort: pass rate not measured. None of the 8 rows separates them. Table: How much of an AI bill is thinking? Reasoning tokens by model and effort · Claude Code · n = 24, 16, 15 per side. Caveat: The calls used one shared Mac and network. Other work and provider load were not controlled. Latency includes those effects.
- Where the seconds go: first text, output speed and prompt size for 6 LLMs: pass rate 100% vs 56%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: Where the seconds go: first text, output speed and prompt size for 6 LLMs · Claude Code · n = 4, 3, 9 per side. Caveat: First text includes CLI start-up and reasoning time. It is not the API’s time to first token; a direct API call would skip the CLI start-up (the CLI reported ready after a median 0.33 s to 0.84 s per cell).
- GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks: pass rate 38% vs 42%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks · Claude Code · n = per side. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.
- No winner where the data shows none. Every row and its reason online.
The decision table
Rate brackets below are 95% intervals. Timing ranges are fastest to slowest, except routing, which shows p50-to-p95 bands. Routing p50 uses the nearest-rank method.
| Job | Our pick | The number | How sure |
|---|---|---|---|
| Typed routing or classification | Sonnet at low effort | 77/82 exact (94%, 86.5% to 97.4%); Haiku 73/82 (80.4% to 94.1%). p50 (nearest-rank) 2.60 s vs 12.54 s; p50-to-p95 bands 2.60–4.30 s vs 12.54–34.48 s | The accuracy intervals overlap; this sample cannot separate them. Sonnet ahead on time |
| Hard single-turn coding | Sonnet | 24/24 (86% to 100%) at $0.01435 per pass (calculation) | Tie with Opus and Fable at a ceiling |
| Short, latency-sensitive calls | Sonnet; test Fable on your calls | Fable median 1.94 s (1.41–9.83), Sonnet 2.31 s (2.17–7.73); n = 15 each | Ranges overlap |
| Exact formats and strict JSON | Sonnet, not Haiku | Haiku 0/10 (0%–28%) and 1/10 (2%–40%); Sonnet 10/10 (72%–100%) on each | Sonnet ahead; n = 10 |
| Work with project memory | Sonnet and a short curated file | 211-line handbook: Haiku 3/10 (11%–60%), Sonnet 15/15 (80%–100%) | Sonnet ahead on that row |
| Reasoning effort | Start at low | 11 of 11 cells 16/16 (81%–100%); Sonnet low 5.82 s median (2.78–19.96 s), n = 16 | Tie at a ceiling |
| Paying for Opus | Only if your validator shows a gain | No supported gain; 2.0x cost per hard pass (calculation) | A gap could hide |
| Using Haiku | Test it first | Half the uncached input/output token price; higher calculated cost per pass in the settings below | Calculation only |
Typed routing or classification: Sonnet at low effort
We gave Sonnet 5.5 and Haiku 4.5 the same 82 routing cases through the Claude Code CLI. Exact means every scored question in a case is right. Sonnet at low effort got 77 of 82 exactly right (94%, 95% interval 86.5% to 97.4%). Haiku got 73 of 82 (89%, 80.4% to 94.1%). The accuracy intervals overlap; this sample cannot separate them.
Time separates them. The p50 (nearest-rank) was 2.60 s per decision for Sonnet and 12.54 s for Haiku. The p50-to-p95 bands do not overlap: 2.60 to 4.30 s against 12.54 to 34.48 s. Sonnet is ahead on these p50-to-p95 bands. The full call ranges also do not overlap: Sonnet 1.99 to 5.58 s, Haiku 5.86 to 51.28 s.
Haiku ran with default thinking and Sonnet at low effort, so the gap mixes model and setting.
Cost per 1,000 decisions (a calculation): $4.996 Sonnet, $8.924 Haiku. The routing calculation prices cache writes at 1.25x input. Sonnet receipts report one-hour writes; their CLI cost estimate is $7.324 per 1,000 decisions. These are different cost assumptions, not bills.
A routing model costs far less. Jev 1.13, from TypeSafe, got 221 of 246 live calls exactly right (90%, case-level 95% interval 81.9% to 95.0%). That pools three repeats of the same 82 cases: 74/82, 73/82 and 74/82. Its interval uses n = 82; these are not 246 independent samples. Its $0.0337 per 1,000 decisions is a calculation from provider-reported input tokens. Its accuracy interval overlaps both Claude intervals, but we revised our cases against Jev's answers: a home advantage. Plain rules in code cost $0 per decision; we did not score their accuracy.
Hard single-turn coding: Sonnet
Eight hard tasks, strict validators, 24 calls per model. Sonnet 5.5, Opus 5.5 and Fable 5.1 each passed 24 of 24 (86% to 100%). Haiku passed 11 of 24 (46%, 28% to 65%).
Cost per strict pass (a calculation): Sonnet $0.01435, Opus $0.02824, Fable $0.09331. That is 2.0x and 6.5x Sonnet (calculations). The set has a ceiling, so this sample cannot rule out a quality gap. The 86% lower bound is for each model, not a confidence interval for their difference.
Short, latency-sensitive calls: Sonnet, then test Fable
On five easy tasks (15 calls per model), Fable 5.1 had the lowest median time: 1.94 s (range 1.41 to 9.83 s). Sonnet's median was 2.31 s (2.17 to 7.73 s). The ranges overlap, so speed does not separate them.
Fable cost $0.02054 per pass and Sonnet $0.00624 (a calculation, 3.3x). Fable passed 15/15 (95% interval 79.6% to 100%). Sonnet passed 12/15 (80%, 55% to 93%), and the intervals overlap. All three misses were one task: the right number plus extra working lines, which the exact-text validator rejects. Test a clearer format rule. We did not measure whether it fixes these misses.
Exact formats and strict JSON: Sonnet, not Haiku
We sent each prompt 10 times to Haiku 4.5 and to Sonnet 5.5. Haiku passed 0 of 10 on the exact-number prompt (95% interval 0% to 27.8%). It gave the same wrong number every time: 289, not 282.
Haiku passed 1 of 10 on the JSON prompt (1.8% to 40.4%). Nine replies held the right JSON in the wrong format, such as a code fence. Sonnet passed 10/10 on both prompts (72.3% to 100%), so Sonnet is ahead. If you use Haiku for strict output, add a validator and a repair step.
Work with project memory: Sonnet, and keep the file short
We gave Claude Code a small repository under eight memory conditions. Team-knowledge checks cover two facts the code does not show: a changelog rule and a late-fee rate. With no memory file, Sonnet passed 6 of 15 (40%, 19.8% to 64.3%).
With an 11-line curated file, Sonnet passed 15 of 15 (79.6% to 100%) and Haiku 8 of 10 (49.0% to 94.3%): a tie. With the same facts in a 211-line handbook, Sonnet still passed 15/15 and Haiku 3 of 10 (10.8% to 60.3%), so Sonnet is ahead. Our reading: keep the file short. Haiku's own 8/10 and 3/10 overlap, so it is a hint.
More: Does CLAUDE.md help agent memory?
Reasoning effort: start low
All 11 cells on the hard set passed 16 of 16 (81% to 100% each). These are Sonnet and Opus at four settings each, plus GPT-6.1 Sol at three. Sonnet at low effort had a median 5.82 s (range 2.78 to 19.96 s, n = 16). Its cost was $0.01219 per pass (calculation). At high effort it had a median 8.81 s (2.93 to 35.81 s, n = 16) and $0.01671 per pass (calculation). The call ranges overlap.
Some ladder cells reuse earlier hard-set calls; these are not independent new trials. The set has a ceiling, so effort may still matter on harder work.
More: Does reasoning effort buy quality?
When to pay for Opus: when your validator says so
These task sets did not show a supported quality gain for Opus. Both models passed 24/24 hard calls (95% interval 86% to 100% each). On short tasks, Opus passed 15/15 (79.6% to 100%), against Sonnet's 12/15 (54.8% to 93.0%); the intervals overlap. Our Sonnet vs Opus comparison holds 56 rows from nine studies: 26 measured metrics and 30 list-price calculations. None separates them: 10 ties and 46 unclear. Calculation rows are derived from recorded counts and prices; they are not new runs.
Opus 5.5 lists at $4 input and $20 output per million tokens, double Sonnet's $2 and $10. That fits the 2.0x per hard pass above.
The calculated price ratio shrinks on our recorded agent-loop tokens. Both models list cache reads at $0.20 per million. Our agent's recorded SWE-bench tokens from 33 attempts cost $143.83 at Opus prices and $87.23 at Sonnet prices. That is 1.6x (a calculation, not a run).
Our rule: run your tasks on both and compare cost per validator pass. More: Sonnet vs Opus: when is Opus worth the price?
When to use Haiku: after you test it
Haiku 4.5 lists at half Sonnet's per-token price: $1 input and $5 output per million tokens. Repricing the same tokens from 33 attempts gives Haiku half the calculated cost ($43.61 against $87.23, a calculation). That holds tokens and cache use fixed; it does not predict Haiku's outcomes. On the hard set it passed 11/24 (95% interval 28% to 65%), against Sonnet's 24/24 (86% to 100%). On the five short tasks Haiku passed 15/15 (79.6% to 100%).
Haiku had a higher calculated cost in the settings below. These calculations describe our samples. They do not establish population cost rankings:
- Short tasks: $0.00836 against $0.00624; n = 15 calls per model.
- Hard tasks: $0.0672 against $0.01435; n = 24 calls per model.
- Routing, per 1,000 decisions: $8.924 against $4.996; n = 82 calls per model. This is cost per decision, not per correct decision.
- Project memory, in all eight conditions (2 to 9 Haiku passes each). Curated file: $0.1095 against $0.0818. This uses full passes: Haiku 7/10 (95% interval 39.7% to 89.2%), Sonnet 15/15 (79.6% to 100%). The team-knowledge counts above use a different metric.
With the CLI default thinking, Haiku reported a median of 4,556 reasoning tokens per call on the hard set. Its median time was 39.01 s (range 15.27 to 75.13 s, n = 24). Sonnet had 7.75 s (2.26 to 34.79 s, n = 24); the ranges overlap. None of the studies above tested Haiku with thinking turned down.
How we measured
The Claude calls in the head-to-head, routing, consistency, effort and memory studies used the Claude Code CLI. Their timings include its start-up. A deterministic validator or test graded each result, and every failed call counts. Rates carry 95% Wilson intervals; times carry a range or a p50-to-p95 band, not an interval. "Ahead" means the intervals or ranges do not overlap. Costs are list-price calculations, not bills. Cost per pass divides the cost of all attempts, including failures, by the number of passes. These cost figures have no uncertainty intervals and do not establish population cost rankings.
Studies: hard tasks, short tasks, consistency, agent memory, routing, routing overhead, effort ladder, cost experiments.
Caveats
- Ceilings. Our tasks are too easy to separate the top models. Opus and Fable may win on work we did not test.
- Small samples. Head-to-head, consistency and effort cells hold 10 to 24 calls. Routing uses 82 cases per model. Memory pools several checks per session, so checks are not independent trials.
- Narrow tasks. Head-to-head, consistency and effort calls were single-turn with tools off. Routing uses typed cases, not general classification. Memory used file and shell tools in one small synthetic repository. Agent loops can differ.
- Prices and picks. List prices are the product price table effective 2026-09-21. Picks are our reading of the data.
Compare: Haiku vs Sonnet, Sonnet vs Opus, Opus vs Fable.
What to read next
- When tasks get hard: Haiku vs Sonnet vs Opus vs Fable
- How to estimate your AI coding bill and the AI cost calculator
Measure your own tasks
Try Agent keeps a receipt for each task: model, tokens, time, cost and validation result.