• Head to head
  • Claude Code
  • Claude Sonnet
  • Claude Opus

Which Claude model should you use? A task-by-task guide from our measurements

Which Claude model fits your job? Sonnet passed 24/24 hard calls (95% interval 86%–100%) at $0.01435 per pass (calculation). The top models hit a ceiling.

TL;DR

  • Start with Claude Sonnet 5.5. It passed 24 of 24 calls on eight hard tasks (95% interval 86% to 100%) at $0.01435 per pass. Sonnet had the lowest recorded cost per pass in both head-to-heads (calculations).
  • Opus 5.5 and Fable 5.1 also passed 24 of 24 (95% interval 86% to 100% each), at $0.02824 and $0.09331 per pass (calculations). Our tasks hit a ceiling, so we found no quality gain.
  • Haiku 4.5 costs half as much per token, not per pass. It passed 11/24 hard calls (95% interval 28% to 65%). Repeated prompts gave 0/10 exact-number passes (0% to 28%) and 1/10 JSON passes (2% to 40%). Sonnet passed 10/10 on both repeated prompts (72% to 100% each).
  • Start at low effort. All 11 effort cells passed 16 of 16 (95% interval 81% to 100% each). Picks are our reading of small samples.
Live story · 101 sClaude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say

Claude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say

66 comparison rows from 10 studies: 0 rows favour Sonnet 5.5, 0 favour Opus 5.5, 66 are ties or unclear. Cost rows are calculations.

Transcript
  1. Comparison · 66 rows · 10 studies. Sonnet 5.5 vs Opus 5.5. A winner only where the 95% intervals or run ranges do not overlap.
  2. 66 comparison rows from 10 studies. None separates them: intervals or run ranges overlap, or none was recorded. Rows where Sonnet 5.5 is ahead: 0 (of 66). Rows where Opus 5.5 is ahead: 0 (of 66). Ties or unclear: 66 (16 ties · 50 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
  3. SWE-bench pairs, interim: pass rate 33% vs 67%, tie: 95% intervals overlap. None of the 3 rows separates them. Table: SWE-bench pairs, interim · Agent · n = 3 per side. Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
  4. Five short tasks: pass rate 80% vs 100%, tie: 95% intervals overlap. None of the 8 rows separates them. Table: Five short tasks · Claude Code · n = 15 per side. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
  5. Eight hard tasks: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 6 rows separates them. Table: Eight hard tasks · Claude Code · n = 24 per side. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  6. Coding agents, hidden tests: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Coding agents, hidden tests · Claude Code · n = 12 per side. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
  7. Effort ladder, default effort: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Effort ladder, default effort · Claude Code · n = 16 per side. Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  8. Caching sessions: pass rate not measured. None of the 4 rows separates them. Table: Caching sessions · Claude Code · n = 15, 3, 12 per side. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  9. Prompt cache break-even: after how many reuses does a cached prefix cost less?: pass rate not measured. None of the 9 rows separates them. Table: Prompt cache break-even: after how many reuses does a cached prefix cost less? · no cache · n = per side. Caveat: The source has 30 attempted Claude turns, 0 failed turns and 0 turns without token usage. Failed turns with usage remain in cost totals. Missing usage cannot be priced. No quality rate or cache-caused speed effect is claimed.
  10. How much of an AI bill is thinking? Reasoning tokens by model and effort: pass rate not measured. None of the 8 rows separates them. Table: How much of an AI bill is thinking? Reasoning tokens by model and effort · Claude Code · n = 24, 16, 15 per side. Caveat: The calls used one shared Mac and network. Other work and provider load were not controlled. Latency includes those effects.
  11. Where the seconds go: first text, output speed and prompt size for 6 LLMs: pass rate 100% vs 56%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: Where the seconds go: first text, output speed and prompt size for 6 LLMs · Claude Code · n = 4, 3, 9 per side. Caveat: First text includes CLI start-up and reasoning time. It is not the API’s time to first token; a direct API call would skip the CLI start-up (the CLI reported ready after a median 0.33 s to 0.84 s per cell).
  12. GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks: pass rate 38% vs 42%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks · Claude Code · n = per side. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.
  13. No winner where the data shows none. Every row and its reason online.

The decision table

Rate brackets below are 95% intervals. Timing ranges are fastest to slowest, except routing, which shows p50-to-p95 bands. Routing p50 uses the nearest-rank method.

JobOur pickThe numberHow sure
Typed routing or classificationSonnet at low effort77/82 exact (94%, 86.5% to 97.4%); Haiku 73/82 (80.4% to 94.1%). p50 (nearest-rank) 2.60 s vs 12.54 s; p50-to-p95 bands 2.60–4.30 s vs 12.54–34.48 sThe accuracy intervals overlap; this sample cannot separate them. Sonnet ahead on time
Hard single-turn codingSonnet24/24 (86% to 100%) at $0.01435 per pass (calculation)Tie with Opus and Fable at a ceiling
Short, latency-sensitive callsSonnet; test Fable on your callsFable median 1.94 s (1.41–9.83), Sonnet 2.31 s (2.17–7.73); n = 15 eachRanges overlap
Exact formats and strict JSONSonnet, not HaikuHaiku 0/10 (0%–28%) and 1/10 (2%–40%); Sonnet 10/10 (72%–100%) on eachSonnet ahead; n = 10
Work with project memorySonnet and a short curated file211-line handbook: Haiku 3/10 (11%–60%), Sonnet 15/15 (80%–100%)Sonnet ahead on that row
Reasoning effortStart at low11 of 11 cells 16/16 (81%–100%); Sonnet low 5.82 s median (2.78–19.96 s), n = 16Tie at a ceiling
Paying for OpusOnly if your validator shows a gainNo supported gain; 2.0x cost per hard pass (calculation)A gap could hide
Using HaikuTest it firstHalf the uncached input/output token price; higher calculated cost per pass in the settings belowCalculation only

Typed routing or classification: Sonnet at low effort

We gave Sonnet 5.5 and Haiku 4.5 the same 82 routing cases through the Claude Code CLI. Exact means every scored question in a case is right. Sonnet at low effort got 77 of 82 exactly right (94%, 95% interval 86.5% to 97.4%). Haiku got 73 of 82 (89%, 80.4% to 94.1%). The accuracy intervals overlap; this sample cannot separate them.

Time separates them. The p50 (nearest-rank) was 2.60 s per decision for Sonnet and 12.54 s for Haiku. The p50-to-p95 bands do not overlap: 2.60 to 4.30 s against 12.54 to 34.48 s. Sonnet is ahead on these p50-to-p95 bands. The full call ranges also do not overlap: Sonnet 1.99 to 5.58 s, Haiku 5.86 to 51.28 s.

Haiku ran with default thinking and Sonnet at low effort, so the gap mixes model and setting.

Cost per 1,000 decisions (a calculation): $4.996 Sonnet, $8.924 Haiku. The routing calculation prices cache writes at 1.25x input. Sonnet receipts report one-hour writes; their CLI cost estimate is $7.324 per 1,000 decisions. These are different cost assumptions, not bills.

A routing model costs far less. Jev 1.13, from TypeSafe, got 221 of 246 live calls exactly right (90%, case-level 95% interval 81.9% to 95.0%). That pools three repeats of the same 82 cases: 74/82, 73/82 and 74/82. Its interval uses n = 82; these are not 246 independent samples. Its $0.0337 per 1,000 decisions is a calculation from provider-reported input tokens. Its accuracy interval overlaps both Claude intervals, but we revised our cases against Jev's answers: a home advantage. Plain rules in code cost $0 per decision; we did not score their accuracy.

Hard single-turn coding: Sonnet

Eight hard tasks, strict validators, 24 calls per model. Sonnet 5.5, Opus 5.5 and Fable 5.1 each passed 24 of 24 (86% to 100%). Haiku passed 11 of 24 (46%, 28% to 65%).

Cost per strict pass (a calculation): Sonnet $0.01435, Opus $0.02824, Fable $0.09331. That is 2.0x and 6.5x Sonnet (calculations). The set has a ceiling, so this sample cannot rule out a quality gap. The 86% lower bound is for each model, not a confidence interval for their difference.

Calculation
  • Claude Code
  • Codex CLI
Better: upper left

Haloed: on the frontier (1 of 7). A point in the shaded area is no better on either axis than a haloed point.

List-price calculation, not a run. 7 points: Strict pass rate against USD per strict pass (list-price calculation). USD per strict pass (list-price calculation) runs from $0.014 to $0.093; Strict pass rate from 46% to 100%. Highlighted: Claude Sonnet 5.5 · Claude Code.

Notesn 16–24 per point

Strict pass rate against list-price cost per strict pass

Upper-left is better. Highlighted points are on the frontier: no other configuration passes at least as often for at most the same cost per pass. Frontier: Claude Sonnet 5.5 · Claude Code. Costs are calculations from tokens. Pass rates with their 95% intervals are in the pass-rate chart.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Short, latency-sensitive calls: Sonnet, then test Fable

On five easy tasks (15 calls per model), Fable 5.1 had the lowest median time: 1.94 s (range 1.41 to 9.83 s). Sonnet's median was 2.31 s (2.17 to 7.73 s). The ranges overlap, so speed does not separate them.

Fable cost $0.02054 per pass and Sonnet $0.00624 (a calculation, 3.3x). Fable passed 15/15 (95% interval 79.6% to 100%). Sonnet passed 12/15 (80%, 55% to 93%), and the intervals overlap. All three misses were one task: the right number plus extra working lines, which the exact-text validator rejects. Test a clearer format rule. We did not measure whether it fixes these misses.

Entrance: medians race at 4.5× real timeMotion reduced: press Replay to animateThe slowest median is 6.3 s. The clock runs at the recorded speed.
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

9 rows. Slowest GPT-6.1 Sol (low) · Codex CLI 6.3 s (range 4.7 s–10.5 s, n 10). Fastest Claude Fable 5.1 · Claude Code 1.9 s (range 1.4 s–9.8 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 10–15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Exact formats and strict JSON: Sonnet, not Haiku

We sent each prompt 10 times to Haiku 4.5 and to Sonnet 5.5. Haiku passed 0 of 10 on the exact-number prompt (95% interval 0% to 27.8%). It gave the same wrong number every time: 289, not 282.

Haiku passed 1 of 10 on the JSON prompt (1.8% to 40.4%). Nine replies held the right JSON in the wrong format, such as a code fence. Sonnet passed 10/10 on both prompts (72.3% to 100%), so Sonnet is ahead. If you use Haiku for strict output, add a validator and a repair step.

  • Exact number
  • JSON object
  • Code fix
Claude Haiku 4.5
Claude Sonnet 5.5
GPT-6.1 Sol (medium)

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 10 per row

One series per prompt; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Work with project memory: Sonnet, and keep the file short

We gave Claude Code a small repository under eight memory conditions. Team-knowledge checks cover two facts the code does not show: a changelog rule and a late-fee rate. With no memory file, Sonnet passed 6 of 15 (40%, 19.8% to 64.3%).

With an 11-line curated file, Sonnet passed 15 of 15 (79.6% to 100%) and Haiku 8 of 10 (49.0% to 94.3%): a tie. With the same facts in a 211-line handbook, Sonnet still passed 15/15 and Haiku 3 of 10 (10.8% to 60.3%), so Sonnet is ahead. Our reading: keep the file short. Haiku's own 8/10 and 3/10 overlap, so it is a hint.

  • Claude Sonnet 5.5
  • Claude Haiku 4.5
No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest No memory 40% (95% interval 20%–64%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest Curated + hook 100% (95% interval 72%–100%, n 10). Lowest No memory 0% (95% interval 0%–28%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row5 of 8 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.

Changelog rule and late-fee rate, pooled · 95% Wilson intervals

Both models had the same memory files. The smaller model followed the team rules less often when the facts sat in long or messy files.

Source: Agent memory study: 8 kinds of project memory on Claude Code

More: Does CLAUDE.md help agent memory?

Reasoning effort: start low

All 11 cells on the hard set passed 16 of 16 (81% to 100% each). These are Sonnet and Opus at four settings each, plus GPT-6.1 Sol at three. Sonnet at low effort had a median 5.82 s (range 2.78 to 19.96 s, n = 16). Its cost was $0.01219 per pass (calculation). At high effort it had a median 8.81 s (2.93 to 35.81 s, n = 16) and $0.01671 per pass (calculation). The call ranges overlap.

Some ladder cells reuse earlier hard-set calls; these are not independent new trials. The set has a ceiling, so effort may still matter on harder work.

Calculation
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (low) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 11 rows. Highest Claude Opus 5.5 (high) · Claude Code $0.034 (n 16). Lowest Claude Sonnet 5.5 (low) · Claude Code $0.012 (n 16).

Notesn = 16 per row

All calls in a configuration divided by its strict passes

Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.

Sources: Effort ladder: the hard task set at each effort level, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

More: Does reasoning effort buy quality?

When to pay for Opus: when your validator says so

These task sets did not show a supported quality gain for Opus. Both models passed 24/24 hard calls (95% interval 86% to 100% each). On short tasks, Opus passed 15/15 (79.6% to 100%), against Sonnet's 12/15 (54.8% to 93.0%); the intervals overlap. Our Sonnet vs Opus comparison holds 56 rows from nine studies: 26 measured metrics and 30 list-price calculations. None separates them: 10 ties and 46 unclear. Calculation rows are derived from recorded counts and prices; they are not new runs.

Opus 5.5 lists at $4 input and $20 output per million tokens, double Sonnet's $2 and $10. That fits the 2.0x per hard pass above.

The calculated price ratio shrinks on our recorded agent-loop tokens. Both models list cache reads at $0.20 per million. Our agent's recorded SWE-bench tokens from 33 attempts cost $143.83 at Opus prices and $87.23 at Sonnet prices. That is 1.6x (a calculation, not a run).

Our rule: run your tasks on both and compare cost per validator pass. More: Sonnet vs Opus: when is Opus worth the price?

When to use Haiku: after you test it

Haiku 4.5 lists at half Sonnet's per-token price: $1 input and $5 output per million tokens. Repricing the same tokens from 33 attempts gives Haiku half the calculated cost ($43.61 against $87.23, a calculation). That holds tokens and cache use fixed; it does not predict Haiku's outcomes. On the hard set it passed 11/24 (95% interval 28% to 65%), against Sonnet's 24/24 (86% to 100%). On the five short tasks Haiku passed 15/15 (79.6% to 100%).

Haiku had a higher calculated cost in the settings below. These calculations describe our samples. They do not establish population cost rankings:

  • Short tasks: $0.00836 against $0.00624; n = 15 calls per model.
  • Hard tasks: $0.0672 against $0.01435; n = 24 calls per model.
  • Routing, per 1,000 decisions: $8.924 against $4.996; n = 82 calls per model. This is cost per decision, not per correct decision.
  • Project memory, in all eight conditions (2 to 9 Haiku passes each). Curated file: $0.1095 against $0.0818. This uses full passes: Haiku 7/10 (95% interval 39.7% to 89.2%), Sonnet 15/15 (79.6% to 100%). The team-knowledge counts above use a different metric.

With the CLI default thinking, Haiku reported a median of 4,556 reasoning tokens per call on the hard set. Its median time was 39.01 s (range 15.27 to 75.13 s, n = 24). Sonnet had 7.75 s (2.26 to 34.79 s, n = 24); the ranges overlap. None of the studies above tested Haiku with thinking turned down.

How we measured

The Claude calls in the head-to-head, routing, consistency, effort and memory studies used the Claude Code CLI. Their timings include its start-up. A deterministic validator or test graded each result, and every failed call counts. Rates carry 95% Wilson intervals; times carry a range or a p50-to-p95 band, not an interval. "Ahead" means the intervals or ranges do not overlap. Costs are list-price calculations, not bills. Cost per pass divides the cost of all attempts, including failures, by the number of passes. These cost figures have no uncertainty intervals and do not establish population cost rankings.

Studies: hard tasks, short tasks, consistency, agent memory, routing, routing overhead, effort ladder, cost experiments.

Caveats

  • Ceilings. Our tasks are too easy to separate the top models. Opus and Fable may win on work we did not test.
  • Small samples. Head-to-head, consistency and effort cells hold 10 to 24 calls. Routing uses 82 cases per model. Memory pools several checks per session, so checks are not independent trials.
  • Narrow tasks. Head-to-head, consistency and effort calls were single-turn with tools off. Routing uses typed cases, not general classification. Memory used file and shell tools in one small synthetic repository. Agent loops can differ.
  • Prices and picks. List prices are the product price table effective 2026-09-21. Picks are our reading of the data.

Compare: Haiku vs Sonnet, Sonnet vs Opus, Opus vs Fable.

Measure your own tasks

Try Agent keeps a receipt for each task: model, tokens, time, cost and validation result.

The data behind this post

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.