Claude Haiku 4.5 vs Sonnet 5.5: all 80 comparison rows, and where the small model loses
Haiku 4.5 vs Sonnet 5.5 on 80 rows: Sonnet ahead on 14, Haiku on none, 31 ties. Hard tasks 11/24 vs 24/24, plus speed, memory and price.
TL;DR
- Sonnet 5.5 was ahead of Haiku 4.5 on 14 of 80 comparison rows. Haiku was ahead on none. 31 rows tie and 35 are unclear. The 14 is 13 rows plus one memory row where Haiku did worse (calculation). 64 rows are measurements and 16 are list-price calculations.
- Hard tasks: Sonnet passed 24 of 24 (95% interval 86% to 100%). Haiku passed 11 of 24 (46%, 28% to 65%). On five easy tasks the two tie.
- Speed: Sonnet was ahead on all 6 speed rows that have a winner. As a router, Haiku's median decision took 12,543 ms and Sonnet's took 2,597 ms.
- Price: Haiku lists at half of Sonnet's price per token (calculation). On the hard set it still cost 4.7 times as much per strict pass (calculation).
- When Haiku fits: no measured row puts it ahead. Test it on your own tasks before you switch.
The side-by-side rows: /compare/claude-haiku-4-5-vs-claude-sonnet-5-5. Model pages: Claude Haiku 4.5 and Claude Sonnet 5.5.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks
139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.
Transcript
- Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
- 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
- Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
- Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
- Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
- Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
- Open benchmarks: intervals, sources and every failure kept.
The 80 rows at a glance
A row goes to one model only when the 95% intervals, ranges or p50 to p95 bands do not overlap. A tie means the sample cannot separate the models. Unclear means the row has no interval, or its ranges overlap. Our count by study:
| Study | Rows | Sonnet ahead | Haiku ahead | Tie | Unclear |
|---|---|---|---|---|---|
| Hard tasks | 6 | 2 | 0 | 0 | 4 |
| Easy tasks | 8 | 0 | 0 | 1 | 7 |
| Same prompt, 10 times | 9 | 3 | 0 | 3 | 3 |
| Agent memory | 40 | 4 | 0 | 20 | 16 |
| Routing (two studies) | 17 | 5 | 0 | 7 | 5 |
| All | 80 | 14 | 0 | 31 | 35 |
The compare page lists one memory row as a Haiku lead because it reads a higher rate as better. That row is a bad outcome for Haiku (section 4), so we count it for Sonnet.
1. Hard tasks: Sonnet 24/24, Haiku 11/24
Each model made 24 calls (3 repeats of 8 tasks). Sonnet passed all 24 (95% interval 86% to 100%). Haiku passed 11 (46%, 28% to 65%). The intervals do not overlap, so Sonnet is ahead.
It is also ahead on the lenient reading, which counts a right answer in the wrong format: Haiku 16/24 (67%, 47% to 82%). Study: /benchmarks/hard-model-head-to-head.
Haiku had 5 format misses (CSV parser 1, room schedule 2, SQLite query 2) and 8 wrong answers. By task, with 3 calls each, it passed 0/3 on event-loop order, room schedule and SQLite query. It passed 1/3 on DST day length, 2/3 on CSV parser and refactor, and 3/3 on interval merge and SemVer regex. Sonnet passed 3/3 on all eight tasks.
Sonnet's 24/24 is this set's ceiling. It shows that Haiku falls short, not how far Sonnet reaches.
2. Easy tasks: a tie
On five short tasks, Haiku passed 15/15 (95% interval 80% to 100%) and Sonnet 12/15 (80%, 55% to 93%). The intervals overlap, so this is a tie. Haiku's 15/15 is the set's ceiling. All three Sonnet misses were format misses on one arithmetic task. Each reply ended on the right answer, 2292, but added working lines that the exact-text validator rejects. Study: /benchmarks/model-head-to-head.
3. The same prompt, 10 times
On the exact-number prompt, Haiku passed 0/10 (0% to 28%). It gave the same wrong answer every time: 289, not the right 282. Sonnet passed 10/10 (72% to 100%).
On the JSON prompt, Haiku passed 1/10 (2% to 40%), with 9 format misses, and Sonnet passed 10/10. Sonnet is ahead on both rows. On the code-fix prompt both passed 10/10: a tie. Study: /benchmarks/caching-consistency.
4. Memory: Sonnet ahead on the handbook and /init files
Both models got the same memory files (Haiku 10 sessions per condition, Sonnet 15). The score is team knowledge followed: a changelog rule and a late-fee rate. Sonnet is ahead on two conditions and ties on the other six.
- 211-line handbook: Haiku 3/10 (11% to 60%), Sonnet 15/15 (80% to 100%).
- /init file: Haiku 1/10 (2% to 40%), Sonnet 10/15 (42% to 85%).
One row reads the other way. This chart shows sessions that ran a stale README test command, which fails on Node 25, so a higher rate is worse. With 60 lines of raw notes, Haiku ran it in 10/10 sessions (72% to 100%) and Sonnet in 0/15 (0% to 20%). Haiku did worse, so we count this row for Sonnet. Study: /benchmarks/agent-memory.
5. Speed: Haiku was slower on every decided row
As a router, Haiku's median decision took 12,543 ms (p50 to p95: 12,543 to 34,481 ms; n = 82). Sonnet at low effort took 2,597 ms (2,597 to 4,298 ms; n = 82). Haiku's median is above Sonnet's p95, so Sonnet is ahead. A p50 to p95 band is not a confidence interval. Study: /benchmarks/routing-overhead.
Six speed rows have a winner, all Sonnet. Five are router rows, which are not independent tests. The sixth is the same-prompt code fix: 5.95 s against 2.67 s (ranges 4.89 to 7.33 s and 2.32 to 4.34 s; n = 10).
Other speed rows are unclear. On the hard set the medians were 39.01 s and 7.75 s, but the ranges overlap (15.27 to 75.13 s and 2.26 to 34.79 s). On the exact-number prompt Haiku's median was lower (5.06 s against 6.89 s), and those ranges overlap too. A range is not an interval.
On the hard set, Haiku's median call wrote 5,064 output tokens, 4,556 of them reasoning. Sonnet wrote 1,050 (585 reasoning). Time and tokens come from the same calls, so they show a link, not a cause.
6. Price: half the rate, not half the cost
List prices per million tokens:
| Model (list price, calculation) | Input | Cache read | Output |
|---|---|---|---|
| Claude Haiku 4.5 | $1 | $0.10 | $5 |
| Claude Sonnet 5.5 | $2 | $0.20 | $10 |
Haiku lists at half of Sonnet's price at all three points (calculation). The charts below price the recorded tokens at these rates. They are calculations, not bills: the calls ran on subscriptions. Study: /benchmarks/cost-thought-experiments.
On the hard set, one strict pass cost $0.0672 with Haiku and $0.01435 with Sonnet (calculation): about 4.7 times (calculation). A failed call still costs, so Haiku's 13 misses raise its figure. The row is unclear: the dataset holds no interval for either cost.
On the easy set, one pass cost $0.00836 with Haiku and $0.00624 with Sonnet (calculation): about 1.3 times (calculation). Sonnet's 3 format misses count against it.
As a router, 1,000 decisions cost $8.924 with Haiku and $4.996 with Sonnet (calculation): about 1.8 times (calculation). Haiku wrote 1,419 output tokens per decision and Sonnet 107. Accuracy tied: Haiku 73/82 exact (89%, 80% to 94%), Sonnet 77/82 (94%, 87% to 97%). Study: /benchmarks/routing-jev-vs-llm.
7. When Haiku fits
No measured row puts Haiku ahead. Half the price per token is its only lead in our data (calculation). All 14 cost rows show a higher figure for Haiku (calculations). All 14 are unclear: one has a range that overlaps, and the other 13 have no interval.
- Test on your own tasks. Compare cost per pass, not price per token.
- Check the output format. Haiku had 5 format misses on the hard set and 9 on the JSON prompt. Sonnet had none on those two, and 3 on the easy arithmetic task.
How we measured
- Task sets: Claude Code, one turn, tools off, deterministic validators. n = 24 (hard), 15 (easy), 10 per prompt. Rates carry 95% Wilson intervals.
- Memory: Claude Code sessions with tools on, one small repository, 8 conditions. Routing: 82 decisions per model, one CLI call each.
- Times: medians with ranges, or p50 to p95 bands. Costs: reported tokens times list price (calculations). Ratios and counts are our arithmetic.
Caveats
- Settings differ. Haiku ran with the CLI's default thinking in this pair. The router rows compare Haiku with thinking on against Sonnet at low effort. CLI timings include start-up and the CLI's own prompt.
- Small samples. n is 10 to 82 per cell (194 for per-question accuracy). A hard-task cell holds 3 calls. The memory cost-per-pass rows divide by as few as 2 passes.
- The 14 rows are not 14 tests. The strict and lenient hard rows share 24 calls. The two handbook rows show the same counts. The five router rows overlap.
- Long agent runs are not in this pair. On 33 SWE-bench Verified tasks, a public Claude 4.5 Haiku run (bash-only harness) resolved 25 (76%, 95% interval 59% to 87%). So did Agent's full pipeline on Sonnet 5.5. These are two systems, so no row compares them: /benchmarks/swe-bench-verified.
What to read next
- Hard tasks: Haiku vs Sonnet vs Opus vs Fable
- Same prompt, ten answers: Claude consistency
- Does CLAUDE.md help agent memory?
- Haiku vs Sonnet as a router
Test the cheaper model on your own work
Agent records the model, the tokens and the result of every step, so you can see where a smaller model fails. Try Agent and compare on your own work.