Is Claude Haiku cheaper than Sonnet? Cost per correct answer, with retries
Haiku 4.5 lists at half the price of Sonnet 5.5, yet calculated cost per correct hard answer was 4.7x higher. Retry maths, two exceptions and every source.
TL;DR
- Per token, yes. Our cost-per-correct-answer estimates usually say no. Claude Haiku 4.5 lists at half the price of Claude Sonnet 5.5. With default thinking, Haiku had a higher cost-per-correct-answer point estimate in the four settings below. These cost gaps have no confidence interval.
- The gaps (list-price calculations): 4.7x on hard tasks, 1.3x on easy tasks, 1.1x to 3.2x across 8 memory setups, 1.9x per correct routing decision under the dataset accounting. Sample sizes and intervals follow below.
- Two observed cost factors. Haiku wrote 3.4x to 13.3x as many output tokens per call (calculation). On hard tasks it also passed 11 of 24 (95% interval 28% to 65%) against 24 of 24 for Sonnet (86% to 100%).
- A simple retry calculation keeps the gap. Under independent tries with a fixed pass probability, retry-until-pass costs about $0.067 per hard task on Haiku and $0.01435 on Sonnet (calculation). Haiku repeated one wrong answer 10 times out of 10.
- Haiku has a lower calculated cost only when its cost per call, as a multiple of Sonnet's, is below its pass rate as a multiple of Sonnet's (calculation). The default-thinking point estimates below do not meet that rule.
- Two lower cost point estimates. With thinking off, Haiku had a lower cost per correct routing decision: $3.89 against $5.32 per 1,000 (calculation, separate run). In a public SWE-bench panel, calculated cost per resolved instance was $0.479 for Haiku 4.5 and $0.913 for Sonnet 4.5. Both resolved 25 of 33 (95% interval 59% to 87% each).
Side by side: Haiku vs Sonnet.
The question
Price pages list cost per token. You want correct answers, so use this:
cost per correct answer = cost of all calls ÷ correct answers
Failed calls stay in the cost, so half the token price can still cost more.
| Per million tokens | Input | Cache read | Output |
|---|---|---|---|
| Claude Haiku 4.5 | $1 | $0.10 | $5 |
| Claude Sonnet 5.5 | $2 | $0.20 | $10 |
The table uses the recorded price snapshot cited by the repricing study. Cost figures below are calculations from reported tokens or CLI list-price estimates. The memory costs use the CLI estimates. The calls ran on a subscription, so no invoice backs them.
The routing dataset assumes the five-minute cache-write rate: $2.50 per million Sonnet tokens. Sonnet's receipts use the one-hour rate: $4 per million. We show both calculations below.
Hard tasks: 4.7x per correct answer
On 8 hard tasks with strict validators, Sonnet passed 24 of 24 and Haiku 11 of 24. The intervals (86% to 100% and 28% to 65%) do not overlap, so Sonnet is ahead on pass rate.
Cost per strict pass: Haiku $0.0672, Sonnet $0.01435. Calculation: 0.0672 ÷ 0.01435 = 4.7x.
Haiku's medians were 5,064 output tokens and 4,556 reasoning tokens per call. Sonnet's medians were 1,050 output tokens and 585 reasoning tokens. These are separate medians from n = 24 calls per model. Calculation: 4.8x the output tokens at half the price is about 2.4x the output cost.
A Haiku call cost on average $0.0672 × 11/24 = $0.0308, against $0.01435 for Sonnet: 2.1x (calculation). Fewer than half of Haiku's calls passed, so per pass the ratio grows to 4.7x.
The retry thought experiment
A common plan: use the cheap model, check the answer, retry on a fail. Each step below is a calculation.
- Expected tries. Retry-until-pass needs 1 ÷ p tries, where p is the pass rate. Haiku: 1 ÷ (11/24) = 2.2 tries (1.5 to 3.6 across its interval). Sonnet: 1 (1.0 to 1.2 across its 86% to 100% interval).
- Expected cost per solved task. Cost per try × tries. Haiku: $0.0308 × 2.18 = $0.067 ($0.047 to $0.110 across the interval). Sonnet: $0.01435 ($0.01435 to $0.01665 across its interval).
These sensitivity ranges hold mean call cost fixed. They are not cost confidence intervals. At Haiku's upper pass-rate bound, the cost is 3.3x Sonnet's point estimate.
- Illustrative time proxy. Tries × median time per call. Haiku: 2.18 × 39.01 s = about 85 s. Sonnet: 7.75 s.
The single-call ranges overlap (Haiku 15.27 s to 75.13 s, Sonnet 2.26 s to 34.79 s). A range is not a confidence interval. This median-based proxy is neither expected elapsed time nor a tested speed difference.
The model assumes independent tries with one fixed pass probability and mean call cost. The pooled pass rate does not predict retries on each task. These runs did not test retry-until-pass. For a retry followed by a switch to Sonnet, see the retry-or-escalate study.
Retries do not fix a repeated error
Some Haiku errors repeat.
- The same wrong answer, 10 of 10. On an exact-number prompt, Haiku passed 0 of 10 (95% interval 0% to 28%). All 10 replies gave 289; the right answer is 282 (details). No try passed in this run. That does not prove that every future retry will fail.
- 0 of 3 on three hard tasks. Haiku passed 0 of 3 on each of the event-loop, room-schedule and SQL tasks (95% interval 0% to 56% each).
- Format misses differ. 5 of Haiku's 24 calls gave the right answer in the wrong format. Counting them gives 16 of 24 (47% to 82%) and $0.046 per pass (calculation: $0.0672 × 11/16), still 3.2x Sonnet.
A retry helps against random errors. When a model fails the same way each time, change the prompt or the model.
Easy tasks: Haiku passed 15/15 and had a higher cost estimate
On five short tasks, Haiku passed 15 of 15 (80% to 100%) and Sonnet 12 of 15 (55% to 93%). The intervals overlap, so this sample does not establish a pass-rate difference. Haiku hit the task set's ceiling. All three Sonnet misses were format misses on one task.
Haiku’s cost estimate was still higher: $0.00836 per pass against $0.00624, 1.3x (calculation). A call cost on average $0.00836 on Haiku and $0.00499 on Sonnet: 1.7x (calculation). The per-call cost ranges overlap.
Haiku used more of both token types:
- Output: medians of 367 output tokens and 297 reasoning tokens for Haiku, against 107 output tokens for Sonnet: 3.4x the tokens at half the price (calculation).
- Input: a mean of 3,790 tokens, none from the cache, against 685 other tokens plus 1,401 cache reads for Sonnet. At list input and cache-read prices, that is about $0.0038 per call for Haiku and $0.0017 for Sonnet (calculation; floors, because cache writes cost more). We do not know why Haiku read nothing from the cache.
Agent sessions with memory: Haiku higher in all 8 setups
In 200 Claude Code sessions (120 Sonnet, 80 Haiku), we tested 8 memory setups. Each setup had n = 15 Sonnet sessions and n = 10 Haiku sessions. Haiku's calculated cost per fully correct result was higher in all 8. These point estimates have no cost interval. The ratio ran from 1.1x (Stop hook only, $0.1419 vs $0.1278) to 3.2x (/init CLAUDE.md, $0.4317 vs $0.1359), calculations. Haiku's median session also read 3.7x to 4.9x as many input tokens, cache reads included, in all 8 (calculation).
The no-memory and curated-plus-hook conditions had different pass rates and costs. This comparison cannot isolate a cause. With no memory, Haiku passed fully in 2 of 10 sessions (6% to 51%) at $0.3786 per full pass, 2.7x Sonnet (calculation). With a curated file plus a hook, it passed 9 of 10 (60% to 98%) at $0.0960, 1.1x Sonnet's $0.0843 (calculation). Sonnet passed 15 of 15 with curated plus hook (80% to 100%), a ceiling. With no memory, Sonnet passed 9 of 15 (36% to 80%).
The no-memory figure rests on 2 full passes.
Routing decisions: 1.9x per correct decision
As a router, Haiku decided 73 of 82 cases exactly (80% to 94%) and Sonnet 77 of 82 (87% to 97%). The intervals overlap; this sample does not establish a difference. Calculated cost under the dataset accounting: $8.924 per 1,000 decisions against $4.996. Per 1,000 exact decisions: $8.924 × 82/73 = $10.02 against $4.996 × 82/77 = $5.32, 1.9x (calculations).
Sonnet's own CLI cost estimate was higher, $7.324. The dataset assumes five-minute cache writes at $2.50 per million tokens. The CLI receipts use one-hour writes at $4 per million. That gives $7.80 per 1,000 exact decisions, and a 1.3x Haiku/Sonnet gap (calculation). Neither cost ratio has a confidence interval.
Haiku wrote a mean 1,419 output tokens per decision, Sonnet 107 (n = 82 each). Haiku ran with the CLI default extended thinking, Sonnet at low effort, as production asks.
What about "Haiku halves the agent bill"?
Our cost thought experiment prices our agent's recorded SWE-bench tokens at list price. That gives $43.61 at Haiku prices and $87.23 at Sonnet prices (a calculation, not a run). It assumes Haiku would use the same tokens and resolve the same tasks. Our other task sets do not validate that assumption: Haiku wrote more output tokens in the easy, hard and routing sets, and passed less often on hard work.
One recorded row goes the other way. In the public SWE-bench panel, Claude 4.5 Haiku resolved 25 of 33 instances (95% interval 59% to 87%) at a calculated $0.479 per resolved instance. Claude 4.5 Sonnet resolved 25 of the same 33 (59% to 87%) at $0.913, 1.9x higher (calculation).
Two limits apply. It is Sonnet 4.5, not 5.5, on a third party's bash-only harness. We calculate cost per resolved instance from the panel's published API costs. Those costs have no interval, so the compare page marks the gap unclear. The answer depends on the setting.
When could Haiku be cheaper?
Cost per correct answer is cost per call ÷ pass rate. So Haiku is cheaper per correct answer only when:
Haiku cost per call ÷ Sonnet cost per call < Haiku pass rate ÷ Sonnet pass rate
Compare cost, not token counts. Input, cache and output tokens have different prices. At half the price and the same token mix, Haiku would need fewer than twice Sonnet's tokens at an equal pass rate. The mix differed.
On the easy set, Haiku used 2.1x Sonnet's mean total tokens (calculation; n = 15 each). The means were 4,704 and 2,283, including cache reads and output. Its calls cost 1.7x as much (calculation). Sonnet read a mean of 1,401 input tokens per call from the cache at a tenth of the input price; Haiku read none.
The rule on our measured costs (calculations from point estimates):
| Setting | Haiku cost per call ÷ Sonnet's | Haiku pass rate ÷ Sonnet's | Lower Haiku cost point estimate? |
|---|---|---|---|
| Hard tasks | 2.1x | 0.46x | No |
| Easy tasks | 1.7x | 1.3x | No |
| Routing, default thinking | 1.8x | 0.95x | No |
| Routing, thinking off | 0.67x | 0.92x | Yes |
| Hard tasks, thinking off | 0.42x | 0.17x | No |
The routing rows use the dataset's Sonnet baseline: five-minute cache writes at $2.50 per million tokens. Sonnet's CLI receipts instead use the one-hour rate of $4 per million.
The two thinking-off rows come from a separate run with MAX_THINKING_TOKENS=0 (details). On routing, Haiku answered 71 of 82 exactly (78% to 92%) at $3.364 per 1,000 decisions, or $3.89 per 1,000 exact decisions (calculation).
It still wrote 3.4x Sonnet's mean output tokens (366 against 107 per decision; calculation, n = 82 each). Sonnet's calls carried 1,552 cache-write tokens per decision on average; Haiku's carried none. The routing pass-rate intervals overlap, and the cost gap has no interval.
On the hard tasks, thinking off passed 4 of 24 (7% to 36%) at $0.0365 per strict pass, 2.5x Sonnet's (calculation).
Best practice
- Compare models on your own tasks by cost per correct answer. Count every failed try.
- Check the effort or thinking setting first.
How we measured
- Hard and easy sets: Claude Code, one turn, tools off, 3 repetitions per task (n = 24 hard, n = 15 easy per model). The hard validators were strict: the whole reply must pass. The easy validators removed one wrapping fence first.
- Memory: 5 tasks, 8 conditions; Sonnet 3 repetitions per cell, Haiku 2. Routing: 82 labelled cases, one call each. Consistency: each prompt 10 times.
- Every ratio, try count and cost figure is a calculation or a CLI list-price estimate. Rate intervals are 95% Wilson intervals. Repeated calls reuse the same tasks; they do not represent that many distinct tasks.
- Disclosure: our recorded SWE-bench work ran on Sonnet 5.5, so we have a stake in this answer.
Caveats
- Thinking and effort. The hard and easy head-to-head rows use default effort; the CLI chose its thinking level. Haiku routing used default thinking; Sonnet routing used low effort. Thinking-off Haiku ran later, so this is not a matched same-time comparison.
- Our runs only. We ran Haiku 4.5 and Sonnet 5.5 in Claude Code. The SWE-bench panel went the other way, with Sonnet 4.5.
- Cost has no interval. The hard-set point estimate is 4.7x. Varying only Haiku's pass rate across its interval gives 3.3x or more against Sonnet's point estimate (calculation). This holds cost fixed and does not establish a cost ranking. The easy, memory and routing gaps rest on small samples, and the compare page marks every cost row unclear. Read them as direction.
- Small samples: 10 to 24 calls per model/condition in the easy, hard and memory comparisons, 3 per hard task. Routing has n = 82 per model. Sonnet passed 24/24 on the hard set, a ceiling.
- A simple retry model: it assumes independent tries.
What to read next
- Cost per correct answer calculator: a fixed independent-retry scenario from the recorded hard-task counts.
- When tasks get hard: Haiku vs Sonnet vs Opus vs Fable
- How to estimate your AI coding bill and the AI cost calculator
- Model: Claude Haiku 4.5. Studies: hard, easy, memory, routing, consistency, repricing, SWE-bench
Pay per correct answer, not per token
Agent records the tokens, cost and validation result of every step, so you can see cost per correct answer on your own work. Try Agent.