• Thought experiment
  • Calculation
  • Claude Haiku
  • Claude Sonnet

Is Claude Haiku cheaper than Sonnet? Cost per correct answer, with retries

Haiku 4.5 lists at half the price of Sonnet 5.5, yet calculated cost per correct hard answer was 4.7x higher. Retry maths, two exceptions and every source.

TL;DR

  • Per token, yes. Our cost-per-correct-answer estimates usually say no. Claude Haiku 4.5 lists at half the price of Claude Sonnet 5.5. With default thinking, Haiku had a higher cost-per-correct-answer point estimate in the four settings below. These cost gaps have no confidence interval.
  • The gaps (list-price calculations): 4.7x on hard tasks, 1.3x on easy tasks, 1.1x to 3.2x across 8 memory setups, 1.9x per correct routing decision under the dataset accounting. Sample sizes and intervals follow below.
  • Two observed cost factors. Haiku wrote 3.4x to 13.3x as many output tokens per call (calculation). On hard tasks it also passed 11 of 24 (95% interval 28% to 65%) against 24 of 24 for Sonnet (86% to 100%).
  • A simple retry calculation keeps the gap. Under independent tries with a fixed pass probability, retry-until-pass costs about $0.067 per hard task on Haiku and $0.01435 on Sonnet (calculation). Haiku repeated one wrong answer 10 times out of 10.
  • Haiku has a lower calculated cost only when its cost per call, as a multiple of Sonnet's, is below its pass rate as a multiple of Sonnet's (calculation). The default-thinking point estimates below do not meet that rule.
  • Two lower cost point estimates. With thinking off, Haiku had a lower cost per correct routing decision: $3.89 against $5.32 per 1,000 (calculation, separate run). In a public SWE-bench panel, calculated cost per resolved instance was $0.479 for Haiku 4.5 and $0.913 for Sonnet 4.5. Both resolved 25 of 33 (95% interval 59% to 87% each).

Side by side: Haiku vs Sonnet.

The question

Price pages list cost per token. You want correct answers, so use this:

cost per correct answer = cost of all calls ÷ correct answers

Failed calls stay in the cost, so half the token price can still cost more.

Per million tokensInputCache readOutput
Claude Haiku 4.5$1$0.10$5
Claude Sonnet 5.5$2$0.20$10

The table uses the recorded price snapshot cited by the repricing study. Cost figures below are calculations from reported tokens or CLI list-price estimates. The memory costs use the CLI estimates. The calls ran on a subscription, so no invoice backs them.

The routing dataset assumes the five-minute cache-write rate: $2.50 per million Sonnet tokens. Sonnet's receipts use the one-hour rate: $4 per million. We show both calculations below.

Hard tasks: 4.7x per correct answer

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

On 8 hard tasks with strict validators, Sonnet passed 24 of 24 and Haiku 11 of 24. The intervals (86% to 100% and 28% to 65%) do not overlap, so Sonnet is ahead on pass rate.

Cost per strict pass: Haiku $0.0672, Sonnet $0.01435. Calculation: 0.0672 ÷ 0.01435 = 4.7x.

  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 5,064 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16). Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 4,556 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 150 (n 16).

Notesn 16–24 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Haiku's medians were 5,064 output tokens and 4,556 reasoning tokens per call. Sonnet's medians were 1,050 output tokens and 585 reasoning tokens. These are separate medians from n = 24 calls per model. Calculation: 4.8x the output tokens at half the price is about 2.4x the output cost.

A Haiku call cost on average $0.0672 × 11/24 = $0.0308, against $0.01435 for Sonnet: 2.1x (calculation). Fewer than half of Haiku's calls passed, so per pass the ratio grows to 4.7x.

The retry thought experiment

A common plan: use the cheap model, check the answer, retry on a fail. Each step below is a calculation.

  1. Expected tries. Retry-until-pass needs 1 ÷ p tries, where p is the pass rate. Haiku: 1 ÷ (11/24) = 2.2 tries (1.5 to 3.6 across its interval). Sonnet: 1 (1.0 to 1.2 across its 86% to 100% interval).
  2. Expected cost per solved task. Cost per try × tries. Haiku: $0.0308 × 2.18 = $0.067 ($0.047 to $0.110 across the interval). Sonnet: $0.01435 ($0.01435 to $0.01665 across its interval).

    These sensitivity ranges hold mean call cost fixed. They are not cost confidence intervals. At Haiku's upper pass-rate bound, the cost is 3.3x Sonnet's point estimate.

  3. Illustrative time proxy. Tries × median time per call. Haiku: 2.18 × 39.01 s = about 85 s. Sonnet: 7.75 s.

    The single-call ranges overlap (Haiku 15.27 s to 75.13 s, Sonnet 2.26 s to 34.79 s). A range is not a confidence interval. This median-based proxy is neither expected elapsed time nor a tested speed difference.

The model assumes independent tries with one fixed pass probability and mean call cost. The pooled pass rate does not predict retries on each task. These runs did not test retry-until-pass. For a retry followed by a switch to Sonnet, see the retry-or-escalate study.

Retries do not fix a repeated error

Some Haiku errors repeat.

  • The same wrong answer, 10 of 10. On an exact-number prompt, Haiku passed 0 of 10 (95% interval 0% to 28%). All 10 replies gave 289; the right answer is 282 (details). No try passed in this run. That does not prove that every future retry will fail.
  • 0 of 3 on three hard tasks. Haiku passed 0 of 3 on each of the event-loop, room-schedule and SQL tasks (95% interval 0% to 56% each).
  • Format misses differ. 5 of Haiku's 24 calls gave the right answer in the wrong format. Counting them gives 16 of 24 (47% to 82%) and $0.046 per pass (calculation: $0.0672 × 11/16), still 3.2x Sonnet.

A retry helps against random errors. When a model fails the same way each time, change the prompt or the model.

Easy tasks: Haiku passed 15/15 and had a higher cost estimate

Calculation
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 9 rows. Highest Claude Fable 5.1 · Claude Code $0.021 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.0062 (n 15).

Notesn 10–15 per row

All calls in a configuration, failures included, divided by its passes

Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.

Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

On five short tasks, Haiku passed 15 of 15 (80% to 100%) and Sonnet 12 of 15 (55% to 93%). The intervals overlap, so this sample does not establish a pass-rate difference. Haiku hit the task set's ceiling. All three Sonnet misses were format misses on one task.

Haiku’s cost estimate was still higher: $0.00836 per pass against $0.00624, 1.3x (calculation). A call cost on average $0.00836 on Haiku and $0.00499 on Sonnet: 1.7x (calculation). The per-call cost ranges overlap.

Haiku used more of both token types:

  • Output: medians of 367 output tokens and 297 reasoning tokens for Haiku, against 107 output tokens for Sonnet: 3.4x the tokens at half the price (calculation).
  • Input: a mean of 3,790 tokens, none from the cache, against 685 other tokens plus 1,401 cache reads for Sonnet. At list input and cache-read prices, that is about $0.0038 per call for Haiku and $0.0017 for Sonnet (calculation; floors, because cache writes cost more). We do not know why Haiku read nothing from the cache.

Agent sessions with memory: Haiku higher in all 8 setups

Calculation
  • Claude Sonnet 5.5
  • Claude Haiku 4.5 (square)
Sorted by gap, largest first.
/init CLAUDE.md
No memory
Handbook, 210 lines
Curated, 11 lines
Dreamed notes
Raw notes, 60 lines
Curated + hook
Stop hook only

Gap labels, Claude Haiku 4.5 vs Claude Sonnet 5.5: Claude Haiku 4.5 is x% higher (+) or lower (−) than Claude Sonnet 5.5, calculated from the two values shown (the change counted from Claude Sonnet 5.5’s value).

List-price calculation, not a run. 8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest No memory $0.14 (n 9). Lowest Curated, 11 lines $0.082 (n 15). Claude Haiku 4.5: highest /init CLAUDE.md $0.43 (n 2). Lowest Curated + hook $0.096 (n 9).

Notesn 2–15 per row

Sum of the CLI's cost estimates for a condition, divided by its full passes

Sessions ran on a subscription; these are the CLI's list-price estimates, not bills. A failed session still costs money, so cost per correct result falls when fewer sessions fail.

Source: Agent memory study: 8 kinds of project memory on Claude Code

In 200 Claude Code sessions (120 Sonnet, 80 Haiku), we tested 8 memory setups. Each setup had n = 15 Sonnet sessions and n = 10 Haiku sessions. Haiku's calculated cost per fully correct result was higher in all 8. These point estimates have no cost interval. The ratio ran from 1.1x (Stop hook only, $0.1419 vs $0.1278) to 3.2x (/init CLAUDE.md, $0.4317 vs $0.1359), calculations. Haiku's median session also read 3.7x to 4.9x as many input tokens, cache reads included, in all 8 (calculation).

The no-memory and curated-plus-hook conditions had different pass rates and costs. This comparison cannot isolate a cause. With no memory, Haiku passed fully in 2 of 10 sessions (6% to 51%) at $0.3786 per full pass, 2.7x Sonnet (calculation). With a curated file plus a hook, it passed 9 of 10 (60% to 98%) at $0.0960, 1.1x Sonnet's $0.0843 (calculation). Sonnet passed 15 of 15 with curated plus hook (80% to 100%), a ceiling. With no memory, Sonnet passed 9 of 15 (36% to 80%).

The no-memory figure rests on 2 full passes.

Routing decisions: 1.9x per correct decision

Calculation
Largest value is 260x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).

Notesn 82–246 per row

List price × reported tokens per decision

List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions

As a router, Haiku decided 73 of 82 cases exactly (80% to 94%) and Sonnet 77 of 82 (87% to 97%). The intervals overlap; this sample does not establish a difference. Calculated cost under the dataset accounting: $8.924 per 1,000 decisions against $4.996. Per 1,000 exact decisions: $8.924 × 82/73 = $10.02 against $4.996 × 82/77 = $5.32, 1.9x (calculations).

Sonnet's own CLI cost estimate was higher, $7.324. The dataset assumes five-minute cache writes at $2.50 per million tokens. The CLI receipts use one-hour writes at $4 per million. That gives $7.80 per 1,000 exact decisions, and a 1.3x Haiku/Sonnet gap (calculation). Neither cost ratio has a confidence interval.

Haiku wrote a mean 1,419 output tokens per decision, Sonnet 107 (n = 82 each). Haiku ran with the CLI default extended thinking, Sonnet at low effort, as production asks.

What about "Haiku halves the agent bill"?

Our cost thought experiment prices our agent's recorded SWE-bench tokens at list price. That gives $43.61 at Haiku prices and $87.23 at Sonnet prices (a calculation, not a run). It assumes Haiku would use the same tokens and resolve the same tasks. Our other task sets do not validate that assumption: Haiku wrote more output tokens in the easy, hard and routing sets, and passed less often on hard work.

One recorded row goes the other way. In the public SWE-bench panel, Claude 4.5 Haiku resolved 25 of 33 instances (95% interval 59% to 87%) at a calculated $0.479 per resolved instance. Claude 4.5 Sonnet resolved 25 of the same 33 (59% to 87%) at $0.913, 1.9x higher (calculation).

Two limits apply. It is Sonnet 4.5, not 5.5, on a third party's bash-only harness. We calculate cost per resolved instance from the panel's published API costs. Those costs have no interval, so the compare page marks the gap unclear. The answer depends on the setting.

When could Haiku be cheaper?

Cost per correct answer is cost per call ÷ pass rate. So Haiku is cheaper per correct answer only when:

Haiku cost per call ÷ Sonnet cost per call < Haiku pass rate ÷ Sonnet pass rate

Compare cost, not token counts. Input, cache and output tokens have different prices. At half the price and the same token mix, Haiku would need fewer than twice Sonnet's tokens at an equal pass rate. The mix differed.

On the easy set, Haiku used 2.1x Sonnet's mean total tokens (calculation; n = 15 each). The means were 4,704 and 2,283, including cache reads and output. Its calls cost 1.7x as much (calculation). Sonnet read a mean of 1,401 input tokens per call from the cache at a tenth of the input price; Haiku read none.

The rule on our measured costs (calculations from point estimates):

SettingHaiku cost per call ÷ Sonnet'sHaiku pass rate ÷ Sonnet'sLower Haiku cost point estimate?
Hard tasks2.1x0.46xNo
Easy tasks1.7x1.3xNo
Routing, default thinking1.8x0.95xNo
Routing, thinking off0.67x0.92xYes
Hard tasks, thinking off0.42x0.17xNo

The routing rows use the dataset's Sonnet baseline: five-minute cache writes at $2.50 per million tokens. Sonnet's CLI receipts instead use the one-hour rate of $4 per million.

The two thinking-off rows come from a separate run with MAX_THINKING_TOKENS=0 (details). On routing, Haiku answered 71 of 82 exactly (78% to 92%) at $3.364 per 1,000 decisions, or $3.89 per 1,000 exact decisions (calculation).

It still wrote 3.4x Sonnet's mean output tokens (366 against 107 per decision; calculation, n = 82 each). Sonnet's calls carried 1,552 cache-write tokens per decision on average; Haiku's carried none. The routing pass-rate intervals overlap, and the cost gap has no interval.

On the hard tasks, thinking off passed 4 of 24 (7% to 36%) at $0.0365 per strict pass, 2.5x Sonnet's (calculation).

Best practice

  1. Compare models on your own tasks by cost per correct answer. Count every failed try.
  2. Check the effort or thinking setting first.

How we measured

  • Hard and easy sets: Claude Code, one turn, tools off, 3 repetitions per task (n = 24 hard, n = 15 easy per model). The hard validators were strict: the whole reply must pass. The easy validators removed one wrapping fence first.
  • Memory: 5 tasks, 8 conditions; Sonnet 3 repetitions per cell, Haiku 2. Routing: 82 labelled cases, one call each. Consistency: each prompt 10 times.
  • Every ratio, try count and cost figure is a calculation or a CLI list-price estimate. Rate intervals are 95% Wilson intervals. Repeated calls reuse the same tasks; they do not represent that many distinct tasks.
  • Disclosure: our recorded SWE-bench work ran on Sonnet 5.5, so we have a stake in this answer.

Caveats

  • Thinking and effort. The hard and easy head-to-head rows use default effort; the CLI chose its thinking level. Haiku routing used default thinking; Sonnet routing used low effort. Thinking-off Haiku ran later, so this is not a matched same-time comparison.
  • Our runs only. We ran Haiku 4.5 and Sonnet 5.5 in Claude Code. The SWE-bench panel went the other way, with Sonnet 4.5.
  • Cost has no interval. The hard-set point estimate is 4.7x. Varying only Haiku's pass rate across its interval gives 3.3x or more against Sonnet's point estimate (calculation). This holds cost fixed and does not establish a cost ranking. The easy, memory and routing gaps rest on small samples, and the compare page marks every cost row unclear. Read them as direction.
  • Small samples: 10 to 24 calls per model/condition in the easy, hard and memory comparisons, 3 per hard task. Routing has n = 82 per model. Sonnet passed 24/24 on the hard set, a ceiling.
  • A simple retry model: it assumes independent tries.

Pay per correct answer, not per token

Agent records the tokens, cost and validation result of every step, so you can see cost per correct answer on your own work. Try Agent.

The data behind this post

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.