• Thought experiment
  • Claude Haiku
  • Claude Sonnet
  • LLM pricing
  • Cost Per Correct Answer
  • Retry
  • Escalation
  • Hard Tasks

Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts

On the 8 hard tasks, what does one correct answer cost, and how long does it take, if you try Claude Haiku 4.5 first and retry or escalate, against Claude Sonnet 5.5 every time?

Published · 6 charts · Download the data or a carousel

46%

95% CI 28%–65% · n = 24

11/24 · Strict passes, Claude Haiku 4.5 · Claude Code, eight hard tasks

The answer

The calculated Haiku-first policies cost more on these 8 hard tasks. These are scenarios from recorded calls, not a measured policy ranking. Sonnet every time costs $0.0132 per correct answer and takes 8.2 s (24/24 passes; 95% interval 86% to 100%). Haiku passes 11/24 calls (95% interval 28% to 65%); Haiku once costs $0.0677 per correct answer. Haiku with one retry, then Sonnet, costs $0.0566 and takes 71.5 s per correct answer. The tables show the other policies and sensitivity ranges; the method states the assumptions.

Key numbers

100% (24/24)

Strict passes, Claude Sonnet 5.5 · Claude Code, eight hard tasks

95% CI 86%–100% · n = 24

100% (16/16)

Strict passes, Claude Sonnet 5.5 (low) · Claude Code, eight hard tasks

95% CI 81%–100% · n = 16

$0.0132

Cost per correct answer, Sonnet 5.5 every time (calculation)

n = 24

8.2s

Wall time per correct answer, Sonnet 5.5 every time (calculation)

n = 24

$0.0677

Cost per correct answer, Haiku once (calculation)

n = 24

91.5s

Wall time per correct answer, Haiku once (calculation)

n = 24

$0.0566

Cost per correct answer, Haiku with one retry then Sonnet (calculation)

n = 48

71.5s

Wall time per correct answer, Haiku with one retry then Sonnet (calculation)

n = 48

4.3x

Haiku with one retry then Sonnet vs Sonnet every time: cost per correct answer (calculation)

n = 48

8.7x

Haiku with one retry then Sonnet vs Sonnet every time: wall time per correct answer (calculation)

n = 48

5.4x

Haiku up to 3 tries then Sonnet vs Sonnet every time: cost per correct answer (calculation)

n = 48

$0.0122

Cost per correct answer, Sonnet 5.5 low effort once (calculation)

n = 16

7.1s

Wall time per correct answer, Sonnet 5.5 low effort once (calculation)

n = 16

8 of 8

Tasks where Haiku’s median call cost more than Sonnet’s (calculation)

n = 8

8 of 8

Tasks where Haiku’s median call took longer than Sonnet’s

n = 8

5,064

Median output tokens per call, Haiku 4.5 (Claude Code)

n = 24

1,050

Median output tokens per call, Sonnet 5.5 (Claude Code)

n = 24

11.8%

Haiku call cost, as a share of measured, at which one retry then Sonnet matches Sonnet every time (calculation)

of measured · n = 8

5.4%

Haiku call time, as a share of measured, at which one retry then Sonnet matches Sonnet every time (calculation)

of measured · n = 8

3.0x

Lowest cost ratio of any Haiku-first policy to Sonnet every time, across the four Haiku pass-rate settings (sensitivity range, calculation)

n = 8

1.3x

Lowest cost ratio of any Haiku-first policy to Sonnet every time, with Haiku’s pass rate at its Wilson upper bound and Sonnet’s at its Wilson lower bound on every task at once (joint extreme, calculation)

n = 8

$0.0459

Cost per correct answer, Haiku with one retry then Sonnet, format misses counted as passes (calculation)

n = 48

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Thought experiment: not a run. These values reprice recorded tokens at list prices. No model was called again.

Calculation
Sonnet 5.5 low effort once (calculation)
Sonnet 5.5 every time (calculation)
Haiku, one retry, then Sonnet (calculation)
Haiku once (calculation)
Haiku up to 3 tries, then Sonnet (calculation)

Hover or focus a bar for its ratio to Sonnet 5.5 every time (calcul… (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 5 rows. Highest Haiku up to 3 tries, then Sonnet (calculation) $0.072 (n 48). Lowest Sonnet 5.5 low effort once (calculation) $0.012 (n 16).

Notesn 16–48 per row

Eight hard tasks, each equally likely. A failed try is paid for. Calls recorded in Claude Code

Calculation, not a run. These are median-input scenarios, not measured mean costs. Total modelled list-price cost of all tries over the 8 tasks, divided by the expected number of correct answers. Per-call cost is each task’s median call at list price (reported tokens × list price); the calls ran on a flat subscription. A try runs only when the earlier tries failed the validator, and tries are independent at each task’s observed pass rate. n is the number of recorded calls behind each bar. A calculation has no confidence interval; the sensitivity chart shows a range. Highlighted: the baseline, Sonnet 5.5 every time.

Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)

Share card (PNG)
Calculation
Largest value is 13x the smallest; Log shows the small bars.
Sonnet 5.5 low effort once (calculation)
Sonnet 5.5 every time (calculation)
Haiku, one retry, then Sonnet (calculation)
Haiku once (calculation)
Haiku up to 3 tries, then Sonnet (calculation)

Hover or focus a bar for its ratio to Sonnet 5.5 every time (calcul… (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 5 rows. Slowest Haiku up to 3 tries, then Sonnet (calculation) 93.4 s (n 48). Fastest Sonnet 5.5 low effort once (calculation) 7.1 s (n 16).

Notesn 16–48 per row

Tries run one after another, so a failed try adds its time

Calculation, not a run. These are median-input scenarios, not measured mean times. Each task’s median total call time (CLI start-up included), summed over the tries a policy expects to run, divided by the expected number of correct answers. Tries are assumed to run one after another. Parallel tries were not tested. One host, one network. A calculation has no confidence interval. Highlighted: the baseline, Sonnet 5.5 every time.

Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)

Share card (PNG)
Calculation
Sonnet 5.5 every time (calculation)
Haiku, one retry, then Sonnet (calculation)
Haiku up to 3 tries, then Sonnet (calculation)
Sonnet 5.5 low effort once (calculation)
Haiku once (calculation)

Hover or focus a bar for its ratio to Sonnet 5.5 every time (calcul… (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 5 rows. Highest Sonnet 5.5 every time (calculation) 100% (95% interval 86%–100%, n 24). Lowest Haiku once (calculation) 46% (95% interval 28%–65%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–48 per row4 of 5 at 100%: this task set cannot separate them.

Share of the eight hard tasks answered correctly under the strict validator

Whiskers are 95% Wilson intervals on the measured strict pass rate and appear only on the three one-try policies (24/24, 11/24 and 16/16). The two escalation policies are calculations without an interval. A value of 100% means no failure was recorded: the model treats Sonnet 5.5 as always right (24/24, 95% interval 86% to 100%), so the escalation policies cannot miss. n is the number of recorded calls behind each bar.

Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)

Share card (PNG)
Interval merge fix
DST day length
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQLite report query

Hover or focus a bar for its ratio to DST day length (the lowest value): a ratio of the two values shown, not a measurement.

8 rows. Highest Interval merge fix 100% (95% interval 44%–100%, n 3). Lowest SQLite report query 0% (95% interval 0%–56%, n 3). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 3 per row

Claude Haiku 4.5 · Claude Code, default effort. 3 calls per task

Measured. Whiskers are 95% Wilson intervals on only 3 calls per task, so they are wide: 0 of 3 allows a true rate up to 56%, and 3 of 3 allows one as low as 44%. A format miss (a right answer in a code fence or with prose) is not a pass. The retry policies take each task’s rate from this chart.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Share card (PNG)
Calculation
  • Claude Haiku 4.5 · Claude Code
  • Claude Sonnet 5.5 · Claude Code (square)
Sorted by gap, largest first.
SemVer regex
Interval merge fix
Room schedule
Event-loop order
SQLite report query
Money refactor
CSV parser
DST day length

Gap labels, Claude Sonnet 5.5 · Claude Code vs Claude Haiku 4.5 · Claude Code: Claude Sonnet 5.5 · Claude Code is x% higher (+) or lower (−) than Claude Haiku 4.5 · Claude Code, calculated from the two values shown (the change counted from Claude Haiku 4.5 · Claude Code’s value).

List-price calculation, not a run. 8 rows, 2 series: Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code. Claude Haiku 4.5 · Claude Code: highest SemVer regex $0.042 (n 3). Lowest Interval merge fix $0.018 (n 3). Claude Sonnet 5.5 · Claude Code: highest DST day length $0.025 (n 3). Lowest SemVer regex $0.0051 (n 3).

Notesn = 3 per row

Median of each task’s calls, default effort, Claude Code

Calculation: reported tokens × list price for each call, then the median per task; the calls ran on a flat subscription. Haiku’s list price per token is lower. Its median calls cost more and contained more output tokens. This does not isolate the effect of effort or thinking. 3 calls per task and configuration; the table shows each call-cost range, not a confidence interval.

Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)

Share card (PNG)
Calculation
Sonnet 5.5 low effort once (calculation)
Sonnet 5.5 every time (calculation)
Haiku, one retry, then Sonnet (calculation)
Haiku once (calculation)
Haiku up to 3 tries, then Sonnet (calculation)

Every interval overlaps every other: this chart does not order these rows.

List-price calculation, not a run. 5 rows. Highest Haiku up to 3 tries, then Sonnet (calculation) $0.072 (range $0.041–$0.09, n 48). Lowest Sonnet 5.5 low effort once (calculation) $0.012 (n 16). All run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 16–48 per row

Dot: Haiku’s pass rate as observed. Whisker: its pass rate on every task set to its Wilson lower and upper bound at once

Calculation. This is a sensitivity range, not a confidence interval: each Haiku policy is recomputed with Haiku’s pass rate on all 8 tasks at once set to the lower and then the upper end of that task’s 95% Wilson interval (3 calls per task, so the ends are far apart), and the whisker spans the lowest and highest result. That is more extreme than a 95% interval on the total. The two Sonnet policies do not depend on Haiku’s pass rate and have no whisker. Highlighted: the baseline, Sonnet 5.5 every time.

Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)

Share card (PNG)

Tables

Retry policies over the eight hard tasks (calculation)

PolicyTries in orderExpected successExpected cost per task (all tries)Cost per correct answerWall time per correct answerCost vs Sonnet every timeTime vs Sonnet every time
Sonnet 5.5 every time (calculation)Sonnet 5.5100%$0.013$0.0138.2 s1x1x
Haiku once (calculation)Haiku 4.546%$0.031$0.06891.5 s5.1x11.1x
Haiku, one retry, then Sonnet (calculation)Haiku 4.5 → Haiku 4.5 → Sonnet 5.5100%$0.057$0.05771.5 s4.3x8.7x
Haiku up to 3 tries, then Sonnet (calculation)Haiku 4.5 → Haiku 4.5 → Haiku 4.5 → Sonnet 5.5100%$0.072$0.07293.4 s5.4x11.3x
Sonnet 5.5 low effort once (calculation)Sonnet 5.5 low100%$0.012$0.0127.1 s0.92x0.86x

Per-task inputs: passes, 95% intervals, median call cost (calculation) and time; ranges are min–max, not confidence intervals

TaskClaude Haiku 4.5 · Claude Code: strict passesHaiku 95% intervalHaiku median call costHaiku call-cost range (calculation)Haiku median call timeHaiku call-time rangeClaude Sonnet 5.5 · Claude Code: strict passesSonnet 95% intervalSonnet median call costSonnet call-cost range (calculation)Sonnet median call timeSonnet call-time rangeClaude Sonnet 5.5 (low) · Claude Code: strict passesSonnet (low) 95% intervalSonnet (low) median call costSonnet (low) call-cost range (calculation)Sonnet (low) median call timeSonnet (low) call-time range
Fix an interval-merge function (off-by-one and edge cases)3/344% to 100%$0.018$0.0148 to $0.018321.2 s18.5 s to 24.7 s3/344% to 100%$0.0056$0.0056 to $0.00912.9 s2.4 s to 3.1 s2/234% to 100%$0.0074$0.0056 to $0.00922.8 s2.8 s to 2.8 s
Fix a time-zone day-length function (DST)1/36% to 79%$0.029$0.0243 to $0.030839 s26.2 s to 40.6 s3/344% to 100%$0.025$0.0231 to $0.042421.6 s19.6 s to 34.8 s2/234% to 100%$0.024$0.0218 to $0.026118.1 s16.3 s to 20.0 s
Write a CSV parser (quoted newlines, strict errors)2/3 (+1 format miss)21% to 94%$0.029$0.0250 to $0.050539 s33.8 s to 75.1 s3/344% to 100%$0.015$0.0146 to $0.01619.6 s8.8 s to 10.6 s2/234% to 100%$0.0093$0.0091 to $0.00964.4 s4.3 s to 4.4 s
Predict JavaScript event-loop output order0/30% to 56%$0.037$0.0231 to $0.039554 s26.0 s to 56.6 s3/344% to 100%$0.016$0.0147 to $0.01619.6 s8.2 s to 9.9 s2/234% to 100%$0.014$0.0139 to $0.01419.3 s8.8 s to 9.7 s
Solve a multi-constraint room schedule0/3 (+2 format misses)0% to 56%$0.036$0.0292 to $0.044454.6 s38.0 s to 68.8 s3/344% to 100%$0.012$0.0118 to $0.01267.7 s7.4 s to 7.8 s2/234% to 100%$0.011$0.0113 to $0.01156.6 s6.5 s to 6.7 s
Write a strict SemVer 2.0.0 regex3/344% to 100%$0.042$0.0373 to $0.044764.5 s54.5 s to 67.0 s3/344% to 100%$0.0051$0.0051 to $0.00792.5 s2.3 s to 3.7 s2/234% to 100%$0.0051$0.0051 to $0.00513.2 s2.9 s to 3.5 s
Refactor to remove duplication, keep 20 tests green2/321% to 94%$0.02$0.0179 to $0.021615.9 s15.3 s to 17.2 s3/344% to 100%$0.0096$0.0094 to $0.01513.6 s3.4 s to 7.2 s2/234% to 100%$0.0095$0.0094 to $0.00954.4 s3.7 s to 5.2 s
Write a SQLite reporting query (fan-out, ties, boundaries)0/3 (+2 format misses)0% to 56%$0.036$0.0289 to $0.040647.1 s38.9 s to 56.1 s3/344% to 100%$0.017$0.0157 to $0.01928.5 s7.6 s to 8.9 s2/234% to 100%$0.017$0.0167 to $0.01727.8 s7.7 s to 7.9 s

Sensitivity of each policy to Haiku’s pass rate (calculation)

PolicyHaiku pass rate set toExpected successCost per correct answerWall time per correct answer
Sonnet 5.5 every time (calculation)Does not depend on Haiku100%$0.0138.2 s
Haiku once (calculation)Haiku pass rate at the Wilson lower bound (every task)17%$0.18248 s
Haiku once (calculation)Haiku pass rate as observed46%$0.06891.5 s
Haiku once (calculation)Haiku pass rate at the Wilson upper bound (every task)79%$0.03952.8 s
Haiku once (calculation)Haiku format misses counted as passes (lenient reading)67%$0.04762.9 s
Haiku, one retry, then Sonnet (calculation)Haiku pass rate at the Wilson lower bound (every task)100%$0.06884.3 s
Haiku, one retry, then Sonnet (calculation)Haiku pass rate as observed100%$0.05771.5 s
Haiku, one retry, then Sonnet (calculation)Haiku pass rate at the Wilson upper bound (every task)100%$0.03952.6 s
Haiku, one retry, then Sonnet (calculation)Haiku format misses counted as passes (lenient reading)100%$0.04659.5 s
Haiku up to 3 tries, then Sonnet (calculation)Haiku pass rate at the Wilson lower bound (every task)100%$0.09115 s
Haiku up to 3 tries, then Sonnet (calculation)Haiku pass rate as observed100%$0.07293.4 s
Haiku up to 3 tries, then Sonnet (calculation)Haiku pass rate at the Wilson upper bound (every task)100%$0.04156.2 s
Haiku up to 3 tries, then Sonnet (calculation)Haiku format misses counted as passes (lenient reading)100%$0.05369.5 s
Sonnet 5.5 low effort once (calculation)Does not depend on Haiku100%$0.0127.1 s

Method

  1. This study makes no new model call and adds no new raw file. It reuses recorded calls from the hard head-to-head (Haiku 4.5 and Sonnet 5.5 at default effort, 3 calls per task) and the effort ladder (Sonnet 5.5 at low effort, 2 calls per task). All calls ran in Claude Code on the same 8 tasks with the same strict validators.
  2. The source controls ran before counted inference. Each controls receipt records 8 passing references, 26 rejected wrong answers and 8 detected wrapped format misses. The hard-set protocol outcome says 27 wrong answers; the controls receipt supports 26. Both Claude batches completed their declared call caps: 120 hard-set calls and 80 ladder calls, with no trims or errors. They used a 300 s timeout and a 16,000-output-token setting. The selected calls stayed below that setting. No Claude probe is recorded in these source folders; the hard-set Codex probe is outside this calculation.
  3. For each task and model, we take the pass rate p. It is strict passes ÷ calls, with its 95% Wilson interval. A format miss is not a pass.
  4. These are median-input scenarios. A median is not an expected mean. Cost and time can depend on whether a call fails. We do not model that link. We take the cost of one call as the median list-price cost of that task’s calls. List price is reported tokens × price. Cache reads and one-hour cache writes are priced as in the hard head-to-head.
  5. We take the time of one call as the median total time of that task’s calls. It includes CLI start-up.
  6. A policy is a list of tries. A try runs only when every earlier try failed. The policies are: Sonnet every time, Haiku once, Haiku with one retry then Sonnet, Haiku with up to 3 tries then Sonnet, and Sonnet at low effort once.
  7. Per task, let q = 1 − p(Haiku). For Haiku with one retry then Sonnet: cost = C(Haiku) × (1 + q) + C(Sonnet) × q². Time = T(Haiku) × (1 + q) + T(Sonnet) × q². Success = 1 − q² × (1 − p(Sonnet)).
  8. For Haiku with up to 3 tries then Sonnet: cost = C(Haiku) × (1 + q + q²) + C(Sonnet) × q³. Success = 1 − q³ × (1 − p(Sonnet)). Time uses the same weights.
  9. Over the 8 tasks, each task is equally likely. Cost per correct answer = total expected cost ÷ total expected correct answers. We compute time the same way, with the tries in sequence.
  10. We assume three things. A validator detects a failed answer, and running it is free. Tries are independent at each task’s observed rate. A failed last Sonnet try is a miss. Every figure that follows is a calculation.
  11. For the sensitivity, we set Haiku’s pass rate on all tasks at once to the lower end of its 95% Wilson interval, then to the upper end. A fourth setting counts Haiku’s format misses as passes. The result is a sensitivity range, not a confidence interval.
  12. A joint extreme moves both models at once: Haiku’s pass rate to the upper end and Sonnet’s to the lower end of each task’s 95% Wilson interval. It is more extreme than a 95% interval on the total, and it is not a forecast.
  13. For the break-even, we scale Haiku’s call cost (or time) by a factor f and keep all pass rates. We solve for the f at which Haiku with one retry then Sonnet equals Sonnet every time: f = (Σ C(Sonnet) − Σ C(Sonnet) × q²) ÷ Σ C(Haiku) × (1 + q). This is arithmetic, not a forecast.

Caveats

  • Advance registration is not verified. The hard-set protocol file birth time is 2026-10-06 04:03:37 UTC; its first counted call started at 03:23:59 UTC. The ladder protocol file birth time is 14:35:11 UTC; its first Claude call started at 14:21:50 UTC. Both files claim advance declaration, but the available file times do not support that claim. Copying could explain the times; we cannot establish it.
  • Wilson intervals treat calls as independent. Repeated calls on eight fixed tasks do not give a population interval for coding work. No task-level paired test or measured retry policy supports a general ranking.
  • This is a calculation, not a run. No policy ran. Costs are list-price calculations from reported tokens. The calls ran on a flat subscription, so no invoice backs them.
  • Retry independence is untested. Calls repeat the same eight prompts. Failures can repeat, and escalation after a Haiku failure may differ from a fresh Sonnet call. The sensitivity does not cover these links.
  • Samples are small: 3 calls per task for Haiku and for Sonnet, and 2 for Sonnet at low effort. A per-task pass rate has a wide interval: 0 of 3 allows up to 56%, and 3 of 3 allows as low as 44%. A median of a few calls has no interval. The sensitivity moves all 8 tasks to a bound at once. That is more extreme than a 95% interval on the total.
  • Sonnet 5.5 passed 24/24, and the model treats it as always right (95% interval 86% to 100%). So every policy that ends in a Sonnet try shows 100% success. A perfect sample does not prove perfect success on future calls. This is the ceiling effect of the task set.
  • Haiku ran at the CLI’s default effort: the effort flag was not passed, so the CLI chose. It wrote a median 5,064 output tokens per call against 1,050 for Sonnet. Output tokens enter the price calculation. The run does not show what caused the timing gap. We did not test Haiku with thinking off. The break-even stats show how far its calls would have to fall.
  • Strict grading counts a reply in a code fence as a failure. 5 of Haiku’s 13 failed calls were such format misses. With them counted as passes (lenient reading), Haiku once costs $0.0466 per correct answer, and Haiku with one retry then Sonnet costs $0.0459. Both stay above Sonnet every time.
  • This study uses each task’s median call for cost and time. The hard head-to-head divides the total cost of all calls by the passes instead. So its figures differ a little: Sonnet $0.0143 and Haiku $0.0672 per strict pass, against $0.0132 and $0.0677 here.
  • This is a hand-built task set, tuned after an earlier set hit a ceiling. It is not a random sample of coding work. This study covers the eight hard tasks only. Haiku may do better on easy tasks, where pass rate hits a ceiling. We did not recalculate the five-task head-to-head here.
  • Haiku and Sonnet at default effort ran in one batch on 2026-10-06 (UTC). Sonnet at low effort ran about 10 hours later in the effort ladder. All calls used one host and one network. We did not control for other load on the host, so wall times may carry some contention. Times include CLI start-up and the CLI’s own system prompt.
  • The sensitivity moves Haiku’s pass rate only, and the Sonnet baseline rests on 24/24. In a joint extreme (Haiku at the upper and Sonnet at the lower end of its Wilson interval on every task), the closest Haiku-first policy costs 1.3 times as much, and the quickest takes 2.8 times as long, as Sonnet every time (calculation). That setting has Sonnet succeed on only 44% of the tasks, which no run showed.

Sources

  • Repricing calculation

    Calculation ·

    Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.

  • Provider head-to-head, hard set: eight hard tasks with strict validators

    Our recorded runs ·

    Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.

    Raw data: provider-h2h-hard/receipts.json

  • Effort ladder: the hard task set at each effort level

    Our recorded runs ·

    The eight hard tasks of the hard head-to-head at low, medium and high effort (Claude Sonnet 5.5 and Claude Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI), 8 tasks × 2 repetitions per cell, declared protocols, every attempt kept. Efforts the hard head-to-head already ran are reused from its receipts as reference cells.

    Raw data: effort-ladder/receipts.json, provider-h2h-hard/receipts.json

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts”, updated October 7, 2026, https://agent.sasid.ai/benchmarks/haiku-retry-or-escalate.

More write-ups that cite this study (1)

Models and comparisons in this study

More studies

All benchmarks
Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Includes calculations
  • Prompt Caching
  • Break Even

Prompt cache break-even: after how many reuses does a cached prefix cost less?

A calculation on list prices: reuses before a cached prompt prefix costs less, per Claude and GPT model, and cost per 1,000 sessions of 1 to 20 turns.

2reuses (the 3rd request) · Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation)

3 chartsUpdated October 6, 2026

Live story
  • Head to head
  • Hard Tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.

91% (139/152)Calls that passed strictly (hard set) · n = 152

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.