Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts
On the 8 hard tasks, what does one correct answer cost, and how long does it take, if you try Claude Haiku 4.5 first and retry or escalate, against Claude Sonnet 5.5 every time?
Published · 6 charts · Download the data or a carousel
46%
The answer
The calculated Haiku-first policies cost more on these 8 hard tasks. These are scenarios from recorded calls, not a measured policy ranking. Sonnet every time costs $0.0132 per correct answer and takes 8.2 s (24/24 passes; 95% interval 86% to 100%). Haiku passes 11/24 calls (95% interval 28% to 65%); Haiku once costs $0.0677 per correct answer. Haiku with one retry, then Sonnet, costs $0.0566 and takes 71.5 s per correct answer. The tables show the other policies and sensitivity ranges; the method states the assumptions.
Key numbers
100% (24/24)
Strict passes, Claude Sonnet 5.5 · Claude Code, eight hard tasks
95% CI 86%–100% · n = 24
100% (16/16)
Strict passes, Claude Sonnet 5.5 (low) · Claude Code, eight hard tasks
95% CI 81%–100% · n = 16
$0.0132
Cost per correct answer, Sonnet 5.5 every time (calculation)
n = 24
8.2s
Wall time per correct answer, Sonnet 5.5 every time (calculation)
n = 24
$0.0677
Cost per correct answer, Haiku once (calculation)
n = 24
91.5s
Wall time per correct answer, Haiku once (calculation)
n = 24
$0.0566
Cost per correct answer, Haiku with one retry then Sonnet (calculation)
n = 48
71.5s
Wall time per correct answer, Haiku with one retry then Sonnet (calculation)
n = 48
4.3x
Haiku with one retry then Sonnet vs Sonnet every time: cost per correct answer (calculation)
n = 48
8.7x
Haiku with one retry then Sonnet vs Sonnet every time: wall time per correct answer (calculation)
n = 48
5.4x
Haiku up to 3 tries then Sonnet vs Sonnet every time: cost per correct answer (calculation)
n = 48
$0.0122
Cost per correct answer, Sonnet 5.5 low effort once (calculation)
n = 16
7.1s
Wall time per correct answer, Sonnet 5.5 low effort once (calculation)
n = 16
8 of 8
Tasks where Haiku’s median call cost more than Sonnet’s (calculation)
n = 8
8 of 8
Tasks where Haiku’s median call took longer than Sonnet’s
n = 8
5,064
Median output tokens per call, Haiku 4.5 (Claude Code)
n = 24
1,050
Median output tokens per call, Sonnet 5.5 (Claude Code)
n = 24
11.8%
Haiku call cost, as a share of measured, at which one retry then Sonnet matches Sonnet every time (calculation)
of measured · n = 8
5.4%
Haiku call time, as a share of measured, at which one retry then Sonnet matches Sonnet every time (calculation)
of measured · n = 8
3.0x
Lowest cost ratio of any Haiku-first policy to Sonnet every time, across the four Haiku pass-rate settings (sensitivity range, calculation)
n = 8
1.3x
Lowest cost ratio of any Haiku-first policy to Sonnet every time, with Haiku’s pass rate at its Wilson upper bound and Sonnet’s at its Wilson lower bound on every task at once (joint extreme, calculation)
n = 8
$0.0459
Cost per correct answer, Haiku with one retry then Sonnet, format misses counted as passes (calculation)
n = 48
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
Hover or focus a bar for its ratio to Sonnet 5.5 every time (calcul… (the highlighted row): a ratio of list-price calculations, not a measurement.
| Item | Cost per correct answer (calculation) | n |
|---|---|---|
| Sonnet 5.5 low effort once (calculation) | $0.012 | 16 |
| Sonnet 5.5 every time (calculation) | $0.013 | 24 |
| Haiku, one retry, then Sonnet (calculation) | $0.057 | 48 |
| Haiku once (calculation) | $0.068 | 24 |
| Haiku up to 3 tries, then Sonnet (calculation) | $0.072 | 48 |
List-price calculation, not a run. 5 rows. Highest Haiku up to 3 tries, then Sonnet (calculation) $0.072 (n 48). Lowest Sonnet 5.5 low effort once (calculation) $0.012 (n 16).
Notesn 16–48 per row
Eight hard tasks, each equally likely. A failed try is paid for. Calls recorded in Claude Code
Calculation, not a run. These are median-input scenarios, not measured mean costs. Total modelled list-price cost of all tries over the 8 tasks, divided by the expected number of correct answers. Per-call cost is each task’s median call at list price (reported tokens × list price); the calls ran on a flat subscription. A try runs only when the earlier tries failed the validator, and tries are independent at each task’s observed pass rate. n is the number of recorded calls behind each bar. A calculation has no confidence interval; the sensitivity chart shows a range. Highlighted: the baseline, Sonnet 5.5 every time.
Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)
Hover or focus a bar for its ratio to Sonnet 5.5 every time (calcul… (the highlighted row): a ratio of list-price calculations, not a measurement.
| Item | Time per correct answer (calculation) | n |
|---|---|---|
| Sonnet 5.5 low effort once (calculation) | 7.1 s | 16 |
| Sonnet 5.5 every time (calculation) | 8.2 s | 24 |
| Haiku, one retry, then Sonnet (calculation) | 71.5 s | 48 |
| Haiku once (calculation) | 91.5 s | 24 |
| Haiku up to 3 tries, then Sonnet (calculation) | 93.4 s | 48 |
List-price calculation, not a run. 5 rows. Slowest Haiku up to 3 tries, then Sonnet (calculation) 93.4 s (n 48). Fastest Sonnet 5.5 low effort once (calculation) 7.1 s (n 16).
Notesn 16–48 per row
Tries run one after another, so a failed try adds its time
Calculation, not a run. These are median-input scenarios, not measured mean times. Each task’s median total call time (CLI start-up included), summed over the tries a policy expects to run, divided by the expected number of correct answers. Tries are assumed to run one after another. Parallel tries were not tested. One host, one network. A calculation has no confidence interval. Highlighted: the baseline, Sonnet 5.5 every time.
Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)
Hover or focus a bar for its ratio to Sonnet 5.5 every time (calcul… (the highlighted row): a ratio of list-price calculations, not a measurement.
| Item | Expected success rate | 95% interval | n |
|---|---|---|---|
| Sonnet 5.5 every time (calculation) | 100% | 86%–100% | 24 |
| Haiku, one retry, then Sonnet (calculation) | 100% | — | 48 |
| Haiku up to 3 tries, then Sonnet (calculation) | 100% | — | 48 |
| Sonnet 5.5 low effort once (calculation) | 100% | 81%–100% | 16 |
| Haiku once (calculation) | 46% | 28%–65% | 24 |
List-price calculation, not a run. 5 rows. Highest Sonnet 5.5 every time (calculation) 100% (95% interval 86%–100%, n 24). Lowest Haiku once (calculation) 46% (95% interval 28%–65%, n 24). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 16–48 per row4 of 5 at 100%: this task set cannot separate them.
Share of the eight hard tasks answered correctly under the strict validator
Whiskers are 95% Wilson intervals on the measured strict pass rate and appear only on the three one-try policies (24/24, 11/24 and 16/16). The two escalation policies are calculations without an interval. A value of 100% means no failure was recorded: the model treats Sonnet 5.5 as always right (24/24, 95% interval 86% to 100%), so the escalation policies cannot miss. n is the number of recorded calls behind each bar.
Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)
Hover or focus a bar for its ratio to DST day length (the lowest value): a ratio of the two values shown, not a measurement.
| Item | Strict pass rate (n = 3 per task) | 95% interval | n |
|---|---|---|---|
| Interval merge fix | 100% | 44%–100% | 3 |
| DST day length | 33% | 6.2%–79% | 3 |
| CSV parser | 67% | 21%–94% | 3 |
| Event-loop order | 0% | 0%–56% | 3 |
| Room schedule | 0% | 0%–56% | 3 |
| SemVer regex | 100% | 44%–100% | 3 |
| Money refactor | 67% | 21%–94% | 3 |
| SQLite report query | 0% | 0%–56% | 3 |
8 rows. Highest Interval merge fix 100% (95% interval 44%–100%, n 3). Lowest SQLite report query 0% (95% interval 0%–56%, n 3). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 3 per row
Claude Haiku 4.5 · Claude Code, default effort. 3 calls per task
Measured. Whiskers are 95% Wilson intervals on only 3 calls per task, so they are wide: 0 of 3 allows a true rate up to 56%, and 3 of 3 allows one as low as 44%. A format miss (a right answer in a code fence or with prose) is not a pass. The retry policies take each task’s rate from this chart.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
- Claude Haiku 4.5 · Claude Code
- Claude Sonnet 5.5 · Claude Code (square)
Gap labels, Claude Sonnet 5.5 · Claude Code vs Claude Haiku 4.5 · Claude Code: Claude Sonnet 5.5 · Claude Code is x% higher (+) or lower (−) than Claude Haiku 4.5 · Claude Code, calculated from the two values shown (the change counted from Claude Haiku 4.5 · Claude Code’s value).
| Item | Claude Haiku 4.5 · Claude Code | Claude Sonnet 5.5 · Claude Code | n |
|---|---|---|---|
| Interval merge fix | $0.018 | $0.0056 | 3 |
| DST day length | $0.029 | $0.025 | 3 |
| CSV parser | $0.029 | $0.015 | 3 |
| Event-loop order | $0.037 | $0.016 | 3 |
| Room schedule | $0.036 | $0.012 | 3 |
| SemVer regex | $0.042 | $0.0051 | 3 |
| Money refactor | $0.02 | $0.0096 | 3 |
| SQLite report query | $0.036 | $0.017 | 3 |
List-price calculation, not a run. 8 rows, 2 series: Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code. Claude Haiku 4.5 · Claude Code: highest SemVer regex $0.042 (n 3). Lowest Interval merge fix $0.018 (n 3). Claude Sonnet 5.5 · Claude Code: highest DST day length $0.025 (n 3). Lowest SemVer regex $0.0051 (n 3).
Notesn = 3 per row
Median of each task’s calls, default effort, Claude Code
Calculation: reported tokens × list price for each call, then the median per task; the calls ran on a flat subscription. Haiku’s list price per token is lower. Its median calls cost more and contained more output tokens. This does not isolate the effect of effort or thinking. 3 calls per task and configuration; the table shows each call-cost range, not a confidence interval.
Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)
Every interval overlaps every other: this chart does not order these rows.
| Item | Cost per correct answer (calculation) | Range (lowest–highest run) | n |
|---|---|---|---|
| Sonnet 5.5 low effort once (calculation) | $0.012 | — | 16 |
| Sonnet 5.5 every time (calculation) | $0.013 | — | 24 |
| Haiku, one retry, then Sonnet (calculation) | $0.057 | $0.039–$0.068 | 48 |
| Haiku once (calculation) | $0.068 | $0.039–$0.18 | 24 |
| Haiku up to 3 tries, then Sonnet (calculation) | $0.072 | $0.041–$0.09 | 48 |
List-price calculation, not a run. 5 rows. Highest Haiku up to 3 tries, then Sonnet (calculation) $0.072 (range $0.041–$0.09, n 48). Lowest Sonnet 5.5 low effort once (calculation) $0.012 (n 16). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 16–48 per row
Dot: Haiku’s pass rate as observed. Whisker: its pass rate on every task set to its Wilson lower and upper bound at once
Calculation. This is a sensitivity range, not a confidence interval: each Haiku policy is recomputed with Haiku’s pass rate on all 8 tasks at once set to the lower and then the upper end of that task’s 95% Wilson interval (3 calls per task, so the ends are far apart), and the whisker spans the lowest and highest result. That is more extreme than a 95% interval on the total. The two Sonnet policies do not depend on Haiku’s pass rate and have no whisker. Highlighted: the baseline, Sonnet 5.5 every time.
Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)
Tables
Retry policies over the eight hard tasks (calculation)
| Policy | Tries in order | Expected success | Expected cost per task (all tries) | Cost per correct answer | Wall time per correct answer | Cost vs Sonnet every time | Time vs Sonnet every time |
|---|---|---|---|---|---|---|---|
| Sonnet 5.5 every time (calculation) | Sonnet 5.5 | 100% | $0.013 | $0.013 | 8.2 s | 1x | 1x |
| Haiku once (calculation) | Haiku 4.5 | 46% | $0.031 | $0.068 | 91.5 s | 5.1x | 11.1x |
| Haiku, one retry, then Sonnet (calculation) | Haiku 4.5 → Haiku 4.5 → Sonnet 5.5 | 100% | $0.057 | $0.057 | 71.5 s | 4.3x | 8.7x |
| Haiku up to 3 tries, then Sonnet (calculation) | Haiku 4.5 → Haiku 4.5 → Haiku 4.5 → Sonnet 5.5 | 100% | $0.072 | $0.072 | 93.4 s | 5.4x | 11.3x |
| Sonnet 5.5 low effort once (calculation) | Sonnet 5.5 low | 100% | $0.012 | $0.012 | 7.1 s | 0.92x | 0.86x |
Per-task inputs: passes, 95% intervals, median call cost (calculation) and time; ranges are min–max, not confidence intervals
| Task | Claude Haiku 4.5 · Claude Code: strict passes | Haiku 95% interval | Haiku median call cost | Haiku call-cost range (calculation) | Haiku median call time | Haiku call-time range | Claude Sonnet 5.5 · Claude Code: strict passes | Sonnet 95% interval | Sonnet median call cost | Sonnet call-cost range (calculation) | Sonnet median call time | Sonnet call-time range | Claude Sonnet 5.5 (low) · Claude Code: strict passes | Sonnet (low) 95% interval | Sonnet (low) median call cost | Sonnet (low) call-cost range (calculation) | Sonnet (low) median call time | Sonnet (low) call-time range |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Fix an interval-merge function (off-by-one and edge cases) | 3/3 | 44% to 100% | $0.018 | $0.0148 to $0.0183 | 21.2 s | 18.5 s to 24.7 s | 3/3 | 44% to 100% | $0.0056 | $0.0056 to $0.0091 | 2.9 s | 2.4 s to 3.1 s | 2/2 | 34% to 100% | $0.0074 | $0.0056 to $0.0092 | 2.8 s | 2.8 s to 2.8 s |
| Fix a time-zone day-length function (DST) | 1/3 | 6% to 79% | $0.029 | $0.0243 to $0.0308 | 39 s | 26.2 s to 40.6 s | 3/3 | 44% to 100% | $0.025 | $0.0231 to $0.0424 | 21.6 s | 19.6 s to 34.8 s | 2/2 | 34% to 100% | $0.024 | $0.0218 to $0.0261 | 18.1 s | 16.3 s to 20.0 s |
| Write a CSV parser (quoted newlines, strict errors) | 2/3 (+1 format miss) | 21% to 94% | $0.029 | $0.0250 to $0.0505 | 39 s | 33.8 s to 75.1 s | 3/3 | 44% to 100% | $0.015 | $0.0146 to $0.0161 | 9.6 s | 8.8 s to 10.6 s | 2/2 | 34% to 100% | $0.0093 | $0.0091 to $0.0096 | 4.4 s | 4.3 s to 4.4 s |
| Predict JavaScript event-loop output order | 0/3 | 0% to 56% | $0.037 | $0.0231 to $0.0395 | 54 s | 26.0 s to 56.6 s | 3/3 | 44% to 100% | $0.016 | $0.0147 to $0.0161 | 9.6 s | 8.2 s to 9.9 s | 2/2 | 34% to 100% | $0.014 | $0.0139 to $0.0141 | 9.3 s | 8.8 s to 9.7 s |
| Solve a multi-constraint room schedule | 0/3 (+2 format misses) | 0% to 56% | $0.036 | $0.0292 to $0.0444 | 54.6 s | 38.0 s to 68.8 s | 3/3 | 44% to 100% | $0.012 | $0.0118 to $0.0126 | 7.7 s | 7.4 s to 7.8 s | 2/2 | 34% to 100% | $0.011 | $0.0113 to $0.0115 | 6.6 s | 6.5 s to 6.7 s |
| Write a strict SemVer 2.0.0 regex | 3/3 | 44% to 100% | $0.042 | $0.0373 to $0.0447 | 64.5 s | 54.5 s to 67.0 s | 3/3 | 44% to 100% | $0.0051 | $0.0051 to $0.0079 | 2.5 s | 2.3 s to 3.7 s | 2/2 | 34% to 100% | $0.0051 | $0.0051 to $0.0051 | 3.2 s | 2.9 s to 3.5 s |
| Refactor to remove duplication, keep 20 tests green | 2/3 | 21% to 94% | $0.02 | $0.0179 to $0.0216 | 15.9 s | 15.3 s to 17.2 s | 3/3 | 44% to 100% | $0.0096 | $0.0094 to $0.0151 | 3.6 s | 3.4 s to 7.2 s | 2/2 | 34% to 100% | $0.0095 | $0.0094 to $0.0095 | 4.4 s | 3.7 s to 5.2 s |
| Write a SQLite reporting query (fan-out, ties, boundaries) | 0/3 (+2 format misses) | 0% to 56% | $0.036 | $0.0289 to $0.0406 | 47.1 s | 38.9 s to 56.1 s | 3/3 | 44% to 100% | $0.017 | $0.0157 to $0.0192 | 8.5 s | 7.6 s to 8.9 s | 2/2 | 34% to 100% | $0.017 | $0.0167 to $0.0172 | 7.8 s | 7.7 s to 7.9 s |
Sensitivity of each policy to Haiku’s pass rate (calculation)
| Policy | Haiku pass rate set to | Expected success | Cost per correct answer | Wall time per correct answer |
|---|---|---|---|---|
| Sonnet 5.5 every time (calculation) | Does not depend on Haiku | 100% | $0.013 | 8.2 s |
| Haiku once (calculation) | Haiku pass rate at the Wilson lower bound (every task) | 17% | $0.18 | 248 s |
| Haiku once (calculation) | Haiku pass rate as observed | 46% | $0.068 | 91.5 s |
| Haiku once (calculation) | Haiku pass rate at the Wilson upper bound (every task) | 79% | $0.039 | 52.8 s |
| Haiku once (calculation) | Haiku format misses counted as passes (lenient reading) | 67% | $0.047 | 62.9 s |
| Haiku, one retry, then Sonnet (calculation) | Haiku pass rate at the Wilson lower bound (every task) | 100% | $0.068 | 84.3 s |
| Haiku, one retry, then Sonnet (calculation) | Haiku pass rate as observed | 100% | $0.057 | 71.5 s |
| Haiku, one retry, then Sonnet (calculation) | Haiku pass rate at the Wilson upper bound (every task) | 100% | $0.039 | 52.6 s |
| Haiku, one retry, then Sonnet (calculation) | Haiku format misses counted as passes (lenient reading) | 100% | $0.046 | 59.5 s |
| Haiku up to 3 tries, then Sonnet (calculation) | Haiku pass rate at the Wilson lower bound (every task) | 100% | $0.09 | 115 s |
| Haiku up to 3 tries, then Sonnet (calculation) | Haiku pass rate as observed | 100% | $0.072 | 93.4 s |
| Haiku up to 3 tries, then Sonnet (calculation) | Haiku pass rate at the Wilson upper bound (every task) | 100% | $0.041 | 56.2 s |
| Haiku up to 3 tries, then Sonnet (calculation) | Haiku format misses counted as passes (lenient reading) | 100% | $0.053 | 69.5 s |
| Sonnet 5.5 low effort once (calculation) | Does not depend on Haiku | 100% | $0.012 | 7.1 s |
Method
- This study makes no new model call and adds no new raw file. It reuses recorded calls from the hard head-to-head (Haiku 4.5 and Sonnet 5.5 at default effort, 3 calls per task) and the effort ladder (Sonnet 5.5 at low effort, 2 calls per task). All calls ran in Claude Code on the same 8 tasks with the same strict validators.
- The source controls ran before counted inference. Each controls receipt records 8 passing references, 26 rejected wrong answers and 8 detected wrapped format misses. The hard-set protocol outcome says 27 wrong answers; the controls receipt supports 26. Both Claude batches completed their declared call caps: 120 hard-set calls and 80 ladder calls, with no trims or errors. They used a 300 s timeout and a 16,000-output-token setting. The selected calls stayed below that setting. No Claude probe is recorded in these source folders; the hard-set Codex probe is outside this calculation.
- For each task and model, we take the pass rate p. It is strict passes ÷ calls, with its 95% Wilson interval. A format miss is not a pass.
- These are median-input scenarios. A median is not an expected mean. Cost and time can depend on whether a call fails. We do not model that link. We take the cost of one call as the median list-price cost of that task’s calls. List price is reported tokens × price. Cache reads and one-hour cache writes are priced as in the hard head-to-head.
- We take the time of one call as the median total time of that task’s calls. It includes CLI start-up.
- A policy is a list of tries. A try runs only when every earlier try failed. The policies are: Sonnet every time, Haiku once, Haiku with one retry then Sonnet, Haiku with up to 3 tries then Sonnet, and Sonnet at low effort once.
- Per task, let q = 1 − p(Haiku). For Haiku with one retry then Sonnet: cost = C(Haiku) × (1 + q) + C(Sonnet) × q². Time = T(Haiku) × (1 + q) + T(Sonnet) × q². Success = 1 − q² × (1 − p(Sonnet)).
- For Haiku with up to 3 tries then Sonnet: cost = C(Haiku) × (1 + q + q²) + C(Sonnet) × q³. Success = 1 − q³ × (1 − p(Sonnet)). Time uses the same weights.
- Over the 8 tasks, each task is equally likely. Cost per correct answer = total expected cost ÷ total expected correct answers. We compute time the same way, with the tries in sequence.
- We assume three things. A validator detects a failed answer, and running it is free. Tries are independent at each task’s observed rate. A failed last Sonnet try is a miss. Every figure that follows is a calculation.
- For the sensitivity, we set Haiku’s pass rate on all tasks at once to the lower end of its 95% Wilson interval, then to the upper end. A fourth setting counts Haiku’s format misses as passes. The result is a sensitivity range, not a confidence interval.
- A joint extreme moves both models at once: Haiku’s pass rate to the upper end and Sonnet’s to the lower end of each task’s 95% Wilson interval. It is more extreme than a 95% interval on the total, and it is not a forecast.
- For the break-even, we scale Haiku’s call cost (or time) by a factor f and keep all pass rates. We solve for the f at which Haiku with one retry then Sonnet equals Sonnet every time: f = (Σ C(Sonnet) − Σ C(Sonnet) × q²) ÷ Σ C(Haiku) × (1 + q). This is arithmetic, not a forecast.
Caveats
- Advance registration is not verified. The hard-set protocol file birth time is 2026-10-06 04:03:37 UTC; its first counted call started at 03:23:59 UTC. The ladder protocol file birth time is 14:35:11 UTC; its first Claude call started at 14:21:50 UTC. Both files claim advance declaration, but the available file times do not support that claim. Copying could explain the times; we cannot establish it.
- Wilson intervals treat calls as independent. Repeated calls on eight fixed tasks do not give a population interval for coding work. No task-level paired test or measured retry policy supports a general ranking.
- This is a calculation, not a run. No policy ran. Costs are list-price calculations from reported tokens. The calls ran on a flat subscription, so no invoice backs them.
- Retry independence is untested. Calls repeat the same eight prompts. Failures can repeat, and escalation after a Haiku failure may differ from a fresh Sonnet call. The sensitivity does not cover these links.
- Samples are small: 3 calls per task for Haiku and for Sonnet, and 2 for Sonnet at low effort. A per-task pass rate has a wide interval: 0 of 3 allows up to 56%, and 3 of 3 allows as low as 44%. A median of a few calls has no interval. The sensitivity moves all 8 tasks to a bound at once. That is more extreme than a 95% interval on the total.
- Sonnet 5.5 passed 24/24, and the model treats it as always right (95% interval 86% to 100%). So every policy that ends in a Sonnet try shows 100% success. A perfect sample does not prove perfect success on future calls. This is the ceiling effect of the task set.
- Haiku ran at the CLI’s default effort: the effort flag was not passed, so the CLI chose. It wrote a median 5,064 output tokens per call against 1,050 for Sonnet. Output tokens enter the price calculation. The run does not show what caused the timing gap. We did not test Haiku with thinking off. The break-even stats show how far its calls would have to fall.
- Strict grading counts a reply in a code fence as a failure. 5 of Haiku’s 13 failed calls were such format misses. With them counted as passes (lenient reading), Haiku once costs $0.0466 per correct answer, and Haiku with one retry then Sonnet costs $0.0459. Both stay above Sonnet every time.
- This study uses each task’s median call for cost and time. The hard head-to-head divides the total cost of all calls by the passes instead. So its figures differ a little: Sonnet $0.0143 and Haiku $0.0672 per strict pass, against $0.0132 and $0.0677 here.
- This is a hand-built task set, tuned after an earlier set hit a ceiling. It is not a random sample of coding work. This study covers the eight hard tasks only. Haiku may do better on easy tasks, where pass rate hits a ceiling. We did not recalculate the five-task head-to-head here.
- Haiku and Sonnet at default effort ran in one batch on 2026-10-06 (UTC). Sonnet at low effort ran about 10 hours later in the effort ladder. All calls used one host and one network. We did not control for other load on the host, so wall times may carry some contention. Times include CLI start-up and the CLI’s own system prompt.
- The sensitivity moves Haiku’s pass rate only, and the Sonnet baseline rests on 24/24. In a joint extreme (Haiku at the upper and Sonnet at the lower end of its Wilson interval on every task), the closest Haiku-first policy costs 1.3 times as much, and the quickest takes 2.8 times as long, as Sonnet every time (calculation). That setting has Sonnet succeed on only 44% of the tasks, which no run showed.
Sources
Repricing calculation
Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.
Provider head-to-head, hard set: eight hard tasks with strict validators
Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.
Effort ladder: the hard task set at each effort level
The eight hard tasks of the hard head-to-head at low, medium and high effort (Claude Sonnet 5.5 and Claude Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI), 8 tasks × 2 repetitions per cell, declared protocols, every attempt kept. Efforts the hard head-to-head already ran are reused from its receipts as reference cells.
Anthropic list prices (Claude models)
Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts”, updated October 7, 2026, https://agent.sasid.ai/benchmarks/haiku-retry-or-escalate.
More write-ups that cite this study (1)
Models and comparisons in this study
Write-ups on this study
Is Claude Haiku cheaper if you retry or escalate to Sonnet? A calculation on real receipts
Haiku 4.5 first, one retry, then Sonnet 5.5? On 8 hard tasks, the calculation gives 4.3x the cost and 8.7x the time per correct answer versus Sonnet alone.
More studies
All benchmarksHow much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.
Prompt cache break-even: after how many reuses does a cached prefix cost less?
A calculation on list prices: reuses before a cached prompt prefix costs less, per Claude and GPT model, and cost per 1,000 sessions of 1 to 20 turns.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.