Is Claude Haiku cheaper if you retry or escalate to Sonnet? A calculation on real receipts
Haiku 4.5 first, one retry, then Sonnet 5.5? On 8 hard tasks, the calculation gives 4.3x the cost and 8.7x the time per correct answer versus Sonnet alone.
TL;DR
- No, not on these tasks. On 8 hard tasks, Sonnet 5.5 every time cost $0.0132 per correct answer and took 8.2 s. Haiku 4.5 with one retry, then Sonnet, cost $0.0566 (4.3 times) and took 71.5 s (8.7 times). Every figure here is a calculation on recorded calls. No policy ran.
- Haiku once costs $0.0677 and takes 91.5 s per correct answer. It passed 11 of 24 calls (46%, 95% interval 28% to 65%). Sonnet passed 24 of 24 (86% to 100%).
- Three Haiku tries do not reach parity in this calculation. Up to three Haiku tries, then Sonnet, cost $0.0719 and took 93.4 s. Their sensitivity ranges overlap, so we do not rank the three Haiku-first policies against each other. Each one stays above Sonnet every time.
- Token price does not tell the whole cost. Haiku lists at half Sonnet's price per token. Yet its median call cost more than Sonnet's on 8 of 8 tasks. Its median output was 5,064 tokens against 1,050 (24 calls each). The call ranges were 1,899 to 9,321 and 176 to 3,895 tokens.
- The best case for Haiku does not reach parity. With Haiku's pass rate on every task set to the top of its interval, the closest Haiku-first case cost 3.0 times as much (a sensitivity range, not an interval). If Sonnet's pass rate is also set to the bottom of its interval on every task, the closest case is 1.3 times.
- What would have to change: for cost parity, Haiku's calls would need to cost about 12% of what they cost here. For time parity, they would need to take about 5% of the time (two separate break-evens, arithmetic, not a forecast).
- Limits. One task set and the CLI's default effort. This calculation excludes the separate thinking-off experiment: 4 of 24 strict passes (17%, 95% interval 7% to 36%). Its strict interval overlaps the thinking-on interval. Read the caveats.
Disclosure: I build Agent. Agent ran these benchmarks, and its default coding model is Sonnet 5.5. The receipts, formulas and code are public, so you can check every number.
Study with all charts, tables and formulas: /benchmarks/haiku-retry-or-escalate.
The plan, and why it sounds good
Many teams plan to start with a cheap model. A check tells them when an answer is wrong. If it is wrong, they retry, or they send the task to a stronger model. On paper this saves money. The price per token is lower, and the cheap model passes some tasks.
We tested that plan on paper, with real numbers. We did not run it. We took recorded calls and calculated what each plan would cost per correct answer. For the short overview of Haiku against Sonnet on cost, read Is Claude Haiku cheaper than Sonnet?. This post goes deeper into the retry and escalate policies.
Five policies
| Policy | What runs |
|---|---|
| Sonnet 5.5 every time | One Sonnet 5.5 try |
| Haiku once | One Haiku 4.5 try, no retry |
| Haiku, one retry, then Sonnet | Haiku. If the check fails, Haiku again. If it fails again, Sonnet |
| Haiku up to 3 tries, then Sonnet | Haiku up to three times, then Sonnet |
| Sonnet 5.5 low effort once | One Sonnet 5.5 try with effort set to low |
A try runs only when every earlier try failed. A failed try is paid for. If the last Sonnet try fails, the task is a miss.
Where the numbers come from
The inputs are 64 recorded calls in Claude Code on the same 8 hard tasks. Each task has a strict validator. Haiku 4.5 and Sonnet 5.5 at default effort made 24 calls each (3 per task). Sonnet 5.5 at low effort made 16 (2 per task). A reply in the wrong format, such as a code fence, counts as a failure.
Per task and model, we take three numbers: the strict pass rate, the median list-price cost of a call, and the median time of a call. Then we compute what each policy expects to spend.
Haiku passed the interval merge and the SemVer regex 3 of 3 each (95% interval 44% to 100%). It passed none of the event-loop order, the room schedule and the SQL query (0 of 3 each; 95% interval 0% to 56%). On the room schedule and the SQL query, 2 of its 3 replies each had the right answer in the wrong format. Strict grading does not count them. With 3 calls per task, each interval is wide: 0 of 3 allows a true rate up to 56%.
Result 1: cost per correct answer
Cost per correct answer, list-price calculation, in order of the point estimate. The inputs are 24 calls for each default-effort model and 16 for low-effort Sonnet. Each escalation scenario uses 48 recorded calls; these are input counts, not executed policy tries:
- Sonnet 5.5 low effort once: $0.0122 (16 calls, 16 of 16 passed; 95% interval 81% to 100%, a ceiling)
- Sonnet 5.5 every time: $0.0132 (the baseline)
- Haiku, one retry, then Sonnet: $0.0566, 4.3 times the baseline
- Haiku once: $0.0677, 5.1 times
- Haiku up to 3 tries, then Sonnet: $0.0719, 5.4 times
Every Haiku-first scenario costs more than Sonnet every time in all four sensitivity settings below. These scenarios do not establish a measured policy ranking. The three Haiku-first policies overlap each other there, and the two Sonnet policies have no interval, so we do not rank within those groups.
One Haiku try costs $0.0310 on average (the mean of the eight task medians). That is 2.3 times one Sonnet try (calculation) at $0.0132. These mean-of-median costs have no interval. And Haiku passes under half the time, so a correct answer costs $0.0677. The scenario gives each retry the same task-specific cost and pass rate as the first try. It charges a Sonnet call only if all Haiku tries fail. Actual retries and follow-up Sonnet calls were not tested.
Sonnet at low effort comes out at $0.0122, a calculated 8% below the default-effort figure. The calculation has no interval, so we do not call that a difference. The effort ladder found that Sonnet's per-call time ranges overlap across effort settings (does reasoning effort buy quality?).
Result 2: wall time per correct answer
This calculation runs tries one after another, so a failed try adds its time. It uses the same 24-call, 16-call and 48-call inputs as the cost scenarios. The sensitivity table shows scenario ranges, not confidence intervals.
- Sonnet 5.5 low effort once: 7.1 s
- Sonnet 5.5 every time: 8.2 s
- Haiku, one retry, then Sonnet: 71.5 s, 8.7 times
- Haiku once: 91.5 s, 11.1 times
- Haiku up to 3 tries, then Sonnet: 93.4 s, 11.3 times
As with cost, this is a scenario comparison, not a measured policy ranking. In all four settings of the sensitivity table, every Haiku-first policy takes longer than Sonnet every time. We do not rank the policies inside each group.
Haiku's median call time exceeded Sonnet's on 8 of 8 tasks. The task medians span 15.9 s to 64.5 s for Haiku and 2.5 s to 21.6 s for Sonnet. These spans are not run ranges. Each row below shows the median and the actual run range (3 calls per model per task). Ranges are not confidence intervals. The DST ranges overlap, so its medians do not establish a speed ranking. Parallel tries were not tested.
| Task | Haiku median; run range | Sonnet median; run range |
|---|---|---|
| Interval merge | 21.2 s; 18.5 s to 24.7 s | 2.9 s; 2.4 s to 3.1 s |
| DST day length | 39.0 s; 26.2 s to 40.6 s | 21.6 s; 19.6 s to 34.8 s |
| CSV parser | 39.0 s; 33.8 s to 75.1 s | 9.6 s; 8.8 s to 10.6 s |
| Event-loop order | 54.0 s; 26.0 s to 56.6 s | 9.6 s; 8.2 s to 9.9 s |
| Room schedule | 54.6 s; 38.0 s to 68.8 s | 7.7 s; 7.4 s to 7.8 s |
| SemVer regex | 64.5 s; 54.5 s to 67.0 s | 2.5 s; 2.3 s to 3.7 s |
| Money refactor | 15.9 s; 15.3 s to 17.2 s | 3.6 s; 3.4 s to 7.2 s |
| SQLite report query | 47.1 s; 38.9 s to 56.1 s | 8.5 s; 7.6 s to 8.9 s |
Result 3: how often each policy is right
Haiku once passed 46% of recorded calls (11 of 24; 95% interval 28% to 65%). The two escalation policies calculate 100% success from 48 recorded calls, with no confidence interval. Sonnet at low effort also passed every call (16 of 16; 95% interval 81% to 100%). That is a ceiling effect, not proof. Sonnet passed 24 of 24, and the model treats it as always right. Its 95% interval is 86% to 100%, which does not prove perfect success on future calls.
Why the Haiku-first calculation costs more
Per token, Haiku is cheaper. It lists at $1 per million input tokens and $5 per million output tokens. Sonnet lists at $2 and $10.
Per call, Haiku cost more on all 8 tasks. Output tokens enter the price calculation. This does not isolate the effect of thinking or effort. At the CLI's default effort, Haiku wrote a median 5,064 output tokens per call. Sonnet wrote 1,050. Each median uses 24 calls. Haiku's call range was 1,899 to 9,321 tokens; Sonnet's was 176 to 3,895. These are ranges, not confidence intervals. That is 4.8 times the tokens at half the price, about 2.4 times the output cost (our arithmetic). Our hard-task post shows that most of Haiku's output was reasoning.
The hard head-to-head calculated cost per pass in a different way: total cost of all calls divided by strict passes. That gives Haiku $0.0672 and Sonnet $0.0143 per strict pass. The retry study uses each task's median call instead, so Sonnet reads $0.0132 here. Both views agree.
How wrong could we be?
Haiku's pass rate rests on 3 calls per task. So we recomputed each Haiku policy with its pass rate on all 8 tasks set to the lower end, then the upper end, of each task's 95% interval. That moves every task to a bound at once. It is a sensitivity range, not a confidence interval. It does not form a confidence interval on the total.
- Haiku, one retry, then Sonnet: $0.0394 to $0.0681 per correct answer. The baseline is $0.0132.
- Haiku once: $0.0391 to $0.1834.
- Haiku up to 3 tries, then Sonnet: $0.0415 to $0.0905.
In every setting, every Haiku-first policy cost more than the baseline. The closest any came was about 3.0 times (Haiku once, with every task at the top of its interval).
That test moves only Haiku. Sonnet's pooled 24 of 24 has a 95% interval of 86% to 100%. Each task's 3 of 3 has a 95% interval of 44% to 100%. The joint extreme uses these per-task bounds. So we also ran a joint extreme: Haiku at the top of its interval and Sonnet at the bottom of its interval, on every task at once. Sonnet every time then succeeds on only 44% of tasks, which no run showed. Even then the closest Haiku-first policy (Haiku once) costs 1.3 times as much and the quickest takes 2.8 times as long as Sonnet every time. It is a calculation on 3 calls per task, not a forecast.
One more reading is fair to Haiku. Five of its 13 failed calls had the right answer in the wrong format. If a harness strips the code fence and accepts them, Haiku once costs $0.0466 per correct answer and Haiku with one retry, then Sonnet, costs $0.0459. Both stay above Sonnet's $0.0132.
What would have to be true
Keep Haiku's pass rates as observed and scale only its call cost. Then Haiku with one retry, then Sonnet, matches Sonnet every time when Haiku's calls cost 11.8% of what they cost here. For time, they would need to take 5.4% of the time. This is arithmetic, not a forecast. These break-evens keep the default-effort pass rates fixed. They do not predict a thinking-off retry policy. The separate thinking-on vs off study recorded 4 of 24 strict passes with thinking off (17%, 95% interval 7% to 36%) on the same 8 hard tasks. Its strict interval overlaps thinking on (11 of 24, 28% to 65%). The arms ran at different times on different accounts, so the study cannot isolate a thinking effect. This calculation excludes those thinking-off calls; no thinking-off retry policy ran.
Repeated errors limit the case for retries
The calculation treats each try as independent. That may favour Haiku if its failures repeat. In our consistency test, Haiku answered 289 on an exact-number prompt in 10 of 10 repeats (wrong-answer rate 100%, 95% interval 72% to 100%). The expected answer was 282. Those repeats found no fix; they do not prove that every future retry will fail. And on three of the 8 hard tasks, Haiku passed 0 of 3 calls each (95% interval 0% to 56%). The calculation already sends those tasks to Sonnet. See same prompt, ten answers.
Best practice
- Price a plan in cost per correct answer. Keep failed tries in the cost.
- Compare cost per call, not price per token. Input, cache and output tokens enter the cost. Haiku's median output was 4.8 times Sonnet's (calculation).
- Repeat a prompt before you add a retry. If the wrong answer repeats, escalate at once or change the prompt.
- A cheap first try pays only if three things hold: the check is cheap, the first try often passes, and its calls cost well below the strong model's. Here, the check was free by assumption. The other two did not hold: Haiku passed 11 of 24 calls (46%, 95% interval 28% to 65%), and its calls cost more.
- Set effort and thinking before you compare models. Haiku ran at the default and thought a lot.
- Count the wait. Tries in a row multiply the wall time. Here, one retry meant a calculated 8.7 times the time.
- Measure on your own tasks. Eight tasks cannot speak for yours.
How we calculated
- Inputs. 64 recorded calls from the hard head-to-head and the effort ladder. No new call was made.
- Per task and model. Pass rate = strict passes ÷ calls. Cost = median of each call's list-price cost, with plain input, cache reads, cache writes and output priced separately. Time = median total call time, CLI start-up included.
- A policy. Expected cost = sum over tries of (chance the try runs) × its cost. For Haiku, one retry, then Sonnet, with q = 1 − Haiku pass rate: cost = C(Haiku) × (1 + q) + C(Sonnet) × q².
- Over the 8 tasks. Each task is equally likely. Cost per correct answer = total expected cost ÷ total expected correct answers. Time is built the same way.
- Sensitivity. Haiku's pass rate on every task at the lower, then the upper, end of its 95% Wilson interval. The joint extreme also sets Sonnet's pass rate to the lower end. Both are sensitivity ranges, not confidence intervals.
- Assumptions. A check detects a failed answer (true for this set, and free here). Tries are independent at the observed rate. A failed last Sonnet try is a miss.
Formulas and every table: /benchmarks/haiku-retry-or-escalate. Price your own mix with the AI cost calculator.
Caveats
- A calculation, not a run. No policy ran. The calls ran on a flat subscription, so costs are list-price estimates, not bills.
- Protocol timing. Advance registration is not verified. Both source protocol files have birth times after their first counted calls. The study lists the exact times. Copying could explain them, but the files cannot establish that.
- Scope and dependence. The hand-built hard tasks followed an earlier ceiling. Repeated calls on these fixed tasks do not give a population interval for coding work. Retry independence is untested. A Sonnet call after a Haiku failure may differ from a fresh Sonnet call. The sensitivity does not cover these links.
- Median inputs. Cost and time scenarios use each task's median call, not measured mean policy costs or times.
- Small samples. 3 calls per task for Haiku and Sonnet, 2 for Sonnet at low effort. A median of 3 calls has no interval.
- Default effort. The effort flag was not passed for Haiku and Sonnet in these inputs. This calculation excludes the separate thinking-off experiment: 4 of 24 strict passes (17%, 95% interval 7% to 36%). Its strict interval overlaps thinking on, and day and account differ. That experiment does not validate these fixed-pass-rate break-evens.
- One task set. These are the 8 hard tasks. Haiku may do better on easy tasks, where pass rate hits a ceiling. The five-task head-to-head is not recalculated here.
- Sonnet is treated as always right. It passed 24 of 24. That is a ceiling, not a guarantee. The joint extreme above lets it fall to 44% and still finds no parity.
- One host. One network. Low-effort Sonnet ran about 10 hours after the default-effort batch. We did not control for other load on the host, so times may carry some contention. Times include CLI start-up.
What to read next
- Is Claude Haiku cheaper than Sonnet? The short version
- When tasks get hard: Haiku vs Sonnet vs Opus vs Fable
- Same prompt, ten answers
- Does reasoning effort buy quality?
- Side by side: Claude Haiku 4.5 vs Claude Sonnet 5.5
Know what each correct answer costs
Agent records tokens, calls, time and cost for every task, and the check that decided pass or fail. That is how we could run this calculation. I build Agent, so weigh that. Try Agent and get the same receipts for your own work.