{"i":24,"study":{"slug":"haiku-retry-or-escalate","title":"Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts","seoTitle":"Haiku retry and escalate: cost per correct answer","description":"Haiku 4.5 first, then retry or escalate to Sonnet 5.5? A calculation on 64 recorded calls across 8 hard tasks: cost and time per correct answer.","question":"On the 8 hard tasks, what does one correct answer cost, and how long does it take, if you try Claude Haiku 4.5 first and retry or escalate, against Claude Sonnet 5.5 every time?","answer":"The calculated Haiku-first policies cost more on these 8 hard tasks. These are scenarios from recorded calls, not a measured policy ranking. Sonnet every time costs $0.0132 per correct answer and takes 8.2 s (24/24 passes; 95% interval 86% to 100%). Haiku passes 11/24 calls (95% interval 28% to 65%); Haiku once costs $0.0677 per correct answer. Haiku with one retry, then Sonnet, costs $0.0566 and takes 71.5 s per correct answer. The tables show the other policies and sensitivity ranges; the method states the assumptions.","date":"2026-10-07","updated":"2026-10-07","tags":["thought-experiment","claude-haiku","claude-sonnet","llm-pricing","cost-per-correct-answer","retry","escalation","hard-tasks"],"caveats":["Advance registration is not verified. The hard-set protocol file birth time is 2026-10-06 04:03:37 UTC; its first counted call started at 03:23:59 UTC. The ladder protocol file birth time is 14:35:11 UTC; its first Claude call started at 14:21:50 UTC. Both files claim advance declaration, but the available file times do not support that claim. Copying could explain the times; we cannot establish it.","Wilson intervals treat calls as independent. Repeated calls on eight fixed tasks do not give a population interval for coding work. No task-level paired test or measured retry policy supports a general ranking.","This is a calculation, not a run. No policy ran. Costs are list-price calculations from reported tokens. The calls ran on a flat subscription, so no invoice backs them.","Retry independence is untested. Calls repeat the same eight prompts. Failures can repeat, and escalation after a Haiku failure may differ from a fresh Sonnet call. The sensitivity does not cover these links.","Samples are small: 3 calls per task for Haiku and for Sonnet, and 2 for Sonnet at low effort. A per-task pass rate has a wide interval: 0 of 3 allows up to 56%, and 3 of 3 allows as low as 44%. A median of a few calls has no interval. The sensitivity moves all 8 tasks to a bound at once. That is more extreme than a 95% interval on the total.","Sonnet 5.5 passed 24/24, and the model treats it as always right (95% interval 86% to 100%). So every policy that ends in a Sonnet try shows 100% success. A perfect sample does not prove perfect success on future calls. This is the ceiling effect of the task set.","Haiku ran at the CLI’s default effort: the effort flag was not passed, so the CLI chose. It wrote a median 5,064 output tokens per call against 1,050 for Sonnet. Output tokens enter the price calculation. The run does not show what caused the timing gap. We did not test Haiku with thinking off. The break-even stats show how far its calls would have to fall.","Strict grading counts a reply in a code fence as a failure. 5 of Haiku’s 13 failed calls were such format misses. With them counted as passes (lenient reading), Haiku once costs $0.0466 per correct answer, and Haiku with one retry then Sonnet costs $0.0459. Both stay above Sonnet every time.","This study uses each task’s median call for cost and time. The hard head-to-head divides the total cost of all calls by the passes instead. So its figures differ a little: Sonnet $0.0143 and Haiku $0.0672 per strict pass, against $0.0132 and $0.0677 here.","This is a hand-built task set, tuned after an earlier set hit a ceiling. It is not a random sample of coding work. This study covers the eight hard tasks only. Haiku may do better on easy tasks, where pass rate hits a ceiling. We did not recalculate the five-task head-to-head here.","Haiku and Sonnet at default effort ran in one batch on 2026-10-06 (UTC). Sonnet at low effort ran about 10 hours later in the effort ladder. All calls used one host and one network. We did not control for other load on the host, so wall times may carry some contention. Times include CLI start-up and the CLI’s own system prompt.","The sensitivity moves Haiku’s pass rate only, and the Sonnet baseline rests on 24/24. In a joint extreme (Haiku at the upper and Sonnet at the lower end of its Wilson interval on every task), the closest Haiku-first policy costs 1.3 times as much, and the quickest takes 2.8 times as long, as Sonnet every time (calculation). That setting has Sonnet succeed on only 44% of the tasks, which no run showed."],"sourceIds":["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["retry-escalate-haiku-passes","Strict passes, Claude Haiku 4.5 · Claude Code, eight hard tasks",0.4583,"rate","46% (11/24)",24,[0.2789,0.6493],"\u0001"],["retry-escalate-sonnet-passes","Strict passes, Claude Sonnet 5.5 · Claude Code, eight hard tasks",1,"rate","100% (24/24)",24,[0.862,1],"\u0001"],["retry-escalate-sonnet-low-passes","Strict passes, Claude Sonnet 5.5 (low) · Claude Code, eight hard tasks",1,"rate","100% (16/16)",16,[0.8064,1],"\u0001"],["retry-escalate-baseline-cost","Cost per correct answer, Sonnet 5.5 every time (calculation)",0.01322,"usd","$0.0132",24,"\u0001","\u0001"],["retry-escalate-baseline-time","Wall time per correct answer, Sonnet 5.5 every time (calculation)",8.24,"seconds","8.2 s",24,"\u0001","\u0001"],["retry-escalate-haiku-once-cost","Cost per correct answer, Haiku once (calculation)",0.06772,"usd","$0.0677",24,"\u0001","\u0001"],["retry-escalate-haiku-once-time","Wall time per correct answer, Haiku once (calculation)",91.47,"seconds","91.5 s",24,"\u0001","\u0001"],["retry-escalate-retry-cost","Cost per correct answer, Haiku with one retry then Sonnet (calculation)",0.05664,"usd","$0.0566",48,"\u0001","\u0001"],["retry-escalate-retry-time","Wall time per correct answer, Haiku with one retry then Sonnet (calculation)",71.54,"seconds","71.5 s",48,"\u0001","\u0001"],["retry-escalate-retry-cost-ratio","Haiku with one retry then Sonnet vs Sonnet every time: cost per correct answer (calculation)",4.28,"ratio","4.3x",48,"\u0001","\u0001"],["retry-escalate-retry-time-ratio","Haiku with one retry then Sonnet vs Sonnet every time: wall time per correct answer (calculation)",8.68,"ratio","8.7x",48,"\u0001","\u0001"],["retry-escalate-three-tries-cost-ratio","Haiku up to 3 tries then Sonnet vs Sonnet every time: cost per correct answer (calculation)",5.44,"ratio","5.4x",48,"\u0001","\u0001"],["retry-escalate-sonnet-low-cost","Cost per correct answer, Sonnet 5.5 low effort once (calculation)",0.01219,"usd","$0.0122",16,"\u0001","\u0001"],["retry-escalate-sonnet-low-time","Wall time per correct answer, Sonnet 5.5 low effort once (calculation)",7.07,"seconds","7.1 s",16,"\u0001","\u0001"],["retry-escalate-haiku-dearer-tasks","Tasks where Haiku’s median call cost more than Sonnet’s (calculation)",8,"count","8 of 8",8,"\u0001","\u0001"],["retry-escalate-haiku-slower-tasks","Tasks where Haiku’s median call took longer than Sonnet’s",8,"count","8 of 8",8,"\u0001","\u0001"],["retry-escalate-haiku-output-tokens","Median output tokens per call, Haiku 4.5 (Claude Code)",5064,"tokens","5,064",24,"\u0001","Call range 1,899 to 9,321 tokens; not a confidence interval."],["retry-escalate-sonnet-output-tokens","Median output tokens per call, Sonnet 5.5 (Claude Code)",1050,"tokens","1,050",24,"\u0001","Call range 176 to 3,895 tokens; not a confidence interval."],["retry-escalate-breakeven-cost","Haiku call cost, as a share of measured, at which one retry then Sonnet matches Sonnet every time (calculation)",0.1177,"ratio","11.8% of measured",8,"\u0001","All pass rates kept as observed. Arithmetic, not a prediction: a Haiku that thinks less may pass at a different rate."],["retry-escalate-breakeven-time","Haiku call time, as a share of measured, at which one retry then Sonnet matches Sonnet every time (calculation)",0.0543,"ratio","5.4% of measured",8,"\u0001","All pass rates kept as observed. Arithmetic, not a prediction."],["retry-escalate-sensitivity-lowest-ratio","Lowest cost ratio of any Haiku-first policy to Sonnet every time, across the four Haiku pass-rate settings (sensitivity range, calculation)",2.96,"ratio","3.0x",8,"\u0001","Haiku once, with Haiku’s pass rate on every task set to its Wilson upper bound."],["retry-escalate-joint-extreme-ratio","Lowest cost ratio of any Haiku-first policy to Sonnet every time, with Haiku’s pass rate at its Wilson upper bound and Sonnet’s at its Wilson lower bound on every task at once (joint extreme, calculation)",1.3,"ratio","1.3x",8,"\u0001","Haiku once. The lowest time ratio is 2.8x (Haiku once). In this setting Sonnet every time succeeds on 44% of the tasks, which no run showed (Sonnet passed 24/24). A joint extreme on 3 calls per task, not a forecast."],["retry-escalate-lenient-retry-cost","Cost per correct answer, Haiku with one retry then Sonnet, format misses counted as passes (calculation)",0.04591,"usd","$0.0459",48,"\u0001","\u0001"]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds","polarity","whisker"],"$r":[["retry-escalate-cost-per-correct","Expected list-price cost per correct answer, by retry policy (calculation)","Eight hard tasks, each equally likely. A failed try is paid for. Calls recorded in Claude Code","bar","usd","USD per correct answer",[{"name":"Cost per correct answer (calculation)","points":{"$k":["label","value","n","highlight"],"$r":[["Sonnet 5.5 low effort once (calculation)",0.01219,16,false],["Sonnet 5.5 every time (calculation)",0.01322,24,true],["Haiku, one retry, then Sonnet (calculation)",0.05664,48,false],["Haiku once (calculation)",0.06772,24,false],["Haiku up to 3 tries, then Sonnet (calculation)",0.07193,48,false]]}}],"Calculation, not a run. These are median-input scenarios, not measured mean costs. Total modelled list-price cost of all tries over the 8 tasks, divided by the expected number of correct answers. Per-call cost is each task’s median call at list price (reported tokens × list price); the calls ran on a flat subscription. A try runs only when the earlier tries failed the validator, and tries are independent at each task’s observed pass rate. n is the number of recorded calls behind each bar. A calculation has no confidence interval; the sensitivity chart shows a range. Highlighted: the baseline, Sonnet 5.5 every time.",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],"\u0001","\u0001"],["retry-escalate-time-per-correct","Expected wall time per correct answer, by retry policy (calculation)","Tries run one after another, so a failed try adds its time","bar","seconds","Seconds per correct answer",[{"name":"Time per correct answer (calculation)","points":{"$k":["label","value","n","highlight"],"$r":[["Sonnet 5.5 low effort once (calculation)",7.07,16,false],["Sonnet 5.5 every time (calculation)",8.24,24,true],["Haiku, one retry, then Sonnet (calculation)",71.54,48,false],["Haiku once (calculation)",91.47,24,false],["Haiku up to 3 tries, then Sonnet (calculation)",93.42,48,false]]}}],"Calculation, not a run. These are median-input scenarios, not measured mean times. Each task’s median total call time (CLI start-up included), summed over the tries a policy expects to run, divided by the expected number of correct answers. Tries are assumed to run one after another. Parallel tries were not tested. One host, one network. A calculation has no confidence interval. Highlighted: the baseline, Sonnet 5.5 every time.",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],"\u0001","\u0001"],["retry-escalate-success","Expected success rate by retry policy (calculation)","Share of the eight hard tasks answered correctly under the strict validator","bar","rate","Correct answers",[{"name":"Expected success rate","points":{"$k":["label","value","n","lo","hi","highlight"],"$r":[["Sonnet 5.5 every time (calculation)",1,24,0.862,1,true],["Haiku, one retry, then Sonnet (calculation)",1,48,"\u0001","\u0001",false],["Haiku up to 3 tries, then Sonnet (calculation)",1,48,"\u0001","\u0001",false],["Sonnet 5.5 low effort once (calculation)",1,16,0.8064,1,false],["Haiku once (calculation)",0.4583,24,0.2789,0.6493,false]]}}],"Whiskers are 95% Wilson intervals on the measured strict pass rate and appear only on the three one-try policies (24/24, 11/24 and 16/16). The two escalation policies are calculations without an interval. A value of 100% means no failure was recorded: the model treats Sonnet 5.5 as always right (24/24, 95% interval 86% to 100%), so the escalation policies cannot miss. n is the number of recorded calls behind each bar.",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],"higher","ci95"],["retry-escalate-haiku-by-task","Strict pass rate of Haiku 4.5 on each hard task","Claude Haiku 4.5 · Claude Code, default effort. 3 calls per task","bar","rate","Strict passes",[{"name":"Strict pass rate (n = 3 per task)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Interval merge fix",1,0.4385,1,3],["DST day length",0.3333,0.0615,0.7923,3],["CSV parser",0.6667,0.2077,0.9385,3],["Event-loop order",0,0,0.5615,3],["Room schedule",0,0,0.5615,3],["SemVer regex",1,0.4385,1,3],["Money refactor",0.6667,0.2077,0.9385,3],["SQLite report query",0,0,0.5615,3]]}}],"Measured. Whiskers are 95% Wilson intervals on only 3 calls per task, so they are wide: 0 of 3 allows a true rate up to 56%, and 3 of 3 allows one as low as 44%. A format miss (a right answer in a code fence or with prose) is not a pass. The retry policies take each task’s rate from this chart.",["agent-provider-h2h-hard"],"higher","ci95"],["retry-escalate-call-cost-by-task","List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation)","Median of each task’s calls, default effort, Claude Code","grouped-bar","usd","USD per call",[{"name":"Claude Haiku 4.5 · Claude Code","points":{"$k":["label","value","n"],"$r":[["Interval merge fix",0.01789,3],["DST day length",0.0293,3],["CSV parser",0.02919,3],["Event-loop order",0.03708,3],["Room schedule",0.03578,3],["SemVer regex",0.04217,3],["Money refactor",0.02039,3],["SQLite report query",0.03648,3]]}},{"name":"Claude Sonnet 5.5 · Claude Code","points":{"$k":["label","value","n"],"$r":[["Interval merge fix",0.00557,3],["DST day length",0.02532,3],["CSV parser",0.01464,3],["Event-loop order",0.01588,3],["Room schedule",0.01243,3],["SemVer regex",0.00514,3],["Money refactor",0.00961,3],["SQLite report query",0.01719,3]]}}],"Calculation: reported tokens × list price for each call, then the median per task; the calls ran on a flat subscription. Haiku’s list price per token is lower. Its median calls cost more and contained more output tokens. This does not isolate the effect of effort or thinking. 3 calls per task and configuration; the table shows each call-cost range, not a confidence interval.",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],"\u0001","\u0001"],["retry-escalate-sensitivity","Sensitivity: cost per correct answer if Haiku’s pass rate is higher or lower (calculation)","Dot: Haiku’s pass rate as observed. Whisker: its pass rate on every task set to its Wilson lower and upper bound at once","dot-range","usd","USD per correct answer",[{"name":"Cost per correct answer (calculation)","points":{"$k":["label","value","n","highlight","lo","hi"],"$r":[["Sonnet 5.5 low effort once (calculation)",0.01219,16,false,"\u0001","\u0001"],["Sonnet 5.5 every time (calculation)",0.01322,24,true,"\u0001","\u0001"],["Haiku, one retry, then Sonnet (calculation)",0.05664,48,false,0.03941,0.06807],["Haiku once (calculation)",0.06772,24,false,0.03908,0.1834],["Haiku up to 3 tries, then Sonnet (calculation)",0.07193,48,false,0.04149,0.09047]]}}],"Calculation. This is a sensitivity range, not a confidence interval: each Haiku policy is recomputed with Haiku’s pass rate on all 8 tasks at once set to the lower and then the upper end of that task’s 95% Wilson interval (3 calls per task, so the ends are far apart), and the whisker spans the lowest and highest result. That is more extreme than a 95% interval on the total. The two Sonnet policies do not depend on Haiku’s pass rate and have no whisker. Highlighted: the baseline, Sonnet 5.5 every time.",["calc-repricing","agent-provider-h2h-hard","agent-effort-ladder","price-anthropic"],"\u0001","minmax"]]},"related":["hard-model-head-to-head","effort-ladder"]}}