• Thought experiment
  • Calculation
  • Claude Haiku
  • Claude Sonnet

Is Claude Haiku cheaper if you retry or escalate to Sonnet? A calculation on real receipts

Haiku 4.5 first, one retry, then Sonnet 5.5? On 8 hard tasks, the calculation gives 4.3x the cost and 8.7x the time per correct answer versus Sonnet alone.

TL;DR

  • No, not on these tasks. On 8 hard tasks, Sonnet 5.5 every time cost $0.0132 per correct answer and took 8.2 s. Haiku 4.5 with one retry, then Sonnet, cost $0.0566 (4.3 times) and took 71.5 s (8.7 times). Every figure here is a calculation on recorded calls. No policy ran.
  • Haiku once costs $0.0677 and takes 91.5 s per correct answer. It passed 11 of 24 calls (46%, 95% interval 28% to 65%). Sonnet passed 24 of 24 (86% to 100%).
  • Three Haiku tries do not reach parity in this calculation. Up to three Haiku tries, then Sonnet, cost $0.0719 and took 93.4 s. Their sensitivity ranges overlap, so we do not rank the three Haiku-first policies against each other. Each one stays above Sonnet every time.
  • Token price does not tell the whole cost. Haiku lists at half Sonnet's price per token. Yet its median call cost more than Sonnet's on 8 of 8 tasks. Its median output was 5,064 tokens against 1,050 (24 calls each). The call ranges were 1,899 to 9,321 and 176 to 3,895 tokens.
  • The best case for Haiku does not reach parity. With Haiku's pass rate on every task set to the top of its interval, the closest Haiku-first case cost 3.0 times as much (a sensitivity range, not an interval). If Sonnet's pass rate is also set to the bottom of its interval on every task, the closest case is 1.3 times.
  • What would have to change: for cost parity, Haiku's calls would need to cost about 12% of what they cost here. For time parity, they would need to take about 5% of the time (two separate break-evens, arithmetic, not a forecast).
  • Limits. One task set and the CLI's default effort. This calculation excludes the separate thinking-off experiment: 4 of 24 strict passes (17%, 95% interval 7% to 36%). Its strict interval overlaps the thinking-on interval. Read the caveats.

Disclosure: I build Agent. Agent ran these benchmarks, and its default coding model is Sonnet 5.5. The receipts, formulas and code are public, so you can check every number.

Study with all charts, tables and formulas: /benchmarks/haiku-retry-or-escalate.

The plan, and why it sounds good

Many teams plan to start with a cheap model. A check tells them when an answer is wrong. If it is wrong, they retry, or they send the task to a stronger model. On paper this saves money. The price per token is lower, and the cheap model passes some tasks.

We tested that plan on paper, with real numbers. We did not run it. We took recorded calls and calculated what each plan would cost per correct answer. For the short overview of Haiku against Sonnet on cost, read Is Claude Haiku cheaper than Sonnet?. This post goes deeper into the retry and escalate policies.

Five policies

PolicyWhat runs
Sonnet 5.5 every timeOne Sonnet 5.5 try
Haiku onceOne Haiku 4.5 try, no retry
Haiku, one retry, then SonnetHaiku. If the check fails, Haiku again. If it fails again, Sonnet
Haiku up to 3 tries, then SonnetHaiku up to three times, then Sonnet
Sonnet 5.5 low effort onceOne Sonnet 5.5 try with effort set to low

A try runs only when every earlier try failed. A failed try is paid for. If the last Sonnet try fails, the task is a miss.

Where the numbers come from

The inputs are 64 recorded calls in Claude Code on the same 8 hard tasks. Each task has a strict validator. Haiku 4.5 and Sonnet 5.5 at default effort made 24 calls each (3 per task). Sonnet 5.5 at low effort made 16 (2 per task). A reply in the wrong format, such as a code fence, counts as a failure.

Per task and model, we take three numbers: the strict pass rate, the median list-price cost of a call, and the median time of a call. Then we compute what each policy expects to spend.

Interval merge fix
DST day length
CSV parser
Event-loop order
Room schedule
SemVer regex
Money refactor
SQLite report query

Hover or focus a bar for its ratio to DST day length (the lowest value): a ratio of the two values shown, not a measurement.

8 rows. Highest Interval merge fix 100% (95% interval 44%–100%, n 3). Lowest SQLite report query 0% (95% interval 0%–56%, n 3). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 3 per row

Claude Haiku 4.5 · Claude Code, default effort. 3 calls per task

Measured. Whiskers are 95% Wilson intervals on only 3 calls per task, so they are wide: 0 of 3 allows a true rate up to 56%, and 3 of 3 allows one as low as 44%. A format miss (a right answer in a code fence or with prose) is not a pass. The retry policies take each task’s rate from this chart.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Haiku passed the interval merge and the SemVer regex 3 of 3 each (95% interval 44% to 100%). It passed none of the event-loop order, the room schedule and the SQL query (0 of 3 each; 95% interval 0% to 56%). On the room schedule and the SQL query, 2 of its 3 replies each had the right answer in the wrong format. Strict grading does not count them. With 3 calls per task, each interval is wide: 0 of 3 allows a true rate up to 56%.

Result 1: cost per correct answer

Calculation
Sonnet 5.5 low effort once (calculation)
Sonnet 5.5 every time (calculation)
Haiku, one retry, then Sonnet (calculation)
Haiku once (calculation)
Haiku up to 3 tries, then Sonnet (calculation)

Hover or focus a bar for its ratio to Sonnet 5.5 every time (calcul… (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 5 rows. Highest Haiku up to 3 tries, then Sonnet (calculation) $0.072 (n 48). Lowest Sonnet 5.5 low effort once (calculation) $0.012 (n 16).

Notesn 16–48 per row

Eight hard tasks, each equally likely. A failed try is paid for. Calls recorded in Claude Code

Calculation, not a run. These are median-input scenarios, not measured mean costs. Total modelled list-price cost of all tries over the 8 tasks, divided by the expected number of correct answers. Per-call cost is each task’s median call at list price (reported tokens × list price); the calls ran on a flat subscription. A try runs only when the earlier tries failed the validator, and tries are independent at each task’s observed pass rate. n is the number of recorded calls behind each bar. A calculation has no confidence interval; the sensitivity chart shows a range. Highlighted: the baseline, Sonnet 5.5 every time.

Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)

Cost per correct answer, list-price calculation, in order of the point estimate. The inputs are 24 calls for each default-effort model and 16 for low-effort Sonnet. Each escalation scenario uses 48 recorded calls; these are input counts, not executed policy tries:

  • Sonnet 5.5 low effort once: $0.0122 (16 calls, 16 of 16 passed; 95% interval 81% to 100%, a ceiling)
  • Sonnet 5.5 every time: $0.0132 (the baseline)
  • Haiku, one retry, then Sonnet: $0.0566, 4.3 times the baseline
  • Haiku once: $0.0677, 5.1 times
  • Haiku up to 3 tries, then Sonnet: $0.0719, 5.4 times

Every Haiku-first scenario costs more than Sonnet every time in all four sensitivity settings below. These scenarios do not establish a measured policy ranking. The three Haiku-first policies overlap each other there, and the two Sonnet policies have no interval, so we do not rank within those groups.

One Haiku try costs $0.0310 on average (the mean of the eight task medians). That is 2.3 times one Sonnet try (calculation) at $0.0132. These mean-of-median costs have no interval. And Haiku passes under half the time, so a correct answer costs $0.0677. The scenario gives each retry the same task-specific cost and pass rate as the first try. It charges a Sonnet call only if all Haiku tries fail. Actual retries and follow-up Sonnet calls were not tested.

Sonnet at low effort comes out at $0.0122, a calculated 8% below the default-effort figure. The calculation has no interval, so we do not call that a difference. The effort ladder found that Sonnet's per-call time ranges overlap across effort settings (does reasoning effort buy quality?).

Result 2: wall time per correct answer

Calculation
Largest value is 13x the smallest; Log shows the small bars.
Sonnet 5.5 low effort once (calculation)
Sonnet 5.5 every time (calculation)
Haiku, one retry, then Sonnet (calculation)
Haiku once (calculation)
Haiku up to 3 tries, then Sonnet (calculation)

Hover or focus a bar for its ratio to Sonnet 5.5 every time (calcul… (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 5 rows. Slowest Haiku up to 3 tries, then Sonnet (calculation) 93.4 s (n 48). Fastest Sonnet 5.5 low effort once (calculation) 7.1 s (n 16).

Notesn 16–48 per row

Tries run one after another, so a failed try adds its time

Calculation, not a run. These are median-input scenarios, not measured mean times. Each task’s median total call time (CLI start-up included), summed over the tries a policy expects to run, divided by the expected number of correct answers. Tries are assumed to run one after another. Parallel tries were not tested. One host, one network. A calculation has no confidence interval. Highlighted: the baseline, Sonnet 5.5 every time.

Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)

This calculation runs tries one after another, so a failed try adds its time. It uses the same 24-call, 16-call and 48-call inputs as the cost scenarios. The sensitivity table shows scenario ranges, not confidence intervals.

  • Sonnet 5.5 low effort once: 7.1 s
  • Sonnet 5.5 every time: 8.2 s
  • Haiku, one retry, then Sonnet: 71.5 s, 8.7 times
  • Haiku once: 91.5 s, 11.1 times
  • Haiku up to 3 tries, then Sonnet: 93.4 s, 11.3 times

As with cost, this is a scenario comparison, not a measured policy ranking. In all four settings of the sensitivity table, every Haiku-first policy takes longer than Sonnet every time. We do not rank the policies inside each group.

Haiku's median call time exceeded Sonnet's on 8 of 8 tasks. The task medians span 15.9 s to 64.5 s for Haiku and 2.5 s to 21.6 s for Sonnet. These spans are not run ranges. Each row below shows the median and the actual run range (3 calls per model per task). Ranges are not confidence intervals. The DST ranges overlap, so its medians do not establish a speed ranking. Parallel tries were not tested.

TaskHaiku median; run rangeSonnet median; run range
Interval merge21.2 s; 18.5 s to 24.7 s2.9 s; 2.4 s to 3.1 s
DST day length39.0 s; 26.2 s to 40.6 s21.6 s; 19.6 s to 34.8 s
CSV parser39.0 s; 33.8 s to 75.1 s9.6 s; 8.8 s to 10.6 s
Event-loop order54.0 s; 26.0 s to 56.6 s9.6 s; 8.2 s to 9.9 s
Room schedule54.6 s; 38.0 s to 68.8 s7.7 s; 7.4 s to 7.8 s
SemVer regex64.5 s; 54.5 s to 67.0 s2.5 s; 2.3 s to 3.7 s
Money refactor15.9 s; 15.3 s to 17.2 s3.6 s; 3.4 s to 7.2 s
SQLite report query47.1 s; 38.9 s to 56.1 s8.5 s; 7.6 s to 8.9 s

Result 3: how often each policy is right

Calculation
Sonnet 5.5 every time (calculation)
Haiku, one retry, then Sonnet (calculation)
Haiku up to 3 tries, then Sonnet (calculation)
Sonnet 5.5 low effort once (calculation)
Haiku once (calculation)

Hover or focus a bar for its ratio to Sonnet 5.5 every time (calcul… (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 5 rows. Highest Sonnet 5.5 every time (calculation) 100% (95% interval 86%–100%, n 24). Lowest Haiku once (calculation) 46% (95% interval 28%–65%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–48 per row4 of 5 at 100%: this task set cannot separate them.

Share of the eight hard tasks answered correctly under the strict validator

Whiskers are 95% Wilson intervals on the measured strict pass rate and appear only on the three one-try policies (24/24, 11/24 and 16/16). The two escalation policies are calculations without an interval. A value of 100% means no failure was recorded: the model treats Sonnet 5.5 as always right (24/24, 95% interval 86% to 100%), so the escalation policies cannot miss. n is the number of recorded calls behind each bar.

Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)

Haiku once passed 46% of recorded calls (11 of 24; 95% interval 28% to 65%). The two escalation policies calculate 100% success from 48 recorded calls, with no confidence interval. Sonnet at low effort also passed every call (16 of 16; 95% interval 81% to 100%). That is a ceiling effect, not proof. Sonnet passed 24 of 24, and the model treats it as always right. Its 95% interval is 86% to 100%, which does not prove perfect success on future calls.

Why the Haiku-first calculation costs more

Calculation
  • Claude Haiku 4.5 · Claude Code
  • Claude Sonnet 5.5 · Claude Code (square)
Sorted by gap, largest first.
SemVer regex
Interval merge fix
Room schedule
Event-loop order
SQLite report query
Money refactor
CSV parser
DST day length

Gap labels, Claude Sonnet 5.5 · Claude Code vs Claude Haiku 4.5 · Claude Code: Claude Sonnet 5.5 · Claude Code is x% higher (+) or lower (−) than Claude Haiku 4.5 · Claude Code, calculated from the two values shown (the change counted from Claude Haiku 4.5 · Claude Code’s value).

List-price calculation, not a run. 8 rows, 2 series: Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code. Claude Haiku 4.5 · Claude Code: highest SemVer regex $0.042 (n 3). Lowest Interval merge fix $0.018 (n 3). Claude Sonnet 5.5 · Claude Code: highest DST day length $0.025 (n 3). Lowest SemVer regex $0.0051 (n 3).

Notesn = 3 per row

Median of each task’s calls, default effort, Claude Code

Calculation: reported tokens × list price for each call, then the median per task; the calls ran on a flat subscription. Haiku’s list price per token is lower. Its median calls cost more and contained more output tokens. This does not isolate the effect of effort or thinking. 3 calls per task and configuration; the table shows each call-cost range, not a confidence interval.

Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)

Per token, Haiku is cheaper. It lists at $1 per million input tokens and $5 per million output tokens. Sonnet lists at $2 and $10.

Per call, Haiku cost more on all 8 tasks. Output tokens enter the price calculation. This does not isolate the effect of thinking or effort. At the CLI's default effort, Haiku wrote a median 5,064 output tokens per call. Sonnet wrote 1,050. Each median uses 24 calls. Haiku's call range was 1,899 to 9,321 tokens; Sonnet's was 176 to 3,895. These are ranges, not confidence intervals. That is 4.8 times the tokens at half the price, about 2.4 times the output cost (our arithmetic). Our hard-task post shows that most of Haiku's output was reasoning.

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

The hard head-to-head calculated cost per pass in a different way: total cost of all calls divided by strict passes. That gives Haiku $0.0672 and Sonnet $0.0143 per strict pass. The retry study uses each task's median call instead, so Sonnet reads $0.0132 here. Both views agree.

How wrong could we be?

Calculation
Sonnet 5.5 low effort once (calculation)
Sonnet 5.5 every time (calculation)
Haiku, one retry, then Sonnet (calculation)
Haiku once (calculation)
Haiku up to 3 tries, then Sonnet (calculation)

Every interval overlaps every other: this chart does not order these rows.

List-price calculation, not a run. 5 rows. Highest Haiku up to 3 tries, then Sonnet (calculation) $0.072 (range $0.041–$0.09, n 48). Lowest Sonnet 5.5 low effort once (calculation) $0.012 (n 16). All run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 16–48 per row

Dot: Haiku’s pass rate as observed. Whisker: its pass rate on every task set to its Wilson lower and upper bound at once

Calculation. This is a sensitivity range, not a confidence interval: each Haiku policy is recomputed with Haiku’s pass rate on all 8 tasks at once set to the lower and then the upper end of that task’s 95% Wilson interval (3 calls per task, so the ends are far apart), and the whisker spans the lowest and highest result. That is more extreme than a 95% interval on the total. The two Sonnet policies do not depend on Haiku’s pass rate and have no whisker. Highlighted: the baseline, Sonnet 5.5 every time.

Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)

Haiku's pass rate rests on 3 calls per task. So we recomputed each Haiku policy with its pass rate on all 8 tasks set to the lower end, then the upper end, of each task's 95% interval. That moves every task to a bound at once. It is a sensitivity range, not a confidence interval. It does not form a confidence interval on the total.

  • Haiku, one retry, then Sonnet: $0.0394 to $0.0681 per correct answer. The baseline is $0.0132.
  • Haiku once: $0.0391 to $0.1834.
  • Haiku up to 3 tries, then Sonnet: $0.0415 to $0.0905.

In every setting, every Haiku-first policy cost more than the baseline. The closest any came was about 3.0 times (Haiku once, with every task at the top of its interval).

That test moves only Haiku. Sonnet's pooled 24 of 24 has a 95% interval of 86% to 100%. Each task's 3 of 3 has a 95% interval of 44% to 100%. The joint extreme uses these per-task bounds. So we also ran a joint extreme: Haiku at the top of its interval and Sonnet at the bottom of its interval, on every task at once. Sonnet every time then succeeds on only 44% of tasks, which no run showed. Even then the closest Haiku-first policy (Haiku once) costs 1.3 times as much and the quickest takes 2.8 times as long as Sonnet every time. It is a calculation on 3 calls per task, not a forecast.

One more reading is fair to Haiku. Five of its 13 failed calls had the right answer in the wrong format. If a harness strips the code fence and accepts them, Haiku once costs $0.0466 per correct answer and Haiku with one retry, then Sonnet, costs $0.0459. Both stay above Sonnet's $0.0132.

What would have to be true

Keep Haiku's pass rates as observed and scale only its call cost. Then Haiku with one retry, then Sonnet, matches Sonnet every time when Haiku's calls cost 11.8% of what they cost here. For time, they would need to take 5.4% of the time. This is arithmetic, not a forecast. These break-evens keep the default-effort pass rates fixed. They do not predict a thinking-off retry policy. The separate thinking-on vs off study recorded 4 of 24 strict passes with thinking off (17%, 95% interval 7% to 36%) on the same 8 hard tasks. Its strict interval overlaps thinking on (11 of 24, 28% to 65%). The arms ran at different times on different accounts, so the study cannot isolate a thinking effect. This calculation excludes those thinking-off calls; no thinking-off retry policy ran.

Repeated errors limit the case for retries

The calculation treats each try as independent. That may favour Haiku if its failures repeat. In our consistency test, Haiku answered 289 on an exact-number prompt in 10 of 10 repeats (wrong-answer rate 100%, 95% interval 72% to 100%). The expected answer was 282. Those repeats found no fix; they do not prove that every future retry will fail. And on three of the 8 hard tasks, Haiku passed 0 of 3 calls each (95% interval 0% to 56%). The calculation already sends those tasks to Sonnet. See same prompt, ten answers.

Best practice

  1. Price a plan in cost per correct answer. Keep failed tries in the cost.
  2. Compare cost per call, not price per token. Input, cache and output tokens enter the cost. Haiku's median output was 4.8 times Sonnet's (calculation).
  3. Repeat a prompt before you add a retry. If the wrong answer repeats, escalate at once or change the prompt.
  4. A cheap first try pays only if three things hold: the check is cheap, the first try often passes, and its calls cost well below the strong model's. Here, the check was free by assumption. The other two did not hold: Haiku passed 11 of 24 calls (46%, 95% interval 28% to 65%), and its calls cost more.
  5. Set effort and thinking before you compare models. Haiku ran at the default and thought a lot.
  6. Count the wait. Tries in a row multiply the wall time. Here, one retry meant a calculated 8.7 times the time.
  7. Measure on your own tasks. Eight tasks cannot speak for yours.

How we calculated

  • Inputs. 64 recorded calls from the hard head-to-head and the effort ladder. No new call was made.
  • Per task and model. Pass rate = strict passes ÷ calls. Cost = median of each call's list-price cost, with plain input, cache reads, cache writes and output priced separately. Time = median total call time, CLI start-up included.
  • A policy. Expected cost = sum over tries of (chance the try runs) × its cost. For Haiku, one retry, then Sonnet, with q = 1 − Haiku pass rate: cost = C(Haiku) × (1 + q) + C(Sonnet) × q².
  • Over the 8 tasks. Each task is equally likely. Cost per correct answer = total expected cost ÷ total expected correct answers. Time is built the same way.
  • Sensitivity. Haiku's pass rate on every task at the lower, then the upper, end of its 95% Wilson interval. The joint extreme also sets Sonnet's pass rate to the lower end. Both are sensitivity ranges, not confidence intervals.
  • Assumptions. A check detects a failed answer (true for this set, and free here). Tries are independent at the observed rate. A failed last Sonnet try is a miss.

Formulas and every table: /benchmarks/haiku-retry-or-escalate. Price your own mix with the AI cost calculator.

Caveats

  • A calculation, not a run. No policy ran. The calls ran on a flat subscription, so costs are list-price estimates, not bills.
  • Protocol timing. Advance registration is not verified. Both source protocol files have birth times after their first counted calls. The study lists the exact times. Copying could explain them, but the files cannot establish that.
  • Scope and dependence. The hand-built hard tasks followed an earlier ceiling. Repeated calls on these fixed tasks do not give a population interval for coding work. Retry independence is untested. A Sonnet call after a Haiku failure may differ from a fresh Sonnet call. The sensitivity does not cover these links.
  • Median inputs. Cost and time scenarios use each task's median call, not measured mean policy costs or times.
  • Small samples. 3 calls per task for Haiku and Sonnet, 2 for Sonnet at low effort. A median of 3 calls has no interval.
  • Default effort. The effort flag was not passed for Haiku and Sonnet in these inputs. This calculation excludes the separate thinking-off experiment: 4 of 24 strict passes (17%, 95% interval 7% to 36%). Its strict interval overlaps thinking on, and day and account differ. That experiment does not validate these fixed-pass-rate break-evens.
  • One task set. These are the 8 hard tasks. Haiku may do better on easy tasks, where pass rate hits a ceiling. The five-task head-to-head is not recalculated here.
  • Sonnet is treated as always right. It passed 24 of 24. That is a ceiling, not a guarantee. The joint extreme above lets it fall to 44% and still finds no parity.
  • One host. One network. Low-effort Sonnet ran about 10 hours after the default-effort batch. We did not control for other load on the host, so times may carry some contention. Times include CLI start-up.

Know what each correct answer costs

Agent records tokens, calls, time and cost for every task, and the check that decided pass or fail. That is how we could run this calculation. I build Agent, so weigh that. Try Agent and get the same receipts for your own work.

The data behind this post

Includes calculations
  • Thought experiment
  • Claude Haiku

Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts

Haiku 4.5 first, then retry or escalate to Sonnet 5.5? A calculation on 64 recorded calls across 8 hard tasks: cost and time per correct answer.

46% (11/24)Strict passes, Claude Haiku 4.5 · Claude Code, eight hard tasks · n = 24

6 chartsUpdated October 7, 2026

  • Claude Haiku
  • Extended Thinking

Does thinking pay for Claude Haiku 4.5? Thinking on vs off

Claude Haiku 4.5 with extended thinking on and off: 82 routing decisions and 8 hard tasks. Accuracy with 95% intervals, time and cost.

87% (71/82)Claude Haiku 4.5 (thinking off): exact routing decisions · n = 82

6 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.