Does extended thinking pay for Claude Haiku 4.5? We turned it off and measured
Claude Haiku 4.5, thinking off: 2.7x faster (a calculation from two medians), no routing accuracy gap shown. Hard tasks: 4/24 passed, thinking on 11/24.
TL;DR
- Question: does extended thinking pay for Claude Haiku 4.5? In our earlier runs it took a median 12.54 s per routing decision through Claude Code and 39.0 s on hard tasks. We turned thinking off and ran both tests again.
- Router, 82 typed decisions: thinking off answered 71 of 82 exactly (87%, 95% interval 78% to 92%). Thinking on answered 73 of 82 (89%, 80% to 94%). On the same cases the paired exact test gives p = 0.75: no difference shown.
- Speed and cost of that router: median wall time 4.66 s with thinking off, 12.54 s with thinking on (p95 8.18 s against 34.48 s). Thinking tokens per decision: 0 against 1,101 (mean). List-price cost per 1,000 decisions: $3.36 against $8.92 (a calculation).
- Hard tasks, 24 calls per cell: thinking off passed 4 of 24 strictly (17%, 7% to 36%). Thinking on passed 11 of 24 (46%, 28% to 65%). The strict intervals overlap. Count the right answers in the wrong format, and it is 4 of 24 against 16 of 24 (67%, 47% to 82%): those intervals do not overlap.
- Time on hard tasks: a median 2.9 s against 39.0 s per call. Cost per strict pass: $0.0365 against $0.0672 (a calculation from 4 and 11 passes, with no interval). Thinking off was cheaper per pass in this sample, but it passed fewer tasks.
- What to do (our reading of this sample): turn thinking off for short typed decisions, where we could not show an accuracy loss (p = 0.75, 10 differing cases). Test thinking on for hard problems. It passed more calls here, but strict intervals overlap. Sonnet 5.5 at low effort had a lower median of 2.60 s through the same CLI. Its range overlaps Haiku with thinking off.
- Limits: the thinking-on arms are recorded runs from other days, so the timing comparison is not concurrent. 82 decisions and 24 hard calls are small samples.
The full study: /benchmarks/haiku-thinking-on-off.
What we tested
Claude Haiku 4.5 reported thinking tokens under the Claude Code default in our reference runs. Those tokens count as billed output at list price. Do they buy anything?
- Thinking off. We set the environment variable
MAX_THINKING_TOKENS=0on the Claude Code process. We did not assume it worked. Two uncounted probes and all 106 counted calls reported 0 thinking tokens in the CLI's own counters. - Test 1, routing. The 82 typed decisions our platform asks a router about: failure class, message intent, is-it-a-rule, context shape. One call per decision. The repository's own decision-eval runner scores each one.
- Test 2, hard tasks. The eight hard tasks of our hard head-to-head, each with a sandboxed validator, three repetitions each. We ran the validator controls first: 8 of 8 reference answers pass and 26 of 26 planted wrong answers fail.
- Reference arms. The thinking-on numbers are our recorded runs: the routing run of 2026-10-05 and the hard-task receipts of 2026-10-06. We reused them. We did not rerun them.
Router accuracy: no difference shown
Exact means every scored question in a case was acceptable. An unanswered question counts as wrong. All intervals below are 95% Wilson intervals. Questions within one case are related, so per-question intervals are descriptive.
| Arm (82 decisions) | Exact | 95% interval | Per question | 95% interval |
|---|---|---|---|---|
| Haiku 4.5, thinking off | 71 (87%) | 78% to 92% | 177 of 194 (91%) | 86% to 94% |
| Haiku 4.5, thinking on | 73 (89%) | 80% to 94% | 183 of 194 (94%) | 90% to 97% |
| Sonnet 5.5, effort low (reference) | 77 (94%) | 87% to 97% | 189 of 194 (97%) | 94% to 99% |
The Haiku intervals overlap, so the data does not rank the two settings. The case-by-case test agrees. Thinking on was right where thinking off was wrong in 6 cases, and the reverse in 4. 67 cases were right both times and 5 wrong both times. The exact McNemar test (two-sided, on those 10 cases) gives p = 0.754.
| Pair, same 82 cases | Both right | Only thinking off right | Only the other arm right | Both wrong | Exact McNemar p |
|---|---|---|---|---|---|
| Thinking off vs thinking on | 67 | 4 | 6 | 5 | 0.754 |
| Thinking off vs Sonnet 5.5 low | 70 | 1 | 7 | 4 | 0.07 |
By decision type, message intent was 20 of 20 in every arm (95% interval 84% to 100%). Is-it-a-rule was 12 of 12 (76% to 100%). These subsets hit a ceiling; they cannot show equal accuracy on harder cases. Failure class scored 15 of 18 with thinking off (61% to 94%) and 17 of 18 with thinking on (74% to 99%). Context shape scored 24 of 32 in both (58% to 87%). Few cases carry the whole difference, so the paired test has little power. Read "no difference shown" as "we cannot tell", not as "equal".
Router speed: about 2.7 times faster (a calculation from two medians)
- Wall time per decision (median, p95): thinking off 4.66 s (8.18 s); thinking on 12.54 s (34.48 s); Sonnet 5.5 at low effort 2.60 s (4.30 s).
- Wall-time ranges (82 calls per arm): thinking off 2.20 s to 11.57 s; thinking on 5.86 s to 51.28 s; Sonnet 1.99 s to 5.58 s. These are single-call ranges, not confidence intervals.
- Model (API) time, median and range: thinking off 3.79 s (1.37 s to 10.83 s); thinking on 10.51 s (4.33 s to 49.49 s); Sonnet 1.60 s (1.06 s to 4.76 s). Each arm has 82 calls; these ranges are not confidence intervals.
The whisker runs from the median to the 95th percentile. It is not a confidence interval. The thinking-on median sits above the thinking-off 95th percentile. The arms ran on different days and accounts, so this gap cannot isolate the effect of thinking. The ranges of single calls overlap: the slowest thinking-off call took 11.57 s and the fastest thinking-on call took 5.86 s.
Haiku with thinking off still had a higher median than Sonnet at low effort through the same CLI: Haiku's thinking-off median (4.66 s) is above Sonnet's 95th percentile (4.30 s). See every measured row for Haiku vs Sonnet for the full list. Their single-call ranges overlap, so these runs do not establish a speed ranking. Where this sits next to rule-based routing and Jev:
That chart shows the thinking-on Haiku and Sonnet medians (12,543 ms and 2,597 ms) next to a deterministic policy that decides in microseconds. The routing-overhead study has the per-task arithmetic. For Jev, see the Jev vs Haiku page.
Router tokens and cost
With thinking on, Haiku wrote a mean 1,101 thinking tokens per decision (range 285 to 4,443). With thinking off it wrote none. The visible output was about the same: 366 tokens with thinking off, 318 with thinking on.
List-price cost per 1,000 decisions, from the tokens the CLI reported: $3.36 with thinking off, $8.92 with thinking on, $7.32 for Sonnet 5.5 at low effort. These are calculations. The calls ran on a flat subscription. Haiku with thinking off had the lowest calculated cost of the three and a higher median time than Sonnet.
Two notes on these costs:
- Sonnet's price basis. Its calls wrote one-hour cache entries, priced at $4 per million tokens. The calculation is $7.32 ($7.324) per 1,000 decisions, matching the CLI's reported cost. The older $4.996 calculation used the 5-minute cache-write rate. It does not match these receipts. The two Haiku calculations also match their CLI-reported costs.
- An input-token gap between the Haiku arms. Each case has the same prompt character count, but the thinking-on calls report about 297 more input tokens per call (1,831 against 1,534, a calculation from the two means). We do not know why, and the CLI version of the thinking-on routing run is not recorded. At Haiku's list price that is about $0.30 of the $5.56 cost gap per 1,000 decisions (a calculation).
Hard tasks: thinking on passed more calls
Eight tasks, three repetitions each, strict validators. A strict pass needs the whole reply to pass as given. Repeats on the same task are related. The call-level intervals are descriptive; they do not measure uncertainty across unseen tasks.
- Thinking off: 4 of 24 strict passes (17%, 95% interval 7% to 36%). 0 format misses, 20 wrong answers.
- Thinking on: 11 of 24 (46%, 28% to 65%). 5 format misses, 8 wrong answers.
The strict intervals overlap, so the strict result alone does not separate the two. The lenient reading counts a correct answer in the wrong wrapping, such as a code fence. It is 4 of 24 against 16 of 24 (67%, 47% to 82%), and those intervals do not overlap. On the lenient reading, thinking on passed more.
By task, thinking off passed the interval-merge fix 3 of 3 (95% interval 44% to 100%) and the refactor 1 of 3 (6% to 79%). Every other task scored 0 of 3 (0% to 56%). Thinking on passed interval merge and strict SemVer 3 of 3 each (44% to 100%). It passed CSV parsing and the refactor 2 of 3 each (21% to 94%), and DST day length 1 of 3 (6% to 79%). Its other three tasks scored 0 of 3 each (0% to 56%). Both arms hit a ceiling on interval merge. This does not establish equal accuracy on harder variants.
Median total time per call: 2.9 s with thinking off (range 1.7 s to 13.0 s) against 39.0 s with thinking on (15.3 s to 75.1 s). The ranges do not overlap and are not confidence intervals. With thinking on, Haiku used a median 4,556 reasoning tokens per call and 5,064 output tokens. With it off, a median 291 output tokens.
At list price, the 24 thinking-off calls cost about $0.15 and the 24 thinking-on calls about $0.74 (calculations, not bills). Per strict pass that is $0.0365 against $0.0672 (a calculation: every call counted, failures included, divided by strict passes). In this sample thinking off was cheaper per strict pass. That figure rests on 4 and 11 passes, has no interval, and the pass-rate intervals overlap, so treat the order as unsettled. Thinking off also passed fewer tasks (4 of 24 against 11 of 24), and that count matters if your task needs a pass rate.
Questions people ask
How do I turn off extended thinking in Claude Code? Set MAX_THINKING_TOKENS=0 in the environment of the claude process. It worked for Haiku 4.5 on Claude Code 2.1.286 (the version the hard-task receipts record): every call reported 0 thinking tokens. We did not test other models, other versions or the API's own settings. Check the thinking-token counter in the CLI's JSON result on your setup.
Is Claude Haiku slow because of thinking? In our runs, the thinking-off router had a lower median time. The thinking-on router took a median 12.54 s and the thinking-off router 4.66 s (the arms ran on different days). The CLI adds its own start-up time on top of the model's time.
Should I keep thinking on for Haiku 4.5? It depends on the task. For short typed decisions we could not show an accuracy loss with it off (p = 0.75, 10 differing cases), and it was about 2.7 times faster at the median (a calculation from two medians). For hard code and logic tasks thinking on passed more calls: the strict intervals overlap, the lenient ones do not.
What to do
- Try thinking off for classification and routing calls. We could not show a difference between thinking off and thinking on across 82 decisions (p = 0.75), and thinking off took less than half the time and cost (calculations from medians and list prices). A paired test on 10 differing cases has little power, so check it on your own cases.
- Test thinking on, or a bigger model, for hard problems. Thinking on passed more calls here, but strict intervals overlap. The lenient intervals do not overlap. Our hard-task results for Haiku, Sonnet, Opus and Fable show the next step up.
- Count cost per pass, not cost per call, and keep its uncertainty in view. Here thinking off was cheaper per strict pass ($0.0365 against $0.0672, a calculation from 4 and 11 passes, with wide uncertainty), but it passed fewer tasks (4 of 24 against 11 of 24). If you need a pass rate, start from that count. See is Haiku cheaper than Sonnet per correct answer.
- If speed is the goal, compare with a larger model at low effort. In the CLI, Sonnet 5.5 at low effort had a lower median time, with overlapping single-call ranges. It scored 77 of 82 exactly (94%, 95% interval 87% to 97%) (the paired test against thinking off gives p = 0.07: not below 0.05, so we cannot rank them).
- Measure your own cases. 82 decisions is a small set. Run yours with the thinking counter on.
More on routing: Jev vs Haiku vs Sonnet as a router, what a router costs you and the routing hub.
How we measured
- Protocol chronology: controls ran at 21:27:56 UTC on 2026-10-06. The protocol file was created at 21:29:01, before the new probes and the first new counted call at 21:29:57. Reused reference calls predate it. Later amendments changed the file; no frozen initial copy verifies its original wording. Two uncounted probes (one per route) showed 0 thinking tokens before the new counted calls.
- Routing: the platform's labelled decision suites at fixed versions, only the cases production asks a router about: 82 decisions, 194 scored questions. Claude Code
-p, tools off, no MCP, the production system text, prompt and JSON schema. One call at a time. Scored by the repository's decision-eval runner. Paired test: exact two-sided McNemar on the cases where the arms differ. - Timing: wall time per call and the CLI's own API time. Percentiles by the nearest-rank rule, the same rule as the routing-overhead study. (The Jev-vs-LLM page interpolates between ranks, so its Haiku median reads 12.67 s.)
- Hard tasks: eight tasks with deterministic validators in a sandbox without network. Three repetitions. 300 s timeout.
- Cost: reported tokens times list price. A calculation, not a bill. Sonnet's figure uses its recorded one-hour cache writes; the calculation and CLI report both give $7.32 per 1,000 decisions.
- Every attempt is kept. 106 counted calls, nothing trimmed, nothing retried, no usage limit hit.
Caveats
- Not concurrent. The thinking-on arms are recorded runs from 2026-10-05 and 2026-10-06. The thinking-off calls ran later on a different Claude subscription account. Day, account, CLI version and input-token differences can affect the results. This comparison cannot isolate the effect of thinking. The shared Mac had no recorded host-load control.
- Small samples. 82 routing decisions in four hand-labelled sets; 24 hard calls per cell. A 4 of 24 result has a 95% interval of 7% to 36%. The case sets and question wording were tuned in fix waves against Jev answers (2026-10-04 to 2026-10-05).
- The Haiku arms differ in input tokens. The thinking-on routing calls report about 297 more input tokens per call than the thinking-off calls, with the same prompt character count. The cause is unknown. It is worth about $0.30 per 1,000 decisions at list price (a calculation).
- The CLI, not the API. Claude Code adds start-up time and tool-schema tokens. A direct API call would skip them. The
MAX_THINKING_TOKENSresult says nothing about the API's thinking settings. - Usage checks. No saved pre-batch usage readings are in the run folder. Later checks do not prove that each pre-batch gate ran.
- Strict format rules decide part of the hard result. We show the lenient reading next to the strict one.
- One model. This is Haiku 4.5 only. Do not read it as a result for other models.
- Disclosure. I build Agent, a product that routes work between models, so I have an interest in cheap, fast routers. The hard-task result goes the other way, and we publish it as measured. The receipts and the method are public, so you can check every number.
What to read next
- Jev vs Claude Haiku vs Claude Sonnet as a router: an honest comparison
- Does reasoning effort buy quality? Claude and Codex on hard tasks
- Why Claude Code is slow: where the seconds go
See which model and effort each step got
Agent records each routing decision with its reason, its cost and its time. Try Agent and see what your work would route to.