• Claude Haiku
  • Extended Thinking
  • Routing
  • Latency

Does extended thinking pay for Claude Haiku 4.5? We turned it off and measured

Claude Haiku 4.5, thinking off: 2.7x faster (a calculation from two medians), no routing accuracy gap shown. Hard tasks: 4/24 passed, thinking on 11/24.

TL;DR

  • Question: does extended thinking pay for Claude Haiku 4.5? In our earlier runs it took a median 12.54 s per routing decision through Claude Code and 39.0 s on hard tasks. We turned thinking off and ran both tests again.
  • Router, 82 typed decisions: thinking off answered 71 of 82 exactly (87%, 95% interval 78% to 92%). Thinking on answered 73 of 82 (89%, 80% to 94%). On the same cases the paired exact test gives p = 0.75: no difference shown.
  • Speed and cost of that router: median wall time 4.66 s with thinking off, 12.54 s with thinking on (p95 8.18 s against 34.48 s). Thinking tokens per decision: 0 against 1,101 (mean). List-price cost per 1,000 decisions: $3.36 against $8.92 (a calculation).
  • Hard tasks, 24 calls per cell: thinking off passed 4 of 24 strictly (17%, 7% to 36%). Thinking on passed 11 of 24 (46%, 28% to 65%). The strict intervals overlap. Count the right answers in the wrong format, and it is 4 of 24 against 16 of 24 (67%, 47% to 82%): those intervals do not overlap.
  • Time on hard tasks: a median 2.9 s against 39.0 s per call. Cost per strict pass: $0.0365 against $0.0672 (a calculation from 4 and 11 passes, with no interval). Thinking off was cheaper per pass in this sample, but it passed fewer tasks.
  • What to do (our reading of this sample): turn thinking off for short typed decisions, where we could not show an accuracy loss (p = 0.75, 10 differing cases). Test thinking on for hard problems. It passed more calls here, but strict intervals overlap. Sonnet 5.5 at low effort had a lower median of 2.60 s through the same CLI. Its range overlaps Haiku with thinking off.
  • Limits: the thinking-on arms are recorded runs from other days, so the timing comparison is not concurrent. 82 decisions and 24 hard calls are small samples.

The full study: /benchmarks/haiku-thinking-on-off.

What we tested

Claude Haiku 4.5 reported thinking tokens under the Claude Code default in our reference runs. Those tokens count as billed output at list price. Do they buy anything?

  • Thinking off. We set the environment variable MAX_THINKING_TOKENS=0 on the Claude Code process. We did not assume it worked. Two uncounted probes and all 106 counted calls reported 0 thinking tokens in the CLI's own counters.
  • Test 1, routing. The 82 typed decisions our platform asks a router about: failure class, message intent, is-it-a-rule, context shape. One call per decision. The repository's own decision-eval runner scores each one.
  • Test 2, hard tasks. The eight hard tasks of our hard head-to-head, each with a sandboxed validator, three repetitions each. We ran the validator controls first: 8 of 8 reference answers pass and 26 of 26 planted wrong answers fail.
  • Reference arms. The thinking-on numbers are our recorded runs: the routing run of 2026-10-05 and the hard-task receipts of 2026-10-06. We reused them. We did not rerun them.

Router accuracy: no difference shown

  • Exact decisions (every scored question right)
  • Per-question accuracy
Claude Haiku 4.5 (thinking off) · Claude Code
Claude Haiku 4.5 (thinking on) · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

3 rows, 2 series: Exact decisions (every scored question right), Per-question accuracy. Exact decisions (every scored question right): highest Claude Sonnet 5.5 (low) · Claude Code 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 (thinking off) · Claude Code 87% (95% interval 78%–92%, n 82). All intervals overlap. Per-question accuracy: highest Claude Sonnet 5.5 (low) · Claude Code 97% (95% interval 94%–99%, n 194). Lowest Claude Haiku 4.5 (thinking off) · Claude Code 91% (95% interval 86%–94%, n 194). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 82–194 per row

Claude Haiku 4.5 and Claude Sonnet 5.5 (low effort); 82 decisions, the same cases for every arm

Whiskers are 95% Wilson intervals. Questions within a decision are related; per-question intervals are descriptive, not an independent-question test. The thinking-on and Sonnet arms are the recorded 2026-10-05 routing run, reused, not rerun; the thinking-off arm ran later on another account. An unanswered question counts as wrong.

Sources: Haiku thinking on vs off, Routing runs: Jev router vs LLM routing

Exact means every scored question in a case was acceptable. An unanswered question counts as wrong. All intervals below are 95% Wilson intervals. Questions within one case are related, so per-question intervals are descriptive.

Arm (82 decisions)Exact95% intervalPer question95% interval
Haiku 4.5, thinking off71 (87%)78% to 92%177 of 194 (91%)86% to 94%
Haiku 4.5, thinking on73 (89%)80% to 94%183 of 194 (94%)90% to 97%
Sonnet 5.5, effort low (reference)77 (94%)87% to 97%189 of 194 (97%)94% to 99%

The Haiku intervals overlap, so the data does not rank the two settings. The case-by-case test agrees. Thinking on was right where thinking off was wrong in 6 cases, and the reverse in 4. 67 cases were right both times and 5 wrong both times. The exact McNemar test (two-sided, on those 10 cases) gives p = 0.754.

Pair, same 82 casesBoth rightOnly thinking off rightOnly the other arm rightBoth wrongExact McNemar p
Thinking off vs thinking on674650.754
Thinking off vs Sonnet 5.5 low701740.07

By decision type, message intent was 20 of 20 in every arm (95% interval 84% to 100%). Is-it-a-rule was 12 of 12 (76% to 100%). These subsets hit a ceiling; they cannot show equal accuracy on harder cases. Failure class scored 15 of 18 with thinking off (61% to 94%) and 17 of 18 with thinking on (74% to 99%). Context shape scored 24 of 32 in both (58% to 87%). Few cases carry the whole difference, so the paired test has little power. Read "no difference shown" as "we cannot tell", not as "equal".

Router speed: about 2.7 times faster (a calculation from two medians)

  • Wall time (CLI)
  • Model time (API)
Entrance: medians race at 9× real timeMotion reduced: press Replay to animateThe slowest median is 12.5 s. The clock runs at the recorded speed.
Claude Haiku 4.5 (thinking off) · Claude Code
Claude Haiku 4.5 (thinking on) · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

3 rows, 2 series: Wall time (CLI), Model time (API). Wall time (CLI): slowest Claude Haiku 4.5 (thinking on) · Claude Code 12.5 s (median to p95 12.5 s–34.5 s, n 82). Fastest Claude Sonnet 5.5 (low) · Claude Code 2.6 s (median to p95 2.6 s–4.3 s, n 82). Not all run ranges overlap. Model time (API): slowest Claude Haiku 4.5 (thinking on) · Claude Code 10.5 s (median to p95 10.5 s–32.1 s, n 82). Fastest Claude Sonnet 5.5 (low) · Claude Code 1.6 s (median to p95 1.6 s–2.6 s, n 82). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n = 82 per row

Median wall time and model (API) time; whisker to the 95th percentile

Whiskers run from p50 to p95, not a confidence interval. Percentiles use the nearest-rank rule of the routing-overhead study, so the thinking-on and Sonnet medians match that study (the Jev-vs-LLM page interpolates between ranks and shows slightly different values). One call at a time through the Claude Code CLI; the thinking-on and Sonnet arms ran on another day.

Sources: Haiku thinking on vs off, Routing runs: Jev router vs LLM routing

  • Wall time per decision (median, p95): thinking off 4.66 s (8.18 s); thinking on 12.54 s (34.48 s); Sonnet 5.5 at low effort 2.60 s (4.30 s).
  • Wall-time ranges (82 calls per arm): thinking off 2.20 s to 11.57 s; thinking on 5.86 s to 51.28 s; Sonnet 1.99 s to 5.58 s. These are single-call ranges, not confidence intervals.
  • Model (API) time, median and range: thinking off 3.79 s (1.37 s to 10.83 s); thinking on 10.51 s (4.33 s to 49.49 s); Sonnet 1.60 s (1.06 s to 4.76 s). Each arm has 82 calls; these ranges are not confidence intervals.

The whisker runs from the median to the 95th percentile. It is not a confidence interval. The thinking-on median sits above the thinking-off 95th percentile. The arms ran on different days and accounts, so this gap cannot isolate the effect of thinking. The ranges of single calls overlap: the slowest thinking-off call took 11.57 s and the fastest thinking-on call took 5.86 s.

Haiku with thinking off still had a higher median than Sonnet at low effort through the same CLI: Haiku's thinking-off median (4.66 s) is above Sonnet's 95th percentile (4.30 s). See every measured row for Haiku vs Sonnet for the full list. Their single-call ranges overlap, so these runs do not establish a speed ranking. Where this sits next to rule-based routing and Jev:

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

Time per decision · log scale: each gridline is 10 times the one before

4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–20000 per row

Median; whiskers = median to 95th percentile

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

That chart shows the thinking-on Haiku and Sonnet medians (12,543 ms and 2,597 ms) next to a deterministic policy that decides in microseconds. The routing-overhead study has the per-task arithmetic. For Jev, see the Jev vs Haiku page.

Router tokens and cost

  • Thinking tokens
  • Visible output tokens
Claude Haiku 4.5 (thinking off) · Claude Code
Claude Haiku 4.5 (thinking on) · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

3 rows, 2 series: Thinking tokens, Visible output tokens. Thinking tokens: highest Claude Haiku 4.5 (thinking on) · Claude Code 1,101 (n 82). Lowest Claude Haiku 4.5 (thinking off) · Claude Code 0 (n 82). Visible output tokens: highest Claude Haiku 4.5 (thinking off) · Claude Code 366 (n 82). Lowest Claude Sonnet 5.5 (low) · Claude Code 105 (n 82).

Notesn = 82 per row

Mean per decision, as the Claude Code CLI reports them

Visible output = output tokens minus thinking tokens. It includes the structured answer the CLI asks for. Thinking tokens are counted by the CLI; their content is never captured. More tokens is not better or worse by itself.

Sources: Haiku thinking on vs off, Routing runs: Jev router vs LLM routing

With thinking on, Haiku wrote a mean 1,101 thinking tokens per decision (range 285 to 4,443). With thinking off it wrote none. The visible output was about the same: 366 tokens with thinking off, 318 with thinking on.

Calculation
Claude Haiku 4.5 (thinking off) · Claude Code
Claude Haiku 4.5 (thinking on) · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

Hover or focus a bar for its ratio to Claude Haiku 4.5 (thinking of… (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 (thinking on) · Claude Code $8.92 (n 82). Lowest Claude Haiku 4.5 (thinking off) · Claude Code $3.36 (n 82).

Notesn = 82 per row

Reported tokens × list price, per 1,000 decisions

Calculation, not a bill: the calls ran on a flat subscription. Reported input, cache and output tokens (thinking tokens are part of output) × list price, with the price table of the Jev-vs-LLM study. Sonnet cache writes use the one-hour rate ($4 per million tokens), as recorded in its receipts. The calculation matches the CLI-reported cost.

Sources: Haiku thinking on vs off, Routing runs: Jev router vs LLM routing, Repricing calculation, Anthropic list prices (Claude models)

List-price cost per 1,000 decisions, from the tokens the CLI reported: $3.36 with thinking off, $8.92 with thinking on, $7.32 for Sonnet 5.5 at low effort. These are calculations. The calls ran on a flat subscription. Haiku with thinking off had the lowest calculated cost of the three and a higher median time than Sonnet.

Two notes on these costs:

  • Sonnet's price basis. Its calls wrote one-hour cache entries, priced at $4 per million tokens. The calculation is $7.32 ($7.324) per 1,000 decisions, matching the CLI's reported cost. The older $4.996 calculation used the 5-minute cache-write rate. It does not match these receipts. The two Haiku calculations also match their CLI-reported costs.
  • An input-token gap between the Haiku arms. Each case has the same prompt character count, but the thinking-on calls report about 297 more input tokens per call (1,831 against 1,534, a calculation from the two means). We do not know why, and the CLI version of the thinking-on routing run is not recorded. At Haiku's list price that is about $0.30 of the $5.56 cost gap per 1,000 decisions (a calculation).

Hard tasks: thinking on passed more calls

  • Strict pass
  • Lenient (format misses counted)
Claude Haiku 4.5 (thinking off) · Claude Code
Claude Haiku 4.5 (thinking on) · Claude Code

2 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Haiku 4.5 (thinking on) · Claude Code 46% (95% interval 28%–65%, n 24). Lowest Claude Haiku 4.5 (thinking off) · Claude Code 17% (95% interval 6.7%–36%, n 24). All intervals overlap. Lenient (format misses counted): highest Claude Haiku 4.5 (thinking on) · Claude Code 67% (95% interval 47%–82%, n 24). Lowest Claude Haiku 4.5 (thinking off) · Claude Code 17% (95% interval 6.7%–36%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 24 per row

Claude Haiku 4.5 in Claude Code; strict and lenient

Whiskers are 95% Wilson intervals over calls. Three repeats per task are related, so these intervals do not measure uncertainty across unseen tasks. A format miss (right answer in the wrong wrapping) never counts as a strict pass. The thinking-on cell is the 24 recorded Haiku receipts of the hard head-to-head, reused; the thinking-off cell ran later on another account.

Sources: Haiku thinking on vs off, Provider head-to-head, hard set: eight hard tasks with strict validators

Eight tasks, three repetitions each, strict validators. A strict pass needs the whole reply to pass as given. Repeats on the same task are related. The call-level intervals are descriptive; they do not measure uncertainty across unseen tasks.

  • Thinking off: 4 of 24 strict passes (17%, 95% interval 7% to 36%). 0 format misses, 20 wrong answers.
  • Thinking on: 11 of 24 (46%, 28% to 65%). 5 format misses, 8 wrong answers.

The strict intervals overlap, so the strict result alone does not separate the two. The lenient reading counts a correct answer in the wrong wrapping, such as a code fence. It is 4 of 24 against 16 of 24 (67%, 47% to 82%), and those intervals do not overlap. On the lenient reading, thinking on passed more.

By task, thinking off passed the interval-merge fix 3 of 3 (95% interval 44% to 100%) and the refactor 1 of 3 (6% to 79%). Every other task scored 0 of 3 (0% to 56%). Thinking on passed interval merge and strict SemVer 3 of 3 each (44% to 100%). It passed CSV parsing and the refactor 2 of 3 each (21% to 94%), and DST day length 1 of 3 (6% to 79%). Its other three tasks scored 0 of 3 each (0% to 56%). Both arms hit a ceiling on interval merge. This does not establish equal accuracy on harder variants.

Entrance: medians race at 28× real time
Claude Haiku 4.5 (thinking off) · Claude Code
Claude Haiku 4.5 (thinking on) · Claude Code

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

2 rows. Slowest Claude Haiku 4.5 (thinking on) · Claude Code 39 s (range 15.3 s–75.1 s, n 24). Fastest Claude Haiku 4.5 (thinking off) · Claude Code 3 s (range 1.7 s–13 s, n 24). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 24 per row

Median per cell; whiskers = fastest and slowest call

Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network; the thinking-on calls ran in an earlier session. Timings include CLI start-up.

Sources: Haiku thinking on vs off, Provider head-to-head, hard set: eight hard tasks with strict validators

Median total time per call: 2.9 s with thinking off (range 1.7 s to 13.0 s) against 39.0 s with thinking on (15.3 s to 75.1 s). The ranges do not overlap and are not confidence intervals. With thinking on, Haiku used a median 4,556 reasoning tokens per call and 5,064 output tokens. With it off, a median 291 output tokens.

At list price, the 24 thinking-off calls cost about $0.15 and the 24 thinking-on calls about $0.74 (calculations, not bills). Per strict pass that is $0.0365 against $0.0672 (a calculation: every call counted, failures included, divided by strict passes). In this sample thinking off was cheaper per strict pass. That figure rests on 4 and 11 passes, has no interval, and the pass-rate intervals overlap, so treat the order as unsettled. Thinking off also passed fewer tasks (4 of 24 against 11 of 24), and that count matters if your task needs a pass rate.

Questions people ask

How do I turn off extended thinking in Claude Code? Set MAX_THINKING_TOKENS=0 in the environment of the claude process. It worked for Haiku 4.5 on Claude Code 2.1.286 (the version the hard-task receipts record): every call reported 0 thinking tokens. We did not test other models, other versions or the API's own settings. Check the thinking-token counter in the CLI's JSON result on your setup.

Is Claude Haiku slow because of thinking? In our runs, the thinking-off router had a lower median time. The thinking-on router took a median 12.54 s and the thinking-off router 4.66 s (the arms ran on different days). The CLI adds its own start-up time on top of the model's time.

Should I keep thinking on for Haiku 4.5? It depends on the task. For short typed decisions we could not show an accuracy loss with it off (p = 0.75, 10 differing cases), and it was about 2.7 times faster at the median (a calculation from two medians). For hard code and logic tasks thinking on passed more calls: the strict intervals overlap, the lenient ones do not.

What to do

  1. Try thinking off for classification and routing calls. We could not show a difference between thinking off and thinking on across 82 decisions (p = 0.75), and thinking off took less than half the time and cost (calculations from medians and list prices). A paired test on 10 differing cases has little power, so check it on your own cases.
  2. Test thinking on, or a bigger model, for hard problems. Thinking on passed more calls here, but strict intervals overlap. The lenient intervals do not overlap. Our hard-task results for Haiku, Sonnet, Opus and Fable show the next step up.
  3. Count cost per pass, not cost per call, and keep its uncertainty in view. Here thinking off was cheaper per strict pass ($0.0365 against $0.0672, a calculation from 4 and 11 passes, with wide uncertainty), but it passed fewer tasks (4 of 24 against 11 of 24). If you need a pass rate, start from that count. See is Haiku cheaper than Sonnet per correct answer.
  4. If speed is the goal, compare with a larger model at low effort. In the CLI, Sonnet 5.5 at low effort had a lower median time, with overlapping single-call ranges. It scored 77 of 82 exactly (94%, 95% interval 87% to 97%) (the paired test against thinking off gives p = 0.07: not below 0.05, so we cannot rank them).
  5. Measure your own cases. 82 decisions is a small set. Run yours with the thinking counter on.

More on routing: Jev vs Haiku vs Sonnet as a router, what a router costs you and the routing hub.

How we measured

  • Protocol chronology: controls ran at 21:27:56 UTC on 2026-10-06. The protocol file was created at 21:29:01, before the new probes and the first new counted call at 21:29:57. Reused reference calls predate it. Later amendments changed the file; no frozen initial copy verifies its original wording. Two uncounted probes (one per route) showed 0 thinking tokens before the new counted calls.
  • Routing: the platform's labelled decision suites at fixed versions, only the cases production asks a router about: 82 decisions, 194 scored questions. Claude Code -p, tools off, no MCP, the production system text, prompt and JSON schema. One call at a time. Scored by the repository's decision-eval runner. Paired test: exact two-sided McNemar on the cases where the arms differ.
  • Timing: wall time per call and the CLI's own API time. Percentiles by the nearest-rank rule, the same rule as the routing-overhead study. (The Jev-vs-LLM page interpolates between ranks, so its Haiku median reads 12.67 s.)
  • Hard tasks: eight tasks with deterministic validators in a sandbox without network. Three repetitions. 300 s timeout.
  • Cost: reported tokens times list price. A calculation, not a bill. Sonnet's figure uses its recorded one-hour cache writes; the calculation and CLI report both give $7.32 per 1,000 decisions.
  • Every attempt is kept. 106 counted calls, nothing trimmed, nothing retried, no usage limit hit.

Caveats

  • Not concurrent. The thinking-on arms are recorded runs from 2026-10-05 and 2026-10-06. The thinking-off calls ran later on a different Claude subscription account. Day, account, CLI version and input-token differences can affect the results. This comparison cannot isolate the effect of thinking. The shared Mac had no recorded host-load control.
  • Small samples. 82 routing decisions in four hand-labelled sets; 24 hard calls per cell. A 4 of 24 result has a 95% interval of 7% to 36%. The case sets and question wording were tuned in fix waves against Jev answers (2026-10-04 to 2026-10-05).
  • The Haiku arms differ in input tokens. The thinking-on routing calls report about 297 more input tokens per call than the thinking-off calls, with the same prompt character count. The cause is unknown. It is worth about $0.30 per 1,000 decisions at list price (a calculation).
  • The CLI, not the API. Claude Code adds start-up time and tool-schema tokens. A direct API call would skip them. The MAX_THINKING_TOKENS result says nothing about the API's thinking settings.
  • Usage checks. No saved pre-batch usage readings are in the run folder. Later checks do not prove that each pre-batch gate ran.
  • Strict format rules decide part of the hard result. We show the lenient reading next to the strict one.
  • One model. This is Haiku 4.5 only. Do not read it as a result for other models.
  • Disclosure. I build Agent, a product that routes work between models, so I have an interest in cheap, fast routers. The hard-task result goes the other way, and we publish it as measured. The receipts and the method are public, so you can check every number.

See which model and effort each step got

Agent records each routing decision with its reason, its cost and its time. Try Agent and see what your work would route to.

The data behind this post

  • Claude Haiku
  • Extended Thinking

Does thinking pay for Claude Haiku 4.5? Thinking on vs off

Claude Haiku 4.5 with extended thinking on and off: 82 routing decisions and 8 hard tasks. Accuracy with 95% intervals, time and cost.

87% (71/82)Claude Haiku 4.5 (thinking off): exact routing decisions · n = 82

6 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.