How to turn off extended thinking in Claude Code, and what we measured for Haiku 4.5
Set MAX_THINKING_TOKENS=0, then check both counters. Our Haiku 4.5 test covers routing and hard tasks, with timings, costs and uncertainty.
TL;DR
- How: set
MAX_THINKING_TOKENS=0in the environment of theclaudeprocess. It held for Haiku 4.5 in Claude Code. All 106 counted calls reported 0 thinking tokens (0/106 with reported thinking; 95% Wilson interval 0% to 3.5%). - Check it: read
usage.output_tokens_details.thinking_tokensandmodelUsage[].thinkingTokensin the--output-format jsonresult. Zero means the CLI reported no thinking tokens. It does not verify the absence of hidden reasoning. Missing counters do not verify the setting. Check both on every call. - Short typed decisions (82 cases): thinking off answered 71 exactly (87%, 95% interval 78% to 92%). Thinking on answered 73 (89%, 80% to 94%). The paired exact test gives p = 0.754: no difference shown. Median time was 4.66 s with thinking off (range 2.20 to 11.57 s) and 12.54 s with it on (5.86 to 51.28 s), n = 82 each. List-price cost per 1,000 decisions fell from $8.92 to $3.36 (a calculation).
- Hard tasks (24 calls per setting): thinking off passed 4 strictly (17%, 7% to 36%). Thinking on passed 11 (46%, 28% to 65%). The strict intervals overlap. Count right answers in the wrong format: 4/24 (17%, 7% to 36%) against 16/24 (67%, 47% to 82%). Those intervals do not overlap.
- Our reading of this sample: try thinking off for short typed decisions. Try thinking on for hard problems, where it led on the lenient reading. Then check both on your own cases.
- Limits: the thinking-on arms reuse earlier sessions, so the timings are not concurrent. The routing reference ran the day before; the hard-task reference ran earlier the same day. The comparison cannot isolate thinking’s effect. The samples are small. This is the Claude Code CLI, not the API.
Full study: /benchmarks/haiku-thinking-on-off. For the results write-up with every chart and the cost detail, read Does extended thinking pay for Claude Haiku 4.5?
Step 1: turn thinking off
Our earlier Haiku 4.5 runs used the Claude Code default and reported thinking tokens. Haiku had the highest router median: 12.54 s (range 5.86 to 51.28 s, n = 82). On the hard tasks it also had the highest median: 39.0 s (15.3 to 75.1 s, n = 24). The ranges overlapped other models, so these medians do not establish a speed ranking. We then tested the thinking-off setting. The ranges are observed calls, not confidence intervals.
Set the variable on the process that starts claude:
MAX_THINKING_TOKENS=0 claude -p --model haiku --output-format json "Your prompt"Our runs also used --setting-sources '' --strict-mcp-config --tools '' --no-session-persistence. Those flags drop settings, MCP servers, tools and session memory. The intended setting change was thinking. The arms also differed in run time and account; the routing CLI versions are not recorded. This comparison cannot isolate the effect of thinking.
If a script or runner starts claude for you, check that it passes the variable through. Our hard-task runner filters the environment with an allowlist. We added MAX_THINKING_TOKENS to the process it starts, because the allowlist alone would not pass it.
We tested the variable only for Haiku 4.5, in print mode (-p). The hard-task receipts record Claude Code 2.1.286. The routing calls ran on the same machine minutes earlier, but their log does not record the version. We did not test other models, interactive sessions, settings files or the API's own thinking setting.
Step 2: check that it held
Do not assume the setting worked. Each JSON result carries two counters. This is the real, trimmed output of our first routing probe in the public receipts:
{
"usage": { "output_tokens": 321, "output_tokens_details": { "thinking_tokens": 0 } },
"modelUsage": { "claude-haiku-4-5-20251001": { "outputTokens": 321, "thinkingTokens": 0 } }
}Read both counters with jq:
MAX_THINKING_TOKENS=0 claude -p --model haiku --output-format json "Your prompt" \
| jq '[.usage.output_tokens_details.thinking_tokens, (.modelUsage[].thinkingTokens)]'Our routing probe printed [0,0]: both counters reported zero. A missing counter (null) does not verify the setting. Before the counted runs, we made one uncounted probe call on each route. The amended protocol states a stop rule if the probe cannot disable thinking. Both probes reported zero thinking tokens.
The counters held on every call. All 82 routing calls and all 24 hard-task calls reported 0 thinking tokens. With thinking on, the 82 recorded routing calls each reported some, a mean of 1,101 per decision (range 285 to 4,443). The CLI counts thinking tokens inside output tokens. Mean visible output (a calculation: output minus thinking) was 366 tokens with thinking off and 318 with thinking on, n = 82 each.
Step 3: the short typed-decision results
The test: the 82 typed decisions our platform asks a router about. They are failure class, message intent, is-it-a-rule and context shape. One Claude Code call per decision. The repository's own decision-eval runner scores each one. "Exact" means every scored question in a case was acceptable.
| Arm (82 decisions) | Exact | 95% interval | Per question | Median wall time (range, s) | p95 | Thinking tokens (mean) | List cost per 1,000 (calculation) |
|---|---|---|---|---|---|---|---|
| Haiku 4.5, thinking off | 71 (87%) | 78% to 92% | 177 of 194 (91%; 86% to 94%) | 4.66 (2.20 to 11.57) | 8.18 s | 0 | $3.36 |
| Haiku 4.5, thinking on | 73 (89%) | 80% to 94% | 183 of 194 (94%; 90% to 97%) | 12.54 (5.86 to 51.28) | 34.48 s | 1,101 | $8.92 |
| Sonnet 5.5, effort low (reference) | 77 (94%) | 87% to 97% | 189 of 194 (97%; 94% to 99%) | 2.60 (1.99 to 5.58) | 4.30 s | 2 | $7.32 |
Per-question intervals are descriptive. Questions within a decision are related, so they are not an independent-question test.
Accuracy: no difference shown. The two Haiku intervals overlap. The paired test agrees. Thinking on was right where thinking off was wrong in 6 cases. The reverse happened in 4 cases. 67 cases were right both times and 5 were wrong both times. The exact McNemar test on those 10 cases gives p = 0.754.
Only 10 cases carry the whole comparison, so the test has little power. Read "no difference shown" as "we cannot tell". Do not read it as "equal". Message intent hit a ceiling in every arm: 20/20 (95% interval 84% to 100%). Is-it-a-rule also hit a ceiling: 12/12 (76% to 100%). Neither set can establish equal accuracy on harder cases. The other two types contain the differences. The cells show exact counts and 95% Wilson intervals:
| Decision type (cases) | Thinking off | Thinking on | Sonnet 5.5 low |
|---|---|---|---|
| Failure class (18) | 15 (61% to 94%) | 17 (74% to 99%) | 18 (82% to 100%) |
| Message intent (20) | 20 (84% to 100%) | 20 (84% to 100%) | 20 (84% to 100%) |
| Is it a rule? (12) | 12 (76% to 100%) | 12 (76% to 100%) | 12 (76% to 100%) |
| Context shape (32) | 24 (58% to 87%) | 24 (58% to 87%) | 27 (68% to 93%) |
Speed: a median-time ratio of about 2.7 (a calculation from two medians). Median wall time was 4.66 s (range 2.20 to 11.57 s) with thinking off and 12.54 s (5.86 to 51.28 s) with thinking on, n = 82 each. Ranges are not confidence intervals. The 95th percentile fell from 34.48 s to 8.18 s. The thinking-on median is above the thinking-off 95th percentile. The ranges of single calls still overlap: the slowest thinking-off call took 11.57 s, and the fastest thinking-on call took 5.86 s. The whiskers on the chart run from the median to the 95th percentile. They are not confidence intervals.
Cost: less than half at list price (a calculation). The calls ran on a flat subscription, so nobody paid these amounts. At list price, 1,000 decisions cost $3.36 with thinking off and $8.92 with thinking on. At 100,000 decisions that is about $336 against $892 (a calculation: 100 times each rounded figure).
Two notes on the cost figures:
- An input-token gap. Each case has the same recorded prompt character count in both Haiku arms. That alone does not prove identical prompt text. Yet the thinking-on calls report about 297 more input tokens per call (1,831 against 1,534, a calculation from two means). We do not know why. The CLI version of the thinking-on routing run is not recorded. That gap is worth about $0.30 of the $5.56 cost gap per 1,000 decisions (a calculation).
- Sonnet's price basis. Its $7.32 per 1,000 decisions uses the one-hour cache-write rate recorded in the receipts. This calculation matches the CLI-reported cost. The older $5.00 figure used the five-minute rate and does not match these receipts.
Sonnet’s reference median was lower than thinking-off Haiku’s. Sonnet 5.5 at low effort took a median 2.60 s (range 1.99 to 5.58 s, n = 82) through the same CLI. It answered 77 of 82 exactly (94%, 87% to 97%). Against thinking-off Haiku, the paired test gives p = 0.07, which is not below 0.05, so we cannot rank the two on accuracy. On time, Haiku's thinking-off median (4.66 s) is above Sonnet's 95th percentile (4.30 s). See every measured row for Haiku against Sonnet. For routers that decide in milliseconds or microseconds, see how fast Jev routes.
Step 4: the hard-task results
The test: the eight hard tasks of our hard head-to-head. Each has a sandboxed validator. Each ran three times per setting. We ran the validator controls before the first new model call, after the reused reference calls: all 8 reference answers pass and all 26 planted wrong answers fail. A strict pass needs the whole reply to pass as written.
| Setting (24 calls) | Strict passes | 95% interval | Format misses | Wrong answers | Median time per call (range) | Median output tokens (range) | List cost per strict pass (calculation) |
|---|---|---|---|---|---|---|---|
| Thinking off | 4 (17%) | 7% to 36% | 0 | 20 | 2.9 s (1.7 to 13.0) | 291 (50 to 1,628) | $0.0365 |
| Thinking on | 11 (46%) | 28% to 65% | 5 | 8 | 39.0 s (15.3 to 75.1) | 5,064 (1,899 to 9,321) | $0.0672 |
The strict intervals overlap, so the strict result alone does not separate the two settings. Three repeats share each task. These call-level intervals describe this sample; they do not measure uncertainty across unseen tasks. A format miss is a right answer in the wrong wrapping, such as a code fence. On the lenient reading, which counts format misses, it is 4/24 (17%, 7% to 36%) against 16/24 (67%, 47% to 82%). Those intervals do not overlap, and thinking on is ahead there. The strict-versus-lenient post explains why we show both.
| Hard task (strict passes of 3; 95% interval) | Thinking off | Thinking on |
|---|---|---|
| Fix an interval-merge function | 3 (44% to 100%) | 3 (44% to 100%) |
| Fix a time-zone day-length function (DST) | 0 (0% to 56%) | 1 (6% to 79%) |
| Write a CSV parser | 0 (0% to 56%) | 2 (21% to 94%) (and 1 format miss) |
| Predict JavaScript event-loop output order | 0 (0% to 56%) | 0 (0% to 56%) |
| Solve a multi-constraint room schedule | 0 (0% to 56%) | 0 (0% to 56%) (and 2 format misses) |
| Write a strict SemVer 2.0.0 regex | 0 (0% to 56%) | 3 (44% to 100%) |
| Refactor to remove duplication, keep 20 tests green | 1 (6% to 79%) | 2 (21% to 94%) |
| Write a SQLite reporting query | 0 (0% to 56%) | 0 (0% to 56%) (and 2 format misses) |
Both arms hit a ceiling on interval merge: 3/3 each (95% interval 44% to 100%). This cannot establish equal accuracy on harder variants.
Thinking off had a lower median: 2.9 s (range 1.7 to 13.0 s) against 39.0 s (15.3 to 75.1 s), n = 24 each. These call ranges do not overlap. They are not confidence intervals, and the runs cannot establish a cause. With thinking on, Haiku used a median 4,556 reasoning tokens per call (range 1,452 to 8,569, n = 24).
Cost per strict pass was lower with thinking off: $0.0365 against $0.0672 (a calculation: the cost of every call, failures included, divided by strict passes). That figure rests on 4 and 11 passes and has no interval. Thinking off had a lower calculated cost per pass in this sample, and it passed fewer calls. The strict pass intervals overlap, so the cost-per-pass order is not settled. If you need a pass rate, start from the pass count.
What to do
- Try thinking off for short typed decisions. That means classification, routing and yes-or-no calls. We could not show an accuracy difference on 82 decisions (p = 0.754). The observed median time was lower with thinking off. The runs cannot isolate a cause.
- Try thinking on for hard problems, or test a larger model. Thinking on led on the lenient reading of this sample. The strict pass intervals overlap. See hard tasks for Haiku, Sonnet, Opus and Fable for the next step up.
- Check the counter on every call. A wrapper can drop the variable without a warning. The counter tells you.
- Run your own cases both ways and pair them by case. Use the exact McNemar test on the cases where the two settings differ. 10 differing cases is a small test. The Wilson interval explainer shows how wide the uncertainty is at this size.
- Count cost per pass, not cost per call. A cheap call that fails more often can cost more per right answer. See Is Claude Haiku cheaper than Sonnet?
- If speed is the goal, test a larger model at low effort. In this CLI, Sonnet 5.5 at low effort had the lower median. What is reasoning effort? explains the setting.
Questions people ask
How do I turn off extended thinking in Claude Code? Set MAX_THINKING_TOKENS=0 in the environment of the claude process, then read both thinking-token counters in the JSON result. Missing counters do not verify the setting. It worked for Haiku 4.5 in print mode. The hard-task receipts record version 2.1.286. Check the counter on your own setup.
Do thinking tokens cost money? In the CLI's counters, thinking tokens are part of the output tokens. With thinking on, the mean output was 1,419 tokens per routing decision: 1,101 thinking and 318 visible. Output tokens cost more than input tokens at list price. Our calls used a subscription, so the dollar figures here are calculations.
Does turning thinking off lower accuracy? We cannot tell on the 82 typed decisions we tested (p = 0.754). On the eight hard tasks, strict passes were 4/24 (95% interval 7% to 36%) off and 11/24 (28% to 65%) on. Those intervals overlap. Lenient passes were 4/24 (7% to 36%) off and 16/24 (47% to 82%) on. Those intervals do not overlap. Repeats share tasks; these intervals do not measure uncertainty across unseen tasks.
Is Claude Haiku slow because of thinking? In our runs, the thinking-off router had a median of 4.66 s (range 2.20 to 11.57 s, n = 82) and the thinking-on router 12.54 s (5.86 to 51.28 s, n = 82). The arms ran on different days. The CLI also adds its own start-up time to every call.
Should I do this for Sonnet or Opus? We do not know. We tested one model. Run the check on yours.
How we measured
- Protocol timing. Controls ran at 21:27:56 UTC on 2026-10-06. The protocol file was created at 21:29:01 UTC. The two probes and first new counted call followed. The reused calls predate it. Later amendments changed the file; no frozen initial copy verifies its original wording.
- Routing: the platform's labelled decision suites at fixed versions. 82 decisions and 194 scored questions. Claude Code
-p, tools off, no MCP, the production system text, prompt and JSON schema. One call at a time. The repository's decision-eval runner scores each call. The paired test is the exact two-sided McNemar test. - Hard tasks: eight tasks, three repetitions, a 300 s timeout, a sandboxed validator for each task.
- Timing: wall time per call. Percentiles use the nearest-rank rule of the routing-overhead study. The Jev-versus-LLM page interpolates between ranks, so it shows 12.67 s for the same thinking-on median.
- Cost: reported tokens times list price. A calculation, not a bill.
- We kept every attempt. 106 counted calls and 2 uncounted probes. We trimmed and retried nothing, and no usage limit stopped a run.
Limits
- Not concurrent. The thinking-on arms and the Sonnet 5.5 arm are recorded runs: the routing run of 2026-10-05 and the hard-task receipts of 2026-10-06. We reused them. The thinking-off calls ran on 2026-10-06, later, on a different Claude subscription account. Day, account, CLI version and input-token differences can affect results. This run cannot isolate the effect of thinking.
- Usage-gate records are missing. No saved pre-batch readings verify that each usage gate ran. Later corroboration does not prove the earlier checks.
- Machine load is unknown. Other jobs may have run on the same Mac. We did not record the load, so wall times may include contention.
- Small samples. 82 routing decisions in four hand-labelled sets and 24 hard calls per setting. A result of 4 of 24 has a 95% interval of 7% to 36%. We tuned the case sets and question wording in fix waves against Jev answers (2026-10-04 to 2026-10-05).
- The Haiku routing arms differ in input tokens. The cause is unknown (see Step 3).
- The CLI, not the API. Claude Code adds start-up time and tool-schema tokens. A direct API call would skip them. Our result says nothing about the API's thinking settings.
- Strict format rules decide part of the hard result. We show the lenient reading next to the strict one.
- One model, one CLI. This is Haiku 4.5 in Claude Code. Only the hard-task receipts record the version (2.1.286).
- Disclosure. I build Agent, a product that routes work between models. That gives me an interest in cheap, fast routers. The hard-task result goes the other way, and we publish it as measured. The receipts and the method are public.
What to read next
- Does extended thinking pay for Claude Haiku 4.5?, the results write-up for this study
- Does reasoning effort buy quality? Claude and Codex on hard tasks
- Why is Claude Code slow? Where the seconds go
- Jev vs Claude Haiku vs Sonnet as a router and the routing hub
See which model and effort each step got
Agent records each routing decision with its reason, its cost and its time. Try Agent and see what your work would route to.