{"i":17,"study":{"slug":"haiku-thinking-on-off","title":"Does thinking pay for Claude Haiku 4.5? Thinking on vs off","seoTitle":"Claude Haiku 4.5 thinking on vs off: benchmark results","description":"Claude Haiku 4.5 with extended thinking on and off: 82 routing decisions and 8 hard tasks. Accuracy with 95% intervals, time and cost.","question":"Does extended thinking pay for Claude Haiku 4.5: what does it buy in accuracy, and what does it cost in time and money, on typed routing decisions and on hard tasks?","answer":"Routing accuracy is unresolved on this sample of 82 paired decisions. Thinking off answered 87% (71/82; 95% interval 78% to 92%) exactly. Thinking on answered 89% (73/82; 95% interval 80% to 94%) exactly. Only thinking on was right in 6 cases; only thinking off in 4. Both were right in 67 cases and wrong in 5. The exact McNemar test gives p = 0.754. The test does not show a difference. This does not establish equal accuracy. The two 95% intervals overlap. Observed median wall time: 4.66 s with thinking off and 12.54 s with thinking on. The thinking-on median is 2.7 times the thinking-off median (a calculation). Wall-time ranges: 2.20 s to 11.57 s off; 5.86 s to 51.28 s on. These ranges overlap and are not confidence intervals. The 95th percentiles are 8.18 s off and 34.48 s on. Mean thinking tokens per decision: 0 off and 1,101 on. List-price cost per 1,000 decisions: $3.36 off and $8.92 on (a calculation). The recorded Sonnet 5.5 low-effort reference answered 94% (77/82; 95% interval 87% to 97%) exactly. Its median wall time was 2.60 s, with range 1.99 s to 5.58 s (not an interval). Only Sonnet was right in 7 cases; only thinking-off Haiku in 1. Exact McNemar p = 0.07 (secondary test). Hard tasks: 8 tasks, each repeated 3 times per arm. Thinking off passed 17% (4/24; 95% interval 7% to 36%) strictly. Thinking on passed 46% (11/24; 95% interval 28% to 65%) strictly. Format misses: 0 off and 5 on. These are right answers in the wrong wrapping, never strict passes. The strict 95% intervals overlap, so strict pass rate does not separate the arms. Lenient passes count format misses: 17% (4/24; 95% interval 7% to 36%) off; 67% (16/24; 95% interval 47% to 82%) on. The lenient intervals do not overlap: thinking on is ahead on this reading. Median total time per call: 2.9 s off and 39.0 s on. Single-call ranges: 1.7 s to 13.0 s off; 15.3 s to 75.1 s on. These ranges do not overlap and are not confidence intervals. List-price cost per strict pass: $0.0365 off and $0.0672 on (a calculation). The calculation divides all call costs by 4 and 11 strict passes. It has no interval. Thinking off was cheaper per strict pass in this sample, but it passed fewer calls. Both thinking-on arms reuse earlier recorded sessions, so the timing comparison is not concurrent.","date":"2026-10-06","updated":"2026-10-06","tags":["claude-haiku","extended-thinking","routing","hard-tasks","latency","cost"],"caveats":["The thinking-on and Sonnet arms are the recorded routing run of 2026-10-05, reused. The thinking-off arm ran on a different day and on a different Claude subscription account, so this is a confounded comparison. Day, account, CLI version and input-token differences can affect the results. The data cannot isolate the effect of thinking.","82 decisions in four small hand-labelled sets: intervals are wide, and a paired test with few discordant cases has little power. The case sets and question wording were tuned in fix waves against Jev answers (2026-10-04 to 2026-10-05).","Thinking off here means MAX_THINKING_TOKENS=0 on the Claude Code process, checked by the CLI’s own thinking-token counters. It says nothing about the API’s thinking settings or about other models.","Claude Code adds its own start-up time and tool-schema tokens to every call; a direct API router would skip them. CLI timings include that overhead.","List-price costs are calculations from reported tokens; the calls used a flat subscription. Sonnet cache writes use the one-hour rate ($4 per million tokens), as recorded in its receipts. The calculation matches the CLI-reported cost.","The Haiku routing arms have the same prompt character count for every case, but thinking-on calls report about 297 more input tokens per call (1,831 against 1,534, a calculation from the two means). The cause is unknown, and the CLI version of the thinking-on run is not recorded. At Haiku’s list price that is about $0.30 of the $5.56 cost gap per 1,000 decisions (a calculation).","Hard tasks: 24 calls per cell (3 per task); a 4/24 result has a 95% interval of 7% to 36%. The thinking-on cell is the recorded hard head-to-head receipts (batch of 2026-10-06, effort flag not passed), reused; the thinking-off cell ran later, on another account.","These are eight distinct hard tasks, repeated three times each. Repeats on the same task are related. Call-level Wilson intervals are descriptive; they do not measure uncertainty across unseen tasks.","Strict format rules decide part of the hard result: a reply in a code fence fails strictly. The lenient reading is shown next to it so the two can be told apart.","Cost per strict pass divides a cell’s cost by its strict passes (4 and 11). It has no interval, and the strict pass-rate intervals overlap, so the order of the two costs per pass is not settled.","Both arms hit a strict-pass ceiling on Fix an interval-merge function (off-by-one and edge cases): 100% (3/3; 95% interval 44% to 100%) in each arm. Repeats on these tasks cannot establish equal accuracy on harder variants.","Controls and reused calls predate protocol creation. The file predates the new probes and counted calls, but later amendments changed it. No frozen initial copy verifies the original wording.","The run folder has no saved pre-batch usage-gate readings. Later corroboration does not prove that each pre-batch gate ran.","The runs used a shared Mac. Host load was not recorded; the run does not prove isolation. Timing gaps do not establish a cause.","Routing has a ceiling on these subsets: Message intent, 100% (20/20; 95% interval 84% to 100%) in every arm; Is it a rule?, 100% (12/12; 95% interval 76% to 100%) in every arm. These results cannot establish equal accuracy on harder cases."],"sourceIds":["agent-haiku-thinking","agent-routing","calc-repricing","price-anthropic","agent-provider-h2h-hard"],"hero":{"statIds":["haiku-thinking-exact-off","haiku-thinking-exact-on"],"testStatId":"haiku-thinking-mcnemar"},"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["haiku-thinking-exact-off","Claude Haiku 4.5 (thinking off): exact routing decisions",0.8659,"rate","87% (71/82)",82,[0.7755,0.9234],"\u0001"],["haiku-thinking-exact-on","Claude Haiku 4.5 (thinking on): exact routing decisions",0.8902,"rate","89% (73/82)",82,[0.8044,0.9412],"\u0001"],["haiku-thinking-keys-off","Claude Haiku 4.5 (thinking off): per-question routing accuracy",0.9124,"rate","91% (177/194)",194,[0.8642,0.9446],"\u0001"],["haiku-thinking-keys-on","Claude Haiku 4.5 (thinking on): per-question routing accuracy",0.9433,"rate","94% (183/194)",194,[0.9013,0.968],"\u0001"],["haiku-thinking-mcnemar","Paired exact test, thinking off vs on (exact McNemar p)",0.754,"score","0.754",82,"\u0001","p = 0.754 (6 only on, 4 only off, 82 cases)"],["haiku-thinking-wall-off","Claude Haiku 4.5 (thinking off): median wall time per routing decision",4.66,"seconds","4.66 s",82,"\u0001","Range 2.20 s to 11.57 s; not a confidence interval. Nearest-rank p50."],["haiku-thinking-wall-on","Claude Haiku 4.5 (thinking on): median wall time per routing decision",12.54,"seconds","12.54 s",82,"\u0001","Range 5.86 s to 51.28 s; not a confidence interval. Nearest-rank p50."],["haiku-thinking-wall-ratio","Median wall time, thinking on ÷ thinking off (calculation)",2.69,"ratio","2.7x (12.54 s ÷ 4.66 s)",82,"\u0001","Calculation from two medians; the arms did not run at the same time."],["haiku-thinking-tokens-off","Claude Haiku 4.5 (thinking off): thinking tokens per routing decision (mean)",0,"tokens","0",82,"\u0001","\u0001"],["haiku-thinking-tokens-on","Claude Haiku 4.5 (thinking on): thinking tokens per routing decision (mean)",1100.8,"tokens","1,101 (range 285 to 4,443)",82,"\u0001","\u0001"],["haiku-thinking-cost-off","Claude Haiku 4.5 (thinking off): list-price cost per 1,000 routing decisions (calculation)",3.364,"usd","$3.364",82,"\u0001","\u0001"],["haiku-thinking-cost-on","Claude Haiku 4.5 (thinking on): list-price cost per 1,000 routing decisions (calculation)",8.924,"usd","$8.924",82,"\u0001","\u0001"],["haiku-thinking-hard-strict-off","Claude Haiku 4.5 (thinking off): strict passes on hard tasks",0.1667,"rate","17% (4/24)",24,[0.0668,0.3585],"\u0001"],["haiku-thinking-hard-strict-on","Claude Haiku 4.5 (thinking on): strict passes on hard tasks",0.4583,"rate","46% (11/24)",24,[0.2789,0.6493],"\u0001"],["haiku-thinking-hard-time-off","Claude Haiku 4.5 (thinking off): median total time per hard-task call",2.95,"seconds","2.9 s",24,"\u0001","Range 1.7 s to 13.0 s; not a confidence interval."],["haiku-thinking-hard-time-on","Claude Haiku 4.5 (thinking on): median total time per hard-task call",39.01,"seconds","39.0 s",24,"\u0001","Range 15.3 s to 75.1 s; not a confidence interval."],["haiku-thinking-hard-reasoning-on","Claude Haiku 4.5 (thinking on): median reasoning tokens per hard-task call",4556,"tokens","4,556",24,"\u0001","\u0001"],["haiku-thinking-hard-cost-per-pass-off","Claude Haiku 4.5 (thinking off): list-price cost per strict pass on hard tasks (calculation)",0.03654,"usd","$0.0365",24,"\u0001","\u0001"],["haiku-thinking-hard-cost-per-pass-on","Claude Haiku 4.5 (thinking on): list-price cost per strict pass on hard tasks (calculation)",0.0672,"usd","$0.0672",24,"\u0001","\u0001"],["haiku-thinking-new-calls","Counted model calls made for this study (thinking off)",106,"calls","106 (82 routing, 24 hard tasks; 2 more uncounted probes)",106,"\u0001","\u0001"],["haiku-thinking-off-check","Thinking-off calls that reported any thinking tokens",0,"count","0 of 106",106,"\u0001","Each call reports its thinking tokens in the CLI result; 0 means the setting held."]]},"charts":{"$k":["id","title","subtitle","kind","unit","polarity","yLabel","series","note","whisker","factContext","sourceIds"],"$r":[["haiku-thinking-router-exact","Haiku thinking study: typed routing decisions answered exactly right","Claude Haiku 4.5 and Claude Sonnet 5.5 (low effort); 82 decisions, the same cases for every arm","dot-range","rate","higher","Correct",[{"name":"Exact decisions (every scored question right)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",0.8659,0.7755,0.9234,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5 (low) · Claude Code",0.939,0.8651,0.9737,82,false]]}},{"name":"Per-question accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",0.9124,0.8642,0.9446,194,true],["Claude Haiku 4.5 (thinking on) · Claude Code",0.9433,0.9013,0.968,194,false],["Claude Sonnet 5.5 (low) · Claude Code",0.9742,0.9411,0.9889,194,false]]}}],"Whiskers are 95% Wilson intervals. Questions within a decision are related; per-question intervals are descriptive, not an independent-question test. The thinking-on and Sonnet arms are the recorded 2026-10-05 routing run, reused, not rerun; the thinking-off arm ran later on another account. An unanswered question counts as wrong.","ci95","typed routing decisions, thinking on vs off",["agent-haiku-thinking","agent-routing"]],["haiku-thinking-router-latency","Haiku thinking study: time per routing decision","Median wall time and model (API) time; whisker to the 95th percentile","dot-range","seconds","\u0001","Seconds",[{"name":"Wall time (CLI)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",4.66,4.66,8.18,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",12.54,12.54,34.48,82,false],["Claude Sonnet 5.5 (low) · Claude Code",2.6,2.6,4.3,82,false]]}},{"name":"Model time (API)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",3.79,3.79,7.43,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",10.51,10.51,32.13,82,false],["Claude Sonnet 5.5 (low) · Claude Code",1.6,1.6,2.58,82,false]]}}],"Whiskers run from p50 to p95, not a confidence interval. Percentiles use the nearest-rank rule of the routing-overhead study, so the thinking-on and Sonnet medians match that study (the Jev-vs-LLM page interpolates between ranks and shows slightly different values). One call at a time through the Claude Code CLI; the thinking-on and Sonnet arms ran on another day.","p50-p95","typed routing decisions, thinking on vs off",["agent-haiku-thinking","agent-routing"]],["haiku-thinking-router-tokens","Haiku thinking study: thinking and visible output tokens per routing decision","Mean per decision, as the Claude Code CLI reports them","grouped-bar","tokens","\u0001","Tokens per decision",[{"name":"Thinking tokens","points":{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",0,82],["Claude Haiku 4.5 (thinking on) · Claude Code",1101,82],["Claude Sonnet 5.5 (low) · Claude Code",2,82]]}},{"name":"Visible output tokens","points":{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",366,82],["Claude Haiku 4.5 (thinking on) · Claude Code",318,82],["Claude Sonnet 5.5 (low) · Claude Code",105,82]]}}],"Visible output = output tokens minus thinking tokens. It includes the structured answer the CLI asks for. Thinking tokens are counted by the CLI; their content is never captured. More tokens is not better or worse by itself.","\u0001","typed routing decisions, thinking on vs off",["agent-haiku-thinking","agent-routing"]],["haiku-thinking-router-cost","Haiku thinking study: list-price cost per 1,000 routing decisions (calculation)","Reported tokens × list price, per 1,000 decisions","bar","usd","\u0001","USD per 1,000 decisions",[{"name":"Cost per 1,000 decisions","points":{"$k":["label","value","n","highlight"],"$r":[["Claude Haiku 4.5 (thinking off) · Claude Code",3.364,82,true],["Claude Haiku 4.5 (thinking on) · Claude Code",8.924,82,false],["Claude Sonnet 5.5 (low) · Claude Code",7.324,82,false]]}}],"Calculation, not a bill: the calls ran on a flat subscription. Reported input, cache and output tokens (thinking tokens are part of output) × list price, with the price table of the Jev-vs-LLM study. Sonnet cache writes use the one-hour rate ($4 per million tokens), as recorded in its receipts. The calculation matches the CLI-reported cost.","\u0001","typed routing decisions, thinking on vs off",["agent-haiku-thinking","agent-routing","calc-repricing","price-anthropic"]],["haiku-thinking-hard-pass","Haiku thinking study: pass rate on eight hard tasks","Claude Haiku 4.5 in Claude Code; strict and lenient","dot-range","rate","higher","Passed",[{"name":"Strict pass","points":[{"label":"Claude Haiku 4.5 (thinking off) · Claude Code","value":0.1667,"lo":0.0668,"hi":0.3585,"n":24,"highlight":true},{"label":"Claude Haiku 4.5 (thinking on) · Claude Code","value":0.4583,"lo":0.2789,"hi":0.6493,"n":24,"highlight":false}]},{"name":"Lenient (format misses counted)","points":[{"label":"Claude Haiku 4.5 (thinking off) · Claude Code","value":0.1667,"lo":0.0668,"hi":0.3585,"n":24,"highlight":true},{"label":"Claude Haiku 4.5 (thinking on) · Claude Code","value":0.6667,"lo":0.4671,"hi":0.8203,"n":24,"highlight":false}]}],"Whiskers are 95% Wilson intervals over calls. Three repeats per task are related, so these intervals do not measure uncertainty across unseen tasks. A format miss (right answer in the wrong wrapping) never counts as a strict pass. The thinking-on cell is the 24 recorded Haiku receipts of the hard head-to-head, reused; the thinking-off cell ran later on another account.","ci95","eight hard validated tasks, thinking on vs off",["agent-haiku-thinking","agent-provider-h2h-hard"]],["haiku-thinking-hard-time","Haiku thinking study: total time per call on hard tasks","Median per cell; whiskers = fastest and slowest call","dot-range","seconds","\u0001","Seconds",[{"name":"Total time per call","points":[{"label":"Claude Haiku 4.5 (thinking off) · Claude Code","value":2.95,"lo":1.7,"hi":13.01,"n":24,"highlight":true},{"label":"Claude Haiku 4.5 (thinking on) · Claude Code","value":39.01,"lo":15.27,"hi":75.13,"n":24,"highlight":false}]}],"Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network; the thinking-on calls ran in an earlier session. Timings include CLI start-up.","minmax","eight hard validated tasks, thinking on vs off",["agent-haiku-thinking","agent-provider-h2h-hard"]]]},"related":["routing-jev-vs-llm","hard-model-head-to-head","routing-overhead"]}}