{"i":22,"study":{"slug":"thinking-token-bill","title":"How much of an AI bill is thinking? Reasoning tokens by model and effort","seoTitle":"Thinking tokens by model and effort: share and cost","description":"Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.","question":"Across 378 recorded calls of Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol, how many output tokens come from reasoning (thinking)? What do they cost at list price per call and per strict pass, and do they track time? The calls cover eight hard tasks, five short tasks and several efforts.","answer":"On eight hard tasks, reasoning tokens were 46% to 92% of the output tokens of a median call (16 to 24 calls per configuration). Haiku 4.5 had the highest median (92%), GPT-6.1 Sol (medium) had the lowest (46%) and the other 5 sat between 54% and 64%. Per-call ranges are wide (Sonnet 5.5 0% to 96%). The ranges of every pair overlap, so this run ranks no configuration. Output share is not bill share. At list price (a calculation) the reasoning part of a call cost a mean $0.0023 (GPT-6.1 Sol (medium)) to $0.0537 (Fable 5.1). That was 9% (GPT-6.1 Sol (medium)) to 80% (Haiku 4.5) of the total list-price cost in each configuration. This pooled share divides summed reasoning cost by summed total cost. Input tokens cost money too. The calls ran on flat subscriptions, so this is not a bill. Fable 5.1's thinking cost 8.1x Sonnet 5.5's per call. The output price explains a factor of 5.0 ($50 against $10 per million tokens). More reasoning tokens explain the rest, a factor of 1.6 (means). Higher effort settings had higher mean recorded reasoning counts in these batches. We paired the same eight tasks. At high effort, mean reasoning tokens per call were higher than at low effort on these tasks: Sonnet 5.5 7 of 8, Opus 5.5 8 of 8, GPT-6.1 Sol 8 of 8. The pooled reasoning share rose from low to high effort. Sonnet 5.5 53% to 73%. Opus 5.5 38% to 69%. GPT-6.1 Sol 29% to 59%. It rose at each step (low, medium, high) in all 3 ladders. The reasoning cost per call rose 2.2x for Sonnet 5.5, 3.6x for Opus 5.5 and 3.4x for GPT-6.1 Sol (a calculation). Every effort cell passed 16/16 strictly. The pass count stayed the same on this set, which has a ceiling (95% Wilson 81% to 100% per cell). Recorded reasoning counts correlated with total time. Within each Claude model, the rank correlation between reasoning tokens and total time was 0.85 to 0.98 (Spearman, 24 to 80 calls each). On the same task, 1,000 more reasoning tokens went with 7.6 to 14.5 s more time (a calculation). For GPT-6.1 Sol in the Codex CLI, the recorded rank correlation was: 0.56 (48 calls). Thinking costs money whether or not the call passes. Haiku 4.5 passed 11/24 strictly (95% Wilson 28% to 65%). The 13 calls that did not pass held 61% of its reasoning cost (a calculation).","date":"2026-10-06","updated":"2026-10-06","tags":["thought-experiment","calculation","reasoning-tokens","thinking-tokens","effort","llm-pricing","claude-haiku","claude-sonnet","claude-opus","claude-fable","gpt-6-1-sol","claude-code","codex-cli","latency"],"caveats":["The hard-set protocol file was created after its first counted call. Both effort-ladder protocol files were created after their batches ended. The short-set file predates its first call, but its top-up amendment timing is unverified. Batch receipts preserve protocol text, but we cannot verify all rules were written before inference. Treat these as exploratory calculations, not preregistered tests.","The calls used one shared Mac and network. Other work and provider load were not controlled. Latency includes those effects.","These are hand-built, tuned case sets, not random workload samples. The short set and most hard-set configurations hit a pass ceiling. Repeats of the same tasks are not independent. Wilson pass intervals describe counted calls under an independence assumption, not performance on new tasks.","Each CLI reports its own reasoning counter, and we never see the reasoning text. So we cannot check what each vendor counts. The sum check shows only that the counters are consistent with inclusion in output; it does not prove their semantics. Shares from Claude Code and Codex CLI do not measure like for like how much each model thinks.","Every prompt asks for a short reply in a strict format (code, JSON, one line or a regex, with no explanation). So the visible answer is short and the reasoning share is high. Longer visible replies could change the share. We did not measure that workload.","Small cells: 16 to 24 calls per configuration over 8 tasks. Per-call ranges are wide and overlap for every pair of configurations. So the medians describe this run and rank nothing. A sentence says one side is ahead only when its ranges do not overlap.","Input includes CLI context that these receipts do not separately count. It moves with cache hits. So the part of the call cost that goes to reasoning depends on the CLI and on the cache, not only on the model. GPT-6.1 Sol at medium and at high effort ran in different batches and show very different input cost per call.","Every effort-ladder cell passed 16/16, so the set has a ceiling. Higher effort had higher mean recorded reasoning cost in these batches. It does not show that more thinking never helps on harder work.","The time link is a correlation, not a cause. Task, effort and batch can affect reasoning, visible output and time together. Total time includes CLI start-up. The within-task slope controls for differences in task means. It does not remove effort or batch effects. Output tokens (reasoning plus visible answer) track time at least as closely as reasoning alone. The Codex CLI slope (36.4 s per 1,000 reasoning tokens) has a calculated ratio of 2.5 to 4.8 times the Claude Code slopes. We did not test why. Calls are repeats of 8 tasks, so they are not independent. The ranges omit one whole task at a time. They show sensitivity to the task mix, not 95% coverage.","Effort levels are not the same scale across vendors. \"Default\" means we did not pass the effort flag, and the CLI chose. The effort-ladder reference cells ran in a different batch and hour than the new cells.","List-price costs are calculations, because the calls used flat subscriptions. A price change moves every cost here. It leaves every token count unchanged."],"sourceIds":["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"],"hero":{"statIds":["thinking-bill-share-highest","thinking-bill-share-lowest"]},"stats":{"$k":["id","label","value","unit","display","n","note"],"$r":[["thinking-bill-calls","Recorded calls analysed (no new calls)",378,"count","378 (152 hard head-to-head, 96 new effort-ladder, 130 five-task head-to-head)",378,"\u0001"],["thinking-bill-zero-reasoning","Calls that reported 0 reasoning tokens (kept as 0, not \"not reported\")",98,"count","98 of 378; 0 calls had no usable count",378,"\u0001"],["thinking-bill-sum-claude","Sum check, Claude Code: output tokens gained per extra reasoning token, same task and model (calculation)",0.993,"ratio","0.99 (leave-one-task-out range 0.98 to 1.01; 290 calls)",290,"About 1 is consistent with reasoning inside output. This correlation does not prove how a CLI counts tokens."],["thinking-bill-sum-codex","Sum check, Codex CLI: output tokens gained per extra reasoning token, same task and model (calculation)",0.967,"ratio","0.97 (leave-one-task-out range 0.97 to 0.99; 88 calls)",88,"About 1 is consistent with reasoning inside output. This correlation does not prove how a CLI counts tokens."],["thinking-bill-share-highest","Highest median reasoning share of output tokens, hard tasks (calculation)",0.9168,"rate","92% (Claude Haiku 4.5 · Claude Code; range 76% to 99%)",24,"\u0001"],["thinking-bill-share-lowest","Lowest median reasoning share of output tokens, hard tasks (calculation)",0.4633,"rate","46% (GPT-6.1 Sol (medium) · Codex CLI; range 11% to 87%)",16,"\u0001"],["thinking-bill-reasoning-cost-highest","Highest mean list-price cost of reasoning per call, hard tasks (calculation)",0.053696,"usd","$0.0537 (Claude Fable 5.1 · Claude Code; call range $0.0037 to $0.2944)",24,"\u0001"],["thinking-bill-reasoning-cost-lowest","Lowest mean list-price cost of reasoning per call, hard tasks (calculation)",0.002273,"usd","$0.0023 (GPT-6.1 Sol (medium) · Codex CLI; call range $0.0006 to $0.0084)",16,"\u0001"],["thinking-bill-cost-share-highest","Highest pooled reasoning share of total list-price cost, hard tasks (calculation)",0.7953,"rate","80% (Claude Haiku 4.5 · Claude Code; call range 40% to 90%)",24,"\u0001"],["thinking-bill-cost-share-lowest","Lowest pooled reasoning share of total list-price cost, hard tasks (calculation)",0.0887,"rate","9% (GPT-6.1 Sol (medium) · Codex CLI; call range 2% to 20%)",16,"\u0001"],["thinking-bill-fable-vs-sonnet","Reasoning cost per call, Fable 5.1 ÷ Sonnet 5.5, hard tasks (calculation, means)",8.056,"ratio","8.1x ($0.0537 vs $0.0067)",24,"\u0001"],["thinking-bill-haiku-failed-share","Share of Haiku 4.5 reasoning cost spent on calls that did not pass (calculation)",0.6114,"rate","61% (13 of 24 calls did not pass strictly)",24,"\u0001"],["thinking-bill-effort-tasks-sonnet-5-5","Tasks where mean reasoning tokens per call were higher at high than at low effort, Claude Sonnet 5.5 · Claude Code (paired, same tasks; calculation)",7,"count","7 of 8 tasks",8,"Each task ran 2 times at each effort. Eight tasks is a small sample; this counts tasks, it is not an interval."],["thinking-bill-effort-tasks-opus-5-5","Tasks where mean reasoning tokens per call were higher at high than at low effort, Claude Opus 5.5 · Claude Code (paired, same tasks; calculation)",8,"count","8 of 8 tasks",8,"Each task ran 2 times at each effort. Eight tasks is a small sample; this counts tasks, it is not an interval."],["thinking-bill-effort-tasks-gpt-6-1-sol","Tasks where mean reasoning tokens per call were higher at high than at low effort, GPT-6.1 Sol · Codex CLI (paired, same tasks; calculation)",8,"count","8 of 8 tasks",8,"Each task ran 2 times at each effort. Eight tasks is a small sample; this counts tasks, it is not an interval."],["thinking-bill-effort-ratio-sonnet-5-5","Reasoning cost per call, high ÷ low effort, Claude Sonnet 5.5 · Claude Code (calculation, means)",2.17,"ratio","2.2x ($0.0043 at low, $0.0094 at high)",16,"\u0001"],["thinking-bill-effort-ratio-opus-5-5","Reasoning cost per call, high ÷ low effort, Claude Opus 5.5 · Claude Code (calculation, means)",3.584,"ratio","3.6x ($0.0050 at low, $0.0180 at high)",16,"\u0001"],["thinking-bill-effort-ratio-gpt-6-1-sol","Reasoning cost per call, high ÷ low effort, GPT-6.1 Sol · Codex CLI (calculation, means)",3.366,"ratio","3.4x ($0.0012 at low, $0.0041 at high)",16,"\u0001"],["thinking-bill-time-rho-claude","Spearman, reasoning tokens vs total time, per Claude model: lowest (calculation)",0.852,"score","0.85 to 0.98 across 4 Claude models",200,"The value is the lowest of the per-model correlations; the display gives the range."],["thinking-bill-time-slope-claude","Seconds per 1,000 reasoning tokens on the same task, per Claude model: lowest (calculation)",7.649,"seconds","7.6 to 14.5 s across 4 Claude models",200,"The value is the lowest of the per-model slopes; the display gives the range."],["thinking-bill-time-rho-codex","Spearman, reasoning tokens vs total time, GPT-6.1 Sol in Codex CLI (calculation)",0.556,"score","0.56 (leave-one-task-out range 0.34 to 0.65; 48 calls)",48,"The range omits one whole task at a time. It is a sensitivity check, not a confidence interval."]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","whisker","polarity","sourceIds","xLabel"],"$r":[["thinking-bill-share","Reasoning share of output tokens per call on hard tasks (calculation)","Median call: reasoning tokens ÷ output tokens. Whiskers: lowest and highest call (16 to 24 calls per configuration)","bar","percent","Reasoning share of output tokens (%)",[{"name":"Median call","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",91.68,76.46,99.27,24],["Claude Fable 5.1 · Claude Code",64.24,23.44,97.19,24],["GPT-6.1 Sol (high) · Codex CLI",57.01,29.19,90.8,16],["Claude Opus 5.5 · Claude Code",54.79,29.92,95.6,24],["Claude Sonnet 5.5 · Claude Code",54.54,0,95.91,24],["Claude Opus 5.5 (high) · Claude Code",54.43,36.14,96.23,24],["GPT-6.1 Sol (medium) · Codex CLI",46.33,11.42,86.85,16]]}}],"Calculation from reported tokens, not a run. Each call gives reasoning ÷ output; the bar is the median of those shares. Whiskers are the lowest and highest call. They are a range, not a confidence interval. They are wide, so the medians describe this run and rank nothing. The pooled share (all reasoning tokens ÷ all output tokens) is in the table. We treat reasoning tokens as part of output tokens; the consistency check supports this accounting assumption. Each CLI reports its own count.","minmax","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"],"\u0001"],["thinking-bill-cost-per-call","List-price cost per call: reasoning, remaining output and input (calculation)","Mean per call on the hard tasks; the three parts add up to the call","stacked-bar","usd","USD per call (list price)",{"$k":["name","points"],"$r":[["Reasoning (output tokens)",{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0.053696,24],["Claude Opus 5.5 (high) · Claude Code",0.017969,24],["Claude Haiku 4.5 · Claude Code",0.024492,24],["Claude Opus 5.5 · Claude Code",0.012528,24],["GPT-6.1 Sol (medium) · Codex CLI",0.002273,16],["GPT-6.1 Sol (high) · Codex CLI",0.004114,16],["Claude Sonnet 5.5 · Claude Code",0.006665,24]]}],["Remaining output (visible-answer estimate)",{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0.018702,24],["Claude Opus 5.5 (high) · Claude Code",0.00799,24],["Claude Haiku 4.5 · Claude Code",0.001795,24],["Claude Opus 5.5 · Claude Code",0.008003,24],["GPT-6.1 Sol (medium) · Codex CLI",0.003025,16],["GPT-6.1 Sol (high) · Codex CLI",0.002905,16],["Claude Sonnet 5.5 · Claude Code",0.003672,24]]}],["Input (prompt, cache priced)",{"$k":["label","value","n"],"$r":[["Claude Fable 5.1 · Claude Code",0.02091,24],["Claude Opus 5.5 (high) · Claude Code",0.007407,24],["Claude Haiku 4.5 · Claude Code",0.00451,24],["Claude Opus 5.5 · Claude Code",0.00771,24],["GPT-6.1 Sol (medium) · Codex CLI",0.020339,16],["GPT-6.1 Sol (high) · Codex CLI",0.008117,16],["Claude Sonnet 5.5 · Claude Code",0.004012,24]]}]]},"Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.","\u0001","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"],"\u0001"],["thinking-bill-by-effort","Reasoning cost per strict pass by effort, with the total (calculation)","List price ÷ strict passes; every cell is 8 tasks × 2 repetitions","grouped-bar","usd","USD per strict pass (list price)",[{"name":"Reasoning cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",0.004317,16],["Claude Sonnet 5.5 (medium) · Claude Code",0.005946,16],["Claude Sonnet 5.5 (high) · Claude Code",0.009369,16],["Claude Sonnet 5.5 · Claude Code",0.006299,16],["Claude Opus 5.5 (low) · Claude Code",0.005031,16],["Claude Opus 5.5 (medium) · Claude Code",0.01344,16],["Claude Opus 5.5 (high) · Claude Code",0.018034,16],["Claude Opus 5.5 · Claude Code",0.013104,16],["GPT-6.1 Sol (low) · Codex CLI",0.001223,16],["GPT-6.1 Sol (medium) · Codex CLI",0.002273,16],["GPT-6.1 Sol (high) · Codex CLI",0.004114,16]]}},{"name":"Total cost per strict pass","points":{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 (low) · Claude Code",0.012191,16],["Claude Sonnet 5.5 (medium) · Claude Code",0.01352,16],["Claude Sonnet 5.5 (high) · Claude Code",0.016705,16],["Claude Sonnet 5.5 · Claude Code",0.013978,16],["Claude Opus 5.5 (low) · Claude Code",0.021152,16],["Claude Opus 5.5 (medium) · Claude Code",0.029475,16],["Claude Opus 5.5 (high) · Claude Code",0.033677,16],["Claude Opus 5.5 · Claude Code",0.028925,16],["GPT-6.1 Sol (low) · Codex CLI",0.012837,16],["GPT-6.1 Sol (medium) · Codex CLI",0.025637,16],["GPT-6.1 Sol (high) · Codex CLI",0.015137,16]]}}],"Calculation, not a bill: reported tokens × list price, divided by the cell's strict passes; the calls ran on flat subscriptions. The effort-ladder cells: new calls plus reference cells reused from the hard head-to-head (Claude repetitions 1-2 only). \"Default\" means the effort flag was not passed. The total is the same value as the effort-ladder cost-per-pass chart. Effort levels are not the same scale across vendors, and the reference cells ran in a different batch and hour.","\u0001","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"],"\u0001"],["thinking-bill-vs-time","Reasoning tokens vs total time per call (calculation)","One point per call: 248 calls from the hard head-to-head and the effort ladder","scatter","seconds","Total time per call (seconds)",{"$k":["name","points"],"$r":[["Claude Haiku 4.5 · Claude Code",{"$k":["label","x","value"],"$r":[["Claude Haiku 4.5 · Claude Code · merge-ranges · rep 1",1977,18.53],["Claude Haiku 4.5 · Claude Code · day-hours · rep 1",4372,39],["Claude Haiku 4.5 · Claude Code · csv-parse · rep 1",3590,33.81],["Claude Haiku 4.5 · Claude Code · event-loop-order · rep 1",3781,25.96],["Claude Haiku 4.5 · Claude Code · talk-schedule · rep 1",6236,54.64],["Claude Haiku 4.5 · Claude Code · semver-regex · rep 1",7498,64.48],["Claude Haiku 4.5 · Claude Code · money-refactor · rep 1",1452,15.91],["Claude Haiku 4.5 · Claude Code · q1-sql · rep 1",4305,38.89],["Claude Haiku 4.5 · Claude Code · merge-ranges · rep 2",2606,24.71],["Claude Haiku 4.5 · Claude Code · day-hours · rep 2",3515,26.15],["Claude Haiku 4.5 · Claude Code · csv-parse · rep 2",4497,39.01],["Claude Haiku 4.5 · Claude Code · event-loop-order · rep 2",6575,54.02],["Claude Haiku 4.5 · Claude Code · talk-schedule · rep 2",7954,68.77],["Claude Haiku 4.5 · Claude Code · semver-regex · rep 2",6538,54.49],["Claude Haiku 4.5 · Claude Code · money-refactor · rep 2",1955,15.27],["Claude Haiku 4.5 · Claude Code · q1-sql · rep 2",6691,56.13],["Claude Haiku 4.5 · Claude Code · merge-ranges · rep 3",2703,21.22],["Claude Haiku 4.5 · Claude Code · day-hours · rep 3",4614,40.56],["Claude Haiku 4.5 · Claude Code · csv-parse · rep 3",8569,75.13],["Claude Haiku 4.5 · Claude Code · event-loop-order · rep 3",7065,56.65],["Claude Haiku 4.5 · Claude Code · talk-schedule · rep 3",4922,37.97],["Claude Haiku 4.5 · Claude Code · semver-regex · rep 3",7994,67.03],["Claude Haiku 4.5 · Claude Code · money-refactor · rep 3",2221,17.16],["Claude Haiku 4.5 · Claude Code · q1-sql · rep 3",5934,47.11]]}],["Claude Sonnet 5.5 · Claude Code",{"$k":["label","x","value"],"$r":[["Claude Sonnet 5.5 · Claude Code · merge-ranges · rep 1",0,2.93],["Claude Sonnet 5.5 · Claude Code · day-hours · rep 1",1249,21.61],["Claude Sonnet 5.5 · Claude Code · csv-parse · rep 1",551,8.85],["Claude Sonnet 5.5 · Claude Code · event-loop-order · rep 1",1064,8.17],["Claude Sonnet 5.5 · Claude Code · talk-schedule · rep 1",716,7.36],["Claude Sonnet 5.5 · Claude Code · semver-regex · rep 1",0,2.26],["Claude Sonnet 5.5 · Claude Code · money-refactor · rep 1",0,3.57],["Claude Sonnet 5.5 · Claude Code · q1-sql · rep 1",799,8.91],["Claude Sonnet 5.5 · Claude Code · merge-ranges · rep 2",0,3.06],["Claude Sonnet 5.5 · Claude Code · day-hours · rep 2",1614,19.62],["Claude Sonnet 5.5 · Claude Code · csv-parse · rep 2",722,9.56],["Claude Sonnet 5.5 · Claude Code · event-loop-order · rep 2",1197,9.87],["Claude Sonnet 5.5 · Claude Code · talk-schedule · rep 2",736,7.76],["Claude Sonnet 5.5 · Claude Code · semver-regex · rep 2",267,3.67],["Claude Sonnet 5.5 · Claude Code · money-refactor · rep 2",545,7.18],["Claude Sonnet 5.5 · Claude Code · q1-sql · rep 2",619,8.47],["Claude Sonnet 5.5 · Claude Code · merge-ranges · rep 3",0,2.38],["Claude Sonnet 5.5 · Claude Code · day-hours · rep 3",3060,34.79],["Claude Sonnet 5.5 · Claude Code · csv-parse · rep 3",547,10.61],["Claude Sonnet 5.5 · Claude Code · event-loop-order · rep 3",1177,9.57],["Claude Sonnet 5.5 · Claude Code · talk-schedule · rep 3",657,7.74],["Claude Sonnet 5.5 · Claude Code · semver-regex · rep 3",0,2.49],["Claude Sonnet 5.5 · Claude Code · money-refactor · rep 3",0,3.44],["Claude Sonnet 5.5 · Claude Code · q1-sql · rep 3",477,7.56],["Claude Sonnet 5.5 (low) · Claude Code · merge-ranges · rep 1",0,2.79],["Claude Sonnet 5.5 (medium) · Claude Code · merge-ranges · rep 1",0,2.71],["Claude Sonnet 5.5 (high) · Claude Code · merge-ranges · rep 1",0,2.93],["Claude Sonnet 5.5 (low) · Claude Code · day-hours · rep 1",1489,19.96],["Claude Sonnet 5.5 (medium) · Claude Code · day-hours · rep 1",1941,22.48],["Claude Sonnet 5.5 (high) · Claude Code · day-hours · rep 1",2597,27.18],["Claude Sonnet 5.5 (low) · Claude Code · csv-parse · rep 1",0,4.32],["Claude Sonnet 5.5 (medium) · Claude Code · csv-parse · rep 1",423,7.83],["Claude Sonnet 5.5 (high) · Claude Code · csv-parse · rep 1",770,10.01],["Claude Sonnet 5.5 (low) · Claude Code · event-loop-order · rep 1",995,8.8],["Claude Sonnet 5.5 (medium) · Claude Code · event-loop-order · rep 1",1200,9.98],["Claude Sonnet 5.5 (high) · Claude Code · event-loop-order · rep 1",1295,10.89],["Claude Sonnet 5.5 (low) · Claude Code · talk-schedule · rep 1",598,6.49],["Claude Sonnet 5.5 (medium) · Claude Code · talk-schedule · rep 1",651,7.9],["Claude Sonnet 5.5 (high) · Claude Code · talk-schedule · rep 1",734,9.07],["Claude Sonnet 5.5 (low) · Claude Code · semver-regex · rep 1",0,3.51],["Claude Sonnet 5.5 (medium) · Claude Code · semver-regex · rep 1",251,4.12],["Claude Sonnet 5.5 (high) · Claude Code · semver-regex · rep 1",252,4.01],["Claude Sonnet 5.5 (low) · Claude Code · money-refactor · rep 1",0,3.72],["Claude Sonnet 5.5 (medium) · Claude Code · money-refactor · rep 1",0,3.89],["Claude Sonnet 5.5 (high) · Claude Code · money-refactor · rep 1",693,8.09],["Claude Sonnet 5.5 (low) · Claude Code · q1-sql · rep 1",545,7.92],["Claude Sonnet 5.5 (medium) · Claude Code · q1-sql · rep 1",421,7.43],["Claude Sonnet 5.5 (high) · Claude Code · q1-sql · rep 1",826,10.79],["Claude Sonnet 5.5 (low) · Claude Code · merge-ranges · rep 2",0,2.78],["Claude Sonnet 5.5 (medium) · Claude Code · merge-ranges · rep 2",0,2.93],["Claude Sonnet 5.5 (high) · Claude Code · merge-ranges · rep 2",0,3.49],["Claude Sonnet 5.5 (low) · Claude Code · day-hours · rep 2",1091,16.28],["Claude Sonnet 5.5 (medium) · Claude Code · day-hours · rep 2",2093,24.01],["Claude Sonnet 5.5 (high) · Claude Code · day-hours · rep 2",3610,35.81],["Claude Sonnet 5.5 (low) · Claude Code · csv-parse · rep 2",0,4.4],["Claude Sonnet 5.5 (medium) · Claude Code · csv-parse · rep 2",0,4.65],["Claude Sonnet 5.5 (high) · Claude Code · csv-parse · rep 2",783,16.1],["Claude Sonnet 5.5 (low) · Claude Code · event-loop-order · rep 2",974,9.73],["Claude Sonnet 5.5 (medium) · Claude Code · event-loop-order · rep 2",1098,8.72],["Claude Sonnet 5.5 (high) · Claude Code · event-loop-order · rep 2",1226,10.77],["Claude Sonnet 5.5 (low) · Claude Code · talk-schedule · rep 2",624,6.71],["Claude Sonnet 5.5 (medium) · Claude Code · talk-schedule · rep 2",624,9.61],["Claude Sonnet 5.5 (high) · Claude Code · talk-schedule · rep 2",755,8.1],["Claude Sonnet 5.5 (low) · Claude Code · semver-regex · rep 2",0,2.86],["Claude Sonnet 5.5 (medium) · Claude Code · semver-regex · rep 2",245,3.82],["Claude Sonnet 5.5 (high) · Claude Code · semver-regex · rep 2",243,3.71],["Claude Sonnet 5.5 (low) · Claude Code · money-refactor · rep 2",0,5.16],["Claude Sonnet 5.5 (medium) · Claude Code · money-refactor · rep 2",0,3.99],["Claude Sonnet 5.5 (high) · Claude Code · money-refactor · rep 2",580,7.61],["Claude Sonnet 5.5 (low) · Claude Code · q1-sql · rep 2",591,7.68],["Claude Sonnet 5.5 (medium) · Claude Code · q1-sql · rep 2",566,9.18],["Claude Sonnet 5.5 (high) · Claude Code · q1-sql · rep 2",626,8.55]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","x","value"],"$r":[["Claude Opus 5.5 · Claude Code · merge-ranges · rep 1",100,4.24],["Claude Opus 5.5 (high) · Claude Code · merge-ranges · rep 1",188,5.36],["Claude Opus 5.5 · Claude Code · day-hours · rep 1",1580,26.3],["Claude Opus 5.5 (high) · Claude Code · day-hours · rep 1",3301,63],["Claude Opus 5.5 · Claude Code · csv-parse · rep 1",543,10.98],["Claude Opus 5.5 (high) · Claude Code · csv-parse · rep 1",528,11.1],["Claude Opus 5.5 · Claude Code · event-loop-order · rep 1",958,11.91],["Claude Opus 5.5 (high) · Claude Code · event-loop-order · rep 1",1274,13.6],["Claude Opus 5.5 · Claude Code · talk-schedule · rep 1",574,7.82],["Claude Opus 5.5 (high) · Claude Code · talk-schedule · rep 1",656,8.9],["Claude Opus 5.5 · Claude Code · semver-regex · rep 1",300,5.35],["Claude Opus 5.5 (high) · Claude Code · semver-regex · rep 1",293,5],["Claude Opus 5.5 · Claude Code · money-refactor · rep 1",429,8.15],["Claude Opus 5.5 (high) · Claude Code · money-refactor · rep 1",528,9.12],["Claude Opus 5.5 · Claude Code · q1-sql · rep 1",796,12.78],["Claude Opus 5.5 (high) · Claude Code · q1-sql · rep 1",920,13.66],["Claude Opus 5.5 · Claude Code · merge-ranges · rep 2",114,4.65],["Claude Opus 5.5 (high) · Claude Code · merge-ranges · rep 2",182,5.52],["Claude Opus 5.5 · Claude Code · day-hours · rep 2",1778,27.21],["Claude Opus 5.5 (high) · Claude Code · day-hours · rep 2",2756,36.44],["Claude Opus 5.5 · Claude Code · csv-parse · rep 2",524,11.47],["Claude Opus 5.5 (high) · Claude Code · csv-parse · rep 2",573,12.22],["Claude Opus 5.5 · Claude Code · event-loop-order · rep 2",1015,10.2],["Claude Opus 5.5 (high) · Claude Code · event-loop-order · rep 2",1300,13.45],["Claude Opus 5.5 · Claude Code · talk-schedule · rep 2",533,7.63],["Claude Opus 5.5 (high) · Claude Code · talk-schedule · rep 2",655,8.54],["Claude Opus 5.5 · Claude Code · semver-regex · rep 2",244,4.75],["Claude Opus 5.5 (high) · Claude Code · semver-regex · rep 2",103,3.63],["Claude Opus 5.5 · Claude Code · money-refactor · rep 2",214,7.2],["Claude Opus 5.5 (high) · Claude Code · money-refactor · rep 2",431,8.84],["Claude Opus 5.5 · Claude Code · q1-sql · rep 2",781,12.44],["Claude Opus 5.5 (high) · Claude Code · q1-sql · rep 2",739,12.46],["Claude Opus 5.5 · Claude Code · merge-ranges · rep 3",122,4.45],["Claude Opus 5.5 (high) · Claude Code · merge-ranges · rep 3",221,5.73],["Claude Opus 5.5 · Claude Code · day-hours · rep 3",1136,19.74],["Claude Opus 5.5 (high) · Claude Code · day-hours · rep 3",3027,38.79],["Claude Opus 5.5 · Claude Code · csv-parse · rep 3",440,11.67],["Claude Opus 5.5 (high) · Claude Code · csv-parse · rep 3",800,12.62],["Claude Opus 5.5 · Claude Code · event-loop-order · rep 3",1109,11.48],["Claude Opus 5.5 (high) · Claude Code · event-loop-order · rep 3",1198,13.5],["Claude Opus 5.5 · Claude Code · talk-schedule · rep 3",444,6.88],["Claude Opus 5.5 (high) · Claude Code · talk-schedule · rep 3",553,7.46],["Claude Opus 5.5 · Claude Code · semver-regex · rep 3",293,5.06],["Claude Opus 5.5 (high) · Claude Code · semver-regex · rep 3",122,3.98],["Claude Opus 5.5 · Claude Code · money-refactor · rep 3",199,6.19],["Claude Opus 5.5 (high) · Claude Code · money-refactor · rep 3",304,10.95],["Claude Opus 5.5 · Claude Code · q1-sql · rep 3",808,12.96],["Claude Opus 5.5 (high) · Claude Code · q1-sql · rep 3",911,12.91],["Claude Opus 5.5 (low) · Claude Code · merge-ranges · rep 1",0,3.34],["Claude Opus 5.5 (medium) · Claude Code · merge-ranges · rep 1",109,5.51],["Claude Opus 5.5 (low) · Claude Code · day-hours · rep 1",394,13.02],["Claude Opus 5.5 (medium) · Claude Code · day-hours · rep 1",1993,31.12],["Claude Opus 5.5 (low) · Claude Code · csv-parse · rep 1",471,9.88],["Claude Opus 5.5 (medium) · Claude Code · csv-parse · rep 1",444,18.26],["Claude Opus 5.5 (low) · Claude Code · event-loop-order · rep 1",684,8.7],["Claude Opus 5.5 (medium) · Claude Code · event-loop-order · rep 1",1120,12.42],["Claude Opus 5.5 (low) · Claude Code · talk-schedule · rep 1",498,7.72],["Claude Opus 5.5 (medium) · Claude Code · talk-schedule · rep 1",471,6.99],["Claude Opus 5.5 (low) · Claude Code · semver-regex · rep 1",88,4.21],["Claude Opus 5.5 (medium) · Claude Code · semver-regex · rep 1",119,6.8],["Claude Opus 5.5 (low) · Claude Code · money-refactor · rep 1",0,5.65],["Claude Opus 5.5 (medium) · Claude Code · money-refactor · rep 1",183,6.87],["Claude Opus 5.5 (low) · Claude Code · q1-sql · rep 1",0,10.64],["Claude Opus 5.5 (medium) · Claude Code · q1-sql · rep 1",793,13.13],["Claude Opus 5.5 (low) · Claude Code · merge-ranges · rep 2",0,3.44],["Claude Opus 5.5 (medium) · Claude Code · merge-ranges · rep 2",115,5.29],["Claude Opus 5.5 (low) · Claude Code · day-hours · rep 2",587,15.82],["Claude Opus 5.5 (medium) · Claude Code · day-hours · rep 2",1846,31.36],["Claude Opus 5.5 (low) · Claude Code · csv-parse · rep 2",0,6.21],["Claude Opus 5.5 (medium) · Claude Code · csv-parse · rep 2",655,10.91],["Claude Opus 5.5 (low) · Claude Code · event-loop-order · rep 2",791,9.16],["Claude Opus 5.5 (medium) · Claude Code · event-loop-order · rep 2",1014,11.19],["Claude Opus 5.5 (low) · Claude Code · talk-schedule · rep 2",426,7.29],["Claude Opus 5.5 (medium) · Claude Code · talk-schedule · rep 2",565,8.53],["Claude Opus 5.5 (low) · Claude Code · semver-regex · rep 2",86,3.79],["Claude Opus 5.5 (medium) · Claude Code · semver-regex · rep 2",243,4.78],["Claude Opus 5.5 (low) · Claude Code · money-refactor · rep 2",0,4.93],["Claude Opus 5.5 (medium) · Claude Code · money-refactor · rep 2",253,7.07],["Claude Opus 5.5 (low) · Claude Code · q1-sql · rep 2",0,9.23],["Claude Opus 5.5 (medium) · Claude Code · q1-sql · rep 2",829,13.8]]}],["Claude Fable 5.1 · Claude Code",{"$k":["label","x","value"],"$r":[["Claude Fable 5.1 · Claude Code · merge-ranges · rep 1",85,4.76],["Claude Fable 5.1 · Claude Code · day-hours · rep 1",1815,33.71],["Claude Fable 5.1 · Claude Code · csv-parse · rep 1",903,16.2],["Claude Fable 5.1 · Claude Code · event-loop-order · rep 1",1765,23.9],["Claude Fable 5.1 · Claude Code · talk-schedule · rep 1",1090,17.27],["Claude Fable 5.1 · Claude Code · semver-regex · rep 1",261,9.6],["Claude Fable 5.1 · Claude Code · money-refactor · rep 1",646,12.91],["Claude Fable 5.1 · Claude Code · q1-sql · rep 1",875,16.07],["Claude Fable 5.1 · Claude Code · merge-ranges · rep 2",75,4.46],["Claude Fable 5.1 · Claude Code · day-hours · rep 2",5889,90],["Claude Fable 5.1 · Claude Code · csv-parse · rep 2",1088,21.38],["Claude Fable 5.1 · Claude Code · event-loop-order · rep 2",1228,16.78],["Claude Fable 5.1 · Claude Code · talk-schedule · rep 2",786,11.23],["Claude Fable 5.1 · Claude Code · semver-regex · rep 2",236,7.52],["Claude Fable 5.1 · Claude Code · money-refactor · rep 2",796,13.68],["Claude Fable 5.1 · Claude Code · q1-sql · rep 2",865,20.31],["Claude Fable 5.1 · Claude Code · merge-ranges · rep 3",197,15.21],["Claude Fable 5.1 · Claude Code · day-hours · rep 3",1427,25.09],["Claude Fable 5.1 · Claude Code · csv-parse · rep 3",1000,15.88],["Claude Fable 5.1 · Claude Code · event-loop-order · rep 3",1606,26.91],["Claude Fable 5.1 · Claude Code · talk-schedule · rep 3",836,12.14],["Claude Fable 5.1 · Claude Code · semver-regex · rep 3",266,5.72],["Claude Fable 5.1 · Claude Code · money-refactor · rep 3",948,21.85],["Claude Fable 5.1 · Claude Code · q1-sql · rep 3",1091,24.58]]}],["GPT-6.1 Sol · Codex CLI",{"$k":["label","x","value"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI · merge-ranges · rep 1",163,13.2],["GPT-6.1 Sol (high) · Codex CLI · merge-ranges · rep 1",192,14.4],["GPT-6.1 Sol (medium) · Codex CLI · day-hours · rep 1",830,61.6],["GPT-6.1 Sol (high) · Codex CLI · day-hours · rep 1",1533,82.99],["GPT-6.1 Sol (medium) · Codex CLI · csv-parse · rep 1",86,14.79],["GPT-6.1 Sol (high) · Codex CLI · csv-parse · rep 1",256,22.91],["GPT-6.1 Sol (medium) · Codex CLI · event-loop-order · rep 1",199,8.54],["GPT-6.1 Sol (high) · Codex CLI · event-loop-order · rep 1",347,15.44],["GPT-6.1 Sol (medium) · Codex CLI · talk-schedule · rep 1",189,12.63],["GPT-6.1 Sol (high) · Codex CLI · talk-schedule · rep 1",189,13.05],["GPT-6.1 Sol (medium) · Codex CLI · semver-regex · rep 1",101,9.63],["GPT-6.1 Sol (high) · Codex CLI · semver-regex · rep 1",306,17.84],["GPT-6.1 Sol (medium) · Codex CLI · money-refactor · rep 1",89,11.74],["GPT-6.1 Sol (high) · Codex CLI · money-refactor · rep 1",224,19.39],["GPT-6.1 Sol (medium) · Codex CLI · q1-sql · rep 1",61,13.76],["GPT-6.1 Sol (high) · Codex CLI · q1-sql · rep 1",181,18.4],["GPT-6.1 Sol (medium) · Codex CLI · merge-ranges · rep 2",137,14.63],["GPT-6.1 Sol (high) · Codex CLI · merge-ranges · rep 2",226,15.12],["GPT-6.1 Sol (medium) · Codex CLI · day-hours · rep 2",839,56.39],["GPT-6.1 Sol (high) · Codex CLI · day-hours · rep 2",1750,92.21],["GPT-6.1 Sol (medium) · Codex CLI · csv-parse · rep 2",112,16.64],["GPT-6.1 Sol (high) · Codex CLI · csv-parse · rep 2",220,20.56],["GPT-6.1 Sol (medium) · Codex CLI · event-loop-order · rep 2",251,12.5],["GPT-6.1 Sol (high) · Codex CLI · event-loop-order · rep 2",375,22.51],["GPT-6.1 Sol (medium) · Codex CLI · talk-schedule · rep 2",197,11.81],["GPT-6.1 Sol (high) · Codex CLI · talk-schedule · rep 2",229,12.21],["GPT-6.1 Sol (medium) · Codex CLI · semver-regex · rep 2",193,13.01],["GPT-6.1 Sol (high) · Codex CLI · semver-regex · rep 2",189,11.67],["GPT-6.1 Sol (medium) · Codex CLI · money-refactor · rep 2",71,11.85],["GPT-6.1 Sol (high) · Codex CLI · money-refactor · rep 2",144,15.43],["GPT-6.1 Sol (medium) · Codex CLI · q1-sql · rep 2",119,17.81],["GPT-6.1 Sol (high) · Codex CLI · q1-sql · rep 2",222,19.62],["GPT-6.1 Sol (low) · Codex CLI · merge-ranges · rep 1",63,17.2],["GPT-6.1 Sol (low) · Codex CLI · day-hours · rep 1",458,44.29],["GPT-6.1 Sol (low) · Codex CLI · csv-parse · rep 1",78,17.66],["GPT-6.1 Sol (low) · Codex CLI · event-loop-order · rep 1",198,14.61],["GPT-6.1 Sol (low) · Codex CLI · talk-schedule · rep 1",189,13.6],["GPT-6.1 Sol (low) · Codex CLI · semver-regex · rep 1",44,8.57],["GPT-6.1 Sol (low) · Codex CLI · money-refactor · rep 1",62,13.12],["GPT-6.1 Sol (low) · Codex CLI · q1-sql · rep 1",0,13.64],["GPT-6.1 Sol (low) · Codex CLI · merge-ranges · rep 2",0,7.94],["GPT-6.1 Sol (low) · Codex CLI · day-hours · rep 2",400,40.71],["GPT-6.1 Sol (low) · Codex CLI · csv-parse · rep 2",0,17.85],["GPT-6.1 Sol (low) · Codex CLI · event-loop-order · rep 2",185,11.1],["GPT-6.1 Sol (low) · Codex CLI · talk-schedule · rep 2",189,12.41],["GPT-6.1 Sol (low) · Codex CLI · semver-regex · rep 2",50,8.9],["GPT-6.1 Sol (low) · Codex CLI · money-refactor · rep 2",0,9.54],["GPT-6.1 Sol (low) · Codex CLI · q1-sql · rep 2",40,14.42]]}]]},"Calculation, not a run: each point is one recorded call. Spearman rank correlations and slopes are in the table. Total time includes CLI start-up and the visible answer. Task, effort and batch can affect both counts and time. The plot shows association, not cause. 248 calls over 8 tasks; calls are not independent.","\u0001","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"],"Reasoning tokens per call"],["thinking-bill-short-vs-hard","Reasoning share on short tasks vs hard tasks (calculation)","Median call per configuration; five short tasks and eight hard tasks","grouped-bar","percent","Reasoning share of output tokens (median call, %)",[{"name":"Eight hard tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",91.68,76.46,99.27,24],["Claude Sonnet 5.5 · Claude Code",54.54,0,95.91,24],["Claude Opus 5.5 · Claude Code",54.79,29.92,95.6,24],["Claude Opus 5.5 (high) · Claude Code",54.43,36.14,96.23,24],["Claude Fable 5.1 · Claude Code",64.24,23.44,97.19,24],["GPT-6.1 Sol (medium) · Codex CLI",46.33,11.42,86.85,16],["GPT-6.1 Sol (high) · Codex CLI",57.01,29.19,90.8,16]]}},{"name":"Five short tasks","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",90.19,73.1,97.59,15],["Claude Sonnet 5.5 · Claude Code",0,0,72.75,15],["Claude Opus 5.5 · Claude Code",0,0,93.33,15],["Claude Opus 5.5 (high) · Claude Code",43.59,0,93.33,15],["Claude Fable 5.1 · Claude Code",0,0,74.01,15],["GPT-6.1 Sol (medium) · Codex CLI",41.05,0,71.43,15],["GPT-6.1 Sol (high) · Codex CLI",58.06,0,75.76,15]]}}],"Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.","minmax","none",["calc-thinking-bill","agent-provider-h2h-hard","agent-effort-ladder","agent-provider-h2h","price-anthropic","price-openai"],"\u0001"]]},"related":["effort-ladder","hard-model-head-to-head"]}}