• Thought experiment
  • Calculation
  • Reasoning Tokens
  • Thinking Tokens
  • Effort
  • LLM pricing
  • Claude Haiku
  • Claude Sonnet
  • Claude Opus
  • Claude Fable
  • GPT-6.1 Sol
  • Claude Code
  • Codex CLI
  • Latency

How much of an AI bill is thinking? Reasoning tokens by model and effort

Across 378 recorded calls of Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol, how many output tokens come from reasoning (thinking)? What do they cost at list price per call and per strict pass, and do they track time? The calls cover eight hard tasks, five short tasks and several efforts.

Published · 5 charts · Download the data or a carousel

92%

Calculation

n = 24

Highest median reasoning share of output tokens, hard tasks (calculation) · (Claude Haiku 4.5 · Claude Code; range 76% to 99%)

46%

Calculation

n = 16

Lowest median reasoning share of output tokens, hard tasks (calculation) · (GPT-6.1 Sol (medium) · Codex CLI; range 11% to 87%)

The answer

On eight hard tasks, reasoning tokens were 46% to 92% of the output tokens of a median call (16 to 24 calls per configuration). Haiku 4.5 had the highest median (92%), GPT-6.1 Sol (medium) had the lowest (46%) and the other 5 sat between 54% and 64%. Per-call ranges are wide (Sonnet 5.5 0% to 96%). The ranges of every pair overlap, so this run ranks no configuration. Output share is not bill share. At list price (a calculation) the reasoning part of a call cost a mean $0.0023 (GPT-6.1 Sol (medium)) to $0.0537 (Fable 5.1). That was 9% (GPT-6.1 Sol (medium)) to 80% (Haiku 4.5) of the total list-price cost in each configuration. This pooled share divides summed reasoning cost by summed total cost. Input tokens cost money too. The calls ran on flat subscriptions, so this is not a bill. Fable 5.1's thinking cost 8.1x Sonnet 5.5's per call. The output price explains a factor of 5.0 ($50 against $10 per million tokens). More reasoning tokens explain the rest, a factor of 1.6 (means). Higher effort settings had higher mean recorded reasoning counts in these batches. We paired the same eight tasks. At high effort, mean reasoning tokens per call were higher than at low effort on these tasks: Sonnet 5.5 7 of 8, Opus 5.5 8 of 8, GPT-6.1 Sol 8 of 8. The pooled reasoning share rose from low to high effort. Sonnet 5.5 53% to 73%. Opus 5.5 38% to 69%. GPT-6.1 Sol 29% to 59%. It rose at each step (low, medium, high) in all 3 ladders. The reasoning cost per call rose 2.2x for Sonnet 5.5, 3.6x for Opus 5.5 and 3.4x for GPT-6.1 Sol (a calculation). Every effort cell passed 16/16 strictly. The pass count stayed the same on this set, which has a ceiling (95% Wilson 81% to 100% per cell). Recorded reasoning counts correlated with total time. Within each Claude model, the rank correlation between reasoning tokens and total time was 0.85 to 0.98 (Spearman, 24 to 80 calls each). On the same task, 1,000 more reasoning tokens went with 7.6 to 14.5 s more time (a calculation). For GPT-6.1 Sol in the Codex CLI, the recorded rank correlation was: 0.56 (48 calls). Thinking costs money whether or not the call passes. Haiku 4.5 passed 11/24 strictly (95% Wilson 28% to 65%). The 13 calls that did not pass held 61% of its reasoning cost (a calculation).

Key numbers

378

Recorded calls analysed (no new calls)

(152 hard head-to-head, 96 new effort-ladder, 130 five-task head-to-head) · n = 378

98

Calls that reported 0 reasoning tokens (kept as 0, not "not reported")

of 378; 0 calls had no usable count · n = 378

0.99

Sum check, Claude Code: output tokens gained per extra reasoning token, same task and model (calculation)

(leave-one-task-out range 0.98 to 1.01; 290 calls) · n = 290

0.97

Sum check, Codex CLI: output tokens gained per extra reasoning token, same task and model (calculation)

(leave-one-task-out range 0.97 to 0.99; 88 calls) · n = 88

$0.0537

Highest mean list-price cost of reasoning per call, hard tasks (calculation)

(Claude Fable 5.1 · Claude Code; call range $0.0037 to $0.2944) · n = 24

$0.0023

Lowest mean list-price cost of reasoning per call, hard tasks (calculation)

(GPT-6.1 Sol (medium) · Codex CLI; call range $0.0006 to $0.0084) · n = 16

80%

Highest pooled reasoning share of total list-price cost, hard tasks (calculation)

(Claude Haiku 4.5 · Claude Code; call range 40% to 90%) · n = 24

9%

Lowest pooled reasoning share of total list-price cost, hard tasks (calculation)

(GPT-6.1 Sol (medium) · Codex CLI; call range 2% to 20%) · n = 16

8.1x

Reasoning cost per call, Fable 5.1 ÷ Sonnet 5.5, hard tasks (calculation, means)

($0.0537 vs $0.0067) · n = 24

61%

Share of Haiku 4.5 reasoning cost spent on calls that did not pass (calculation)

(13 of 24 calls did not pass strictly) · n = 24

7

Tasks where mean reasoning tokens per call were higher at high than at low effort, Claude Sonnet 5.5 · Claude Code (paired, same tasks; calculation)

of 8 tasks · n = 8

8

Tasks where mean reasoning tokens per call were higher at high than at low effort, Claude Opus 5.5 · Claude Code (paired, same tasks; calculation)

of 8 tasks · n = 8

8

Tasks where mean reasoning tokens per call were higher at high than at low effort, GPT-6.1 Sol · Codex CLI (paired, same tasks; calculation)

of 8 tasks · n = 8

2.2x

Reasoning cost per call, high ÷ low effort, Claude Sonnet 5.5 · Claude Code (calculation, means)

($0.0043 at low, $0.0094 at high) · n = 16

3.6x

Reasoning cost per call, high ÷ low effort, Claude Opus 5.5 · Claude Code (calculation, means)

($0.0050 at low, $0.0180 at high) · n = 16

3.4x

Reasoning cost per call, high ÷ low effort, GPT-6.1 Sol · Codex CLI (calculation, means)

($0.0012 at low, $0.0041 at high) · n = 16

0.85

Spearman, reasoning tokens vs total time, per Claude model: lowest (calculation)

to 0.98 across 4 Claude models · n = 200

7.6

Seconds per 1,000 reasoning tokens on the same task, per Claude model: lowest (calculation)

to 14.5 s across 4 Claude models · n = 200

0.56

Spearman, reasoning tokens vs total time, GPT-6.1 Sol in Codex CLI (calculation)

(leave-one-task-out range 0.34 to 0.65; 48 calls) · n = 48

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Thought experiment: not a run. These values reprice recorded tokens at list prices. No model was called again.

Calculation
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Hover or focus a bar for its ratio to GPT-6.1 Sol (medium) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Haiku 4.5 · Claude Code 92% (range 76%–99%, n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 46% (range 11%–87%, n 16). All run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 16–24 per row

Median call: reasoning tokens ÷ output tokens. Whiskers: lowest and highest call (16 to 24 calls per configuration)

Calculation from reported tokens, not a run. Each call gives reasoning ÷ output; the bar is the median of those shares. Whiskers are the lowest and highest call. They are a range, not a confidence interval. They are wide, so the medians describe this run and rank nothing. The pooled share (all reasoning tokens ÷ all output tokens) is in the table. We treat reasoning tokens as part of output tokens; the consistency check supports this accounting assumption. Each CLI reports its own count.

Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)
Calculation
  • Reasoning (output tokens)
  • Remaining output (visible-answer estimate)
  • Input (prompt, cache priced)
Bar length is the total; segments are its parts.
Claude Fable 5.1 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI
Claude Sonnet 5.5 · Claude Code

Totals are the sum of the parts shown. Shares are calculated from the same values.

List-price calculation, not a run. 7 rows, 3 series: Reasoning (output tokens), Remaining output (visible-answer estimate), Input (prompt, cache priced). Reasoning (output tokens): highest Claude Fable 5.1 · Claude Code $0.054 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.0023 (n 16). Remaining output (visible-answer estimate): highest Claude Fable 5.1 · Claude Code $0.019 (n 24). Lowest Claude Haiku 4.5 · Claude Code $0.0018 (n 24).

Notesn 16–24 per row

Mean per call on the hard tasks; the three parts add up to the call

Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.

Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)
Calculation
  • Reasoning cost per strict pass
  • Total cost per strict pass (square)
Sorted by gap, largest first.
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI
Claude Opus 5.5 (low) · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code

Gap labels, Total cost per strict pass vs Reasoning cost per strict pass: Total cost per strict pass is x% higher (+) or lower (−) than Reasoning cost per strict pass, calculated from the two values shown (the change counted from Reasoning cost per strict pass’s value).

List-price calculation, not a run. 11 rows, 2 series: Reasoning cost per strict pass, Total cost per strict pass. Reasoning cost per strict pass: highest Claude Opus 5.5 (high) · Claude Code $0.018 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI $0.0012 (n 16). Total cost per strict pass: highest Claude Opus 5.5 (high) · Claude Code $0.034 (n 16). Lowest Claude Sonnet 5.5 (low) · Claude Code $0.012 (n 16).

Notesn = 16 per row

List price ÷ strict passes; every cell is 8 tasks × 2 repetitions

Calculation, not a bill: reported tokens × list price, divided by the cell's strict passes; the calls ran on flat subscriptions. The effort-ladder cells: new calls plus reference cells reused from the hard head-to-head (Claude repetitions 1-2 only). "Default" means the effort flag was not passed. The total is the same value as the effort-ladder cost-per-pass chart. Effort levels are not the same scale across vendors, and the reference cells ran in a different batch and hour.

Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)
Calculation
  • Claude Haiku 4.5 · Claude Code
  • Claude Sonnet 5.5 · Claude Code
  • Claude Opus 5.5 · Claude Code
  • Claude Fable 5.1 · Claude Code
  • GPT-6.1 Sol · Codex CLI

List-price calculation, not a run. 248 points: Total time per call (seconds) against Reasoning tokens per call. Reasoning tokens per call runs from 0 to 8,569; Total time per call (seconds) from 2.3 s to 92.2 s.

Notes

One point per call: 248 calls from the hard head-to-head and the effort ladder

Calculation, not a run: each point is one recorded call. Spearman rank correlations and slopes are in the table. Total time includes CLI start-up and the visible answer. Task, effort and batch can affect both counts and time. The plot shows association, not cause. 248 calls over 8 tasks; calls are not independent.

Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)
Calculation
  • Eight hard tasks
  • Five short tasks (square)
Sorted by gap, largest first.
Claude Fable 5.1 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI

Gap labels, Five short tasks vs Eight hard tasks: Five short tasks is x percentage points higher (+) or lower (−) than Eight hard tasks, calculated from the two values shown; lines are the lowest–highest run (not an interval).

List-price calculation, not a run. 7 rows, 2 series: Eight hard tasks, Five short tasks. Eight hard tasks: highest Claude Haiku 4.5 · Claude Code 92% (range 76%–99%, n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 46% (range 11%–87%, n 16). All run ranges overlap. Five short tasks: highest Claude Haiku 4.5 · Claude Code 90% (range 73%–98%, n 15). Lowest Claude Fable 5.1 · Claude Code 0% (range 0%–74%, n 15). Not all run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 15–24 per row

Median call per configuration; five short tasks and eight hard tasks

Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.

Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)

Tables

Reasoning tokens and list-price cost per call, hard tasks (calculation)

ConfigurationCallsStrict passesMedian output tokensMedian reasoning tokensOutput / reasoning tokens, lowest to highest callReasoning / total cost, lowest to highest call (USD, calculation)Reasoning cost share, lowest to highest call (calculation)Calls with 0 reasoning tokensMedian call: reasoning share of outputShare, lowest to highest callPooled share (all reasoning ÷ all output)Reasoning cost per call (USD, calculation)Remaining output cost per call (USD, calculation)Input cost per call (USD)Total cost per call (USD)Pooled reasoning share of list-price cost (calculation)Reasoning cost per strict pass (USD)Total cost per strict pass (USD)
Claude Haiku 4.5 · Claude Code2411/24 (95% Wilson 28% to 65%)5,0644,5561899 to 9321 / 1452 to 8569$0.0073 to $0.0428 / $0.0148 to $0.050540% to 90%092%76% to 99%93%$0.024$0.0018$0.0045$0.03180%$0.053$0.067
Claude Sonnet 5.5 · Claude Code2424/24 (95% Wilson 86% to 100%)1,050585176 to 3895 / 0 to 3060$0.0000 to $0.0306 / $0.0051 to $0.04240% to 74%755%0% to 96%64%$0.0067$0.0037$0.004$0.01446%$0.0067$0.014
Claude Opus 5.5 · Claude Code2424/24 (95% Wilson 86% to 100%)945529323 to 2531 / 100 to 1778$0.0020 to $0.0356 / $0.0131 to $0.057210% to 74%055%30% to 96%61%$0.013$0.008$0.0077$0.02844%$0.013$0.028
Claude Opus 5.5 (high) · Claude Code2424/24 (95% Wilson 86% to 100%)1,052614285 to 4052 / 103 to 3301$0.0021 to $0.0660 / $0.0121 to $0.087717% to 77%054%36% to 96%69%$0.018$0.008$0.0074$0.03354%$0.018$0.033
Claude Fable 5.1 · Claude Code2424/24 (95% Wilson 86% to 100%)1,366889318 to 6465 / 75 to 5889$0.0037 to $0.2944 / $0.0322 to $0.33995% to 87%064%23% to 97%74%$0.054$0.019$0.021$0.09358%$0.054$0.093
GPT-6.1 Sol (medium) · Codex CLI1616/16 (95% Wilson 81% to 100%)335150237 to 1766 / 61 to 839$0.0006 to $0.0084 / $0.0099 to $0.04222% to 20%046%11% to 87%43%$0.0023$0.003$0.02$0.0268.9%$0.0023$0.026
GPT-6.1 Sol (high) · Codex CLI1616/16 (95% Wilson 81% to 100%)436225284 to 2569 / 144 to 1750$0.0014 to $0.0175 / $0.0092 to $0.03078% to 60%057%29% to 91%59%$0.0041$0.0029$0.0081$0.01527%$0.0041$0.015

Reasoning tokens and list-price cost by effort, every effort-ladder cell (calculation)

ConfigurationCallsStrict passesMedian output tokensMedian reasoning tokensOutput / reasoning tokens, lowest to highest callReasoning / total cost, lowest to highest call (USD, calculation)Reasoning cost share, lowest to highest call (calculation)Calls with 0 reasoning tokensMedian call: reasoning share of outputShare, lowest to highest callPooled share (all reasoning ÷ all output)Reasoning cost per call (USD, calculation)Remaining output cost per call (USD, calculation)Input cost per call (USD)Total cost per call (USD)Pooled reasoning share of list-price cost (calculation)Reasoning cost per strict pass (USD)Total cost per strict pass (USD)
Claude Sonnet 5.5 (low) · Claude Code1616/16 (95% Wilson 81% to 100%)667273176 to 2263 / 0 to 1489$0.0000 to $0.0149 / $0.0051 to $0.02610% to 71%821%0% to 95%53%$0.0043$0.0038$0.0041$0.01235%$0.0043$0.012
Claude Sonnet 5.5 (medium) · Claude Code1616/16 (95% Wilson 81% to 100%)770422224 to 2836 / 0 to 2093$0.0000 to $0.0209 / $0.0056 to $0.03180% to 75%552%0% to 96%62%$0.0059$0.0037$0.0039$0.01444%$0.0059$0.014
Claude Sonnet 5.5 (high) · Claude Code1616/16 (95% Wilson 81% to 100%)1,192745220 to 4187 / 0 to 3610$0.0000 to $0.0361 / $0.0056 to $0.04530% to 80%261%0% to 96%73%$0.0094$0.0035$0.0039$0.01756%$0.0094$0.017
Claude Sonnet 5.5 · Claude Code1616/16 (95% Wilson 81% to 100%)1,054668176 to 2185 / 0 to 1614$0.0000 to $0.0161 / $0.0051 to $0.02530% to 74%456%0% to 96%64%$0.0063$0.0036$0.0041$0.01445%$0.0063$0.014
Claude Opus 5.5 (low) · Claude Code1616/16 (95% Wilson 81% to 100%)59487220 to 1569 / 0 to 791$0.0000 to $0.0158 / $0.0108 to $0.03800% to 67%731%0% to 94%38%$0.005$0.0083$0.0079$0.02124%$0.005$0.021
Claude Opus 5.5 (medium) · Claude Code1616/16 (95% Wilson 81% to 100%)853518301 to 3075 / 109 to 1993$0.0022 to $0.0399 / $0.0125 to $0.068116% to 74%054%28% to 96%61%$0.013$0.0086$0.0074$0.02946%$0.013$0.029
Claude Opus 5.5 (high) · Claude Code1616/16 (95% Wilson 81% to 100%)1,052614285 to 4052 / 103 to 3301$0.0021 to $0.0660 / $0.0121 to $0.087717% to 77%054%36% to 96%69%$0.018$0.0082$0.0074$0.03454%$0.018$0.034
Claude Opus 5.5 · Claude Code1616/16 (95% Wilson 81% to 100%)945538323 to 2531 / 100 to 1778$0.0020 to $0.0356 / $0.0131 to $0.057210% to 72%054%31% to 95%62%$0.013$0.008$0.0079$0.02945%$0.013$0.029
GPT-6.1 Sol (low) · Codex CLI1616/16 (95% Wilson 81% to 100%)28463172 to 1310 / 0 to 458$0.0000 to $0.0046 / $0.0092 to $0.02640% to 22%424%0% to 84%29%$0.0012$0.003$0.0086$0.0139.5%$0.0012$0.013
GPT-6.1 Sol (medium) · Codex CLI1616/16 (95% Wilson 81% to 100%)335150237 to 1766 / 61 to 839$0.0006 to $0.0084 / $0.0099 to $0.04222% to 20%046%11% to 87%43%$0.0023$0.003$0.02$0.0268.9%$0.0023$0.026
GPT-6.1 Sol (high) · Codex CLI1616/16 (95% Wilson 81% to 100%)436225284 to 2569 / 144 to 1750$0.0014 to $0.0175 / $0.0092 to $0.03078% to 60%057%29% to 91%59%$0.0041$0.0029$0.0081$0.01527%$0.0041$0.015

Reasoning tokens and list-price cost per call, five short tasks (calculation)

ConfigurationCallsStrict passesMedian output tokensMedian reasoning tokensOutput / reasoning tokens, lowest to highest callReasoning / total cost, lowest to highest call (USD, calculation)Reasoning cost share, lowest to highest call (calculation)Calls with 0 reasoning tokensMedian call: reasoning share of outputShare, lowest to highest callPooled share (all reasoning ÷ all output)Reasoning cost per call (USD, calculation)Remaining output cost per call (USD, calculation)Input cost per call (USD)Total cost per call (USD)Pooled reasoning share of list-price cost (calculation)Reasoning cost per strict pass (USD)Total cost per strict pass (USD)
Claude Haiku 4.5 · Claude Code1515/15 (95% Wilson 80% to 100%)367297275 to 2851 / 212 to 2643$0.0011 to $0.0132 / $0.0051 to $0.018020% to 74%090%73% to 98%90%$0.0041$0.00043$0.0038$0.008449%$0.0041$0.0084
Claude Sonnet 5.5 · Claude Code1512/15 (95% Wilson 55% to 93%)107044 to 745 / 0 to 542$0.0000 to $0.0054 / $0.0034 to $0.01020% to 53%120%0% to 73%45%$0.00088$0.0011$0.003$0.00518%$0.0011$0.0062
Claude Opus 5.5 · Claude Code1515/15 (95% Wilson 80% to 100%)64044 to 853 / 0 to 616$0.0000 to $0.0123 / $0.0059 to $0.02230% to 55%90%0% to 93%59%$0.0026$0.0018$0.0057$0.0126%$0.0026$0.01
Claude Opus 5.5 (low) · Claude Code1515/15 (95% Wilson 80% to 100%)64044 to 637 / 0 to 405$0.0000 to $0.0081 / $0.0058 to $0.01790% to 45%90%0% to 93%41%$0.0013$0.0018$0.0052$0.008315%$0.0013$0.0083
Claude Opus 5.5 (high) · Claude Code1515/15 (95% Wilson 80% to 100%)783444 to 1094 / 0 to 857$0.0000 to $0.0171 / $0.0059 to $0.02710% to 63%744%0% to 93%66%$0.0034$0.0018$0.0052$0.0133%$0.0034$0.01
Claude Fable 5.1 · Claude Code1515/15 (95% Wilson 80% to 100%)6404 to 908 / 0 to 672$0.0000 to $0.0336 / $0.0049 to $0.05840% to 67%120%0% to 74%57%$0.006$0.0044$0.01$0.02129%$0.006$0.021
GPT-6.1 Sol (low) · Codex CLI1010/10 (95% Wilson 72% to 100%)422028 to 225 / 0 to 103$0.0000 to $0.0010 / $0.0074 to $0.02650% to 11%425%0% to 71%31%$0.00033$0.00073$0.0089$0.013.3%$0.00033$0.01
GPT-6.1 Sol (medium) · Codex CLI1515/15 (95% Wilson 80% to 100%)422027 to 295 / 0 to 137$0.0000 to $0.0014 / $0.0054 to $0.02690% to 23%641%0% to 71%41%$0.00051$0.00072$0.014$0.0163.3%$0.00051$0.016
GPT-6.1 Sol (high) · Codex CLI1515/15 (95% Wilson 80% to 100%)422128 to 457 / 0 to 289$0.0000 to $0.0029 / $0.0066 to $0.02810% to 29%658%0% to 76%58%$0.001$0.00072$0.011$0.0137.6%$0.001$0.013

Consistency check for reasoning within output tokens (calculation)

CLI routeCallsCalls with more than 0 reasoning tokensCalls that reported 0 (kept as 0)Calls with reasoning above outputOutput tokens gained per extra reasoning token (within task)Leave-one-task-out range (not a 95% interval)Characters per token of output minus reasoningCharacters per output token, calls with 0 reasoningCharacters per output token, calls with reasoningReasoning is part of output
Claude Code2902127800.99x0.982 to 1.0091.941.980.45consistent with inclusion; not proof
Codex CLI88682000.97x0.965 to 0.9862.992.761.36consistent with inclusion; not proof

Method

  1. This study is a calculation over receipts that already exist. We made no new model call. The inputs are three recorded runs. The hard head-to-head gives 152 completed calls. We leave its 30 pre-inference blocked attempts out of token calculations. They remain failures of the original route attempt, not model answers. The effort ladder gives 96 new calls. The five-task head-to-head gives 130 calls. All current counted calls completed. Recorded errors count as failures when their usage is known. We keep the separate blocked batch on record. The original Codex route was blocked. A later batch ran after a successful uncounted probe. It resumed after two completed calls and skipped them.
  2. Sum check. Before we compute any share, we test one point. Do the output tokens include the reasoning tokens? We use this accounting assumption and check its consistency with the receipts. We run three tests on each CLI route. Test 1: reasoning never exceeds output. Test 2: within one task and model, one more reasoning token adds about one output token. A slope near 1 supports the assumption. Correlation alone cannot prove the counter semantics. Test 3: we compare characters per remaining token with characters per output token on calls whose reasoning counter is zero. Claude Code: consistent with inclusion, not proof. Reasoning never exceeded output (0 of 290 calls). Within one task and model, each extra reasoning token went with 0.99 extra output tokens (leave-one-task-out range 0.98 to 1.01, 290 calls). Output minus reasoning has 1.94 characters per token. Calls with 0 reasoning have 1.98. For comparison, output including reasoning has only 0.45. Codex CLI: consistent with inclusion, not proof. Reasoning never exceeded output (0 of 88 calls). Within one task and model, each extra reasoning token went with 0.97 extra output tokens (leave-one-task-out range 0.97 to 0.99, 88 calls). Output minus reasoning has 2.99 characters per token. Calls with 0 reasoning have 2.76. For comparison, output including reasoning has only 1.36.
  3. Not reported is not zero. A route reports reasoning when at least one of its calls reports more than 0. Both routes do (Claude Code 212 of 290 calls, Codex CLI 68 of 88 calls). So a reported 0 is the recorded counter, not proof of no internal reasoning. 98 calls reported 0, and their replies look like visible text only (test 3). If any model attempt lacks usable token counts, this builder returns no study. It never prices unknown usage as zero. 0 calls had no usable count.
  4. Reasoning share of one call = reasoning tokens ÷ output tokens. A configuration gets two numbers. One is the median of its per-call shares. The other is the pooled share (all reasoning tokens ÷ all output tokens). Ranges are the lowest and highest call. They are not intervals.
  5. Hard tasks: we use all calls of each configuration in the hard head-to-head. That is 16 to 24 calls: 8 tasks × 3 repetitions for Claude Code and 8 × 2 for Codex CLI. Effort ladder: 11 cells of 8 tasks × 2 repetitions, the design of the effort-ladder study. We reuse its reference cells from the hard head-to-head (Claude repetitions 1-2 only). Short tasks: the five validated tasks (10 to 15 calls per configuration).
  6. Cost is a calculation at list price, not a bill. The calls ran on flat subscriptions. Reasoning cost = reasoning tokens × the model's output price. Remaining output cost = (output tokens − reasoning tokens) × the same price. We use the rest as a visible-answer estimate. Input cost covers the whole prompt. Cache writes use the one-hour list-price assumption; the receipts do not state the cache lifetime. Prices per million output tokens: Claude Haiku 4.5 $5, Claude Sonnet 5.5 $10, Claude Opus 5.5 $20, Claude Fable 5.1 $50, GPT-6.1 Sol $10. Sources: Anthropic list prices of 2026-09-21, OpenAI of 2026-10-03.
  7. Cost per strict pass = the list-price cost of all calls in the cell, failures included, ÷ the cell's strict passes. Hard and effort cells use the hard-set strict rule. Short-task cells use their original rule, which can strip a wrapping fence.
  8. Time link: we use all calls of the hard head-to-head and the effort ladder. For each model and route, we compute the Spearman rank correlation of reasoning tokens with total time (with a leave-one-task-out sensitivity range, not a confidence interval). We also compute the least-squares slope in seconds per 1,000 reasoning tokens. Then we compute the slope within task. We centre each task on its own mean. This controls for differences in task means, not effort or batch effects.
  9. Isolation, tasks, validators and flags follow the hard head-to-head and the five-task head-to-head. Each call ran in a fresh empty folder with tools off and one turn. The timeout was 300 s per call (180 s for the short tasks). One call ran at a time per account.
  10. Claude Code calls ran with an output cap of 16,000 tokens in all three runs. The highest output of any analysed Claude call was 9,321, so no call hit the cap. The Codex CLI has no cap setting. Its highest output was 2,569.

Caveats

  • The hard-set protocol file was created after its first counted call. Both effort-ladder protocol files were created after their batches ended. The short-set file predates its first call, but its top-up amendment timing is unverified. Batch receipts preserve protocol text, but we cannot verify all rules were written before inference. Treat these as exploratory calculations, not preregistered tests.
  • The calls used one shared Mac and network. Other work and provider load were not controlled. Latency includes those effects.
  • These are hand-built, tuned case sets, not random workload samples. The short set and most hard-set configurations hit a pass ceiling. Repeats of the same tasks are not independent. Wilson pass intervals describe counted calls under an independence assumption, not performance on new tasks.
  • Each CLI reports its own reasoning counter, and we never see the reasoning text. So we cannot check what each vendor counts. The sum check shows only that the counters are consistent with inclusion in output; it does not prove their semantics. Shares from Claude Code and Codex CLI do not measure like for like how much each model thinks.
  • Every prompt asks for a short reply in a strict format (code, JSON, one line or a regex, with no explanation). So the visible answer is short and the reasoning share is high. Longer visible replies could change the share. We did not measure that workload.
  • Small cells: 16 to 24 calls per configuration over 8 tasks. Per-call ranges are wide and overlap for every pair of configurations. So the medians describe this run and rank nothing. A sentence says one side is ahead only when its ranges do not overlap.
  • Input includes CLI context that these receipts do not separately count. It moves with cache hits. So the part of the call cost that goes to reasoning depends on the CLI and on the cache, not only on the model. GPT-6.1 Sol at medium and at high effort ran in different batches and show very different input cost per call.
  • Every effort-ladder cell passed 16/16, so the set has a ceiling. Higher effort had higher mean recorded reasoning cost in these batches. It does not show that more thinking never helps on harder work.
  • The time link is a correlation, not a cause. Task, effort and batch can affect reasoning, visible output and time together. Total time includes CLI start-up. The within-task slope controls for differences in task means. It does not remove effort or batch effects. Output tokens (reasoning plus visible answer) track time at least as closely as reasoning alone. The Codex CLI slope (36.4 s per 1,000 reasoning tokens) has a calculated ratio of 2.5 to 4.8 times the Claude Code slopes. We did not test why. Calls are repeats of 8 tasks, so they are not independent. The ranges omit one whole task at a time. They show sensitivity to the task mix, not 95% coverage.
  • Effort levels are not the same scale across vendors. "Default" means we did not pass the effort flag, and the CLI chose. The effort-ladder reference cells ran in a different batch and hour than the new cells.
  • List-price costs are calculations, because the calls used flat subscriptions. A price change moves every cost here. It leaves every token count unchanged.

Sources

  • Reasoning token bill (calculation)

    Calculation ·

    Reported reasoning tokens priced at the recorded list prices. A calculation, not a new run.

    Raw data: provider-h2h-hard/receipts.json, effort-ladder/receipts.json, provider-h2h/receipts.json

  • Provider head-to-head, hard set: eight hard tasks with strict validators

    Our recorded runs ·

    Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.

    Raw data: provider-h2h-hard/receipts.json

  • Effort ladder: the hard task set at each effort level

    Our recorded runs ·

    The eight hard tasks of the hard head-to-head at low, medium and high effort (Claude Sonnet 5.5 and Claude Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI), 8 tasks × 2 repetitions per cell, declared protocols, every attempt kept. Efforts the hard head-to-head already ran are reused from its receipts as reference cells.

    Raw data: effort-ladder/receipts.json, provider-h2h-hard/receipts.json

  • Provider head-to-head: Claude Code models vs Codex efforts

    Our recorded runs ·

    Five short tasks with deterministic validators, declared protocol, every attempt kept.

    Raw data: provider-h2h/receipts.json

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • OpenAI list prices

    Vendor price list ·

    Token prices as listed by the vendor on 2026-10-03.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “How much of an AI bill is thinking? Reasoning tokens by model and effort”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/thinking-token-bill.

More comparisons based on this study (18)

These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.

Models and comparisons in this study

More studies

All benchmarks
Includes calculations
  • Prompt Caching
  • Break Even

Prompt cache break-even: after how many reuses does a cached prefix cost less?

A calculation on list prices: reuses before a cached prompt prefix costs less, per Claude and GPT model, and cost per 1,000 sessions of 1 to 20 turns.

2reuses (the 3rd request) · Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation)

3 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.