• Thought experiment
  • Reasoning Tokens
  • Thinking Tokens
  • Effort

How much of your AI bill is thinking tokens? Claude and GPT-6.1 Sol, measured

Thinking tokens: 46% to 92% of output per median call, 9% to 80% of pooled list-price cost. Calculation over recorded Claude and GPT-6.1 Sol calls.

Live story · 48 sDoes more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.

Transcript
  1. Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
  2. 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  3. All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  4. Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
  5. More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
  6. List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
  7. 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  8. Start at low effort and measure. Every call, interval and cost online.

TL;DR

  • This is a calculation, not a new run. We re-read 378 calls that we had already recorded: Claude Haiku 4.5, Sonnet 5.5, Opus 5.5 and Fable 5.1 in Claude Code, and GPT-6.1 Sol in the Codex CLI. The calls ran on flat subscriptions, so every dollar figure is list-price arithmetic.
  • We treat thinking tokens as part of output. Within one task and model, each extra reasoning token went with 0.99 extra output tokens in Claude Code (290 calls). In the Codex CLI it was 0.97 (88 calls). Leave-one-task-out ranges were 0.98 to 1.01 and 0.97 to 0.99; these are not 95% intervals. This calculation supports the accounting assumption; it does not prove how the CLIs count tokens.
  • Share of output (calculation): on eight hard tasks (16 to 24 calls per configuration), reasoning was 46% to 92% of the output tokens of a median call. Haiku 4.5 had the highest median (92%). The other five configurations sat between 54% and 64%. The per-call ranges overlap, so this ranks no model.
  • Share of the bill (calculation): reasoning was 9% to 80% of pooled list-price cost per configuration (16 to 24 calls each). Input tokens cost money too. The reasoning part of a call cost a mean $0.0023 (GPT-6.1 Sol, medium effort) to $0.0537 (Fable 5.1).
  • Effort (calculation, 16 calls per cell): from low to high effort, the reasoning cost per call rose 2.2x for Sonnet, 3.6x for Opus and 3.4x for GPT-6.1 Sol. Every effort cell passed 16 of 16 (95% Wilson interval 81% to 100% each). The pass count stayed the same on this ceiling-limited set.
  • Time (calculation): for each Claude model, reasoning tokens and total time had a rank correlation of 0.85 to 0.98 (24 to 80 calls each). On the same task, 1,000 more reasoning tokens went with 7.6 to 14.5 s more time.
  • Failed calls pay too: 13 of 24 Haiku 4.5 calls did not pass strictly (54%, 95% Wilson interval 35% to 72%). They held 61% of Haiku's reasoning cost (calculation).

The study with every input: /benchmarks/thinking-token-bill.

What are thinking tokens, and are they extra?

A reasoning model can use internal reasoning tokens before it gives the answer. Vendors call these reasoning tokens or thinking tokens. This post uses both words for the same thing. Some tools report the count and hide the text. New to the setting that controls how much a model thinks? Read what reasoning effort is.

The bill question is simple. Do thinking tokens come on top of the output count, or inside it? If they are inside, you pay the output price for them. We check that assumption first, because every cost split in this post depends on it.

Are thinking tokens part of the output tokens?

The receipts are consistent with that accounting assumption. They do not prove how the CLIs count tokens. We checked three things on each CLI route. First, reasoning never exceeded output. Second, within one task and model, one more reasoning token went with about one more output token. Third, output minus reasoning had a similar characters-per-token ratio to replies with a zero reasoning counter. The table gives the first two results. The method section gives the third.

CLICallsCalls with reasoning above outputExtra output tokens per extra reasoning token (task sensitivity range)
Claude Code29000.99 (0.982 to 1.009)
Codex CLI8800.97 (0.965 to 0.986)

Both CLIs report reasoning. 212 of 290 Claude Code calls and 68 of 88 Codex CLI calls reported more than 0 reasoning tokens. A reported 0 is the recorded counter, not proof of no internal reasoning. 98 calls reported 0. For this calculation, we keep each reported zero. A missing count is unknown, not zero. No call had a missing count.

How much of a call's output is thinking tokens?

Calculation
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Hover or focus a bar for its ratio to GPT-6.1 Sol (medium) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Haiku 4.5 · Claude Code 92% (range 76%–99%, n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 46% (range 11%–87%, n 16). All run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 16–24 per row

Median call: reasoning tokens ÷ output tokens. Whiskers: lowest and highest call (16 to 24 calls per configuration)

Calculation from reported tokens, not a run. Each call gives reasoning ÷ output; the bar is the median of those shares. Whiskers are the lowest and highest call. They are a range, not a confidence interval. They are wide, so the medians describe this run and rank nothing. The pooled share (all reasoning tokens ÷ all output tokens) is in the table. We treat reasoning tokens as part of output tokens; the consistency check supports this accounting assumption. Each CLI reports its own count.

Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices

On eight hard tasks, reasoning was 46% to 92% of the output tokens of a median call. The tasks are bug fixes, a CSV parser, an event-loop output order, a constraint schedule, a strict SemVer regex and more. Each task has a deterministic validator.

ConfigurationCallsMedian output tokensMedian reasoning tokensMedian shareLowest to highest callOutput / reasoning tokens, lowest to highest call
Claude Haiku 4.5 · Claude Code245,0644,55692%76% to 99%1899 to 9321 / 1452 to 8569
Claude Sonnet 5.5 · Claude Code241,05058555%0% to 96%176 to 3895 / 0 to 3060
Claude Opus 5.5 · Claude Code2494552955%30% to 96%323 to 2531 / 100 to 1778
Claude Opus 5.5 (high) · Claude Code241,05261454%36% to 96%285 to 4052 / 103 to 3301
Claude Fable 5.1 · Claude Code241,36688964%23% to 97%318 to 6465 / 75 to 5889
GPT-6.1 Sol (medium) · Codex CLI1633515046%11% to 87%237 to 1766 / 61 to 839
GPT-6.1 Sol (high) · Codex CLI1643622557%29% to 91%284 to 2569 / 144 to 1750

Haiku 4.5 had the highest median share in this run. Its separate medians were 4,556 reasoning tokens and 5,064 output tokens. The median share is not the ratio of those medians. It reported reasoning tokens on all 24 calls under the CLI default (we did not pass an effort flag). The other five configurations sat between 54% and 64%.

All output shares are calculations. The ranges are lowest to highest call, not confidence intervals.

Per-call ranges are wide. Sonnet 5.5 ran from 0% (7 calls reported zero reasoning) to 96%. Every pair of ranges overlaps. So read these medians as a description of this run, not as a ranking.

On short tasks, many calls report zero thinking tokens

Calculation
  • Eight hard tasks
  • Five short tasks (square)
Sorted by gap, largest first.
Claude Fable 5.1 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI

Gap labels, Five short tasks vs Eight hard tasks: Five short tasks is x percentage points higher (+) or lower (−) than Eight hard tasks, calculated from the two values shown; lines are the lowest–highest run (not an interval).

List-price calculation, not a run. 7 rows, 2 series: Eight hard tasks, Five short tasks. Eight hard tasks: highest Claude Haiku 4.5 · Claude Code 92% (range 76%–99%, n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 46% (range 11%–87%, n 16). All run ranges overlap. Five short tasks: highest Claude Haiku 4.5 · Claude Code 90% (range 73%–98%, n 15). Lowest Claude Fable 5.1 · Claude Code 0% (range 0%–74%, n 15). Not all run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 15–24 per row

Median call per configuration; five short tasks and eight hard tasks

Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.

Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices

The share table gives calculations and call ranges, not 95% intervals. Each median uses 10 to 15 calls.

The five short tasks are a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification. Here the picture splits.

ConfigurationCallsCalls with 0 reasoning tokensMedian shareLowest to highest call
Claude Haiku 4.5 · Claude Code15090%73% to 98%
Claude Sonnet 5.5 · Claude Code15120%0% to 73%
Claude Opus 5.5 · Claude Code1590%0% to 93%
Claude Opus 5.5 (low) · Claude Code1590%0% to 93%
Claude Opus 5.5 (high) · Claude Code15744%0% to 93%
Claude Fable 5.1 · Claude Code15120%0% to 74%
GPT-6.1 Sol (low) · Codex CLI10425%0% to 71%
GPT-6.1 Sol (medium) · Codex CLI15641%0% to 71%
GPT-6.1 Sol (high) · Codex CLI15658%0% to 76%

Sonnet 5.5, Opus 5.5 and Fable 5.1 reported 0 reasoning tokens on at least half of their calls (12, 9 and 12 of 15). So their median share is 0%. Haiku 4.5 reported reasoning on all 15 calls (median share 90%). GPT-6.1 Sol reported 0 on 4 to 6 calls per configuration. These counts vary with the task, the effort and the CLI default. They do not reveal all internal reasoning.

How much do Claude thinking tokens cost?

Calculation
  • Reasoning (output tokens)
  • Remaining output (visible-answer estimate)
  • Input (prompt, cache priced)
Bar length is the total; segments are its parts.
Claude Fable 5.1 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI
Claude Sonnet 5.5 · Claude Code

Totals are the sum of the parts shown. Shares are calculated from the same values.

List-price calculation, not a run. 7 rows, 3 series: Reasoning (output tokens), Remaining output (visible-answer estimate), Input (prompt, cache priced). Reasoning (output tokens): highest Claude Fable 5.1 · Claude Code $0.054 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.0023 (n 16). Remaining output (visible-answer estimate): highest Claude Fable 5.1 · Claude Code $0.019 (n 24). Lowest Claude Haiku 4.5 · Claude Code $0.0018 (n 24).

Notesn 16–24 per row

Mean per call on the hard tasks; the three parts add up to the call

Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.

Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices

All cost rows below are mean USD per call, a list-price calculation. Remaining output estimates the visible answer; the counters do not prove its exact meaning. Pooled cost share is summed reasoning cost ÷ summed total cost. It is not a median or a per-call range.

Thinking is not always most of the bill. Output share is not bill share. A call also sends input tokens, and they cost money. Input includes each CLI's system prompt and tool schemas. These receipts do not count CLI context apart from task input. Figures may not add exactly due to rounding.

At list price (a calculation), the reasoning part of a call cost a mean $0.0023 for GPT-6.1 Sol at medium effort. For Fable 5.1 it cost $0.0537. As a pooled share of list-price cost per configuration, reasoning ran from 9% (GPT-6.1 Sol, medium effort) to 80% (Haiku 4.5).

ConfigurationCallsReasoningRemaining outputInputTotalPooled reasoning share of cost
Claude Haiku 4.5 · Claude Code24$0.0245$0.0018$0.0045$0.030880%
Claude Sonnet 5.5 · Claude Code24$0.0067$0.0037$0.0040$0.014346%
Claude Opus 5.5 · Claude Code24$0.0125$0.0080$0.0077$0.028244%
Claude Opus 5.5 (high) · Claude Code24$0.0180$0.0080$0.0074$0.033454%
Claude Fable 5.1 · Claude Code24$0.0537$0.0187$0.0209$0.093358%
GPT-6.1 Sol (medium) · Codex CLI16$0.0023$0.0030$0.0203$0.02569%
GPT-6.1 Sol (high) · Codex CLI16$0.0041$0.0029$0.0081$0.015127%
ConfigurationReasoning / total cost, lowest to highest call (USD, calculation)Reasoning cost share, lowest to highest call
Claude Haiku 4.5 · Claude Code$0.0073 to $0.0428 / $0.0148 to $0.050540% to 90%
Claude Sonnet 5.5 · Claude Code$0.0000 to $0.0306 / $0.0051 to $0.04240% to 74%
Claude Opus 5.5 · Claude Code$0.0020 to $0.0356 / $0.0131 to $0.057210% to 74%
Claude Opus 5.5 (high) · Claude Code$0.0021 to $0.0660 / $0.0121 to $0.087717% to 77%
Claude Fable 5.1 · Claude Code$0.0037 to $0.2944 / $0.0322 to $0.33995% to 87%
GPT-6.1 Sol (medium) · Codex CLI$0.0006 to $0.0084 / $0.0099 to $0.04222% to 20%
GPT-6.1 Sol (high) · Codex CLI$0.0014 to $0.0175 / $0.0092 to $0.03078% to 60%

These are call ranges, not confidence intervals; sample sizes match the cost table.

Two things move the share.

The CLI and its input. GPT-6.1 Sol at medium effort sent $0.0203 of input per call. That dwarfs its $0.0023 of reasoning. The same model at high effort sent $0.0081. Input cost depends on each CLI's prompt and on cache hits, and these batches ran at different hours. We did not test why the two cells differ.

The output price. Fable 5.1's thinking cost 8.1x Sonnet 5.5's per call (a calculation from means). The output price explains a factor of 5.0 ($50 against $10 per million tokens). More reasoning tokens explain the rest, a factor of 1.6 (means).

Does more effort mean more thinking tokens?

  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

11 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 (high) · Claude Code 1,192 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 284 (n 16). Reasoning tokens: highest Claude Sonnet 5.5 (high) · Claude Code 745 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 63 (n 16).

Notesn = 16 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean "not reported". Their content is never captured. More tokens is not better or worse by itself.

Source: Effort ladder: the hard task set at each effort level

In these batches, higher effort had higher mean recorded reasoning counts. We paired the same eight tasks and compared mean reasoning tokens per call at low and at high effort. High was higher on 7 of 8 tasks for Sonnet, and on 8 of 8 tasks for Opus and for GPT-6.1 Sol. The pooled reasoning share of output rose at each step (low, medium, high) in all three ladders.

Model and routeMedian reasoning tokensPooled reasoning share of outputMean reasoning cost per callTotal cost per strict pass
Claude Sonnet 5.5 · Claude Code273 → 422 → 74553% → 62% → 73%$0.0043 → $0.0059 → $0.0094$0.0122 → $0.0135 → $0.0167
Claude Opus 5.5 · Claude Code87 → 518 → 61438% → 61% → 69%$0.0050 → $0.0134 → $0.0180$0.0212 → $0.0295 → $0.0337
GPT-6.1 Sol · Codex CLI63 → 150 → 22529% → 43% → 59%$0.0012 → $0.0023 → $0.0041$0.0128 → $0.0256 → $0.0151
Calculation
  • Reasoning cost per strict pass
  • Total cost per strict pass (square)
Sorted by gap, largest first.
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI
Claude Opus 5.5 (low) · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code

Gap labels, Total cost per strict pass vs Reasoning cost per strict pass: Total cost per strict pass is x% higher (+) or lower (−) than Reasoning cost per strict pass, calculated from the two values shown (the change counted from Reasoning cost per strict pass’s value).

List-price calculation, not a run. 11 rows, 2 series: Reasoning cost per strict pass, Total cost per strict pass. Reasoning cost per strict pass: highest Claude Opus 5.5 (high) · Claude Code $0.018 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI $0.0012 (n 16). Total cost per strict pass: highest Claude Opus 5.5 (high) · Claude Code $0.034 (n 16). Lowest Claude Sonnet 5.5 (low) · Claude Code $0.012 (n 16).

Notesn = 16 per row

List price ÷ strict passes; every cell is 8 tasks × 2 repetitions

Calculation, not a bill: reported tokens × list price, divided by the cell's strict passes; the calls ran on flat subscriptions. The effort-ladder cells: new calls plus reference cells reused from the hard head-to-head (Claude repetitions 1-2 only). "Default" means the effort flag was not passed. The total is the same value as the effort-ladder cost-per-pass chart. Effort levels are not the same scale across vendors, and the reference cells ran in a different batch and hour.

Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices

The arrows in the table mean low → medium → high. Each cell has 16 calls over eight tasks. Costs and pooled shares are calculations.

From low to high effort, the reasoning cost per call rose 2.2x for Sonnet, 3.6x for Opus and 3.4x for GPT-6.1 Sol (a calculation).

The total cost per strict pass for GPT-6.1 Sol is not a straight line: its medium cell cost the most. That is input cost, not thinking. The medium cell sent $0.0203 of input per call, as the cost table above shows.

Every effort cell passed 16 of 16 strictly (95% Wilson interval 81% to 100% each). The pass count stayed the same. The set has a ceiling. Higher effort had higher mean recorded reasoning cost in these batches. It does not show that more thinking never helps on harder work. The effort post has the pass rates and intervals.

Does thinking make a call slower?

Calculation
  • Claude Haiku 4.5 · Claude Code
  • Claude Sonnet 5.5 · Claude Code
  • Claude Opus 5.5 · Claude Code
  • Claude Fable 5.1 · Claude Code
  • GPT-6.1 Sol · Codex CLI

List-price calculation, not a run. 248 points: Total time per call (seconds) against Reasoning tokens per call. Reasoning tokens per call runs from 0 to 8,569; Total time per call (seconds) from 2.3 s to 92.2 s.

Notes

One point per call: 248 calls from the hard head-to-head and the effort ladder

Calculation, not a run: each point is one recorded call. Spearman rank correlations and slopes are in the table. Total time includes CLI start-up and the visible answer. Task, effort and batch can affect both counts and time. The plot shows association, not cause. 248 calls over 8 tasks; calls are not independent.

Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices

The recorded counts track time closely; they do not show cause. For each Claude model, reasoning tokens and total time had a rank correlation of 0.85 to 0.98 (24 to 80 calls each). On the same task, 1,000 more reasoning tokens went with 7.6 to 14.5 s more time (a calculation). Haiku 4.5's median call took 39.0 s (n = 24; call range 15.27 to 75.13 s, not an interval).

Model and routeCallsRank correlation (task sensitivity range)Seconds per 1,000 reasoning tokens (task sensitivity range)
Claude Haiku 4.5 · Claude Code240.98 (0.98 to 0.99)8.8 (8.6 to 9.1)
Claude Sonnet 5.5 · Claude Code720.92 (0.88 to 0.94)7.6 (7.4 to 7.8)
Claude Opus 5.5 · Claude Code800.85 (0.82 to 0.89)11.8 (5.7 to 12.7)
Claude Fable 5.1 · Claude Code240.92 (0.89 to 0.93)14.5 (14.4 to 22.6)
GPT-6.1 Sol · Codex CLI480.56 (0.34 to 0.65)36.4 (31.8 to 36.7)

The table ranges omit one whole task at a time. They are sensitivity ranges, not 95% confidence intervals.

For GPT-6.1 Sol in the Codex CLI, the recorded point estimate was lower: 0.56 (48 calls). Its within-task slope was 36.4 s per 1,000 reasoning tokens. That is 2.5 to 4.8 times the Claude Code point estimates (calculation). The task sensitivity ranges do not overlap with any Claude range. They are not confidence intervals and do not establish a tested ranking. Model, CLI, effort and batch effects remain. We did not test why.

This is a correlation, not a cause. Task, effort and batch effects may affect reasoning, visible output and time. Total time includes CLI start-up. Output tokens track time at least as closely as reasoning alone. For the Claude models, the rank correlation of output tokens with time was 0.94 to 0.99.

Do failed calls still cost thinking tokens?

Yes. Haiku 4.5 passed 11 of 24 hard calls strictly (46%, 95% Wilson interval 28% to 65%). The 13 calls that did not pass (54%, 95% Wilson interval 35% to 72%) held 61% of Haiku's reasoning cost (a calculation). Five were format misses with correct answers; eight had wrong answers.

At list price, Haiku 4.5 cost $0.0672 per strict pass. Of that, $0.0534 went to reasoning. Sonnet 5.5 passed 24 of 24 (95% Wilson interval 86% to 100%) at $0.0143 per strict pass. Both figures are calculations without an interval. The pass counts differ clearly: see the hard head-to-head for the intervals.

What this means for your settings

  1. Log reasoning tokens apart from output tokens. Both CLIs here reported them. A zero counter does not prove that no internal reasoning took place. If your tool does not, you cannot see this part of the bill. A reported 0 and a missing count are not the same.
  2. Price thinking at the output rate, then compare it with input. Reasoning cost = reasoning tokens × output price. In our hard-set configurations it was 9% to 80% of pooled list-price cost (calculation). Measure your own mix.
  3. Start at low effort and measure. On these tasks, each low-effort cell passed 16 of 16 (95% Wilson interval 81% to 100%). Low had the lowest mean recorded reasoning count. Raise it when your own checks fail at low.
  4. Check a model's default thinking before you call it cheap. Haiku 4.5 reported reasoning on every call under the CLI default. The list-price calculation was $0.0672 per strict pass here (11/24 passes), against $0.0143 for Sonnet 5.5 (24/24).
  5. Count failures in your cost per answer. On a metered plan, thinking on a failed call is still billed.
  6. Budget time and money. In Claude Code, 1,000 reasoning tokens went with 7.6 to 14.5 s on the same task.

Long, open agent work may differ. Our prompts ask for short replies in a strict format. We did not measure agent tasks with long visible output.

How we measured

  • Calls: 378 recorded calls from three runs. The hard head-to-head gave 152 completed calls (30 blocked Codex attempts never reached a model, so we leave them out). The effort ladder gave 96 new calls. The five-task head-to-head gave 130 calls. We made no new call and we retried none.
  • Sum check (calculation): the three consistency checks above. The ranges omit one whole task at a time. They show sensitivity to task mix, not 95% coverage. The median characters-per-token ratios were 1.94 for output minus reasoning in Claude Code and 2.99 in the Codex CLI. Calls with 0 reasoning had median ratios of 1.98 and 2.76. Among calls with reasoning, output including reasoning had median ratios of 0.45 and 1.36.
  • Share: reasoning ÷ output for each call. A cell gets the median of its call shares and a pooled share (all reasoning ÷ all output). Whiskers are the lowest and highest call, not intervals.
  • Cost: reported tokens × list price. Output prices per million tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50, GPT-6.1 Sol $10. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. Cache writes assume the one-hour price; the receipts do not state cache lifetime. Prices come from the recorded Anthropic table dated 2026-09-21 and OpenAI table dated 2026-10-03. A calculation, not a bill.
  • Time: the Spearman rank correlation of reasoning tokens with total time, and the least-squares slope within each task, over the hard head-to-head and effort-ladder calls. The ranges omit one whole task at a time. The within-task slope controls task means, but does not remove effort or batch effects.
  • Isolation: fresh empty folder, tools off, one turn, 300 s timeout (180 s for the short tasks), one call at a time per account. Claude Code ran with a 16,000-token output cap. The highest Claude call used 9,321 tokens, so no call hit it.

Caveats

  • Exploratory protocols. The hard-set protocol file was created after its first counted call. Both effort-ladder protocol files were created after their respective batches ended. The short-set file predates its first call, but its top-up amendment timing is unverified. Receipts preserve protocol text, but we cannot verify that all rules preceded inference.
  • Workload and host. These hand-built, tuned task sets are not random workload samples. The calls used one shared Mac and network; other work and provider load were not controlled. Wilson pass intervals assume independent calls. Repeats of eight tasks do not measure performance on new tasks.
  • Each CLI counts reasoning in its own way. We never see the thinking text, so we cannot check what a vendor counts. The sum check supports treating the counter as part of output. It does not prove counter semantics. Do not read the shares as a like-for-like measure of how much each model thinks.
  • Our prompts ask for short, strict replies. The measured shares apply to those replies. Long visible output (diffs, explanations) could change the share. We did not measure it.
  • Small cells. Each hard-set configuration has 16 to 24 calls over 8 tasks. Short-set cells have 10 to 15 calls over 5 tasks. On the hard tasks, every pair of output-share ranges overlaps, so those medians rank nothing.
  • Input depends on the CLI and the cache. The part of the cost that goes to reasoning is not a property of the model alone.
  • Ceiling. Every effort cell passed 16 of 16 (95% Wilson interval 81% to 100% each), so the data cannot say whether more thinking helps on harder work.
  • Correlation. The time link is not a cause. The calls repeat 8 tasks, so they are not independent.
  • Effort scales differ across vendors. The effort-ladder reference cells also ran in a different batch and hour than the new cells.
  • List price is not your bill. The calls ran on flat subscriptions. A price change moves every cost here and no token count.

Measure the thinking you pay for

Agent records reported token usage, calls, time and cost estimates for task steps, so you can inspect the available usage data. Try Agent.

Disclosure: I build Agent, the product behind these benchmarks. The calls analysed here went straight to Claude Code and the Codex CLI, not through Agent.

The data behind this post

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.