Explainer · Reasoning tokens

Reasoning tokens, explained: the hidden output you pay for

Definition

Reasoning tokens are the tokens a reasoning model writes to think before it writes its visible answer. Vendors generally bill them as output tokens, and a tool may report their count without showing their text. A short answer can hide a long chain of thought that you pay for.

Agent team · · 5 min read · Every number is from the public studies

Calculation
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Hover or focus a bar for its ratio to GPT-6.1 Sol (medium) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Haiku 4.5 · Claude Code 92% (range 76%–99%, n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 46% (range 11%–87%, n 16). All run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 16–24 per row

Median call: reasoning tokens ÷ output tokens. Whiskers: lowest and highest call (16 to 24 calls per configuration)

Calculation from reported tokens, not a run. Each call gives reasoning ÷ output; the bar is the median of those shares. Whiskers are the lowest and highest call. They are a range, not a confidence interval. They are wide, so the medians describe this run and rank nothing. The pooled share (all reasoning tokens ÷ all output tokens) is in the table. We treat reasoning tokens as part of output tokens; the consistency check supports this accounting assumption. Each CLI reports its own count.

Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices

Where reasoning tokens show up

Our receipts include reported reasoning tokens in the output-token count. We never capture their text.

Two limits apply to every number on this page:

  • A reported 0 can mean "not reported", not "none". So our model and leaderboard pages do not list reasoning tokens as a fact of a model. Each count is what one route reported on one run. Token ranges below show the smallest and largest reported counts, not confidence intervals.
  • Effort levels are not one scale across vendors. Default effort means we did not pass the flag, so the CLI chose a level it does not publish.

How many reasoning tokens we measured

On 8 hard tasks, the hard head-to-head recorded these median counts per call. Each Claude Code configuration has n = 24. Each Codex CLI configuration has n = 16.

  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 5,064 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16). Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 4,556 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 150 (n 16).

Notesn 16–24 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

  • Claude Haiku 4.5: median 4,556 reasoning tokens (range 1,452 to 8,569). Median output: 5,064 tokens (1,899 to 9,321).
  • Claude Sonnet 5.5: median 585 reasoning tokens (range 0 to 3,060). Median output: 1,050 tokens (176 to 3,895).
  • GPT-6.1 Sol (high), Codex CLI: median 225 reasoning tokens (range 144 to 1,750). Median output: 436 tokens (284 to 2,569). At medium: 150 reasoning (61 to 839) and 335 output (237 to 1,766).

For Haiku, the reasoning median is about 90% of the output median (calculation: 4,556 / 5,064). For Sonnet it is about 56% (calculation: 585 / 1,050). These are ratios of medians, not the median share per call.

How effort changes the count

The effort ladder used the same 8 tasks, with n = 16 per configuration (2 calls per task). Five configurations reuse hard head-to-head calls from different batches and hours; provider load can differ. Claude reference cells use repetitions 1 and 2. Median reasoning tokens per call, with reported ranges:

  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

11 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 (high) · Claude Code 1,192 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 284 (n 16). Reasoning tokens: highest Claude Sonnet 5.5 (high) · Claude Code 745 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 63 (n 16).

Notesn = 16 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean "not reported". Their content is never captured. More tokens is not better or worse by itself.

Source: Effort ladder: the hard task set at each effort level

  • Sonnet 5.5: 273 at low (0 to 1,489), 422 at medium (0 to 2,093), 745 at high (0 to 3,610). Default: 668 (0 to 1,614).
  • Opus 5.5: 87 at low (0 to 791), 518 at medium (109 to 1,993), 614 at high (103 to 3,301). Default: 538 (100 to 1,778).
  • GPT-6.1 Sol, Codex CLI: 63 at low (0 to 458), 150 at medium (61 to 839), 225 at high (144 to 1,750).

Medians rose from low to medium to high for all three models, but their ranges overlap. Default is not a known step on that scale. High-to-low ratios were 2.7 for Sonnet, 7.1 for Opus and 3.6 for GPT-6.1 Sol. These are calculations: ratios of medians.

What they cost in time and money

Claude rows ran in Claude Code and GPT-6.1 Sol rows through the Codex CLI. Each row has n = 16.

ConfigurationReasoning tokens (median)Total time (median, range)List-price cost per strict pass
Sonnet 5.5, low2735.8 s (2.8 to 20.0)$0.0122
Sonnet 5.5, high7458.8 s (2.9 to 35.8)$0.0167
Opus 5.5, low877.5 s (3.3 to 15.8)$0.0212
Opus 5.5, high61410.1 s (3.6 to 63.0)$0.0337
GPT-6.1 Sol, low6313.6 s (7.9 to 44.3)$0.0128
GPT-6.1 Sol, high22518.1 s (11.7 to 92.2)$0.0151

All six rows passed 16 of 16 strict calls (95% interval 81% to 100%). Each cost is a calculation: all reported tokens times their list rates, divided by strict passes. It includes input and output, with separate cache-read and cache-write rates. It is not the cost of reasoning alone. The calls ran on flat subscriptions; these figures are not per-call bills.

A range is not a confidence interval. For each model the low and high ranges overlap, so the medians describe this run and are not a tested ranking. Time did not always rise with effort: GPT-6.1 Sol took 13.6 s at low (range 7.9 to 44.3 s) and 13.1 s at medium (8.5 to 61.6 s).

Haiku's median total time on the hard tasks was 39.0 s (range 15.3 to 75.1 s, n = 24). Sonnet 5.5 took 7.7 s (2.3 to 34.8 s, n = 24). The ranges overlap, so this is a gap in medians.

Haiku's cost per strict pass was $0.0672 against $0.0143 for Sonnet (calculation, n = 24 each). Haiku passed 11 of 24 strict calls (46%, 95% interval 28% to 65%). Failed calls still cost; we did not separate their cost from reasoning cost.

How to see and control them

Read your CLI or API usage report. In the scheduler repair run, both GPT-6.1 Sol routes used medium effort, n = 3 each:

  • Codex CLI: median 156 reasoning tokens (range 135 to 254).
  • OpenAI API: median 267 (range 214 to 349).
  • Claude Code: 3 calls, none with a separate reasoning count (reported-count n = 0). This means "not reported", not zero reasoning. Reporting differs by run.
  • Output tokens
  • of which reasoning tokens (reported) (inner bar)
Claude Code CLI · Sonnet 5.5 · medium
Codex CLI · GPT-6.1 Sol · medium
OpenAI API · GPT-6.1 Sol · medium

3 rows, 2 series: Output tokens, Reasoning tokens (reported). Output tokens: highest Claude Code CLI · Sonnet 5.5 · medium 2,227 (n 3). Lowest Codex CLI · GPT-6.1 Sol · medium 1,181 (n 3). Reasoning tokens (reported): highest OpenAI API · GPT-6.1 Sol · medium 267 (n 3). Lowest Claude Code CLI · Sonnet 5.5 · medium 0 (n 0).

Notesn 0–3 per row

Median per run; reasoning tokens shown separately where reported

The Claude CLI does not report reasoning tokens separately; 0 there means "not reported", not "none".

Source: Provider explorer receipts: CLI vs API

  • Control them with effort. Claude Code takes an effort flag and the Codex CLI takes a thread effort. See what reasoning effort is.
  • Know the limits. This set hits a ceiling, so it cannot show that high effort never helps. These studies did not switch thinking off or cap a thinking budget. They support no claim about those settings.
  • Measure on your own tasks. Log reasoning tokens for every call and compare cost per pass. The AI cost calculator turns token counts into list-price cost.

Frequently asked questions

Are reasoning tokens billed?

Vendors generally bill them at the output-token rate. Our subscription calls have no per-call bill; the costs above are list-price calculations. Check your vendor's exact rule.

Can I see a model's reasoning tokens?

Often you can see the count, not the text. The scheduler example above shows how reporting differs by route. We never captured reasoning text.

Why did Haiku take so long on hard tasks?

Haiku reported more reasoning tokens and a higher median time than Sonnet. The time ranges overlap, so they do not establish a speed ranking. These observations do not show cause; we did not test with thinking off.

Do more reasoning tokens mean better answers?

More tokens do not prove better answers. All 11 effort configurations passed 16 of 16 strict calls (95% interval 81% to 100% each). That ceiling hides possible quality differences. Sonnet passed 24 of 24 hard head-to-head calls (95% interval 86% to 100%). Its interval does not overlap Haiku's, but the models differ; token count does not explain the gap.

Watch the data

Live story · 48 sDoes more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.

Transcript
  1. Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
  2. 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  3. All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  4. Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
  5. More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
  6. List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
  7. 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  8. Start at low effort and measure. Every call, interval and cost online.

The data behind this explainer

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.