Explainer · Reasoning tokens
Reasoning tokens, explained: the hidden output you pay for
Definition
Reasoning tokens are the tokens a reasoning model writes to think before it writes its visible answer. Vendors generally bill them as output tokens, and a tool may report their count without showing their text. A short answer can hide a long chain of thought that you pay for.
Agent team · · 5 min read · Every number is from the public studies
Where reasoning tokens show up
Our receipts include reported reasoning tokens in the output-token count. We never capture their text.
Two limits apply to every number on this page:
- A reported 0 can mean "not reported", not "none". So our model and leaderboard pages do not list reasoning tokens as a fact of a model. Each count is what one route reported on one run. Token ranges below show the smallest and largest reported counts, not confidence intervals.
- Effort levels are not one scale across vendors. Default effort means we did not pass the flag, so the CLI chose a level it does not publish.
How many reasoning tokens we measured
On 8 hard tasks, the hard head-to-head recorded these median counts per call. Each Claude Code configuration has n = 24. Each Codex CLI configuration has n = 16.
- Claude Haiku 4.5: median 4,556 reasoning tokens (range 1,452 to 8,569). Median output: 5,064 tokens (1,899 to 9,321).
- Claude Sonnet 5.5: median 585 reasoning tokens (range 0 to 3,060). Median output: 1,050 tokens (176 to 3,895).
- GPT-6.1 Sol (high), Codex CLI: median 225 reasoning tokens (range 144 to 1,750). Median output: 436 tokens (284 to 2,569). At medium: 150 reasoning (61 to 839) and 335 output (237 to 1,766).
For Haiku, the reasoning median is about 90% of the output median (calculation: 4,556 / 5,064). For Sonnet it is about 56% (calculation: 585 / 1,050). These are ratios of medians, not the median share per call.
How effort changes the count
The effort ladder used the same 8 tasks, with n = 16 per configuration (2 calls per task). Five configurations reuse hard head-to-head calls from different batches and hours; provider load can differ. Claude reference cells use repetitions 1 and 2. Median reasoning tokens per call, with reported ranges:
- Sonnet 5.5: 273 at low (0 to 1,489), 422 at medium (0 to 2,093), 745 at high (0 to 3,610). Default: 668 (0 to 1,614).
- Opus 5.5: 87 at low (0 to 791), 518 at medium (109 to 1,993), 614 at high (103 to 3,301). Default: 538 (100 to 1,778).
- GPT-6.1 Sol, Codex CLI: 63 at low (0 to 458), 150 at medium (61 to 839), 225 at high (144 to 1,750).
Medians rose from low to medium to high for all three models, but their ranges overlap. Default is not a known step on that scale. High-to-low ratios were 2.7 for Sonnet, 7.1 for Opus and 3.6 for GPT-6.1 Sol. These are calculations: ratios of medians.
What they cost in time and money
Claude rows ran in Claude Code and GPT-6.1 Sol rows through the Codex CLI. Each row has n = 16.
| Configuration | Reasoning tokens (median) | Total time (median, range) | List-price cost per strict pass |
|---|---|---|---|
| Sonnet 5.5, low | 273 | 5.8 s (2.8 to 20.0) | $0.0122 |
| Sonnet 5.5, high | 745 | 8.8 s (2.9 to 35.8) | $0.0167 |
| Opus 5.5, low | 87 | 7.5 s (3.3 to 15.8) | $0.0212 |
| Opus 5.5, high | 614 | 10.1 s (3.6 to 63.0) | $0.0337 |
| GPT-6.1 Sol, low | 63 | 13.6 s (7.9 to 44.3) | $0.0128 |
| GPT-6.1 Sol, high | 225 | 18.1 s (11.7 to 92.2) | $0.0151 |
All six rows passed 16 of 16 strict calls (95% interval 81% to 100%). Each cost is a calculation: all reported tokens times their list rates, divided by strict passes. It includes input and output, with separate cache-read and cache-write rates. It is not the cost of reasoning alone. The calls ran on flat subscriptions; these figures are not per-call bills.
A range is not a confidence interval. For each model the low and high ranges overlap, so the medians describe this run and are not a tested ranking. Time did not always rise with effort: GPT-6.1 Sol took 13.6 s at low (range 7.9 to 44.3 s) and 13.1 s at medium (8.5 to 61.6 s).
Haiku's median total time on the hard tasks was 39.0 s (range 15.3 to 75.1 s, n = 24). Sonnet 5.5 took 7.7 s (2.3 to 34.8 s, n = 24). The ranges overlap, so this is a gap in medians.
Haiku's cost per strict pass was $0.0672 against $0.0143 for Sonnet (calculation, n = 24 each). Haiku passed 11 of 24 strict calls (46%, 95% interval 28% to 65%). Failed calls still cost; we did not separate their cost from reasoning cost.
How to see and control them
Read your CLI or API usage report. In the scheduler repair run, both GPT-6.1 Sol routes used medium effort, n = 3 each:
- Codex CLI: median 156 reasoning tokens (range 135 to 254).
- OpenAI API: median 267 (range 214 to 349).
- Claude Code: 3 calls, none with a separate reasoning count (reported-count n = 0). This means "not reported", not zero reasoning. Reporting differs by run.
- Control them with effort. Claude Code takes an effort flag and the Codex CLI takes a thread effort. See what reasoning effort is.
- Know the limits. This set hits a ceiling, so it cannot show that high effort never helps. These studies did not switch thinking off or cap a thinking budget. They support no claim about those settings.
- Measure on your own tasks. Log reasoning tokens for every call and compare cost per pass. The AI cost calculator turns token counts into list-price cost.
Frequently asked questions
Are reasoning tokens billed?
Vendors generally bill them at the output-token rate. Our subscription calls have no per-call bill; the costs above are list-price calculations. Check your vendor's exact rule.
Can I see a model's reasoning tokens?
Often you can see the count, not the text. The scheduler example above shows how reporting differs by route. We never captured reasoning text.
Why did Haiku take so long on hard tasks?
Haiku reported more reasoning tokens and a higher median time than Sonnet. The time ranges overlap, so they do not establish a speed ranking. These observations do not show cause; we did not test with thinking off.
Do more reasoning tokens mean better answers?
More tokens do not prove better answers. All 11 effort configurations passed 16 of 16 strict calls (95% interval 81% to 100% each). That ceiling hides possible quality differences. Sonnet passed 24 of 24 hard head-to-head calls (95% interval 86% to 100%). Its interval does not overlap Haiku's, but the models differ; token count does not explain the gap.
Watch the data
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks
All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.
Transcript
- Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
- 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
- All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
- More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
- List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
- 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Start at low effort and measure. Every call, interval and cost online.