How much of your AI bill is thinking tokens? Claude and GPT-6.1 Sol, measured
Thinking tokens: 46% to 92% of output per median call, 9% to 80% of pooled list-price cost. Calculation over recorded Claude and GPT-6.1 Sol calls.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks
All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.
Transcript
- Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
- 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
- All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
- More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
- List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
- 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Start at low effort and measure. Every call, interval and cost online.
TL;DR
- This is a calculation, not a new run. We re-read 378 calls that we had already recorded: Claude Haiku 4.5, Sonnet 5.5, Opus 5.5 and Fable 5.1 in Claude Code, and GPT-6.1 Sol in the Codex CLI. The calls ran on flat subscriptions, so every dollar figure is list-price arithmetic.
- We treat thinking tokens as part of output. Within one task and model, each extra reasoning token went with 0.99 extra output tokens in Claude Code (290 calls). In the Codex CLI it was 0.97 (88 calls). Leave-one-task-out ranges were 0.98 to 1.01 and 0.97 to 0.99; these are not 95% intervals. This calculation supports the accounting assumption; it does not prove how the CLIs count tokens.
- Share of output (calculation): on eight hard tasks (16 to 24 calls per configuration), reasoning was 46% to 92% of the output tokens of a median call. Haiku 4.5 had the highest median (92%). The other five configurations sat between 54% and 64%. The per-call ranges overlap, so this ranks no model.
- Share of the bill (calculation): reasoning was 9% to 80% of pooled list-price cost per configuration (16 to 24 calls each). Input tokens cost money too. The reasoning part of a call cost a mean $0.0023 (GPT-6.1 Sol, medium effort) to $0.0537 (Fable 5.1).
- Effort (calculation, 16 calls per cell): from low to high effort, the reasoning cost per call rose 2.2x for Sonnet, 3.6x for Opus and 3.4x for GPT-6.1 Sol. Every effort cell passed 16 of 16 (95% Wilson interval 81% to 100% each). The pass count stayed the same on this ceiling-limited set.
- Time (calculation): for each Claude model, reasoning tokens and total time had a rank correlation of 0.85 to 0.98 (24 to 80 calls each). On the same task, 1,000 more reasoning tokens went with 7.6 to 14.5 s more time.
- Failed calls pay too: 13 of 24 Haiku 4.5 calls did not pass strictly (54%, 95% Wilson interval 35% to 72%). They held 61% of Haiku's reasoning cost (calculation).
The study with every input: /benchmarks/thinking-token-bill.
What are thinking tokens, and are they extra?
A reasoning model can use internal reasoning tokens before it gives the answer. Vendors call these reasoning tokens or thinking tokens. This post uses both words for the same thing. Some tools report the count and hide the text. New to the setting that controls how much a model thinks? Read what reasoning effort is.
The bill question is simple. Do thinking tokens come on top of the output count, or inside it? If they are inside, you pay the output price for them. We check that assumption first, because every cost split in this post depends on it.
Are thinking tokens part of the output tokens?
The receipts are consistent with that accounting assumption. They do not prove how the CLIs count tokens. We checked three things on each CLI route. First, reasoning never exceeded output. Second, within one task and model, one more reasoning token went with about one more output token. Third, output minus reasoning had a similar characters-per-token ratio to replies with a zero reasoning counter. The table gives the first two results. The method section gives the third.
| CLI | Calls | Calls with reasoning above output | Extra output tokens per extra reasoning token (task sensitivity range) |
|---|---|---|---|
| Claude Code | 290 | 0 | 0.99 (0.982 to 1.009) |
| Codex CLI | 88 | 0 | 0.97 (0.965 to 0.986) |
Both CLIs report reasoning. 212 of 290 Claude Code calls and 68 of 88 Codex CLI calls reported more than 0 reasoning tokens. A reported 0 is the recorded counter, not proof of no internal reasoning. 98 calls reported 0. For this calculation, we keep each reported zero. A missing count is unknown, not zero. No call had a missing count.
How much of a call's output is thinking tokens?
On eight hard tasks, reasoning was 46% to 92% of the output tokens of a median call. The tasks are bug fixes, a CSV parser, an event-loop output order, a constraint schedule, a strict SemVer regex and more. Each task has a deterministic validator.
| Configuration | Calls | Median output tokens | Median reasoning tokens | Median share | Lowest to highest call | Output / reasoning tokens, lowest to highest call |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 24 | 5,064 | 4,556 | 92% | 76% to 99% | 1899 to 9321 / 1452 to 8569 |
| Claude Sonnet 5.5 · Claude Code | 24 | 1,050 | 585 | 55% | 0% to 96% | 176 to 3895 / 0 to 3060 |
| Claude Opus 5.5 · Claude Code | 24 | 945 | 529 | 55% | 30% to 96% | 323 to 2531 / 100 to 1778 |
| Claude Opus 5.5 (high) · Claude Code | 24 | 1,052 | 614 | 54% | 36% to 96% | 285 to 4052 / 103 to 3301 |
| Claude Fable 5.1 · Claude Code | 24 | 1,366 | 889 | 64% | 23% to 97% | 318 to 6465 / 75 to 5889 |
| GPT-6.1 Sol (medium) · Codex CLI | 16 | 335 | 150 | 46% | 11% to 87% | 237 to 1766 / 61 to 839 |
| GPT-6.1 Sol (high) · Codex CLI | 16 | 436 | 225 | 57% | 29% to 91% | 284 to 2569 / 144 to 1750 |
Haiku 4.5 had the highest median share in this run. Its separate medians were 4,556 reasoning tokens and 5,064 output tokens. The median share is not the ratio of those medians. It reported reasoning tokens on all 24 calls under the CLI default (we did not pass an effort flag). The other five configurations sat between 54% and 64%.
All output shares are calculations. The ranges are lowest to highest call, not confidence intervals.
Per-call ranges are wide. Sonnet 5.5 ran from 0% (7 calls reported zero reasoning) to 96%. Every pair of ranges overlaps. So read these medians as a description of this run, not as a ranking.
On short tasks, many calls report zero thinking tokens
The share table gives calculations and call ranges, not 95% intervals. Each median uses 10 to 15 calls.
The five short tasks are a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification. Here the picture splits.
| Configuration | Calls | Calls with 0 reasoning tokens | Median share | Lowest to highest call |
|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 15 | 0 | 90% | 73% to 98% |
| Claude Sonnet 5.5 · Claude Code | 15 | 12 | 0% | 0% to 73% |
| Claude Opus 5.5 · Claude Code | 15 | 9 | 0% | 0% to 93% |
| Claude Opus 5.5 (low) · Claude Code | 15 | 9 | 0% | 0% to 93% |
| Claude Opus 5.5 (high) · Claude Code | 15 | 7 | 44% | 0% to 93% |
| Claude Fable 5.1 · Claude Code | 15 | 12 | 0% | 0% to 74% |
| GPT-6.1 Sol (low) · Codex CLI | 10 | 4 | 25% | 0% to 71% |
| GPT-6.1 Sol (medium) · Codex CLI | 15 | 6 | 41% | 0% to 71% |
| GPT-6.1 Sol (high) · Codex CLI | 15 | 6 | 58% | 0% to 76% |
Sonnet 5.5, Opus 5.5 and Fable 5.1 reported 0 reasoning tokens on at least half of their calls (12, 9 and 12 of 15). So their median share is 0%. Haiku 4.5 reported reasoning on all 15 calls (median share 90%). GPT-6.1 Sol reported 0 on 4 to 6 calls per configuration. These counts vary with the task, the effort and the CLI default. They do not reveal all internal reasoning.
How much do Claude thinking tokens cost?
All cost rows below are mean USD per call, a list-price calculation. Remaining output estimates the visible answer; the counters do not prove its exact meaning. Pooled cost share is summed reasoning cost ÷ summed total cost. It is not a median or a per-call range.
Thinking is not always most of the bill. Output share is not bill share. A call also sends input tokens, and they cost money. Input includes each CLI's system prompt and tool schemas. These receipts do not count CLI context apart from task input. Figures may not add exactly due to rounding.
At list price (a calculation), the reasoning part of a call cost a mean $0.0023 for GPT-6.1 Sol at medium effort. For Fable 5.1 it cost $0.0537. As a pooled share of list-price cost per configuration, reasoning ran from 9% (GPT-6.1 Sol, medium effort) to 80% (Haiku 4.5).
| Configuration | Calls | Reasoning | Remaining output | Input | Total | Pooled reasoning share of cost |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 24 | $0.0245 | $0.0018 | $0.0045 | $0.0308 | 80% |
| Claude Sonnet 5.5 · Claude Code | 24 | $0.0067 | $0.0037 | $0.0040 | $0.0143 | 46% |
| Claude Opus 5.5 · Claude Code | 24 | $0.0125 | $0.0080 | $0.0077 | $0.0282 | 44% |
| Claude Opus 5.5 (high) · Claude Code | 24 | $0.0180 | $0.0080 | $0.0074 | $0.0334 | 54% |
| Claude Fable 5.1 · Claude Code | 24 | $0.0537 | $0.0187 | $0.0209 | $0.0933 | 58% |
| GPT-6.1 Sol (medium) · Codex CLI | 16 | $0.0023 | $0.0030 | $0.0203 | $0.0256 | 9% |
| GPT-6.1 Sol (high) · Codex CLI | 16 | $0.0041 | $0.0029 | $0.0081 | $0.0151 | 27% |
| Configuration | Reasoning / total cost, lowest to highest call (USD, calculation) | Reasoning cost share, lowest to highest call |
|---|---|---|
| Claude Haiku 4.5 · Claude Code | $0.0073 to $0.0428 / $0.0148 to $0.0505 | 40% to 90% |
| Claude Sonnet 5.5 · Claude Code | $0.0000 to $0.0306 / $0.0051 to $0.0424 | 0% to 74% |
| Claude Opus 5.5 · Claude Code | $0.0020 to $0.0356 / $0.0131 to $0.0572 | 10% to 74% |
| Claude Opus 5.5 (high) · Claude Code | $0.0021 to $0.0660 / $0.0121 to $0.0877 | 17% to 77% |
| Claude Fable 5.1 · Claude Code | $0.0037 to $0.2944 / $0.0322 to $0.3399 | 5% to 87% |
| GPT-6.1 Sol (medium) · Codex CLI | $0.0006 to $0.0084 / $0.0099 to $0.0422 | 2% to 20% |
| GPT-6.1 Sol (high) · Codex CLI | $0.0014 to $0.0175 / $0.0092 to $0.0307 | 8% to 60% |
These are call ranges, not confidence intervals; sample sizes match the cost table.
Two things move the share.
The CLI and its input. GPT-6.1 Sol at medium effort sent $0.0203 of input per call. That dwarfs its $0.0023 of reasoning. The same model at high effort sent $0.0081. Input cost depends on each CLI's prompt and on cache hits, and these batches ran at different hours. We did not test why the two cells differ.
The output price. Fable 5.1's thinking cost 8.1x Sonnet 5.5's per call (a calculation from means). The output price explains a factor of 5.0 ($50 against $10 per million tokens). More reasoning tokens explain the rest, a factor of 1.6 (means).
Does more effort mean more thinking tokens?
In these batches, higher effort had higher mean recorded reasoning counts. We paired the same eight tasks and compared mean reasoning tokens per call at low and at high effort. High was higher on 7 of 8 tasks for Sonnet, and on 8 of 8 tasks for Opus and for GPT-6.1 Sol. The pooled reasoning share of output rose at each step (low, medium, high) in all three ladders.
| Model and route | Median reasoning tokens | Pooled reasoning share of output | Mean reasoning cost per call | Total cost per strict pass |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 273 → 422 → 745 | 53% → 62% → 73% | $0.0043 → $0.0059 → $0.0094 | $0.0122 → $0.0135 → $0.0167 |
| Claude Opus 5.5 · Claude Code | 87 → 518 → 614 | 38% → 61% → 69% | $0.0050 → $0.0134 → $0.0180 | $0.0212 → $0.0295 → $0.0337 |
| GPT-6.1 Sol · Codex CLI | 63 → 150 → 225 | 29% → 43% → 59% | $0.0012 → $0.0023 → $0.0041 | $0.0128 → $0.0256 → $0.0151 |
The arrows in the table mean low → medium → high. Each cell has 16 calls over eight tasks. Costs and pooled shares are calculations.
From low to high effort, the reasoning cost per call rose 2.2x for Sonnet, 3.6x for Opus and 3.4x for GPT-6.1 Sol (a calculation).
The total cost per strict pass for GPT-6.1 Sol is not a straight line: its medium cell cost the most. That is input cost, not thinking. The medium cell sent $0.0203 of input per call, as the cost table above shows.
Every effort cell passed 16 of 16 strictly (95% Wilson interval 81% to 100% each). The pass count stayed the same. The set has a ceiling. Higher effort had higher mean recorded reasoning cost in these batches. It does not show that more thinking never helps on harder work. The effort post has the pass rates and intervals.
Does thinking make a call slower?
The recorded counts track time closely; they do not show cause. For each Claude model, reasoning tokens and total time had a rank correlation of 0.85 to 0.98 (24 to 80 calls each). On the same task, 1,000 more reasoning tokens went with 7.6 to 14.5 s more time (a calculation). Haiku 4.5's median call took 39.0 s (n = 24; call range 15.27 to 75.13 s, not an interval).
| Model and route | Calls | Rank correlation (task sensitivity range) | Seconds per 1,000 reasoning tokens (task sensitivity range) |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 24 | 0.98 (0.98 to 0.99) | 8.8 (8.6 to 9.1) |
| Claude Sonnet 5.5 · Claude Code | 72 | 0.92 (0.88 to 0.94) | 7.6 (7.4 to 7.8) |
| Claude Opus 5.5 · Claude Code | 80 | 0.85 (0.82 to 0.89) | 11.8 (5.7 to 12.7) |
| Claude Fable 5.1 · Claude Code | 24 | 0.92 (0.89 to 0.93) | 14.5 (14.4 to 22.6) |
| GPT-6.1 Sol · Codex CLI | 48 | 0.56 (0.34 to 0.65) | 36.4 (31.8 to 36.7) |
The table ranges omit one whole task at a time. They are sensitivity ranges, not 95% confidence intervals.
For GPT-6.1 Sol in the Codex CLI, the recorded point estimate was lower: 0.56 (48 calls). Its within-task slope was 36.4 s per 1,000 reasoning tokens. That is 2.5 to 4.8 times the Claude Code point estimates (calculation). The task sensitivity ranges do not overlap with any Claude range. They are not confidence intervals and do not establish a tested ranking. Model, CLI, effort and batch effects remain. We did not test why.
This is a correlation, not a cause. Task, effort and batch effects may affect reasoning, visible output and time. Total time includes CLI start-up. Output tokens track time at least as closely as reasoning alone. For the Claude models, the rank correlation of output tokens with time was 0.94 to 0.99.
Do failed calls still cost thinking tokens?
Yes. Haiku 4.5 passed 11 of 24 hard calls strictly (46%, 95% Wilson interval 28% to 65%). The 13 calls that did not pass (54%, 95% Wilson interval 35% to 72%) held 61% of Haiku's reasoning cost (a calculation). Five were format misses with correct answers; eight had wrong answers.
At list price, Haiku 4.5 cost $0.0672 per strict pass. Of that, $0.0534 went to reasoning. Sonnet 5.5 passed 24 of 24 (95% Wilson interval 86% to 100%) at $0.0143 per strict pass. Both figures are calculations without an interval. The pass counts differ clearly: see the hard head-to-head for the intervals.
What this means for your settings
- Log reasoning tokens apart from output tokens. Both CLIs here reported them. A zero counter does not prove that no internal reasoning took place. If your tool does not, you cannot see this part of the bill. A reported 0 and a missing count are not the same.
- Price thinking at the output rate, then compare it with input. Reasoning cost = reasoning tokens × output price. In our hard-set configurations it was 9% to 80% of pooled list-price cost (calculation). Measure your own mix.
- Start at low effort and measure. On these tasks, each low-effort cell passed 16 of 16 (95% Wilson interval 81% to 100%). Low had the lowest mean recorded reasoning count. Raise it when your own checks fail at low.
- Check a model's default thinking before you call it cheap. Haiku 4.5 reported reasoning on every call under the CLI default. The list-price calculation was $0.0672 per strict pass here (11/24 passes), against $0.0143 for Sonnet 5.5 (24/24).
- Count failures in your cost per answer. On a metered plan, thinking on a failed call is still billed.
- Budget time and money. In Claude Code, 1,000 reasoning tokens went with 7.6 to 14.5 s on the same task.
Long, open agent work may differ. Our prompts ask for short replies in a strict format. We did not measure agent tasks with long visible output.
How we measured
- Calls: 378 recorded calls from three runs. The hard head-to-head gave 152 completed calls (30 blocked Codex attempts never reached a model, so we leave them out). The effort ladder gave 96 new calls. The five-task head-to-head gave 130 calls. We made no new call and we retried none.
- Sum check (calculation): the three consistency checks above. The ranges omit one whole task at a time. They show sensitivity to task mix, not 95% coverage. The median characters-per-token ratios were 1.94 for output minus reasoning in Claude Code and 2.99 in the Codex CLI. Calls with 0 reasoning had median ratios of 1.98 and 2.76. Among calls with reasoning, output including reasoning had median ratios of 0.45 and 1.36.
- Share: reasoning ÷ output for each call. A cell gets the median of its call shares and a pooled share (all reasoning ÷ all output). Whiskers are the lowest and highest call, not intervals.
- Cost: reported tokens × list price. Output prices per million tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50, GPT-6.1 Sol $10. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. Cache writes assume the one-hour price; the receipts do not state cache lifetime. Prices come from the recorded Anthropic table dated 2026-09-21 and OpenAI table dated 2026-10-03. A calculation, not a bill.
- Time: the Spearman rank correlation of reasoning tokens with total time, and the least-squares slope within each task, over the hard head-to-head and effort-ladder calls. The ranges omit one whole task at a time. The within-task slope controls task means, but does not remove effort or batch effects.
- Isolation: fresh empty folder, tools off, one turn, 300 s timeout (180 s for the short tasks), one call at a time per account. Claude Code ran with a 16,000-token output cap. The highest Claude call used 9,321 tokens, so no call hit it.
Caveats
- Exploratory protocols. The hard-set protocol file was created after its first counted call. Both effort-ladder protocol files were created after their respective batches ended. The short-set file predates its first call, but its top-up amendment timing is unverified. Receipts preserve protocol text, but we cannot verify that all rules preceded inference.
- Workload and host. These hand-built, tuned task sets are not random workload samples. The calls used one shared Mac and network; other work and provider load were not controlled. Wilson pass intervals assume independent calls. Repeats of eight tasks do not measure performance on new tasks.
- Each CLI counts reasoning in its own way. We never see the thinking text, so we cannot check what a vendor counts. The sum check supports treating the counter as part of output. It does not prove counter semantics. Do not read the shares as a like-for-like measure of how much each model thinks.
- Our prompts ask for short, strict replies. The measured shares apply to those replies. Long visible output (diffs, explanations) could change the share. We did not measure it.
- Small cells. Each hard-set configuration has 16 to 24 calls over 8 tasks. Short-set cells have 10 to 15 calls over 5 tasks. On the hard tasks, every pair of output-share ranges overlaps, so those medians rank nothing.
- Input depends on the CLI and the cache. The part of the cost that goes to reasoning is not a property of the model alone.
- Ceiling. Every effort cell passed 16 of 16 (95% Wilson interval 81% to 100% each), so the data cannot say whether more thinking helps on harder work.
- Correlation. The time link is not a cause. The calls repeat 8 tasks, so they are not independent.
- Effort scales differ across vendors. The effort-ladder reference cells also ran in a different batch and hour than the new cells.
- List price is not your bill. The calls ran on flat subscriptions. A price change moves every cost here and no token count.
What to read next
- Does reasoning effort buy quality? for pass rates, time and cost per pass by effort.
- Effort ladder and hard head-to-head: the receipts behind this post.
- What reasoning effort is, in plain words.
- How to estimate your AI coding bill
- What if every call ran on Opus? for list-price repricing of recorded tokens.
Measure the thinking you pay for
Agent records reported token usage, calls, time and cost estimates for task steps, so you can inspect the available usage data. Try Agent.
Disclosure: I build Agent, the product behind these benchmarks. The calls analysed here went straight to Claude Code and the Codex CLI, not through Agent.