• Reasoning Effort
  • Tokens
  • LLM pricing
  • Calculation

What thinking costs: reasoning tokens per call for Claude and GPT-6.1 Sol

Haiku 4.5 spent a median 4,556 reasoning tokens per hard call, Sonnet 5.5 585, GPT-6.1 Sol 150. List price per 1,000 calls: $22.78, $5.85, $1.50 (calculation).

TL;DR

  • Thinking takes a large share of reported output. On eight hard tasks, the median call spent 150 (GPT-6.1 Sol, medium effort) to 4,556 (Claude Haiku 4.5) reasoning tokens. The ratio of median reasoning to median output is about 45% to 90% (a calculation). The per-call ranges overlap, so this ranks no model.
  • At list price, median reported thinking costs $0.0015 to $0.0445 per call. Repeating those medians 1,000 times costs $1.50 to $44.45 (a calculation, not a forecast).
  • The effort ladder found no extra passes at higher effort. In the separate hard head-to-head, Sonnet, Opus and Fable passed 24/24 (95% interval 86% to 100%). Sol passed 16/16 (81% to 100%). Haiku had the highest median and passed 11/24 (28% to 65%).
  • Haiku had the higher calculated routing cost in these setups. Haiku cost $8.92 per 1,000 decisions and Sonnet $7.32 (a calculation). Haiku wrote 1,419 output tokens per decision and Sonnet wrote 107.
  • Our reading: start at low effort. Log reasoning tokens. Price them at the output rate.

How many reasoning tokens does each model spend?

A reasoning token is a token that a model spends on thinking before it answers. The CLI counts it inside the output tokens. We never capture its content. See what reasoning effort is.

  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 5,064 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16). Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 4,556 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 150 (n 16).

Notesn 16–24 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

The chart and table cover the hard head-to-head: eight tasks with strict validators. Each Claude cell has n = 24 calls and each Sol cell n = 16. The lowest-to-highest column shows reasoning-token ranges, not confidence intervals.

Route + modelMedian outputMedian reasoningLowest to highest callShare of output (calculation)
Haiku 4.5 (Claude Code)5,0644,5561,452 to 8,56990%
Fable 5.1 (Claude Code)1,36688975 to 5,88965%
Opus 5.5, high (Claude Code)1,052614103 to 3,30158%
Sonnet 5.5 (Claude Code)1,0505850 to 3,06056%
Opus 5.5 (Claude Code)945529100 to 1,77856%
Sol, high (Codex CLI)436225144 to 1,75052%
Sol, medium (Codex CLI)33515061 to 83945%

The Haiku median is 7.8 times the Sonnet median (calculation: 4,556 ÷ 585). The per-call ranges overlap, so this data does not rank the two. Vendors count reasoning differently, so compare inside one vendor.

On five short tasks (five-task head-to-head, n = 15 each), Haiku's median was 297 reasoning tokens (range 212 to 2,643). Sonnet's median was 0 (range 0 to 542). A reported zero can mean "not reported".

What do those tokens cost?

We price reasoning at each model's list output rate, in dollars per million tokens: Haiku $5, Sonnet $10, Opus $20, Fable $50 and Sol $10. Every figure below is a calculation. It covers reported reasoning tokens only. Answer, input and cache tokens cost extra.

Each row uses the same n and reasoning-token range as the table above. The 1,000-call columns scale this sample; they do not predict future calls.

Route + modelPer call, medianPer 1,000 calls, medianPer 1,000 calls, mean
Fable 5.1 (Claude Code)$0.0445$44.45$53.70
Haiku 4.5 (Claude Code)$0.0228$22.78$24.49
Opus 5.5, high (Claude Code)$0.0123$12.28$17.97
Opus 5.5 (Claude Code)$0.0106$10.57$12.53
Sonnet 5.5 (Claude Code)$0.0059$5.85$6.67
Sol, high (Codex CLI)$0.0023$2.25$4.11
Sol, medium (Codex CLI)$0.0015$1.50$2.27

Haiku costs half as much as Sonnet per output token. Yet the Haiku reasoning alone costs $0.0228 per call. That is 2.2 times the whole median output of a Sonnet call, $0.0105 (calculation).

The mean is higher than the median in every row, by 8% to 83%, because a few long calls pull it up. A bill adds up every call, so budget with the mean.

Did the thinking buy more passes?

The effort ladder found no pass-rate gain. The hard head-to-head cannot test what thinking caused: it changes the model and route too. Six of its seven cells passed every call (four at 24/24, two at 16/16). They hit the ceiling, so pass rate cannot separate them.

Haiku had the highest median reasoning and passed 11/24 strictly (46%, 28% to 65%). That interval does not overlap the others (86% to 100% and 81% to 100%). We cannot say if thinking helped or hurt Haiku. These runs have no thinking-off arm.

The effort ladder varies effort on the same eight tasks (n = 16 per cell). It reuses some cells from another batch, so batch conditions can also differ. The ranges below are reasoning tokens per call, not confidence intervals.

  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

11 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 (high) · Claude Code 1,192 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 284 (n 16). Reasoning tokens: highest Claude Sonnet 5.5 (high) · Claude Code 745 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 63 (n 16).

Notesn = 16 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean "not reported". Their content is never captured. More tokens is not better or worse by itself.

Source: Effort ladder: the hard task set at each effort level

Route + modelMedian reasoning, low → highRange, low; highHigh ÷ low (calculation)Per 1,000 calls at each median (calculation)
Sonnet 5.5 (Claude Code)273 → 7450 to 1,489; 0 to 3,6102.7 times$2.73 → $7.45
Opus 5.5 (Claude Code)87 → 6140 to 791; 103 to 3,3017.1 times$1.74 → $12.28
Sol (Codex CLI)63 → 2250 to 458; 144 to 1,7503.6 times$0.63 → $2.25

All three models passed 16/16 at low and at high effort (95% interval 81% to 100% each). None of the 130 metric rows in the 16 effort comparisons has a winner: 31 are ties and 99 are unclear. A perfect 16/16 has a lower interval bound about 19 percentage points below 100% (calculation). That is not an interval for the difference between two arms. Low effort passed every observed call. Harder work may need more.

Why did the smaller model cost more per routing decision?

The routing study gave each router 82 typed decisions. Haiku answered 73 exactly (89%, 80% to 94%). Sonnet answered 77 (94%, 87% to 97%). The intervals overlap, so accuracy does not separate them. The calculated cost totals differ; they have no uncertainty interval.

Haiku cost $8.92 per 1,000 decisions and Sonnet cost $7.32 (a calculation), although Haiku has half the price per token. Per decision, Haiku wrote a mean of 1,419 output tokens and Sonnet wrote 107. That is 13.3 times as many (calculation from unrounded receipts; the displayed means are rounded). The output ranges were 453 to 4,731 for Haiku and 52 to 221 for Sonnet.

The receipts show a mean of 1,101 thinking tokens in the Haiku output (78%, a calculation), against 1.5 for Sonnet. The thinking-token ranges were 285 to 4,443 for Haiku and 0 to 63 for Sonnet. At the Haiku output price, those thinking tokens alone cost $5.50 per 1,000 decisions (calculation).

The two arms differ. Haiku ran with the CLI default thinking. Sonnet ran at low effort, as production asks. This routing study did not test Haiku with thinking off. The two arms also reported different cache counters. Sonnet wrote to the one-hour cache, priced at twice the input rate.

The dataset’s $5.00 Sonnet figure assumes five-minute writes at 1.25 times the input rate. We use the recorded one-hour writes for the $7.32 calculation here. The total cost gap cannot isolate the effect of thinking or model size. These runs cannot show what caused the latency gap.

What to do (our reading)

  1. Start at low effort. On the hard set, low effort passed 16/16 for Sonnet, Opus and Sol with lower median reasoning counts than high effort. Their per-call reasoning ranges overlap. Raise it when your own checks fail.
  2. Log reasoning tokens for every call. Track the median and the mean. Price them at the output rate. The AI cost calculator takes your own counts.
  3. Compare tokens per call, not price per token. The half-price model wrote 13.3 times the output tokens.
  4. Test a thinking limit on small models for your own tasks. These four studies did not test one. This is a test to run, not a result here.

How we measured

  • Data: the four studies in the front matter. Claude Code ran the Claude models and the Codex CLI ran Sol, so each row is a route + model pair.
  • Setup: each hard-set call used a fresh empty folder, tools off, one turn and a 300 s timeout. We retried and trimmed nothing. Thirty Codex attempts never reached a model and are not counted.
  • Tokens: as the CLI reports them. Hard-set medians come from the dataset. Per-call ranges and means come from the hard head-to-head receipts. Routing token means and the CLI-reported cost come from the routing receipts. Effort ranges also use the effort-ladder receipts.
  • Calculations: share = median reasoning ÷ median output. Cost = tokens × list output price ÷ 1,000,000 (cost thought experiments). We calculate costs and ratios from unrounded receipt values, then round half up. Displayed token medians use whole tokens. Shares are ratios of medians, not median per-call shares. Intervals are Wilson 95% intervals.

Caveats

  • Ceiling. Six of seven hard-set cells passed every call. These sets cannot separate thinking levels by quality.
  • Zeros. Seven of Sonnet's 24 hard calls reported 0 reasoning tokens. A 0 can mean "not reported".
  • Reused cells. In the effort ladder, the Opus high and Sol medium/high cells come from the hard head-to-head. Claude reference cells use repetitions 1 and 2 only. They ran in another batch than the low cells.
  • Routing. The costs have no interval. The corrected Sonnet calculation uses the recorded one-hour cache writes. The routing chart still assumes five-minute writes, so we omit it here.
  • Small samples. Each task cell has 2 to 3 calls. Costs are calculations, not bills: the calls ran on subscriptions.

See what thinking costs in your own work

Agent records the model, the effort, the tokens and the validation result of every step, so you can compare token use with validation results. Try Agent. Then measure it on your own work.

The data behind this post

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.