How much of an AI bill is thinking? Reasoning tokens by model and effort
Across 378 recorded calls of Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol, how many output tokens come from reasoning (thinking)? What do they cost at list price per call and per strict pass, and do they track time? The calls cover eight hard tasks, five short tasks and several efforts.
Published · 5 charts · Download the data or a carousel
92%
Calculationn = 24
Highest median reasoning share of output tokens, hard tasks (calculation) · (Claude Haiku 4.5 · Claude Code; range 76% to 99%)
46%
Calculationn = 16
Lowest median reasoning share of output tokens, hard tasks (calculation) · (GPT-6.1 Sol (medium) · Codex CLI; range 11% to 87%)
The answer
On eight hard tasks, reasoning tokens were 46% to 92% of the output tokens of a median call (16 to 24 calls per configuration). Haiku 4.5 had the highest median (92%), GPT-6.1 Sol (medium) had the lowest (46%) and the other 5 sat between 54% and 64%. Per-call ranges are wide (Sonnet 5.5 0% to 96%). The ranges of every pair overlap, so this run ranks no configuration. Output share is not bill share. At list price (a calculation) the reasoning part of a call cost a mean $0.0023 (GPT-6.1 Sol (medium)) to $0.0537 (Fable 5.1). That was 9% (GPT-6.1 Sol (medium)) to 80% (Haiku 4.5) of the total list-price cost in each configuration. This pooled share divides summed reasoning cost by summed total cost. Input tokens cost money too. The calls ran on flat subscriptions, so this is not a bill. Fable 5.1's thinking cost 8.1x Sonnet 5.5's per call. The output price explains a factor of 5.0 ($50 against $10 per million tokens). More reasoning tokens explain the rest, a factor of 1.6 (means). Higher effort settings had higher mean recorded reasoning counts in these batches. We paired the same eight tasks. At high effort, mean reasoning tokens per call were higher than at low effort on these tasks: Sonnet 5.5 7 of 8, Opus 5.5 8 of 8, GPT-6.1 Sol 8 of 8. The pooled reasoning share rose from low to high effort. Sonnet 5.5 53% to 73%. Opus 5.5 38% to 69%. GPT-6.1 Sol 29% to 59%. It rose at each step (low, medium, high) in all 3 ladders. The reasoning cost per call rose 2.2x for Sonnet 5.5, 3.6x for Opus 5.5 and 3.4x for GPT-6.1 Sol (a calculation). Every effort cell passed 16/16 strictly. The pass count stayed the same on this set, which has a ceiling (95% Wilson 81% to 100% per cell). Recorded reasoning counts correlated with total time. Within each Claude model, the rank correlation between reasoning tokens and total time was 0.85 to 0.98 (Spearman, 24 to 80 calls each). On the same task, 1,000 more reasoning tokens went with 7.6 to 14.5 s more time (a calculation). For GPT-6.1 Sol in the Codex CLI, the recorded rank correlation was: 0.56 (48 calls). Thinking costs money whether or not the call passes. Haiku 4.5 passed 11/24 strictly (95% Wilson 28% to 65%). The 13 calls that did not pass held 61% of its reasoning cost (a calculation).
Key numbers
378
Recorded calls analysed (no new calls)
(152 hard head-to-head, 96 new effort-ladder, 130 five-task head-to-head) · n = 378
98
Calls that reported 0 reasoning tokens (kept as 0, not "not reported")
of 378; 0 calls had no usable count · n = 378
0.99
Sum check, Claude Code: output tokens gained per extra reasoning token, same task and model (calculation)
(leave-one-task-out range 0.98 to 1.01; 290 calls) · n = 290
0.97
Sum check, Codex CLI: output tokens gained per extra reasoning token, same task and model (calculation)
(leave-one-task-out range 0.97 to 0.99; 88 calls) · n = 88
$0.0537
Highest mean list-price cost of reasoning per call, hard tasks (calculation)
(Claude Fable 5.1 · Claude Code; call range $0.0037 to $0.2944) · n = 24
$0.0023
Lowest mean list-price cost of reasoning per call, hard tasks (calculation)
(GPT-6.1 Sol (medium) · Codex CLI; call range $0.0006 to $0.0084) · n = 16
80%
Highest pooled reasoning share of total list-price cost, hard tasks (calculation)
(Claude Haiku 4.5 · Claude Code; call range 40% to 90%) · n = 24
9%
Lowest pooled reasoning share of total list-price cost, hard tasks (calculation)
(GPT-6.1 Sol (medium) · Codex CLI; call range 2% to 20%) · n = 16
8.1x
Reasoning cost per call, Fable 5.1 ÷ Sonnet 5.5, hard tasks (calculation, means)
($0.0537 vs $0.0067) · n = 24
61%
Share of Haiku 4.5 reasoning cost spent on calls that did not pass (calculation)
(13 of 24 calls did not pass strictly) · n = 24
7
Tasks where mean reasoning tokens per call were higher at high than at low effort, Claude Sonnet 5.5 · Claude Code (paired, same tasks; calculation)
of 8 tasks · n = 8
8
Tasks where mean reasoning tokens per call were higher at high than at low effort, Claude Opus 5.5 · Claude Code (paired, same tasks; calculation)
of 8 tasks · n = 8
8
Tasks where mean reasoning tokens per call were higher at high than at low effort, GPT-6.1 Sol · Codex CLI (paired, same tasks; calculation)
of 8 tasks · n = 8
2.2x
Reasoning cost per call, high ÷ low effort, Claude Sonnet 5.5 · Claude Code (calculation, means)
($0.0043 at low, $0.0094 at high) · n = 16
3.6x
Reasoning cost per call, high ÷ low effort, Claude Opus 5.5 · Claude Code (calculation, means)
($0.0050 at low, $0.0180 at high) · n = 16
3.4x
Reasoning cost per call, high ÷ low effort, GPT-6.1 Sol · Codex CLI (calculation, means)
($0.0012 at low, $0.0041 at high) · n = 16
0.85
Spearman, reasoning tokens vs total time, per Claude model: lowest (calculation)
to 0.98 across 4 Claude models · n = 200
7.6
Seconds per 1,000 reasoning tokens on the same task, per Claude model: lowest (calculation)
to 14.5 s across 4 Claude models · n = 200
0.56
Spearman, reasoning tokens vs total time, GPT-6.1 Sol in Codex CLI (calculation)
(leave-one-task-out range 0.34 to 0.65; 48 calls) · n = 48
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
- Reasoning (output tokens)
- Remaining output (visible-answer estimate)
- Input (prompt, cache priced)
Totals are the sum of the parts shown. Shares are calculated from the same values.
| Item | Reasoning (output tokens) | Remaining output (visible-answer estimate) | Input (prompt, cache priced) | n |
|---|---|---|---|---|
| Claude Fable 5.1 · Claude Code | $0.054 | $0.019 | $0.021 | 24 |
| Claude Opus 5.5 (high) · Claude Code | $0.018 | $0.008 | $0.0074 | 24 |
| Claude Haiku 4.5 · Claude Code | $0.024 | $0.0018 | $0.0045 | 24 |
| Claude Opus 5.5 · Claude Code | $0.013 | $0.008 | $0.0077 | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.0023 | $0.003 | $0.02 | 16 |
| GPT-6.1 Sol (high) · Codex CLI | $0.0041 | $0.0029 | $0.0081 | 16 |
| Claude Sonnet 5.5 · Claude Code | $0.0067 | $0.0037 | $0.004 | 24 |
List-price calculation, not a run. 7 rows, 3 series: Reasoning (output tokens), Remaining output (visible-answer estimate), Input (prompt, cache priced). Reasoning (output tokens): highest Claude Fable 5.1 · Claude Code $0.054 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.0023 (n 16). Remaining output (visible-answer estimate): highest Claude Fable 5.1 · Claude Code $0.019 (n 24). Lowest Claude Haiku 4.5 · Claude Code $0.0018 (n 24).
Notesn 16–24 per row
Mean per call on the hard tasks; the three parts add up to the call
Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
- Reasoning cost per strict pass
- Total cost per strict pass (square)
Gap labels, Total cost per strict pass vs Reasoning cost per strict pass: Total cost per strict pass is x% higher (+) or lower (−) than Reasoning cost per strict pass, calculated from the two values shown (the change counted from Reasoning cost per strict pass’s value).
| Item | Reasoning cost per strict pass | Total cost per strict pass | n |
|---|---|---|---|
| Claude Sonnet 5.5 (low) · Claude Code | $0.0043 | $0.012 | 16 |
| Claude Sonnet 5.5 (medium) · Claude Code | $0.0059 | $0.014 | 16 |
| Claude Sonnet 5.5 (high) · Claude Code | $0.0094 | $0.017 | 16 |
| Claude Sonnet 5.5 · Claude Code | $0.0063 | $0.014 | 16 |
| Claude Opus 5.5 (low) · Claude Code | $0.005 | $0.021 | 16 |
| Claude Opus 5.5 (medium) · Claude Code | $0.013 | $0.029 | 16 |
| Claude Opus 5.5 (high) · Claude Code | $0.018 | $0.034 | 16 |
| Claude Opus 5.5 · Claude Code | $0.013 | $0.029 | 16 |
| GPT-6.1 Sol (low) · Codex CLI | $0.0012 | $0.013 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.0023 | $0.026 | 16 |
| GPT-6.1 Sol (high) · Codex CLI | $0.0041 | $0.015 | 16 |
List-price calculation, not a run. 11 rows, 2 series: Reasoning cost per strict pass, Total cost per strict pass. Reasoning cost per strict pass: highest Claude Opus 5.5 (high) · Claude Code $0.018 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI $0.0012 (n 16). Total cost per strict pass: highest Claude Opus 5.5 (high) · Claude Code $0.034 (n 16). Lowest Claude Sonnet 5.5 (low) · Claude Code $0.012 (n 16).
Notesn = 16 per row
List price ÷ strict passes; every cell is 8 tasks × 2 repetitions
Calculation, not a bill: reported tokens × list price, divided by the cell's strict passes; the calls ran on flat subscriptions. The effort-ladder cells: new calls plus reference cells reused from the hard head-to-head (Claude repetitions 1-2 only). "Default" means the effort flag was not passed. The total is the same value as the effort-ladder cost-per-pass chart. Effort levels are not the same scale across vendors, and the reference cells ran in a different batch and hour.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
- Claude Haiku 4.5 · Claude Code
- Claude Sonnet 5.5 · Claude Code
- Claude Opus 5.5 · Claude Code
- Claude Fable 5.1 · Claude Code
- GPT-6.1 Sol · Codex CLI
| Point | Series | Reasoning tokens per call | Total time per call (seconds) |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code · merge-ranges · rep 1 | Claude Haiku 4.5 · Claude Code | 1,977 | 18.5 s |
| Claude Haiku 4.5 · Claude Code · day-hours · rep 1 | Claude Haiku 4.5 · Claude Code | 4,372 | 39 s |
| Claude Haiku 4.5 · Claude Code · csv-parse · rep 1 | Claude Haiku 4.5 · Claude Code | 3,590 | 33.8 s |
| Claude Haiku 4.5 · Claude Code · event-loop-order · rep 1 | Claude Haiku 4.5 · Claude Code | 3,781 | 26 s |
| Claude Haiku 4.5 · Claude Code · talk-schedule · rep 1 | Claude Haiku 4.5 · Claude Code | 6,236 | 54.6 s |
| Claude Haiku 4.5 · Claude Code · semver-regex · rep 1 | Claude Haiku 4.5 · Claude Code | 7,498 | 64.5 s |
| Claude Haiku 4.5 · Claude Code · money-refactor · rep 1 | Claude Haiku 4.5 · Claude Code | 1,452 | 15.9 s |
| Claude Haiku 4.5 · Claude Code · q1-sql · rep 1 | Claude Haiku 4.5 · Claude Code | 4,305 | 38.9 s |
| Claude Haiku 4.5 · Claude Code · merge-ranges · rep 2 | Claude Haiku 4.5 · Claude Code | 2,606 | 24.7 s |
| Claude Haiku 4.5 · Claude Code · day-hours · rep 2 | Claude Haiku 4.5 · Claude Code | 3,515 | 26.2 s |
| Claude Haiku 4.5 · Claude Code · csv-parse · rep 2 | Claude Haiku 4.5 · Claude Code | 4,497 | 39 s |
| Claude Haiku 4.5 · Claude Code · event-loop-order · rep 2 | Claude Haiku 4.5 · Claude Code | 6,575 | 54 s |
| Claude Haiku 4.5 · Claude Code · talk-schedule · rep 2 | Claude Haiku 4.5 · Claude Code | 7,954 | 68.8 s |
| Claude Haiku 4.5 · Claude Code · semver-regex · rep 2 | Claude Haiku 4.5 · Claude Code | 6,538 | 54.5 s |
| Claude Haiku 4.5 · Claude Code · money-refactor · rep 2 | Claude Haiku 4.5 · Claude Code | 1,955 | 15.3 s |
| Claude Haiku 4.5 · Claude Code · q1-sql · rep 2 | Claude Haiku 4.5 · Claude Code | 6,691 | 56.1 s |
| Claude Haiku 4.5 · Claude Code · merge-ranges · rep 3 | Claude Haiku 4.5 · Claude Code | 2,703 | 21.2 s |
| Claude Haiku 4.5 · Claude Code · day-hours · rep 3 | Claude Haiku 4.5 · Claude Code | 4,614 | 40.6 s |
| Claude Haiku 4.5 · Claude Code · csv-parse · rep 3 | Claude Haiku 4.5 · Claude Code | 8,569 | 75.1 s |
| Claude Haiku 4.5 · Claude Code · event-loop-order · rep 3 | Claude Haiku 4.5 · Claude Code | 7,065 | 56.7 s |
| Claude Haiku 4.5 · Claude Code · talk-schedule · rep 3 | Claude Haiku 4.5 · Claude Code | 4,922 | 38 s |
| Claude Haiku 4.5 · Claude Code · semver-regex · rep 3 | Claude Haiku 4.5 · Claude Code | 7,994 | 67 s |
| Claude Haiku 4.5 · Claude Code · money-refactor · rep 3 | Claude Haiku 4.5 · Claude Code | 2,221 | 17.2 s |
| Claude Haiku 4.5 · Claude Code · q1-sql · rep 3 | Claude Haiku 4.5 · Claude Code | 5,934 | 47.1 s |
| Claude Sonnet 5.5 · Claude Code · merge-ranges · rep 1 | Claude Sonnet 5.5 · Claude Code | 0 | 2.9 s |
| Claude Sonnet 5.5 · Claude Code · day-hours · rep 1 | Claude Sonnet 5.5 · Claude Code | 1,249 | 21.6 s |
| Claude Sonnet 5.5 · Claude Code · csv-parse · rep 1 | Claude Sonnet 5.5 · Claude Code | 551 | 8.9 s |
| Claude Sonnet 5.5 · Claude Code · event-loop-order · rep 1 | Claude Sonnet 5.5 · Claude Code | 1,064 | 8.2 s |
| Claude Sonnet 5.5 · Claude Code · talk-schedule · rep 1 | Claude Sonnet 5.5 · Claude Code | 716 | 7.4 s |
| Claude Sonnet 5.5 · Claude Code · semver-regex · rep 1 | Claude Sonnet 5.5 · Claude Code | 0 | 2.3 s |
| Claude Sonnet 5.5 · Claude Code · money-refactor · rep 1 | Claude Sonnet 5.5 · Claude Code | 0 | 3.6 s |
| Claude Sonnet 5.5 · Claude Code · q1-sql · rep 1 | Claude Sonnet 5.5 · Claude Code | 799 | 8.9 s |
| Claude Sonnet 5.5 · Claude Code · merge-ranges · rep 2 | Claude Sonnet 5.5 · Claude Code | 0 | 3.1 s |
| Claude Sonnet 5.5 · Claude Code · day-hours · rep 2 | Claude Sonnet 5.5 · Claude Code | 1,614 | 19.6 s |
| Claude Sonnet 5.5 · Claude Code · csv-parse · rep 2 | Claude Sonnet 5.5 · Claude Code | 722 | 9.6 s |
| Claude Sonnet 5.5 · Claude Code · event-loop-order · rep 2 | Claude Sonnet 5.5 · Claude Code | 1,197 | 9.9 s |
| Claude Sonnet 5.5 · Claude Code · talk-schedule · rep 2 | Claude Sonnet 5.5 · Claude Code | 736 | 7.8 s |
| Claude Sonnet 5.5 · Claude Code · semver-regex · rep 2 | Claude Sonnet 5.5 · Claude Code | 267 | 3.7 s |
| Claude Sonnet 5.5 · Claude Code · money-refactor · rep 2 | Claude Sonnet 5.5 · Claude Code | 545 | 7.2 s |
| Claude Sonnet 5.5 · Claude Code · q1-sql · rep 2 | Claude Sonnet 5.5 · Claude Code | 619 | 8.5 s |
| Claude Sonnet 5.5 · Claude Code · merge-ranges · rep 3 | Claude Sonnet 5.5 · Claude Code | 0 | 2.4 s |
| Claude Sonnet 5.5 · Claude Code · day-hours · rep 3 | Claude Sonnet 5.5 · Claude Code | 3,060 | 34.8 s |
| Claude Sonnet 5.5 · Claude Code · csv-parse · rep 3 | Claude Sonnet 5.5 · Claude Code | 547 | 10.6 s |
| Claude Sonnet 5.5 · Claude Code · event-loop-order · rep 3 | Claude Sonnet 5.5 · Claude Code | 1,177 | 9.6 s |
| Claude Sonnet 5.5 · Claude Code · talk-schedule · rep 3 | Claude Sonnet 5.5 · Claude Code | 657 | 7.7 s |
| Claude Sonnet 5.5 · Claude Code · semver-regex · rep 3 | Claude Sonnet 5.5 · Claude Code | 0 | 2.5 s |
| Claude Sonnet 5.5 · Claude Code · money-refactor · rep 3 | Claude Sonnet 5.5 · Claude Code | 0 | 3.4 s |
| Claude Sonnet 5.5 · Claude Code · q1-sql · rep 3 | Claude Sonnet 5.5 · Claude Code | 477 | 7.6 s |
| Claude Sonnet 5.5 (low) · Claude Code · merge-ranges · rep 1 | Claude Sonnet 5.5 · Claude Code | 0 | 2.8 s |
| Claude Sonnet 5.5 (medium) · Claude Code · merge-ranges · rep 1 | Claude Sonnet 5.5 · Claude Code | 0 | 2.7 s |
| Claude Sonnet 5.5 (high) · Claude Code · merge-ranges · rep 1 | Claude Sonnet 5.5 · Claude Code | 0 | 2.9 s |
| Claude Sonnet 5.5 (low) · Claude Code · day-hours · rep 1 | Claude Sonnet 5.5 · Claude Code | 1,489 | 20 s |
| Claude Sonnet 5.5 (medium) · Claude Code · day-hours · rep 1 | Claude Sonnet 5.5 · Claude Code | 1,941 | 22.5 s |
| Claude Sonnet 5.5 (high) · Claude Code · day-hours · rep 1 | Claude Sonnet 5.5 · Claude Code | 2,597 | 27.2 s |
| Claude Sonnet 5.5 (low) · Claude Code · csv-parse · rep 1 | Claude Sonnet 5.5 · Claude Code | 0 | 4.3 s |
| Claude Sonnet 5.5 (medium) · Claude Code · csv-parse · rep 1 | Claude Sonnet 5.5 · Claude Code | 423 | 7.8 s |
| Claude Sonnet 5.5 (high) · Claude Code · csv-parse · rep 1 | Claude Sonnet 5.5 · Claude Code | 770 | 10 s |
| Claude Sonnet 5.5 (low) · Claude Code · event-loop-order · rep 1 | Claude Sonnet 5.5 · Claude Code | 995 | 8.8 s |
| Claude Sonnet 5.5 (medium) · Claude Code · event-loop-order · rep 1 | Claude Sonnet 5.5 · Claude Code | 1,200 | 10 s |
| Claude Sonnet 5.5 (high) · Claude Code · event-loop-order · rep 1 | Claude Sonnet 5.5 · Claude Code | 1,295 | 10.9 s |
| Claude Sonnet 5.5 (low) · Claude Code · talk-schedule · rep 1 | Claude Sonnet 5.5 · Claude Code | 598 | 6.5 s |
| Claude Sonnet 5.5 (medium) · Claude Code · talk-schedule · rep 1 | Claude Sonnet 5.5 · Claude Code | 651 | 7.9 s |
| Claude Sonnet 5.5 (high) · Claude Code · talk-schedule · rep 1 | Claude Sonnet 5.5 · Claude Code | 734 | 9.1 s |
| Claude Sonnet 5.5 (low) · Claude Code · semver-regex · rep 1 | Claude Sonnet 5.5 · Claude Code | 0 | 3.5 s |
| Claude Sonnet 5.5 (medium) · Claude Code · semver-regex · rep 1 | Claude Sonnet 5.5 · Claude Code | 251 | 4.1 s |
| Claude Sonnet 5.5 (high) · Claude Code · semver-regex · rep 1 | Claude Sonnet 5.5 · Claude Code | 252 | 4 s |
| Claude Sonnet 5.5 (low) · Claude Code · money-refactor · rep 1 | Claude Sonnet 5.5 · Claude Code | 0 | 3.7 s |
| Claude Sonnet 5.5 (medium) · Claude Code · money-refactor · rep 1 | Claude Sonnet 5.5 · Claude Code | 0 | 3.9 s |
| Claude Sonnet 5.5 (high) · Claude Code · money-refactor · rep 1 | Claude Sonnet 5.5 · Claude Code | 693 | 8.1 s |
| Claude Sonnet 5.5 (low) · Claude Code · q1-sql · rep 1 | Claude Sonnet 5.5 · Claude Code | 545 | 7.9 s |
| Claude Sonnet 5.5 (medium) · Claude Code · q1-sql · rep 1 | Claude Sonnet 5.5 · Claude Code | 421 | 7.4 s |
| Claude Sonnet 5.5 (high) · Claude Code · q1-sql · rep 1 | Claude Sonnet 5.5 · Claude Code | 826 | 10.8 s |
| Claude Sonnet 5.5 (low) · Claude Code · merge-ranges · rep 2 | Claude Sonnet 5.5 · Claude Code | 0 | 2.8 s |
| Claude Sonnet 5.5 (medium) · Claude Code · merge-ranges · rep 2 | Claude Sonnet 5.5 · Claude Code | 0 | 2.9 s |
| Claude Sonnet 5.5 (high) · Claude Code · merge-ranges · rep 2 | Claude Sonnet 5.5 · Claude Code | 0 | 3.5 s |
| Claude Sonnet 5.5 (low) · Claude Code · day-hours · rep 2 | Claude Sonnet 5.5 · Claude Code | 1,091 | 16.3 s |
| Claude Sonnet 5.5 (medium) · Claude Code · day-hours · rep 2 | Claude Sonnet 5.5 · Claude Code | 2,093 | 24 s |
| Claude Sonnet 5.5 (high) · Claude Code · day-hours · rep 2 | Claude Sonnet 5.5 · Claude Code | 3,610 | 35.8 s |
| Claude Sonnet 5.5 (low) · Claude Code · csv-parse · rep 2 | Claude Sonnet 5.5 · Claude Code | 0 | 4.4 s |
| Claude Sonnet 5.5 (medium) · Claude Code · csv-parse · rep 2 | Claude Sonnet 5.5 · Claude Code | 0 | 4.7 s |
| Claude Sonnet 5.5 (high) · Claude Code · csv-parse · rep 2 | Claude Sonnet 5.5 · Claude Code | 783 | 16.1 s |
| Claude Sonnet 5.5 (low) · Claude Code · event-loop-order · rep 2 | Claude Sonnet 5.5 · Claude Code | 974 | 9.7 s |
| Claude Sonnet 5.5 (medium) · Claude Code · event-loop-order · rep 2 | Claude Sonnet 5.5 · Claude Code | 1,098 | 8.7 s |
| Claude Sonnet 5.5 (high) · Claude Code · event-loop-order · rep 2 | Claude Sonnet 5.5 · Claude Code | 1,226 | 10.8 s |
| Claude Sonnet 5.5 (low) · Claude Code · talk-schedule · rep 2 | Claude Sonnet 5.5 · Claude Code | 624 | 6.7 s |
| Claude Sonnet 5.5 (medium) · Claude Code · talk-schedule · rep 2 | Claude Sonnet 5.5 · Claude Code | 624 | 9.6 s |
| Claude Sonnet 5.5 (high) · Claude Code · talk-schedule · rep 2 | Claude Sonnet 5.5 · Claude Code | 755 | 8.1 s |
| Claude Sonnet 5.5 (low) · Claude Code · semver-regex · rep 2 | Claude Sonnet 5.5 · Claude Code | 0 | 2.9 s |
| Claude Sonnet 5.5 (medium) · Claude Code · semver-regex · rep 2 | Claude Sonnet 5.5 · Claude Code | 245 | 3.8 s |
| Claude Sonnet 5.5 (high) · Claude Code · semver-regex · rep 2 | Claude Sonnet 5.5 · Claude Code | 243 | 3.7 s |
| Claude Sonnet 5.5 (low) · Claude Code · money-refactor · rep 2 | Claude Sonnet 5.5 · Claude Code | 0 | 5.2 s |
| Claude Sonnet 5.5 (medium) · Claude Code · money-refactor · rep 2 | Claude Sonnet 5.5 · Claude Code | 0 | 4 s |
| Claude Sonnet 5.5 (high) · Claude Code · money-refactor · rep 2 | Claude Sonnet 5.5 · Claude Code | 580 | 7.6 s |
| Claude Sonnet 5.5 (low) · Claude Code · q1-sql · rep 2 | Claude Sonnet 5.5 · Claude Code | 591 | 7.7 s |
| Claude Sonnet 5.5 (medium) · Claude Code · q1-sql · rep 2 | Claude Sonnet 5.5 · Claude Code | 566 | 9.2 s |
| Claude Sonnet 5.5 (high) · Claude Code · q1-sql · rep 2 | Claude Sonnet 5.5 · Claude Code | 626 | 8.6 s |
| Claude Opus 5.5 · Claude Code · merge-ranges · rep 1 | Claude Opus 5.5 · Claude Code | 100 | 4.2 s |
| Claude Opus 5.5 (high) · Claude Code · merge-ranges · rep 1 | Claude Opus 5.5 · Claude Code | 188 | 5.4 s |
| Claude Opus 5.5 · Claude Code · day-hours · rep 1 | Claude Opus 5.5 · Claude Code | 1,580 | 26.3 s |
| Claude Opus 5.5 (high) · Claude Code · day-hours · rep 1 | Claude Opus 5.5 · Claude Code | 3,301 | 63 s |
| Claude Opus 5.5 · Claude Code · csv-parse · rep 1 | Claude Opus 5.5 · Claude Code | 543 | 11 s |
| Claude Opus 5.5 (high) · Claude Code · csv-parse · rep 1 | Claude Opus 5.5 · Claude Code | 528 | 11.1 s |
| Claude Opus 5.5 · Claude Code · event-loop-order · rep 1 | Claude Opus 5.5 · Claude Code | 958 | 11.9 s |
| Claude Opus 5.5 (high) · Claude Code · event-loop-order · rep 1 | Claude Opus 5.5 · Claude Code | 1,274 | 13.6 s |
| Claude Opus 5.5 · Claude Code · talk-schedule · rep 1 | Claude Opus 5.5 · Claude Code | 574 | 7.8 s |
| Claude Opus 5.5 (high) · Claude Code · talk-schedule · rep 1 | Claude Opus 5.5 · Claude Code | 656 | 8.9 s |
| Claude Opus 5.5 · Claude Code · semver-regex · rep 1 | Claude Opus 5.5 · Claude Code | 300 | 5.4 s |
| Claude Opus 5.5 (high) · Claude Code · semver-regex · rep 1 | Claude Opus 5.5 · Claude Code | 293 | 5 s |
| Claude Opus 5.5 · Claude Code · money-refactor · rep 1 | Claude Opus 5.5 · Claude Code | 429 | 8.2 s |
| Claude Opus 5.5 (high) · Claude Code · money-refactor · rep 1 | Claude Opus 5.5 · Claude Code | 528 | 9.1 s |
| Claude Opus 5.5 · Claude Code · q1-sql · rep 1 | Claude Opus 5.5 · Claude Code | 796 | 12.8 s |
| Claude Opus 5.5 (high) · Claude Code · q1-sql · rep 1 | Claude Opus 5.5 · Claude Code | 920 | 13.7 s |
| Claude Opus 5.5 · Claude Code · merge-ranges · rep 2 | Claude Opus 5.5 · Claude Code | 114 | 4.7 s |
| Claude Opus 5.5 (high) · Claude Code · merge-ranges · rep 2 | Claude Opus 5.5 · Claude Code | 182 | 5.5 s |
| Claude Opus 5.5 · Claude Code · day-hours · rep 2 | Claude Opus 5.5 · Claude Code | 1,778 | 27.2 s |
| Claude Opus 5.5 (high) · Claude Code · day-hours · rep 2 | Claude Opus 5.5 · Claude Code | 2,756 | 36.4 s |
| Claude Opus 5.5 · Claude Code · csv-parse · rep 2 | Claude Opus 5.5 · Claude Code | 524 | 11.5 s |
| Claude Opus 5.5 (high) · Claude Code · csv-parse · rep 2 | Claude Opus 5.5 · Claude Code | 573 | 12.2 s |
| Claude Opus 5.5 · Claude Code · event-loop-order · rep 2 | Claude Opus 5.5 · Claude Code | 1,015 | 10.2 s |
| Claude Opus 5.5 (high) · Claude Code · event-loop-order · rep 2 | Claude Opus 5.5 · Claude Code | 1,300 | 13.5 s |
| Claude Opus 5.5 · Claude Code · talk-schedule · rep 2 | Claude Opus 5.5 · Claude Code | 533 | 7.6 s |
| Claude Opus 5.5 (high) · Claude Code · talk-schedule · rep 2 | Claude Opus 5.5 · Claude Code | 655 | 8.5 s |
| Claude Opus 5.5 · Claude Code · semver-regex · rep 2 | Claude Opus 5.5 · Claude Code | 244 | 4.8 s |
| Claude Opus 5.5 (high) · Claude Code · semver-regex · rep 2 | Claude Opus 5.5 · Claude Code | 103 | 3.6 s |
| Claude Opus 5.5 · Claude Code · money-refactor · rep 2 | Claude Opus 5.5 · Claude Code | 214 | 7.2 s |
| Claude Opus 5.5 (high) · Claude Code · money-refactor · rep 2 | Claude Opus 5.5 · Claude Code | 431 | 8.8 s |
| Claude Opus 5.5 · Claude Code · q1-sql · rep 2 | Claude Opus 5.5 · Claude Code | 781 | 12.4 s |
| Claude Opus 5.5 (high) · Claude Code · q1-sql · rep 2 | Claude Opus 5.5 · Claude Code | 739 | 12.5 s |
| Claude Opus 5.5 · Claude Code · merge-ranges · rep 3 | Claude Opus 5.5 · Claude Code | 122 | 4.5 s |
| Claude Opus 5.5 (high) · Claude Code · merge-ranges · rep 3 | Claude Opus 5.5 · Claude Code | 221 | 5.7 s |
| Claude Opus 5.5 · Claude Code · day-hours · rep 3 | Claude Opus 5.5 · Claude Code | 1,136 | 19.7 s |
| Claude Opus 5.5 (high) · Claude Code · day-hours · rep 3 | Claude Opus 5.5 · Claude Code | 3,027 | 38.8 s |
| Claude Opus 5.5 · Claude Code · csv-parse · rep 3 | Claude Opus 5.5 · Claude Code | 440 | 11.7 s |
| Claude Opus 5.5 (high) · Claude Code · csv-parse · rep 3 | Claude Opus 5.5 · Claude Code | 800 | 12.6 s |
| Claude Opus 5.5 · Claude Code · event-loop-order · rep 3 | Claude Opus 5.5 · Claude Code | 1,109 | 11.5 s |
| Claude Opus 5.5 (high) · Claude Code · event-loop-order · rep 3 | Claude Opus 5.5 · Claude Code | 1,198 | 13.5 s |
| Claude Opus 5.5 · Claude Code · talk-schedule · rep 3 | Claude Opus 5.5 · Claude Code | 444 | 6.9 s |
| Claude Opus 5.5 (high) · Claude Code · talk-schedule · rep 3 | Claude Opus 5.5 · Claude Code | 553 | 7.5 s |
| Claude Opus 5.5 · Claude Code · semver-regex · rep 3 | Claude Opus 5.5 · Claude Code | 293 | 5.1 s |
| Claude Opus 5.5 (high) · Claude Code · semver-regex · rep 3 | Claude Opus 5.5 · Claude Code | 122 | 4 s |
| Claude Opus 5.5 · Claude Code · money-refactor · rep 3 | Claude Opus 5.5 · Claude Code | 199 | 6.2 s |
| Claude Opus 5.5 (high) · Claude Code · money-refactor · rep 3 | Claude Opus 5.5 · Claude Code | 304 | 11 s |
| Claude Opus 5.5 · Claude Code · q1-sql · rep 3 | Claude Opus 5.5 · Claude Code | 808 | 13 s |
| Claude Opus 5.5 (high) · Claude Code · q1-sql · rep 3 | Claude Opus 5.5 · Claude Code | 911 | 12.9 s |
| Claude Opus 5.5 (low) · Claude Code · merge-ranges · rep 1 | Claude Opus 5.5 · Claude Code | 0 | 3.3 s |
| Claude Opus 5.5 (medium) · Claude Code · merge-ranges · rep 1 | Claude Opus 5.5 · Claude Code | 109 | 5.5 s |
| Claude Opus 5.5 (low) · Claude Code · day-hours · rep 1 | Claude Opus 5.5 · Claude Code | 394 | 13 s |
| Claude Opus 5.5 (medium) · Claude Code · day-hours · rep 1 | Claude Opus 5.5 · Claude Code | 1,993 | 31.1 s |
| Claude Opus 5.5 (low) · Claude Code · csv-parse · rep 1 | Claude Opus 5.5 · Claude Code | 471 | 9.9 s |
| Claude Opus 5.5 (medium) · Claude Code · csv-parse · rep 1 | Claude Opus 5.5 · Claude Code | 444 | 18.3 s |
| Claude Opus 5.5 (low) · Claude Code · event-loop-order · rep 1 | Claude Opus 5.5 · Claude Code | 684 | 8.7 s |
| Claude Opus 5.5 (medium) · Claude Code · event-loop-order · rep 1 | Claude Opus 5.5 · Claude Code | 1,120 | 12.4 s |
| Claude Opus 5.5 (low) · Claude Code · talk-schedule · rep 1 | Claude Opus 5.5 · Claude Code | 498 | 7.7 s |
| Claude Opus 5.5 (medium) · Claude Code · talk-schedule · rep 1 | Claude Opus 5.5 · Claude Code | 471 | 7 s |
| Claude Opus 5.5 (low) · Claude Code · semver-regex · rep 1 | Claude Opus 5.5 · Claude Code | 88 | 4.2 s |
| Claude Opus 5.5 (medium) · Claude Code · semver-regex · rep 1 | Claude Opus 5.5 · Claude Code | 119 | 6.8 s |
| Claude Opus 5.5 (low) · Claude Code · money-refactor · rep 1 | Claude Opus 5.5 · Claude Code | 0 | 5.7 s |
| Claude Opus 5.5 (medium) · Claude Code · money-refactor · rep 1 | Claude Opus 5.5 · Claude Code | 183 | 6.9 s |
| Claude Opus 5.5 (low) · Claude Code · q1-sql · rep 1 | Claude Opus 5.5 · Claude Code | 0 | 10.6 s |
| Claude Opus 5.5 (medium) · Claude Code · q1-sql · rep 1 | Claude Opus 5.5 · Claude Code | 793 | 13.1 s |
| Claude Opus 5.5 (low) · Claude Code · merge-ranges · rep 2 | Claude Opus 5.5 · Claude Code | 0 | 3.4 s |
| Claude Opus 5.5 (medium) · Claude Code · merge-ranges · rep 2 | Claude Opus 5.5 · Claude Code | 115 | 5.3 s |
| Claude Opus 5.5 (low) · Claude Code · day-hours · rep 2 | Claude Opus 5.5 · Claude Code | 587 | 15.8 s |
| Claude Opus 5.5 (medium) · Claude Code · day-hours · rep 2 | Claude Opus 5.5 · Claude Code | 1,846 | 31.4 s |
| Claude Opus 5.5 (low) · Claude Code · csv-parse · rep 2 | Claude Opus 5.5 · Claude Code | 0 | 6.2 s |
| Claude Opus 5.5 (medium) · Claude Code · csv-parse · rep 2 | Claude Opus 5.5 · Claude Code | 655 | 10.9 s |
| Claude Opus 5.5 (low) · Claude Code · event-loop-order · rep 2 | Claude Opus 5.5 · Claude Code | 791 | 9.2 s |
| Claude Opus 5.5 (medium) · Claude Code · event-loop-order · rep 2 | Claude Opus 5.5 · Claude Code | 1,014 | 11.2 s |
| Claude Opus 5.5 (low) · Claude Code · talk-schedule · rep 2 | Claude Opus 5.5 · Claude Code | 426 | 7.3 s |
| Claude Opus 5.5 (medium) · Claude Code · talk-schedule · rep 2 | Claude Opus 5.5 · Claude Code | 565 | 8.5 s |
| Claude Opus 5.5 (low) · Claude Code · semver-regex · rep 2 | Claude Opus 5.5 · Claude Code | 86 | 3.8 s |
| Claude Opus 5.5 (medium) · Claude Code · semver-regex · rep 2 | Claude Opus 5.5 · Claude Code | 243 | 4.8 s |
| Claude Opus 5.5 (low) · Claude Code · money-refactor · rep 2 | Claude Opus 5.5 · Claude Code | 0 | 4.9 s |
| Claude Opus 5.5 (medium) · Claude Code · money-refactor · rep 2 | Claude Opus 5.5 · Claude Code | 253 | 7.1 s |
| Claude Opus 5.5 (low) · Claude Code · q1-sql · rep 2 | Claude Opus 5.5 · Claude Code | 0 | 9.2 s |
| Claude Opus 5.5 (medium) · Claude Code · q1-sql · rep 2 | Claude Opus 5.5 · Claude Code | 829 | 13.8 s |
| Claude Fable 5.1 · Claude Code · merge-ranges · rep 1 | Claude Fable 5.1 · Claude Code | 85 | 4.8 s |
| Claude Fable 5.1 · Claude Code · day-hours · rep 1 | Claude Fable 5.1 · Claude Code | 1,815 | 33.7 s |
| Claude Fable 5.1 · Claude Code · csv-parse · rep 1 | Claude Fable 5.1 · Claude Code | 903 | 16.2 s |
| Claude Fable 5.1 · Claude Code · event-loop-order · rep 1 | Claude Fable 5.1 · Claude Code | 1,765 | 23.9 s |
| Claude Fable 5.1 · Claude Code · talk-schedule · rep 1 | Claude Fable 5.1 · Claude Code | 1,090 | 17.3 s |
| Claude Fable 5.1 · Claude Code · semver-regex · rep 1 | Claude Fable 5.1 · Claude Code | 261 | 9.6 s |
| Claude Fable 5.1 · Claude Code · money-refactor · rep 1 | Claude Fable 5.1 · Claude Code | 646 | 12.9 s |
| Claude Fable 5.1 · Claude Code · q1-sql · rep 1 | Claude Fable 5.1 · Claude Code | 875 | 16.1 s |
| Claude Fable 5.1 · Claude Code · merge-ranges · rep 2 | Claude Fable 5.1 · Claude Code | 75 | 4.5 s |
| Claude Fable 5.1 · Claude Code · day-hours · rep 2 | Claude Fable 5.1 · Claude Code | 5,889 | 90 s |
| Claude Fable 5.1 · Claude Code · csv-parse · rep 2 | Claude Fable 5.1 · Claude Code | 1,088 | 21.4 s |
| Claude Fable 5.1 · Claude Code · event-loop-order · rep 2 | Claude Fable 5.1 · Claude Code | 1,228 | 16.8 s |
| Claude Fable 5.1 · Claude Code · talk-schedule · rep 2 | Claude Fable 5.1 · Claude Code | 786 | 11.2 s |
| Claude Fable 5.1 · Claude Code · semver-regex · rep 2 | Claude Fable 5.1 · Claude Code | 236 | 7.5 s |
| Claude Fable 5.1 · Claude Code · money-refactor · rep 2 | Claude Fable 5.1 · Claude Code | 796 | 13.7 s |
| Claude Fable 5.1 · Claude Code · q1-sql · rep 2 | Claude Fable 5.1 · Claude Code | 865 | 20.3 s |
| Claude Fable 5.1 · Claude Code · merge-ranges · rep 3 | Claude Fable 5.1 · Claude Code | 197 | 15.2 s |
| Claude Fable 5.1 · Claude Code · day-hours · rep 3 | Claude Fable 5.1 · Claude Code | 1,427 | 25.1 s |
| Claude Fable 5.1 · Claude Code · csv-parse · rep 3 | Claude Fable 5.1 · Claude Code | 1,000 | 15.9 s |
| Claude Fable 5.1 · Claude Code · event-loop-order · rep 3 | Claude Fable 5.1 · Claude Code | 1,606 | 26.9 s |
| Claude Fable 5.1 · Claude Code · talk-schedule · rep 3 | Claude Fable 5.1 · Claude Code | 836 | 12.1 s |
| Claude Fable 5.1 · Claude Code · semver-regex · rep 3 | Claude Fable 5.1 · Claude Code | 266 | 5.7 s |
| Claude Fable 5.1 · Claude Code · money-refactor · rep 3 | Claude Fable 5.1 · Claude Code | 948 | 21.9 s |
| Claude Fable 5.1 · Claude Code · q1-sql · rep 3 | Claude Fable 5.1 · Claude Code | 1,091 | 24.6 s |
| GPT-6.1 Sol (medium) · Codex CLI · merge-ranges · rep 1 | GPT-6.1 Sol · Codex CLI | 163 | 13.2 s |
| GPT-6.1 Sol (high) · Codex CLI · merge-ranges · rep 1 | GPT-6.1 Sol · Codex CLI | 192 | 14.4 s |
| GPT-6.1 Sol (medium) · Codex CLI · day-hours · rep 1 | GPT-6.1 Sol · Codex CLI | 830 | 61.6 s |
| GPT-6.1 Sol (high) · Codex CLI · day-hours · rep 1 | GPT-6.1 Sol · Codex CLI | 1,533 | 83 s |
| GPT-6.1 Sol (medium) · Codex CLI · csv-parse · rep 1 | GPT-6.1 Sol · Codex CLI | 86 | 14.8 s |
| GPT-6.1 Sol (high) · Codex CLI · csv-parse · rep 1 | GPT-6.1 Sol · Codex CLI | 256 | 22.9 s |
| GPT-6.1 Sol (medium) · Codex CLI · event-loop-order · rep 1 | GPT-6.1 Sol · Codex CLI | 199 | 8.5 s |
| GPT-6.1 Sol (high) · Codex CLI · event-loop-order · rep 1 | GPT-6.1 Sol · Codex CLI | 347 | 15.4 s |
| GPT-6.1 Sol (medium) · Codex CLI · talk-schedule · rep 1 | GPT-6.1 Sol · Codex CLI | 189 | 12.6 s |
| GPT-6.1 Sol (high) · Codex CLI · talk-schedule · rep 1 | GPT-6.1 Sol · Codex CLI | 189 | 13.1 s |
| GPT-6.1 Sol (medium) · Codex CLI · semver-regex · rep 1 | GPT-6.1 Sol · Codex CLI | 101 | 9.6 s |
| GPT-6.1 Sol (high) · Codex CLI · semver-regex · rep 1 | GPT-6.1 Sol · Codex CLI | 306 | 17.8 s |
| GPT-6.1 Sol (medium) · Codex CLI · money-refactor · rep 1 | GPT-6.1 Sol · Codex CLI | 89 | 11.7 s |
| GPT-6.1 Sol (high) · Codex CLI · money-refactor · rep 1 | GPT-6.1 Sol · Codex CLI | 224 | 19.4 s |
| GPT-6.1 Sol (medium) · Codex CLI · q1-sql · rep 1 | GPT-6.1 Sol · Codex CLI | 61 | 13.8 s |
| GPT-6.1 Sol (high) · Codex CLI · q1-sql · rep 1 | GPT-6.1 Sol · Codex CLI | 181 | 18.4 s |
| GPT-6.1 Sol (medium) · Codex CLI · merge-ranges · rep 2 | GPT-6.1 Sol · Codex CLI | 137 | 14.6 s |
| GPT-6.1 Sol (high) · Codex CLI · merge-ranges · rep 2 | GPT-6.1 Sol · Codex CLI | 226 | 15.1 s |
| GPT-6.1 Sol (medium) · Codex CLI · day-hours · rep 2 | GPT-6.1 Sol · Codex CLI | 839 | 56.4 s |
| GPT-6.1 Sol (high) · Codex CLI · day-hours · rep 2 | GPT-6.1 Sol · Codex CLI | 1,750 | 92.2 s |
| GPT-6.1 Sol (medium) · Codex CLI · csv-parse · rep 2 | GPT-6.1 Sol · Codex CLI | 112 | 16.6 s |
| GPT-6.1 Sol (high) · Codex CLI · csv-parse · rep 2 | GPT-6.1 Sol · Codex CLI | 220 | 20.6 s |
| GPT-6.1 Sol (medium) · Codex CLI · event-loop-order · rep 2 | GPT-6.1 Sol · Codex CLI | 251 | 12.5 s |
| GPT-6.1 Sol (high) · Codex CLI · event-loop-order · rep 2 | GPT-6.1 Sol · Codex CLI | 375 | 22.5 s |
| GPT-6.1 Sol (medium) · Codex CLI · talk-schedule · rep 2 | GPT-6.1 Sol · Codex CLI | 197 | 11.8 s |
| GPT-6.1 Sol (high) · Codex CLI · talk-schedule · rep 2 | GPT-6.1 Sol · Codex CLI | 229 | 12.2 s |
| GPT-6.1 Sol (medium) · Codex CLI · semver-regex · rep 2 | GPT-6.1 Sol · Codex CLI | 193 | 13 s |
| GPT-6.1 Sol (high) · Codex CLI · semver-regex · rep 2 | GPT-6.1 Sol · Codex CLI | 189 | 11.7 s |
| GPT-6.1 Sol (medium) · Codex CLI · money-refactor · rep 2 | GPT-6.1 Sol · Codex CLI | 71 | 11.9 s |
| GPT-6.1 Sol (high) · Codex CLI · money-refactor · rep 2 | GPT-6.1 Sol · Codex CLI | 144 | 15.4 s |
| GPT-6.1 Sol (medium) · Codex CLI · q1-sql · rep 2 | GPT-6.1 Sol · Codex CLI | 119 | 17.8 s |
| GPT-6.1 Sol (high) · Codex CLI · q1-sql · rep 2 | GPT-6.1 Sol · Codex CLI | 222 | 19.6 s |
| GPT-6.1 Sol (low) · Codex CLI · merge-ranges · rep 1 | GPT-6.1 Sol · Codex CLI | 63 | 17.2 s |
| GPT-6.1 Sol (low) · Codex CLI · day-hours · rep 1 | GPT-6.1 Sol · Codex CLI | 458 | 44.3 s |
| GPT-6.1 Sol (low) · Codex CLI · csv-parse · rep 1 | GPT-6.1 Sol · Codex CLI | 78 | 17.7 s |
| GPT-6.1 Sol (low) · Codex CLI · event-loop-order · rep 1 | GPT-6.1 Sol · Codex CLI | 198 | 14.6 s |
| GPT-6.1 Sol (low) · Codex CLI · talk-schedule · rep 1 | GPT-6.1 Sol · Codex CLI | 189 | 13.6 s |
| GPT-6.1 Sol (low) · Codex CLI · semver-regex · rep 1 | GPT-6.1 Sol · Codex CLI | 44 | 8.6 s |
| GPT-6.1 Sol (low) · Codex CLI · money-refactor · rep 1 | GPT-6.1 Sol · Codex CLI | 62 | 13.1 s |
| GPT-6.1 Sol (low) · Codex CLI · q1-sql · rep 1 | GPT-6.1 Sol · Codex CLI | 0 | 13.6 s |
| GPT-6.1 Sol (low) · Codex CLI · merge-ranges · rep 2 | GPT-6.1 Sol · Codex CLI | 0 | 7.9 s |
| GPT-6.1 Sol (low) · Codex CLI · day-hours · rep 2 | GPT-6.1 Sol · Codex CLI | 400 | 40.7 s |
| GPT-6.1 Sol (low) · Codex CLI · csv-parse · rep 2 | GPT-6.1 Sol · Codex CLI | 0 | 17.9 s |
| GPT-6.1 Sol (low) · Codex CLI · event-loop-order · rep 2 | GPT-6.1 Sol · Codex CLI | 185 | 11.1 s |
| GPT-6.1 Sol (low) · Codex CLI · talk-schedule · rep 2 | GPT-6.1 Sol · Codex CLI | 189 | 12.4 s |
| GPT-6.1 Sol (low) · Codex CLI · semver-regex · rep 2 | GPT-6.1 Sol · Codex CLI | 50 | 8.9 s |
| GPT-6.1 Sol (low) · Codex CLI · money-refactor · rep 2 | GPT-6.1 Sol · Codex CLI | 0 | 9.5 s |
| GPT-6.1 Sol (low) · Codex CLI · q1-sql · rep 2 | GPT-6.1 Sol · Codex CLI | 40 | 14.4 s |
List-price calculation, not a run. 248 points: Total time per call (seconds) against Reasoning tokens per call. Reasoning tokens per call runs from 0 to 8,569; Total time per call (seconds) from 2.3 s to 92.2 s.
Notes
One point per call: 248 calls from the hard head-to-head and the effort ladder
Calculation, not a run: each point is one recorded call. Spearman rank correlations and slopes are in the table. Total time includes CLI start-up and the visible answer. Task, effort and batch can affect both counts and time. The plot shows association, not cause. 248 calls over 8 tasks; calls are not independent.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
- Eight hard tasks
- Five short tasks (square)
Gap labels, Five short tasks vs Eight hard tasks: Five short tasks is x percentage points higher (+) or lower (−) than Eight hard tasks, calculated from the two values shown; lines are the lowest–highest run (not an interval).
| Item | Eight hard tasks | Five short tasks | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 92% | 90% | Eight hard tasks: 76%–99%; Five short tasks: 73%–98% | 24 |
| Claude Sonnet 5.5 · Claude Code | 55% | 0% | Eight hard tasks: 0%–96%; Five short tasks: 0%–73% | 24 |
| Claude Opus 5.5 · Claude Code | 55% | 0% | Eight hard tasks: 30%–96%; Five short tasks: 0%–93% | 24 |
| Claude Opus 5.5 (high) · Claude Code | 54% | 44% | Eight hard tasks: 36%–96%; Five short tasks: 0%–93% | 24 |
| Claude Fable 5.1 · Claude Code | 64% | 0% | Eight hard tasks: 23%–97%; Five short tasks: 0%–74% | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 46% | 41% | Eight hard tasks: 11%–87%; Five short tasks: 0%–71% | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 57% | 58% | Eight hard tasks: 29%–91%; Five short tasks: 0%–76% | 16 |
List-price calculation, not a run. 7 rows, 2 series: Eight hard tasks, Five short tasks. Eight hard tasks: highest Claude Haiku 4.5 · Claude Code 92% (range 76%–99%, n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 46% (range 11%–87%, n 16). All run ranges overlap. Five short tasks: highest Claude Haiku 4.5 · Claude Code 90% (range 73%–98%, n 15). Lowest Claude Fable 5.1 · Claude Code 0% (range 0%–74%, n 15). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 15–24 per row
Median call per configuration; five short tasks and eight hard tasks
Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
Tables
Reasoning tokens and list-price cost per call, hard tasks (calculation)
| Configuration | Calls | Strict passes | Median output tokens | Median reasoning tokens | Output / reasoning tokens, lowest to highest call | Reasoning / total cost, lowest to highest call (USD, calculation) | Reasoning cost share, lowest to highest call (calculation) | Calls with 0 reasoning tokens | Median call: reasoning share of output | Share, lowest to highest call | Pooled share (all reasoning ÷ all output) | Reasoning cost per call (USD, calculation) | Remaining output cost per call (USD, calculation) | Input cost per call (USD) | Total cost per call (USD) | Pooled reasoning share of list-price cost (calculation) | Reasoning cost per strict pass (USD) | Total cost per strict pass (USD) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 24 | 11/24 (95% Wilson 28% to 65%) | 5,064 | 4,556 | 1899 to 9321 / 1452 to 8569 | $0.0073 to $0.0428 / $0.0148 to $0.0505 | 40% to 90% | 0 | 92% | 76% to 99% | 93% | $0.024 | $0.0018 | $0.0045 | $0.031 | 80% | $0.053 | $0.067 |
| Claude Sonnet 5.5 · Claude Code | 24 | 24/24 (95% Wilson 86% to 100%) | 1,050 | 585 | 176 to 3895 / 0 to 3060 | $0.0000 to $0.0306 / $0.0051 to $0.0424 | 0% to 74% | 7 | 55% | 0% to 96% | 64% | $0.0067 | $0.0037 | $0.004 | $0.014 | 46% | $0.0067 | $0.014 |
| Claude Opus 5.5 · Claude Code | 24 | 24/24 (95% Wilson 86% to 100%) | 945 | 529 | 323 to 2531 / 100 to 1778 | $0.0020 to $0.0356 / $0.0131 to $0.0572 | 10% to 74% | 0 | 55% | 30% to 96% | 61% | $0.013 | $0.008 | $0.0077 | $0.028 | 44% | $0.013 | $0.028 |
| Claude Opus 5.5 (high) · Claude Code | 24 | 24/24 (95% Wilson 86% to 100%) | 1,052 | 614 | 285 to 4052 / 103 to 3301 | $0.0021 to $0.0660 / $0.0121 to $0.0877 | 17% to 77% | 0 | 54% | 36% to 96% | 69% | $0.018 | $0.008 | $0.0074 | $0.033 | 54% | $0.018 | $0.033 |
| Claude Fable 5.1 · Claude Code | 24 | 24/24 (95% Wilson 86% to 100%) | 1,366 | 889 | 318 to 6465 / 75 to 5889 | $0.0037 to $0.2944 / $0.0322 to $0.3399 | 5% to 87% | 0 | 64% | 23% to 97% | 74% | $0.054 | $0.019 | $0.021 | $0.093 | 58% | $0.054 | $0.093 |
| GPT-6.1 Sol (medium) · Codex CLI | 16 | 16/16 (95% Wilson 81% to 100%) | 335 | 150 | 237 to 1766 / 61 to 839 | $0.0006 to $0.0084 / $0.0099 to $0.0422 | 2% to 20% | 0 | 46% | 11% to 87% | 43% | $0.0023 | $0.003 | $0.02 | $0.026 | 8.9% | $0.0023 | $0.026 |
| GPT-6.1 Sol (high) · Codex CLI | 16 | 16/16 (95% Wilson 81% to 100%) | 436 | 225 | 284 to 2569 / 144 to 1750 | $0.0014 to $0.0175 / $0.0092 to $0.0307 | 8% to 60% | 0 | 57% | 29% to 91% | 59% | $0.0041 | $0.0029 | $0.0081 | $0.015 | 27% | $0.0041 | $0.015 |
Reasoning tokens and list-price cost by effort, every effort-ladder cell (calculation)
| Configuration | Calls | Strict passes | Median output tokens | Median reasoning tokens | Output / reasoning tokens, lowest to highest call | Reasoning / total cost, lowest to highest call (USD, calculation) | Reasoning cost share, lowest to highest call (calculation) | Calls with 0 reasoning tokens | Median call: reasoning share of output | Share, lowest to highest call | Pooled share (all reasoning ÷ all output) | Reasoning cost per call (USD, calculation) | Remaining output cost per call (USD, calculation) | Input cost per call (USD) | Total cost per call (USD) | Pooled reasoning share of list-price cost (calculation) | Reasoning cost per strict pass (USD) | Total cost per strict pass (USD) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 (low) · Claude Code | 16 | 16/16 (95% Wilson 81% to 100%) | 667 | 273 | 176 to 2263 / 0 to 1489 | $0.0000 to $0.0149 / $0.0051 to $0.0261 | 0% to 71% | 8 | 21% | 0% to 95% | 53% | $0.0043 | $0.0038 | $0.0041 | $0.012 | 35% | $0.0043 | $0.012 |
| Claude Sonnet 5.5 (medium) · Claude Code | 16 | 16/16 (95% Wilson 81% to 100%) | 770 | 422 | 224 to 2836 / 0 to 2093 | $0.0000 to $0.0209 / $0.0056 to $0.0318 | 0% to 75% | 5 | 52% | 0% to 96% | 62% | $0.0059 | $0.0037 | $0.0039 | $0.014 | 44% | $0.0059 | $0.014 |
| Claude Sonnet 5.5 (high) · Claude Code | 16 | 16/16 (95% Wilson 81% to 100%) | 1,192 | 745 | 220 to 4187 / 0 to 3610 | $0.0000 to $0.0361 / $0.0056 to $0.0453 | 0% to 80% | 2 | 61% | 0% to 96% | 73% | $0.0094 | $0.0035 | $0.0039 | $0.017 | 56% | $0.0094 | $0.017 |
| Claude Sonnet 5.5 · Claude Code | 16 | 16/16 (95% Wilson 81% to 100%) | 1,054 | 668 | 176 to 2185 / 0 to 1614 | $0.0000 to $0.0161 / $0.0051 to $0.0253 | 0% to 74% | 4 | 56% | 0% to 96% | 64% | $0.0063 | $0.0036 | $0.0041 | $0.014 | 45% | $0.0063 | $0.014 |
| Claude Opus 5.5 (low) · Claude Code | 16 | 16/16 (95% Wilson 81% to 100%) | 594 | 87 | 220 to 1569 / 0 to 791 | $0.0000 to $0.0158 / $0.0108 to $0.0380 | 0% to 67% | 7 | 31% | 0% to 94% | 38% | $0.005 | $0.0083 | $0.0079 | $0.021 | 24% | $0.005 | $0.021 |
| Claude Opus 5.5 (medium) · Claude Code | 16 | 16/16 (95% Wilson 81% to 100%) | 853 | 518 | 301 to 3075 / 109 to 1993 | $0.0022 to $0.0399 / $0.0125 to $0.0681 | 16% to 74% | 0 | 54% | 28% to 96% | 61% | $0.013 | $0.0086 | $0.0074 | $0.029 | 46% | $0.013 | $0.029 |
| Claude Opus 5.5 (high) · Claude Code | 16 | 16/16 (95% Wilson 81% to 100%) | 1,052 | 614 | 285 to 4052 / 103 to 3301 | $0.0021 to $0.0660 / $0.0121 to $0.0877 | 17% to 77% | 0 | 54% | 36% to 96% | 69% | $0.018 | $0.0082 | $0.0074 | $0.034 | 54% | $0.018 | $0.034 |
| Claude Opus 5.5 · Claude Code | 16 | 16/16 (95% Wilson 81% to 100%) | 945 | 538 | 323 to 2531 / 100 to 1778 | $0.0020 to $0.0356 / $0.0131 to $0.0572 | 10% to 72% | 0 | 54% | 31% to 95% | 62% | $0.013 | $0.008 | $0.0079 | $0.029 | 45% | $0.013 | $0.029 |
| GPT-6.1 Sol (low) · Codex CLI | 16 | 16/16 (95% Wilson 81% to 100%) | 284 | 63 | 172 to 1310 / 0 to 458 | $0.0000 to $0.0046 / $0.0092 to $0.0264 | 0% to 22% | 4 | 24% | 0% to 84% | 29% | $0.0012 | $0.003 | $0.0086 | $0.013 | 9.5% | $0.0012 | $0.013 |
| GPT-6.1 Sol (medium) · Codex CLI | 16 | 16/16 (95% Wilson 81% to 100%) | 335 | 150 | 237 to 1766 / 61 to 839 | $0.0006 to $0.0084 / $0.0099 to $0.0422 | 2% to 20% | 0 | 46% | 11% to 87% | 43% | $0.0023 | $0.003 | $0.02 | $0.026 | 8.9% | $0.0023 | $0.026 |
| GPT-6.1 Sol (high) · Codex CLI | 16 | 16/16 (95% Wilson 81% to 100%) | 436 | 225 | 284 to 2569 / 144 to 1750 | $0.0014 to $0.0175 / $0.0092 to $0.0307 | 8% to 60% | 0 | 57% | 29% to 91% | 59% | $0.0041 | $0.0029 | $0.0081 | $0.015 | 27% | $0.0041 | $0.015 |
Reasoning tokens and list-price cost per call, five short tasks (calculation)
| Configuration | Calls | Strict passes | Median output tokens | Median reasoning tokens | Output / reasoning tokens, lowest to highest call | Reasoning / total cost, lowest to highest call (USD, calculation) | Reasoning cost share, lowest to highest call (calculation) | Calls with 0 reasoning tokens | Median call: reasoning share of output | Share, lowest to highest call | Pooled share (all reasoning ÷ all output) | Reasoning cost per call (USD, calculation) | Remaining output cost per call (USD, calculation) | Input cost per call (USD) | Total cost per call (USD) | Pooled reasoning share of list-price cost (calculation) | Reasoning cost per strict pass (USD) | Total cost per strict pass (USD) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 15 | 15/15 (95% Wilson 80% to 100%) | 367 | 297 | 275 to 2851 / 212 to 2643 | $0.0011 to $0.0132 / $0.0051 to $0.0180 | 20% to 74% | 0 | 90% | 73% to 98% | 90% | $0.0041 | $0.00043 | $0.0038 | $0.0084 | 49% | $0.0041 | $0.0084 |
| Claude Sonnet 5.5 · Claude Code | 15 | 12/15 (95% Wilson 55% to 93%) | 107 | 0 | 44 to 745 / 0 to 542 | $0.0000 to $0.0054 / $0.0034 to $0.0102 | 0% to 53% | 12 | 0% | 0% to 73% | 45% | $0.00088 | $0.0011 | $0.003 | $0.005 | 18% | $0.0011 | $0.0062 |
| Claude Opus 5.5 · Claude Code | 15 | 15/15 (95% Wilson 80% to 100%) | 64 | 0 | 44 to 853 / 0 to 616 | $0.0000 to $0.0123 / $0.0059 to $0.0223 | 0% to 55% | 9 | 0% | 0% to 93% | 59% | $0.0026 | $0.0018 | $0.0057 | $0.01 | 26% | $0.0026 | $0.01 |
| Claude Opus 5.5 (low) · Claude Code | 15 | 15/15 (95% Wilson 80% to 100%) | 64 | 0 | 44 to 637 / 0 to 405 | $0.0000 to $0.0081 / $0.0058 to $0.0179 | 0% to 45% | 9 | 0% | 0% to 93% | 41% | $0.0013 | $0.0018 | $0.0052 | $0.0083 | 15% | $0.0013 | $0.0083 |
| Claude Opus 5.5 (high) · Claude Code | 15 | 15/15 (95% Wilson 80% to 100%) | 78 | 34 | 44 to 1094 / 0 to 857 | $0.0000 to $0.0171 / $0.0059 to $0.0271 | 0% to 63% | 7 | 44% | 0% to 93% | 66% | $0.0034 | $0.0018 | $0.0052 | $0.01 | 33% | $0.0034 | $0.01 |
| Claude Fable 5.1 · Claude Code | 15 | 15/15 (95% Wilson 80% to 100%) | 64 | 0 | 4 to 908 / 0 to 672 | $0.0000 to $0.0336 / $0.0049 to $0.0584 | 0% to 67% | 12 | 0% | 0% to 74% | 57% | $0.006 | $0.0044 | $0.01 | $0.021 | 29% | $0.006 | $0.021 |
| GPT-6.1 Sol (low) · Codex CLI | 10 | 10/10 (95% Wilson 72% to 100%) | 42 | 20 | 28 to 225 / 0 to 103 | $0.0000 to $0.0010 / $0.0074 to $0.0265 | 0% to 11% | 4 | 25% | 0% to 71% | 31% | $0.00033 | $0.00073 | $0.0089 | $0.01 | 3.3% | $0.00033 | $0.01 |
| GPT-6.1 Sol (medium) · Codex CLI | 15 | 15/15 (95% Wilson 80% to 100%) | 42 | 20 | 27 to 295 / 0 to 137 | $0.0000 to $0.0014 / $0.0054 to $0.0269 | 0% to 23% | 6 | 41% | 0% to 71% | 41% | $0.00051 | $0.00072 | $0.014 | $0.016 | 3.3% | $0.00051 | $0.016 |
| GPT-6.1 Sol (high) · Codex CLI | 15 | 15/15 (95% Wilson 80% to 100%) | 42 | 21 | 28 to 457 / 0 to 289 | $0.0000 to $0.0029 / $0.0066 to $0.0281 | 0% to 29% | 6 | 58% | 0% to 76% | 58% | $0.001 | $0.00072 | $0.011 | $0.013 | 7.6% | $0.001 | $0.013 |
Does reasoning track time? Rank correlation and slope per model (calculation)
| Model and route | Calls | Spearman: reasoning tokens vs time | Leave-one-task-out range (not a 95% interval) | Spearman: output tokens vs time | Seconds per 1,000 reasoning tokens (all calls) | Seconds per 1,000 reasoning tokens (within task) | Within-task slope, leave-one-task-out range | Median reasoning tokens | Median total time (s) | Total time, lowest to highest call (s; not an interval) |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 24 | 0.98 | 0.98 to 0.99 | 0.99 | 8.4 s | 8.8 s | 8.6 to 9.1 | 4,556 | 39 s | 15.27 to 75.13 |
| Claude Sonnet 5.5 · Claude Code | 72 | 0.92 | 0.88 to 0.94 | 0.94 | 9.3 s | 7.6 s | 7.4 to 7.8 | 586 | 7.8 s | 2.26 to 35.81 |
| Claude Opus 5.5 · Claude Code | 80 | 0.85 | 0.82 to 0.89 | 0.97 | 13.1 s | 11.8 s | 5.7 to 12.7 | 511 | 9 s | 3.34 to 63.00 |
| Claude Fable 5.1 · Claude Code | 24 | 0.92 | 0.89 to 0.93 | 0.94 | 14.4 s | 14.5 s | 14.4 to 22.6 | 889 | 16.1 s | 4.46 to 90.00 |
| GPT-6.1 Sol · Codex CLI | 48 | 0.56 | 0.34 to 0.65 | 0.85 | 50.4 s | 36.4 s | 31.8 to 36.7 | 189 | 14.5 s | 7.94 to 92.20 |
Consistency check for reasoning within output tokens (calculation)
| CLI route | Calls | Calls with more than 0 reasoning tokens | Calls that reported 0 (kept as 0) | Calls with reasoning above output | Output tokens gained per extra reasoning token (within task) | Leave-one-task-out range (not a 95% interval) | Characters per token of output minus reasoning | Characters per output token, calls with 0 reasoning | Characters per output token, calls with reasoning | Reasoning is part of output |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Code | 290 | 212 | 78 | 0 | 0.99x | 0.982 to 1.009 | 1.94 | 1.98 | 0.45 | consistent with inclusion; not proof |
| Codex CLI | 88 | 68 | 20 | 0 | 0.97x | 0.965 to 0.986 | 2.99 | 2.76 | 1.36 | consistent with inclusion; not proof |
Method
- This study is a calculation over receipts that already exist. We made no new model call. The inputs are three recorded runs. The hard head-to-head gives 152 completed calls. We leave its 30 pre-inference blocked attempts out of token calculations. They remain failures of the original route attempt, not model answers. The effort ladder gives 96 new calls. The five-task head-to-head gives 130 calls. All current counted calls completed. Recorded errors count as failures when their usage is known. We keep the separate blocked batch on record. The original Codex route was blocked. A later batch ran after a successful uncounted probe. It resumed after two completed calls and skipped them.
- Sum check. Before we compute any share, we test one point. Do the output tokens include the reasoning tokens? We use this accounting assumption and check its consistency with the receipts. We run three tests on each CLI route. Test 1: reasoning never exceeds output. Test 2: within one task and model, one more reasoning token adds about one output token. A slope near 1 supports the assumption. Correlation alone cannot prove the counter semantics. Test 3: we compare characters per remaining token with characters per output token on calls whose reasoning counter is zero. Claude Code: consistent with inclusion, not proof. Reasoning never exceeded output (0 of 290 calls). Within one task and model, each extra reasoning token went with 0.99 extra output tokens (leave-one-task-out range 0.98 to 1.01, 290 calls). Output minus reasoning has 1.94 characters per token. Calls with 0 reasoning have 1.98. For comparison, output including reasoning has only 0.45. Codex CLI: consistent with inclusion, not proof. Reasoning never exceeded output (0 of 88 calls). Within one task and model, each extra reasoning token went with 0.97 extra output tokens (leave-one-task-out range 0.97 to 0.99, 88 calls). Output minus reasoning has 2.99 characters per token. Calls with 0 reasoning have 2.76. For comparison, output including reasoning has only 1.36.
- Not reported is not zero. A route reports reasoning when at least one of its calls reports more than 0. Both routes do (Claude Code 212 of 290 calls, Codex CLI 68 of 88 calls). So a reported 0 is the recorded counter, not proof of no internal reasoning. 98 calls reported 0, and their replies look like visible text only (test 3). If any model attempt lacks usable token counts, this builder returns no study. It never prices unknown usage as zero. 0 calls had no usable count.
- Reasoning share of one call = reasoning tokens ÷ output tokens. A configuration gets two numbers. One is the median of its per-call shares. The other is the pooled share (all reasoning tokens ÷ all output tokens). Ranges are the lowest and highest call. They are not intervals.
- Hard tasks: we use all calls of each configuration in the hard head-to-head. That is 16 to 24 calls: 8 tasks × 3 repetitions for Claude Code and 8 × 2 for Codex CLI. Effort ladder: 11 cells of 8 tasks × 2 repetitions, the design of the effort-ladder study. We reuse its reference cells from the hard head-to-head (Claude repetitions 1-2 only). Short tasks: the five validated tasks (10 to 15 calls per configuration).
- Cost is a calculation at list price, not a bill. The calls ran on flat subscriptions. Reasoning cost = reasoning tokens × the model's output price. Remaining output cost = (output tokens − reasoning tokens) × the same price. We use the rest as a visible-answer estimate. Input cost covers the whole prompt. Cache writes use the one-hour list-price assumption; the receipts do not state the cache lifetime. Prices per million output tokens: Claude Haiku 4.5 $5, Claude Sonnet 5.5 $10, Claude Opus 5.5 $20, Claude Fable 5.1 $50, GPT-6.1 Sol $10. Sources: Anthropic list prices of 2026-09-21, OpenAI of 2026-10-03.
- Cost per strict pass = the list-price cost of all calls in the cell, failures included, ÷ the cell's strict passes. Hard and effort cells use the hard-set strict rule. Short-task cells use their original rule, which can strip a wrapping fence.
- Time link: we use all calls of the hard head-to-head and the effort ladder. For each model and route, we compute the Spearman rank correlation of reasoning tokens with total time (with a leave-one-task-out sensitivity range, not a confidence interval). We also compute the least-squares slope in seconds per 1,000 reasoning tokens. Then we compute the slope within task. We centre each task on its own mean. This controls for differences in task means, not effort or batch effects.
- Isolation, tasks, validators and flags follow the hard head-to-head and the five-task head-to-head. Each call ran in a fresh empty folder with tools off and one turn. The timeout was 300 s per call (180 s for the short tasks). One call ran at a time per account.
- Claude Code calls ran with an output cap of 16,000 tokens in all three runs. The highest output of any analysed Claude call was 9,321, so no call hit the cap. The Codex CLI has no cap setting. Its highest output was 2,569.
Caveats
- The hard-set protocol file was created after its first counted call. Both effort-ladder protocol files were created after their batches ended. The short-set file predates its first call, but its top-up amendment timing is unverified. Batch receipts preserve protocol text, but we cannot verify all rules were written before inference. Treat these as exploratory calculations, not preregistered tests.
- The calls used one shared Mac and network. Other work and provider load were not controlled. Latency includes those effects.
- These are hand-built, tuned case sets, not random workload samples. The short set and most hard-set configurations hit a pass ceiling. Repeats of the same tasks are not independent. Wilson pass intervals describe counted calls under an independence assumption, not performance on new tasks.
- Each CLI reports its own reasoning counter, and we never see the reasoning text. So we cannot check what each vendor counts. The sum check shows only that the counters are consistent with inclusion in output; it does not prove their semantics. Shares from Claude Code and Codex CLI do not measure like for like how much each model thinks.
- Every prompt asks for a short reply in a strict format (code, JSON, one line or a regex, with no explanation). So the visible answer is short and the reasoning share is high. Longer visible replies could change the share. We did not measure that workload.
- Small cells: 16 to 24 calls per configuration over 8 tasks. Per-call ranges are wide and overlap for every pair of configurations. So the medians describe this run and rank nothing. A sentence says one side is ahead only when its ranges do not overlap.
- Input includes CLI context that these receipts do not separately count. It moves with cache hits. So the part of the call cost that goes to reasoning depends on the CLI and on the cache, not only on the model. GPT-6.1 Sol at medium and at high effort ran in different batches and show very different input cost per call.
- Every effort-ladder cell passed 16/16, so the set has a ceiling. Higher effort had higher mean recorded reasoning cost in these batches. It does not show that more thinking never helps on harder work.
- The time link is a correlation, not a cause. Task, effort and batch can affect reasoning, visible output and time together. Total time includes CLI start-up. The within-task slope controls for differences in task means. It does not remove effort or batch effects. Output tokens (reasoning plus visible answer) track time at least as closely as reasoning alone. The Codex CLI slope (36.4 s per 1,000 reasoning tokens) has a calculated ratio of 2.5 to 4.8 times the Claude Code slopes. We did not test why. Calls are repeats of 8 tasks, so they are not independent. The ranges omit one whole task at a time. They show sensitivity to the task mix, not 95% coverage.
- Effort levels are not the same scale across vendors. "Default" means we did not pass the effort flag, and the CLI chose. The effort-ladder reference cells ran in a different batch and hour than the new cells.
- List-price costs are calculations, because the calls used flat subscriptions. A price change moves every cost here. It leaves every token count unchanged.
Sources
Reasoning token bill (calculation)
Reported reasoning tokens priced at the recorded list prices. A calculation, not a new run.
Provider head-to-head, hard set: eight hard tasks with strict validators
Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.
Effort ladder: the hard task set at each effort level
The eight hard tasks of the hard head-to-head at low, medium and high effort (Claude Sonnet 5.5 and Claude Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI), 8 tasks × 2 repetitions per cell, declared protocols, every attempt kept. Efforts the hard head-to-head already ran are reused from its receipts as reference cells.
Provider head-to-head: Claude Code models vs Codex efforts
Five short tasks with deterministic validators, declared protocol, every attempt kept.
Anthropic list prices (Claude models)
Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.
Token prices as listed by the vendor on 2026-10-03.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “How much of an AI bill is thinking? Reasoning tokens by model and effort”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/thinking-token-bill.
More comparisons based on this study (18)
These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.
- Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI)
- Claude Haiku 4.5 vs Claude Fable 5.1
- Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)
- Claude Opus 5.5: high vs default effort
- Claude Opus 5.5: low vs default effort
- Claude Opus 5.5: low vs high effort
- Claude Opus 5.5: low vs medium effort
- Claude Opus 5.5: medium vs default effort
- Claude Opus 5.5: medium vs high effort
- Claude Sonnet 5.5: high vs default effort
- Claude Sonnet 5.5: low vs default effort
- Claude Sonnet 5.5: low vs high effort
- Claude Sonnet 5.5: low vs medium effort
- Claude Sonnet 5.5: medium vs default effort
- Claude Sonnet 5.5: medium vs high effort
- GPT-6.1 Sol (Codex CLI): low vs high effort
- GPT-6.1 Sol (Codex CLI): low vs medium effort
- GPT-6.1 Sol (Codex CLI): medium vs high effort
Models and comparisons in this study
- Claude Sonnet 5.5
- Claude Opus 5.5
- Claude Haiku 4.5
- Claude Fable 5.1
- GPT-6.1 Sol (Codex CLI)
- Claude Code
- Codex CLI
- Claude Sonnet 5.5 vs Claude Opus 5.5
- Claude Haiku 4.5 vs Claude Sonnet 5.5
- Claude Opus 5.5 vs Claude Fable 5.1
- Claude Sonnet 5.5 vs Claude Fable 5.1
- Claude Haiku 4.5 vs Claude Opus 5.5
- Claude Code vs Codex CLI
- Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)
- Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI)
- All comparisons
Write-ups on this study
How much of your AI bill is thinking tokens? Claude and GPT-6.1 Sol, measured
Thinking tokens: 46% to 92% of output per median call, 9% to 80% of pooled list-price cost. Calculation over recorded Claude and GPT-6.1 Sol calls.
More studies
All benchmarksWhere the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.
Prompt cache break-even: after how many reuses does a cached prefix cost less?
A calculation on list prices: reuses before a cached prompt prefix costs less, per Claude and GPT model, and cost per 1,000 sessions of 1 to 20 turns.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
176 calls on 8 hard tasks at low, medium, high and default effort: Sonnet, Opus and GPT-6.1 Sol. Pass rate with 95% intervals, time, tokens, cost per pass.