Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
On 8 hard tasks with strict validators, does a higher effort setting buy a higher pass rate for Sonnet, Opus and GPT-6.1 Sol, and what does it cost in time, tokens and list price per pass?
Published · 5 charts · Download the data or a carousel
100%
The answer
Not on this set. Every one of the 11 configurations (Sonnet at low, medium, high and default, Opus at low, medium, high and default and GPT-6.1 Sol through Codex CLI at low, medium and high) passed all 16 calls strictly (95% interval 81% to 100% each), so pass rate does not separate any effort level. The hard set has a ceiling for these models: with 16 calls per cell it cannot rule out a difference of up to about 19 points. What more effort did change is output and time. Median total time per call by effort: Sonnet: low 5.8 s, medium 7.6 s, high 8.8 s, default 8.0 s (median output tokens 667 / 770 / 1,192 / 1,054); Opus: low 7.5 s, medium 9.7 s, high 10.1 s, default 9.2 s (median output tokens 594 / 853 / 1,052 / 945); GPT-6.1 Sol (Codex CLI): low 13.6 s, medium 13.1 s, high 18.1 s (median output tokens 284 / 335 / 436). Within each model, the fastest and slowest calls of every effort overlap, so these medians describe this run; they are not a tested ranking. At list price (a calculation; the calls ran on subscriptions), cost per strict pass went from $0.0122 at low to $0.0167 at high for Sonnet, $0.0212 at low to $0.0337 at high for Opus and $0.0128 at low to $0.0151 at high for GPT-6.1 Sol (Codex CLI).
Every metric on one effort axis: hover or focus a rung for all four values at once, or pick a model to isolate it.
Strict pass rate by effort on eight hard tasks · 95% Wilson interval
Median total time per call, by effort
Output tokens per call by effort on hard tasks
List-price cost per strict pass by effort (calculation)
* default: the effort flag was not passed; its level is not known, so no line joins it.
Strict pass rate by effort on eight hard tasks
| Item | Strict pass | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 (low) · Claude Code | 100% | 81%–100% | 16 |
| Claude Sonnet 5.5 (medium) · Claude Code | 100% | 81%–100% | 16 |
| Claude Sonnet 5.5 (high) · Claude Code | 100% | 81%–100% | 16 |
| Claude Sonnet 5.5 · Claude Code | 100% | 81%–100% | 16 |
| Claude Opus 5.5 (low) · Claude Code | 100% | 81%–100% | 16 |
| Claude Opus 5.5 (medium) · Claude Code | 100% | 81%–100% | 16 |
| Claude Opus 5.5 (high) · Claude Code | 100% | 81%–100% | 16 |
| Claude Opus 5.5 · Claude Code | 100% | 81%–100% | 16 |
| GPT-6.1 Sol (low) · Codex CLI | 100% | 81%–100% | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | 100% | 81%–100% | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 100% | 81%–100% | 16 |
Median total time per call, by effort
| Effort | Claude Sonnet 5.5 · Claude Code | Claude Opus 5.5 · Claude Code | GPT-6.1 Sol · Codex CLI | n |
|---|---|---|---|---|
| low | 5.8 s | 7.5 s | 13.6 s | 16 |
| medium | 7.6 s | 9.7 s | 13.1 s | 16 |
| high | 8.8 s | 10.1 s | 18.1 s | 16 |
| default | 8 s | 9.2 s | — | 16 |
Output tokens per call by effort on hard tasks
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Sonnet 5.5 (low) · Claude Code | 667 | 273 | 16 |
| Claude Sonnet 5.5 (medium) · Claude Code | 770 | 422 | 16 |
| Claude Sonnet 5.5 (high) · Claude Code | 1,192 | 745 | 16 |
| Claude Sonnet 5.5 · Claude Code | 1,054 | 668 | 16 |
| Claude Opus 5.5 (low) · Claude Code | 594 | 87 | 16 |
| Claude Opus 5.5 (medium) · Claude Code | 853 | 518 | 16 |
| Claude Opus 5.5 (high) · Claude Code | 1,052 | 614 | 16 |
| Claude Opus 5.5 · Claude Code | 945 | 538 | 16 |
| GPT-6.1 Sol (low) · Codex CLI | 284 | 63 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | 335 | 150 | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 436 | 225 | 16 |
List-price cost per strict pass by effort (calculation)
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Sonnet 5.5 (low) · Claude Code | $0.012 | 16 |
| Claude Sonnet 5.5 (medium) · Claude Code | $0.014 | 16 |
| Claude Sonnet 5.5 (high) · Claude Code | $0.017 | 16 |
| Claude Sonnet 5.5 · Claude Code | $0.014 | 16 |
| Claude Opus 5.5 (low) · Claude Code | $0.021 | 16 |
| Claude Opus 5.5 (medium) · Claude Code | $0.029 | 16 |
| Claude Opus 5.5 (high) · Claude Code | $0.034 | 16 |
| Claude Opus 5.5 · Claude Code | $0.029 | 16 |
| GPT-6.1 Sol (low) · Codex CLI | $0.013 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.026 | 16 |
| GPT-6.1 Sol (high) · Codex CLI | $0.015 | 16 |
11 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 16 per row11 of 11 at 100%: this task set cannot separate them.
Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.
Source: Effort ladder: the hard task set at each effort level
Live story
Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks
All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.
Transcript
- Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
- 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
- All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
- More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
- List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
- 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Start at low effort and measure. Every call, interval and cost online.
Key numbers
100% (176/176)
All effort-ladder calls that passed strictly, reference cells included
95% CI 98%–100% · n = 176
11
Configurations on the ladder (new + reference)
(6 new, 5 reference) · n = 176
0
Format misses and wrong answers on the ladder
format misses, 0 wrong answers · n = 176
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call by effort on hard tasks | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 (low) · Claude Code | 5.8 s | 2.8 s–20 s | 16 |
| Claude Sonnet 5.5 (medium) · Claude Code | 7.6 s | 2.7 s–24 s | 16 |
| Claude Sonnet 5.5 (high) · Claude Code | 8.8 s | 2.9 s–35.8 s | 16 |
| Claude Sonnet 5.5 · Claude Code | 8 s | 2.3 s–21.6 s | 16 |
| Claude Opus 5.5 (low) · Claude Code | 7.5 s | 3.3 s–15.8 s | 16 |
| Claude Opus 5.5 (medium) · Claude Code | 9.7 s | 4.8 s–31.4 s | 16 |
| Claude Opus 5.5 (high) · Claude Code | 10.1 s | 3.6 s–63 s | 16 |
| Claude Opus 5.5 · Claude Code | 9.2 s | 4.2 s–27.2 s | 16 |
| GPT-6.1 Sol (low) · Codex CLI | 13.6 s | 7.9 s–44.3 s | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | 13.1 s | 8.5 s–61.6 s | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 18.1 s | 11.7 s–92.2 s | 16 |
11 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 18.1 s (range 11.7 s–92.2 s, n 16). Fastest Claude Sonnet 5.5 (low) · Claude Code 5.8 s (range 2.8 s–20 s, n 16). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 16 per row
Median per configuration; whiskers = fastest and slowest call
Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.
Source: Effort ladder: the hard task set at each effort level
The ladder’s panels as single charts (3)
Seconds (median)
* default: the effort flag was not passed; its level is not known, so no line joins it.
| Effort | Claude Sonnet 5.5 · Claude Code | Claude Opus 5.5 · Claude Code | GPT-6.1 Sol · Codex CLI | n |
|---|---|---|---|---|
| low | 5.8 s | 7.5 s | 13.6 s | 16 |
| medium | 7.6 s | 9.7 s | 13.1 s | 16 |
| high | 8.8 s | 10.1 s | 18.1 s | 16 |
| default | 8 s | 9.2 s | — | 16 |
4 efforts, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol · Codex CLI. Claude Sonnet 5.5 · Claude Code: slowest high 8.8 s (n 16). Fastest low 5.8 s (n 16). Claude Opus 5.5 · Claude Code: slowest high 10.1 s (n 16). Fastest low 7.5 s (n 16).
Notesn = 16 per row
One line per model and route; default = the effort flag was not passed
Medians only; the per-call ranges are in the total-time chart and they overlap. "default" is placed last because its level is not known: the CLI chose it.
Source: Effort ladder: the hard task set at each effort level
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Sonnet 5.5 (low) · Claude Code | 667 | 273 | 16 |
| Claude Sonnet 5.5 (medium) · Claude Code | 770 | 422 | 16 |
| Claude Sonnet 5.5 (high) · Claude Code | 1,192 | 745 | 16 |
| Claude Sonnet 5.5 · Claude Code | 1,054 | 668 | 16 |
| Claude Opus 5.5 (low) · Claude Code | 594 | 87 | 16 |
| Claude Opus 5.5 (medium) · Claude Code | 853 | 518 | 16 |
| Claude Opus 5.5 (high) · Claude Code | 1,052 | 614 | 16 |
| Claude Opus 5.5 · Claude Code | 945 | 538 | 16 |
| GPT-6.1 Sol (low) · Codex CLI | 284 | 63 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | 335 | 150 | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 436 | 225 | 16 |
11 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 (high) · Claude Code 1,192 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 284 (n 16). Reasoning tokens: highest Claude Sonnet 5.5 (high) · Claude Code 745 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 63 (n 16).
Notesn = 16 per row
Median per configuration; reasoning tokens as the CLI reports them
Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean "not reported". Their content is never captured. More tokens is not better or worse by itself.
Source: Effort ladder: the hard task set at each effort level
Hover or focus a bar for its ratio to Claude Sonnet 5.5 (low) (the lowest value): a ratio of list-price calculations, not a measurement.
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Sonnet 5.5 (low) · Claude Code | $0.012 | 16 |
| Claude Sonnet 5.5 (medium) · Claude Code | $0.014 | 16 |
| Claude Sonnet 5.5 (high) · Claude Code | $0.017 | 16 |
| Claude Sonnet 5.5 · Claude Code | $0.014 | 16 |
| Claude Opus 5.5 (low) · Claude Code | $0.021 | 16 |
| Claude Opus 5.5 (medium) · Claude Code | $0.029 | 16 |
| Claude Opus 5.5 (high) · Claude Code | $0.034 | 16 |
| Claude Opus 5.5 · Claude Code | $0.029 | 16 |
| GPT-6.1 Sol (low) · Codex CLI | $0.013 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.026 | 16 |
| GPT-6.1 Sol (high) · Codex CLI | $0.015 | 16 |
List-price calculation, not a run. 11 rows. Highest Claude Opus 5.5 (high) · Claude Code $0.034 (n 16). Lowest Claude Sonnet 5.5 (low) · Claude Code $0.012 (n 16).
Notesn = 16 per row
All calls in a configuration divided by its strict passes
Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.
Sources: Effort ladder: the hard task set at each effort level, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
Tables
Every effort-ladder cell
| Configuration | Effort | Cell | Strict passes | 95% interval | Format misses | Wrong answers | Median total (s) | Fastest to slowest (s) | Median output tokens | Mean output tokens | Median reasoning tokens | USD per strict pass (calculation) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 (low) · Claude Code | low | new run | 16/16 | 81% to 100% | 0 | 0 | 5.8 s | 2.8 to 20 | 667 | 810 | 273 | $0.012 |
| Claude Sonnet 5.5 (medium) · Claude Code | medium | new run | 16/16 | 81% to 100% | 0 | 0 | 7.6 s | 2.7 to 24 | 770 | 966 | 422 | $0.014 |
| Claude Sonnet 5.5 (high) · Claude Code | high | new run | 16/16 | 81% to 100% | 0 | 0 | 8.8 s | 2.9 to 35.8 | 1,192 | 1,284 | 745 | $0.017 |
| Claude Sonnet 5.5 · Claude Code | default | reference (hard head-to-head) | 16/16 | 81% to 100% | 0 | 0 | 8 s | 2.3 to 21.6 | 1,054 | 989 | 668 | $0.014 |
| Claude Opus 5.5 (low) · Claude Code | low | new run | 16/16 | 81% to 100% | 0 | 0 | 7.5 s | 3.3 to 15.8 | 594 | 665 | 87 | $0.021 |
| Claude Opus 5.5 (medium) · Claude Code | medium | new run | 16/16 | 81% to 100% | 0 | 0 | 9.7 s | 4.8 to 31.4 | 853 | 1,104 | 518 | $0.029 |
| Claude Opus 5.5 (high) · Claude Code | high | reference (hard head-to-head) | 16/16 | 81% to 100% | 0 | 0 | 10.1 s | 3.6 to 63 | 1,052 | 1,314 | 614 | $0.034 |
| Claude Opus 5.5 · Claude Code | default | reference (hard head-to-head) | 16/16 | 81% to 100% | 0 | 0 | 9.2 s | 4.2 to 27.2 | 945 | 1,053 | 538 | $0.029 |
| GPT-6.1 Sol (low) · Codex CLI | low | new run | 16/16 | 81% to 100% | 0 | 0 | 13.6 s | 7.9 to 44.3 | 284 | 420 | 63 | $0.013 |
| GPT-6.1 Sol (medium) · Codex CLI | medium | reference (hard head-to-head) | 16/16 | 81% to 100% | 0 | 0 | 13.1 s | 8.5 to 61.6 | 335 | 530 | 150 | $0.026 |
| GPT-6.1 Sol (high) · Codex CLI | high | reference (hard head-to-head) | 16/16 | 81% to 100% | 0 | 0 | 18.1 s | 11.7 to 92.2 | 436 | 702 | 225 | $0.015 |
Method
- Protocols declared before the first call, one per route. A follow-up to the hard head-to-head (/benchmarks/hard-model-head-to-head), on the same 8 tasks, validators, controls and CLI flags.
- New cells: Claude Sonnet 5.5 (low) · Claude Code; Claude Sonnet 5.5 (medium) · Claude Code; Claude Sonnet 5.5 (high) · Claude Code; Claude Opus 5.5 (low) · Claude Code; Claude Opus 5.5 (medium) · Claude Code; GPT-6.1 Sol (low) · Codex CLI. 96 calls (8 tasks × 2 repetitions per cell), one call at a time per account; order rep-major, then task, then configuration.
- Reference cells (5): Claude Sonnet 5.5 · Claude Code; Claude Opus 5.5 (high) · Claude Code; Claude Opus 5.5 · Claude Code; GPT-6.1 Sol (medium) · Codex CLI; GPT-6.1 Sol (high) · Codex CLI, from the hard head-to-head. Reused, not rerun. Claude reference cells keep repetitions 1-2 (n = 16) so every cell has the same design; all 24 of their calls passed in the hard study.
- Effort: "--effort <level>" for Claude Code and the thread effort for Codex CLI. "default" means the flag was not passed and the CLI chose the level.
- Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 wrapped references are flagged as format misses.
- Strict pass, format miss and wrong answer as in the hard head-to-head. Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn, 300 s timeout per call.
- Stop rules: stop at the first usage-limit or rate-limit text. No batch stopped early, nothing was trimmed or retried, and no call failed.
- Cost per strict pass: list price × reported tokens for every call in the cell, divided by its strict passes. A calculation.
Caveats
- Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
- Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
- Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
- List-price costs are calculations; the calls used flat subscriptions.
Sources
Effort ladder: the hard task set at each effort level
The eight hard tasks of the hard head-to-head at low, medium and high effort (Claude Sonnet 5.5 and Claude Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI), 8 tasks × 2 repetitions per cell, declared protocols, every attempt kept. Efforts the hard head-to-head already ran are reused from its receipts as reference cells.
Repricing calculation
Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.
Anthropic list prices (Claude models)
Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.
Token prices as listed by the vendor on 2026-10-03.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/effort-ladder.
Explainers that cite this study
Read the methods and terms in the context of these recorded results.
- Benchmark saturation: when every model scores 100%
- Cost per correct answer: the LLM price that counts failures
- How many runs does an LLM eval need? Sample size, with real intervals
- Reasoning tokens, explained: the hidden output you pay for
- Time to first token (TTFT), explained with CLI and API timings
- What is reasoning effort?
- Wilson confidence intervals for AI benchmarks
More comparisons based on this study (15)
These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.
- Claude Opus 5.5: high vs default effort
- Claude Opus 5.5: low vs default effort
- Claude Opus 5.5: low vs high effort
- Claude Opus 5.5: low vs medium effort
- Claude Opus 5.5: medium vs default effort
- Claude Opus 5.5: medium vs high effort
- Claude Sonnet 5.5: high vs default effort
- Claude Sonnet 5.5: low vs default effort
- Claude Sonnet 5.5: low vs high effort
- Claude Sonnet 5.5: low vs medium effort
- Claude Sonnet 5.5: medium vs default effort
- Claude Sonnet 5.5: medium vs high effort
- GPT-6.1 Sol (Codex CLI): low vs high effort
- GPT-6.1 Sol (Codex CLI): low vs medium effort
- GPT-6.1 Sol (Codex CLI): medium vs high effort
More write-ups that cite this study (1)
Models and comparisons in this study
Write-ups on this study
AI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
Format misses vs wrong answers: your LLM eval may be failing right answers
5 of 13 failed calls on our hard set were right answers in the wrong format. How to grade LLM output strictly, test validators first and report both numbers.
GPT-6.1 Sol vs Claude Sonnet 5.5 vs Opus 5.5: every row we measured
Of 70 comparison rows for GPT-6.1 Sol, Claude Sonnet 5.5 and Opus 5.5, only 2 have a winner (speed). All 15 pass-rate rows tie. Tokens, price and route differ.
How long does an AI coding agent take per task? Minutes, calls and where the time goes
Agent took a median 9.6 minutes per SWE-bench attempt (n = 33, range 1.6 to 54.1). Compare single calls, repairs and full tasks with limits.
How many runs do you need to compare two AI models? A sample-size table
Calculation: 100 runs can separate an observed 8-point gap at a perfect score. See Wilson intervals, paired tests and limits for LLM eval sample sizes.
How much of your AI bill is thinking tokens? Claude and GPT-6.1 Sol, measured
Thinking tokens: 46% to 92% of output per median call, 9% to 80% of pooled list-price cost. Calculation over recorded Claude and GPT-6.1 Sol calls.
Most AI model comparisons are ties: 44 of 1,196 rows show a clear gap
44 of 1,196 AI model comparison rows show a gap under our overlap rules. Most gaps are timing rows. No effort row separates quality.
Opus at low effort or Sonnet at high effort? A bigger model that thinks less, tested
Opus 5.5 at low effort and Sonnet 5.5 at high effort both passed 16/16. Opus low cost 1.3x as much per pass (calculation). Sonnet at low effort cost least.
Plan for p95, not the median: LLM tail latency in our runs
4.30 s p95 against a 2.60 s median for Claude Sonnet 5.5; 34.5 s against 12.5 s for Haiku 4.5. Measured LLM tail latency and what to do about it.
The cheapest way to run an AI coding agent: 7 levers from measured runs
7 levers that may cut an AI coding agent's bill, sized from our data: prompt cache 3.9x, Fable/Sonnet cost per pass 6.5x, and 5 more. List-price calculations.
What is the best AI model for coding? Our data says four models tie
4 models tied at the top of our hard coding set: Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol. Only Haiku 4.5 separated. A tier list built from intervals.
What thinking costs: reasoning tokens per call for Claude and GPT-6.1 Sol
Haiku 4.5 spent a median 4,556 reasoning tokens per hard call, Sonnet 5.5 585, GPT-6.1 Sol 150. List price per 1,000 calls: $22.78, $5.85, $1.50 (calculation).
Which Claude model is fastest? It depends on the task, and on thinking
1.94 s was the lowest median on short calls (Fable 5.1). 7.75 s on hard calls (Sonnet 5.5). Haiku 4.5 took 4.43 s and 39.01 s with default thinking.
Which Claude model should you use? A task-by-task guide from our measurements
Which Claude model fits your job? Sonnet passed 24/24 hard calls (95% interval 86%–100%) at $0.01435 per pass (calculation). The top models hit a ceiling.
Why is Claude Code slow? Where the seconds go in a coding CLI call, and what to change
A one-word Claude Code call took a median 2.5 s, with 1.7 s outside the model (n = 5). Where the rest goes: thinking, effort, route. Measured splits and fixes.
AI coding benchmarks roundup, October 2026: sixteen studies, every number in one place
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
An AI model leaderboard without a composite score: why, and how to read ours
Our leaderboard lists 46 models, CLIs, routers and providers with their best-supported facts, each with n and an interval. No single score. Here is why.
Does reasoning effort buy quality? Claude and Codex on hard tasks, low to high
176 calls on 8 hard tasks: Sonnet, Opus and GPT-6.1 Sol at low, medium and high effort. Every cell passed 16/16. Higher effort cost more tokens and money.
More studies
All benchmarksHow much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.