• Effort
  • Reasoning Effort
  • Hard Tasks
  • Claude Sonnet
  • Claude Opus
  • GPT-6.1 Sol
  • Claude Code
  • Codex CLI
  • Latency
  • Tokens

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks

On 8 hard tasks with strict validators, does a higher effort setting buy a higher pass rate for Sonnet, Opus and GPT-6.1 Sol, and what does it cost in time, tokens and list price per pass?

Published · 5 charts · Download the data or a carousel

100%

95% CI 96%–100% · n = 96

96/96 · New effort-ladder calls that passed strictly

The answer

Not on this set. Every one of the 11 configurations (Sonnet at low, medium, high and default, Opus at low, medium, high and default and GPT-6.1 Sol through Codex CLI at low, medium and high) passed all 16 calls strictly (95% interval 81% to 100% each), so pass rate does not separate any effort level. The hard set has a ceiling for these models: with 16 calls per cell it cannot rule out a difference of up to about 19 points. What more effort did change is output and time. Median total time per call by effort: Sonnet: low 5.8 s, medium 7.6 s, high 8.8 s, default 8.0 s (median output tokens 667 / 770 / 1,192 / 1,054); Opus: low 7.5 s, medium 9.7 s, high 10.1 s, default 9.2 s (median output tokens 594 / 853 / 1,052 / 945); GPT-6.1 Sol (Codex CLI): low 13.6 s, medium 13.1 s, high 18.1 s (median output tokens 284 / 335 / 436). Within each model, the fastest and slowest calls of every effort overlap, so these medians describe this run; they are not a tested ranking. At list price (a calculation; the calls ran on subscriptions), cost per strict pass went from $0.0122 at low to $0.0167 at high for Sonnet, $0.0212 at low to $0.0337 at high for Opus and $0.0128 at low to $0.0151 at high for GPT-6.1 Sol (Codex CLI).

Every metric on one effort axis: hover or focus a rung for all four values at once, or pick a model to isolate it.

Strict pass rate by effort on eight hard tasks · 95% Wilson interval

Median total time per call, by effort

Output tokens per call by effort on hard tasks

List-price cost per strict pass by effort (calculation)

* default: the effort flag was not passed; its level is not known, so no line joins it.

11 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 16 per row11 of 11 at 100%: this task set cannot separate them.

Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.

Source: Effort ladder: the hard task set at each effort level

Share card (PNG)

Live story

Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.

Live story · 48 sDoes more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.

Transcript
  1. Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
  2. 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  3. All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  4. Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
  5. More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
  6. List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
  7. 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  8. Start at low effort and measure. Every call, interval and cost online.

Key numbers

100% (176/176)

All effort-ladder calls that passed strictly, reference cells included

95% CI 98%–100% · n = 176

11

Configurations on the ladder (new + reference)

(6 new, 5 reference) · n = 176

0

Format misses and wrong answers on the ladder

format misses, 0 wrong answers · n = 176

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Entrance: medians race at 13× real timeMotion reduced: press Replay to animateThe slowest median is 18.1 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

11 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 18.1 s (range 11.7 s–92.2 s, n 16). Fastest Claude Sonnet 5.5 (low) · Claude Code 5.8 s (range 2.8 s–20 s, n 16). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 16 per row

Median per configuration; whiskers = fastest and slowest call

Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.

Source: Effort ladder: the hard task set at each effort level

Share card (PNG)
The ladder’s panels as single charts (3)

Seconds (median)

* default: the effort flag was not passed; its level is not known, so no line joins it.

4 efforts, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol · Codex CLI. Claude Sonnet 5.5 · Claude Code: slowest high 8.8 s (n 16). Fastest low 5.8 s (n 16). Claude Opus 5.5 · Claude Code: slowest high 10.1 s (n 16). Fastest low 7.5 s (n 16).

Notesn = 16 per row

One line per model and route; default = the effort flag was not passed

Medians only; the per-call ranges are in the total-time chart and they overlap. "default" is placed last because its level is not known: the CLI chose it.

Source: Effort ladder: the hard task set at each effort level

Share card (PNG)
  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

11 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 (high) · Claude Code 1,192 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 284 (n 16). Reasoning tokens: highest Claude Sonnet 5.5 (high) · Claude Code 745 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 63 (n 16).

Notesn = 16 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean "not reported". Their content is never captured. More tokens is not better or worse by itself.

Source: Effort ladder: the hard task set at each effort level

Share card (PNG)
Calculation
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (low) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 11 rows. Highest Claude Opus 5.5 (high) · Claude Code $0.034 (n 16). Lowest Claude Sonnet 5.5 (low) · Claude Code $0.012 (n 16).

Notesn = 16 per row

All calls in a configuration divided by its strict passes

Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.

Sources: Effort ladder: the hard task set at each effort level, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)

Tables

Every effort-ladder cell

ConfigurationEffortCellStrict passes95% intervalFormat missesWrong answersMedian total (s)Fastest to slowest (s)Median output tokensMean output tokensMedian reasoning tokensUSD per strict pass (calculation)
Claude Sonnet 5.5 (low) · Claude Codelownew run16/1681% to 100%005.8 s2.8 to 20667810273$0.012
Claude Sonnet 5.5 (medium) · Claude Codemediumnew run16/1681% to 100%007.6 s2.7 to 24770966422$0.014
Claude Sonnet 5.5 (high) · Claude Codehighnew run16/1681% to 100%008.8 s2.9 to 35.81,1921,284745$0.017
Claude Sonnet 5.5 · Claude Codedefaultreference (hard head-to-head)16/1681% to 100%008 s2.3 to 21.61,054989668$0.014
Claude Opus 5.5 (low) · Claude Codelownew run16/1681% to 100%007.5 s3.3 to 15.859466587$0.021
Claude Opus 5.5 (medium) · Claude Codemediumnew run16/1681% to 100%009.7 s4.8 to 31.48531,104518$0.029
Claude Opus 5.5 (high) · Claude Codehighreference (hard head-to-head)16/1681% to 100%0010.1 s3.6 to 631,0521,314614$0.034
Claude Opus 5.5 · Claude Codedefaultreference (hard head-to-head)16/1681% to 100%009.2 s4.2 to 27.29451,053538$0.029
GPT-6.1 Sol (low) · Codex CLIlownew run16/1681% to 100%0013.6 s7.9 to 44.328442063$0.013
GPT-6.1 Sol (medium) · Codex CLImediumreference (hard head-to-head)16/1681% to 100%0013.1 s8.5 to 61.6335530150$0.026
GPT-6.1 Sol (high) · Codex CLIhighreference (hard head-to-head)16/1681% to 100%0018.1 s11.7 to 92.2436702225$0.015

Method

  1. Protocols declared before the first call, one per route. A follow-up to the hard head-to-head (/benchmarks/hard-model-head-to-head), on the same 8 tasks, validators, controls and CLI flags.
  2. New cells: Claude Sonnet 5.5 (low) · Claude Code; Claude Sonnet 5.5 (medium) · Claude Code; Claude Sonnet 5.5 (high) · Claude Code; Claude Opus 5.5 (low) · Claude Code; Claude Opus 5.5 (medium) · Claude Code; GPT-6.1 Sol (low) · Codex CLI. 96 calls (8 tasks × 2 repetitions per cell), one call at a time per account; order rep-major, then task, then configuration.
  3. Reference cells (5): Claude Sonnet 5.5 · Claude Code; Claude Opus 5.5 (high) · Claude Code; Claude Opus 5.5 · Claude Code; GPT-6.1 Sol (medium) · Codex CLI; GPT-6.1 Sol (high) · Codex CLI, from the hard head-to-head. Reused, not rerun. Claude reference cells keep repetitions 1-2 (n = 16) so every cell has the same design; all 24 of their calls passed in the hard study.
  4. Effort: "--effort <level>" for Claude Code and the thread effort for Codex CLI. "default" means the flag was not passed and the CLI chose the level.
  5. Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 wrapped references are flagged as format misses.
  6. Strict pass, format miss and wrong answer as in the hard head-to-head. Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn, 300 s timeout per call.
  7. Stop rules: stop at the first usage-limit or rate-limit text. No batch stopped early, nothing was trimmed or retried, and no call failed.
  8. Cost per strict pass: list price × reported tokens for every call in the cell, divided by its strict passes. A calculation.

Caveats

  • Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  • Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  • Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
  • Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
  • List-price costs are calculations; the calls used flat subscriptions.

Sources

  • Effort ladder: the hard task set at each effort level

    Our recorded runs ·

    The eight hard tasks of the hard head-to-head at low, medium and high effort (Claude Sonnet 5.5 and Claude Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI), 8 tasks × 2 repetitions per cell, declared protocols, every attempt kept. Efforts the hard head-to-head already ran are reused from its receipts as reference cells.

    Raw data: effort-ladder/receipts.json, provider-h2h-hard/receipts.json

  • Repricing calculation

    Calculation ·

    Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • OpenAI list prices

    Vendor price list ·

    Token prices as listed by the vendor on 2026-10-03.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/effort-ladder.

Explainers that cite this study

Read the methods and terms in the context of these recorded results.

More comparisons based on this study (15)

These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.

More write-ups that cite this study (1)

Models and comparisons in this study

More studies

All benchmarks
Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Live story
  • Head to head
  • Hard Tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.

91% (139/152)Calls that passed strictly (hard set) · n = 152

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.