• Effort
  • Reasoning Effort
  • Hard Tasks
  • Claude Sonnet

Does reasoning effort buy quality? Claude and Codex on hard tasks, low to high

176 calls on 8 hard tasks: Sonnet, Opus and GPT-6.1 Sol at low, medium and high effort. Every cell passed 16/16. Higher effort cost more tokens and money.

TL;DR

  • We ran Claude Sonnet 5.5, Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) at low, medium and high effort on 8 hard tasks with strict validators: 11 configurations, 176 calls, 16 per cell.
  • Every cell passed 16 of 16 strictly (95% interval 81% to 100%), with 0 format misses and 0 wrong answers. On this set, more effort bought no extra passes.
  • The set has a ceiling. With 16 calls per cell, it cannot rule out a gap of up to about 19 points. We do not claim that effort never matters.
  • More effort raised median output tokens, and for both Claude models the median time. Median time per call went from 5.8 s to 8.8 s for Sonnet and from 7.5 s to 10.1 s for Opus, low to high. The ranges overlap, so this is not a tested ranking.
  • List-price cost per strict pass (a calculation) rose from low to high: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151.
  • All 16 effort comparisons in the dataset are tie or unclear. No effort level wins a row.

Study: /benchmarks/effort-ladder.

The question

Claude Code has an --effort flag. The Codex CLI has a reasoning effort per thread. Both promise the same trade: let the model think longer, and it gets more answers right. New to the setting? Read what reasoning effort is.

Higher effort costs something you can see at once: more output tokens and a longer wait. The gain is harder to see. So we asked one narrow question:

On 8 hard tasks with strict validators, does a higher effort setting buy a higher pass rate, and what does it cost in time, tokens and list price per pass?

What we ran

The effort ladder is a follow-up to our hard head-to-head. It uses the same 8 tasks, the same validators, the same controls and the same CLI flags. The only change is the effort setting.

  • New cells (6): Sonnet at low, medium and high; Opus at low and medium; GPT-6.1 Sol at low. That is 96 new calls: 8 tasks × 2 repetitions per cell.
  • Reference cells (5): Sonnet and Opus at default effort, Opus at high, and GPT-6.1 Sol at medium and high, read from the hard head-to-head. We reused them; we did not run them again. The Claude reference cells keep repetitions 1 and 2 only, so every cell has n = 16.
  • "Default" means we did not pass the flag, and the CLI chose the level.
  • Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 wrapped references are flagged as format misses.
  • Isolation: fresh empty folder, tools off, no MCP servers, one turn, 300 s timeout. Nothing was retried, trimmed or stopped early, and no call failed.

Quality: 176 of 176

Every rate is 95% or more
Claude Sonnet 5.5 (low)
Claude Sonnet 5.5 (medium)
Claude Sonnet 5.5 (high)
Claude Sonnet 5.5
Claude Opus 5.5 (low)
Claude Opus 5.5 (medium)
Claude Opus 5.5 (high)
Claude Opus 5.5
GPT-6.1 Sol (low)
GPT-6.1 Sol (medium)
GPT-6.1 Sol (high)

Every interval overlaps every other: this chart does not order these rows.

11 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 16 per row11 of 11 at 100%: this task set cannot separate them.

Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.

Source: Effort ladder: the hard task set at each effort level

All 11 cells passed all 16 calls. Across the ladder, 176 of 176 calls passed strictly (95% interval 98% to 100%).

This is a clear result, but read it the right way. A perfect 16/16 has a 95% interval of 81% to 100%. Two cells that both score 16/16 can still differ by up to about 19 points in their true pass rate. These 8 tasks do not need more than low effort for these three models. That does not show that harder or longer work would not need more.

The comparison pages say the same thing in a fixed format. The dataset holds 16 effort comparisons, one per model and effort pair, for example Sonnet low vs high. Both sides of each row come from the same chart, with the same route and task set; only the effort differs. Every one of the 16 records is tie or unclear.

Time: more effort, longer waits (probably)

Seconds (median)

* default: the effort flag was not passed; its level is not known, so no line joins it.

4 efforts, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol · Codex CLI. Claude Sonnet 5.5 · Claude Code: slowest high 8.8 s (n 16). Fastest low 5.8 s (n 16). Claude Opus 5.5 · Claude Code: slowest high 10.1 s (n 16). Fastest low 7.5 s (n 16).

Notesn = 16 per row

One line per model and route; default = the effort flag was not passed

Medians only; the per-call ranges are in the total-time chart and they overlap. "default" is placed last because its level is not known: the CLI chose it.

Source: Effort ladder: the hard task set at each effort level

Median total time per call:

ConfigurationLowMediumHighDefault
Claude Sonnet 5.5 · Claude Code5.8 s7.6 s8.8 s8.0 s
Claude Opus 5.5 · Claude Code7.5 s9.7 s10.1 s9.2 s
GPT-6.1 Sol · Codex CLI13.6 s13.1 s18.1 s(not run)

For both Claude models, the median goes up with each step from low to high. For GPT-6.1 Sol, low and medium are close, and high is slower.

But each median rests on 16 calls, and the single-call ranges are wide. Sonnet at low ran from 2.8 s to 20.0 s; at high, from 2.9 s to 35.8 s. Within each model, the fastest-to-slowest ranges of every effort overlap. So these medians describe this run. They are not a tested ranking.

Entrance: medians race at 13× real timeMotion reduced: press Replay to animateThe slowest median is 18.1 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

11 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 18.1 s (range 11.7 s–92.2 s, n 16). Fastest Claude Sonnet 5.5 (low) · Claude Code 5.8 s (range 2.8 s–20 s, n 16). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 16 per row

Median per configuration; whiskers = fastest and slowest call

Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.

Source: Effort ladder: the hard task set at each effort level

Tokens: where the time goes

  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

11 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 (high) · Claude Code 1,192 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 284 (n 16). Reasoning tokens: highest Claude Sonnet 5.5 (high) · Claude Code 745 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 63 (n 16).

Notesn = 16 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean "not reported". Their content is never captured. More tokens is not better or worse by itself.

Source: Effort ladder: the hard task set at each effort level

Median output tokens per call, low → medium → high:

  • Sonnet: 667 → 770 → 1,192 (default 1,054)
  • Opus: 594 → 853 → 1,052 (default 945)
  • GPT-6.1 Sol: 284 → 335 → 436

Reasoning tokens are part of those figures where the CLI reports them. They grow faster than the total: Sonnet 273 → 422 → 745, Opus 87 → 518 → 614, GPT-6.1 Sol 63 → 150 → 225.

More tokens is not better or worse by itself. Here it describes behaviour: the model wrote more, and the answer did not change from pass to fail or back.

Cost per strict pass

Calculation
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (low) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 11 rows. Highest Claude Opus 5.5 (high) · Claude Code $0.034 (n 16). Lowest Claude Sonnet 5.5 (low) · Claude Code $0.012 (n 16).

Notesn = 16 per row

All calls in a configuration divided by its strict passes

Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.

Sources: Effort ladder: the hard task set at each effort level, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

This is a calculation: list price × reported tokens for every call in the cell, divided by its strict passes. The calls ran on flat subscriptions.

ConfigurationLowMediumHighDefault
Claude Sonnet 5.5 · Claude Code$0.0122$0.0135$0.0167$0.0140
Claude Opus 5.5 · Claude Code$0.0212$0.0295$0.0337$0.0289
GPT-6.1 Sol · Codex CLI$0.0128$0.0256$0.0151(not run)

Our arithmetic on those values: high effort cost 1.4x low per pass for Sonnet, 1.6x for Opus and 1.2x for GPT-6.1 Sol. Since every cell passed everything, that extra spend bought no extra passes on this set.

One value does not follow the pattern. GPT-6.1 Sol at medium cost more per pass than at high ($0.0256 against $0.0151). Medium is a reference cell from a different batch. The output-token medians do not explain it (335 against 436), and we did not test the cause. These costs have no interval, so read small gaps as a direction only.

Two more ratios from the table, also our arithmetic: Opus at low cost 1.7x Sonnet at low per pass, and Opus at high cost 2.8x Sonnet at low. On this set, the cheapest setting that passed everything was Sonnet at low effort, with GPT-6.1 Sol at low close behind.

What this means for your settings

  1. Start low, then measure. On tasks of this size, low effort passed everything for all three models and was the cheapest per pass. Raise the effort only when your own validators show failures at low.
  2. Effort is not a quality dial you can trust blind. On this set, the extra tokens did not change one result.
  3. Effort levels do not match across vendors. "High" in Claude Code and "high" in the Codex CLI are different scales. Compare within one model.
  4. Long, open work may differ. Our tasks are single-turn and short. A multi-step agent task may gain from effort where these did not. We have not measured that.

How we measured

  • Tasks: the 8 hard tasks of the hard head-to-head: bug fixes, a CSV parser, event-loop output order, a constraint schedule, a strict SemVer regex and more, each with a deterministic validator.
  • Outcomes: strict pass, format miss (right answer, wrong format) and wrong answer. Every attempt is counted.
  • Intervals: Wilson 95% intervals for rates; fastest-to-slowest ranges for times. A range is not a confidence interval.
  • Costs: reported tokens × list price per call, cache reads and writes priced apart. A calculation, not a bill.

Caveats

  • Ceiling. Every cell scored 16/16, so the set cannot separate effort levels by quality.
  • Different batches. The 5 reference cells ran at a different time from the 6 new cells, on the same host, CLI versions, cases and flags. Provider load can differ.
  • Small cells. 2 calls per task per cell. A few slow calls move a median of 16.
  • Route + model. Each row pairs a CLI with a model. Each CLI adds its own system prompt and start-up time.

Pay for effort only where it helps

Agent records the model, the effort, the tokens and the validation result of every step, so you can see where a higher setting changes the outcome. Try Agent and measure it on your own work.

The data behind this post

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.