Does reasoning effort buy quality? Claude and Codex on hard tasks, low to high
176 calls on 8 hard tasks: Sonnet, Opus and GPT-6.1 Sol at low, medium and high effort. Every cell passed 16/16. Higher effort cost more tokens and money.
TL;DR
- We ran Claude Sonnet 5.5, Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) at low, medium and high effort on 8 hard tasks with strict validators: 11 configurations, 176 calls, 16 per cell.
- Every cell passed 16 of 16 strictly (95% interval 81% to 100%), with 0 format misses and 0 wrong answers. On this set, more effort bought no extra passes.
- The set has a ceiling. With 16 calls per cell, it cannot rule out a gap of up to about 19 points. We do not claim that effort never matters.
- More effort raised median output tokens, and for both Claude models the median time. Median time per call went from 5.8 s to 8.8 s for Sonnet and from 7.5 s to 10.1 s for Opus, low to high. The ranges overlap, so this is not a tested ranking.
- List-price cost per strict pass (a calculation) rose from low to high: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151.
- All 16 effort comparisons in the dataset are tie or unclear. No effort level wins a row.
Study: /benchmarks/effort-ladder.
The question
Claude Code has an --effort flag. The Codex CLI has a reasoning effort per thread. Both promise the same trade: let the model think longer, and it gets more answers right. New to the setting? Read what reasoning effort is.
Higher effort costs something you can see at once: more output tokens and a longer wait. The gain is harder to see. So we asked one narrow question:
On 8 hard tasks with strict validators, does a higher effort setting buy a higher pass rate, and what does it cost in time, tokens and list price per pass?
What we ran
The effort ladder is a follow-up to our hard head-to-head. It uses the same 8 tasks, the same validators, the same controls and the same CLI flags. The only change is the effort setting.
- New cells (6): Sonnet at low, medium and high; Opus at low and medium; GPT-6.1 Sol at low. That is 96 new calls: 8 tasks × 2 repetitions per cell.
- Reference cells (5): Sonnet and Opus at default effort, Opus at high, and GPT-6.1 Sol at medium and high, read from the hard head-to-head. We reused them; we did not run them again. The Claude reference cells keep repetitions 1 and 2 only, so every cell has n = 16.
- "Default" means we did not pass the flag, and the CLI chose the level.
- Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 wrapped references are flagged as format misses.
- Isolation: fresh empty folder, tools off, no MCP servers, one turn, 300 s timeout. Nothing was retried, trimmed or stopped early, and no call failed.
Quality: 176 of 176
All 11 cells passed all 16 calls. Across the ladder, 176 of 176 calls passed strictly (95% interval 98% to 100%).
This is a clear result, but read it the right way. A perfect 16/16 has a 95% interval of 81% to 100%. Two cells that both score 16/16 can still differ by up to about 19 points in their true pass rate. These 8 tasks do not need more than low effort for these three models. That does not show that harder or longer work would not need more.
The comparison pages say the same thing in a fixed format. The dataset holds 16 effort comparisons, one per model and effort pair, for example Sonnet low vs high. Both sides of each row come from the same chart, with the same route and task set; only the effort differs. Every one of the 16 records is tie or unclear.
Time: more effort, longer waits (probably)
Median total time per call:
| Configuration | Low | Medium | High | Default |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 5.8 s | 7.6 s | 8.8 s | 8.0 s |
| Claude Opus 5.5 · Claude Code | 7.5 s | 9.7 s | 10.1 s | 9.2 s |
| GPT-6.1 Sol · Codex CLI | 13.6 s | 13.1 s | 18.1 s | (not run) |
For both Claude models, the median goes up with each step from low to high. For GPT-6.1 Sol, low and medium are close, and high is slower.
But each median rests on 16 calls, and the single-call ranges are wide. Sonnet at low ran from 2.8 s to 20.0 s; at high, from 2.9 s to 35.8 s. Within each model, the fastest-to-slowest ranges of every effort overlap. So these medians describe this run. They are not a tested ranking.
Tokens: where the time goes
Median output tokens per call, low → medium → high:
- Sonnet: 667 → 770 → 1,192 (default 1,054)
- Opus: 594 → 853 → 1,052 (default 945)
- GPT-6.1 Sol: 284 → 335 → 436
Reasoning tokens are part of those figures where the CLI reports them. They grow faster than the total: Sonnet 273 → 422 → 745, Opus 87 → 518 → 614, GPT-6.1 Sol 63 → 150 → 225.
More tokens is not better or worse by itself. Here it describes behaviour: the model wrote more, and the answer did not change from pass to fail or back.
Cost per strict pass
This is a calculation: list price × reported tokens for every call in the cell, divided by its strict passes. The calls ran on flat subscriptions.
| Configuration | Low | Medium | High | Default |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.0122 | $0.0135 | $0.0167 | $0.0140 |
| Claude Opus 5.5 · Claude Code | $0.0212 | $0.0295 | $0.0337 | $0.0289 |
| GPT-6.1 Sol · Codex CLI | $0.0128 | $0.0256 | $0.0151 | (not run) |
Our arithmetic on those values: high effort cost 1.4x low per pass for Sonnet, 1.6x for Opus and 1.2x for GPT-6.1 Sol. Since every cell passed everything, that extra spend bought no extra passes on this set.
One value does not follow the pattern. GPT-6.1 Sol at medium cost more per pass than at high ($0.0256 against $0.0151). Medium is a reference cell from a different batch. The output-token medians do not explain it (335 against 436), and we did not test the cause. These costs have no interval, so read small gaps as a direction only.
Two more ratios from the table, also our arithmetic: Opus at low cost 1.7x Sonnet at low per pass, and Opus at high cost 2.8x Sonnet at low. On this set, the cheapest setting that passed everything was Sonnet at low effort, with GPT-6.1 Sol at low close behind.
What this means for your settings
- Start low, then measure. On tasks of this size, low effort passed everything for all three models and was the cheapest per pass. Raise the effort only when your own validators show failures at low.
- Effort is not a quality dial you can trust blind. On this set, the extra tokens did not change one result.
- Effort levels do not match across vendors. "High" in Claude Code and "high" in the Codex CLI are different scales. Compare within one model.
- Long, open work may differ. Our tasks are single-turn and short. A multi-step agent task may gain from effort where these did not. We have not measured that.
How we measured
- Tasks: the 8 hard tasks of the hard head-to-head: bug fixes, a CSV parser, event-loop output order, a constraint schedule, a strict SemVer regex and more, each with a deterministic validator.
- Outcomes: strict pass, format miss (right answer, wrong format) and wrong answer. Every attempt is counted.
- Intervals: Wilson 95% intervals for rates; fastest-to-slowest ranges for times. A range is not a confidence interval.
- Costs: reported tokens × list price per call, cache reads and writes priced apart. A calculation, not a bill.
Caveats
- Ceiling. Every cell scored 16/16, so the set cannot separate effort levels by quality.
- Different batches. The 5 reference cells ran at a different time from the 6 new cells, on the same host, CLI versions, cases and flags. Provider load can differ.
- Small cells. 2 calls per task per cell. A few slow calls move a median of 16.
- Route + model. Each row pairs a CLI with a model. Each CLI adds its own system prompt and start-up time.
What to read next
- Claude vs Codex on hard tasks: GPT-6.1 Sol joins the hard set
- Claude Sonnet vs Opus: when is Opus worth the price?
- Same prompt, ten answers: how consistent are Claude and Codex?
- Pairs: Sonnet vs Opus and Sonnet vs GPT-6.1 Sol
Pay for effort only where it helps
Agent records the model, the effort, the tokens and the validation result of every step, so you can see where a higher setting changes the outcome. Try Agent and measure it on your own work.