Explainer · Reasoning effort
What is reasoning effort?
Definition
Reasoning effort is a setting on a reasoning model that controls how much internal work, usually in the form of hidden thinking tokens, the model does before it gives its answer. Typical levels are low, medium and high. A higher effort can help on hard problems, but it costs more output tokens, more time and more money per call, so the useful question is whether a higher effort changes the result on your tasks.
Agent team · · 4 min read · Every number is from the public studies
Seconds (median)
* default: the effort flag was not passed; its level is not known, so no line joins it.
| Effort | Claude Sonnet 5.5 · Claude Code | Claude Opus 5.5 · Claude Code | GPT-6.1 Sol · Codex CLI | n |
|---|---|---|---|---|
| low | 5.8 s | 7.5 s | 13.6 s | 16 |
| medium | 7.6 s | 9.7 s | 13.1 s | 16 |
| high | 8.8 s | 10.1 s | 18.1 s | 16 |
| default | 8 s | 9.2 s | — | 16 |
4 efforts, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol · Codex CLI. Claude Sonnet 5.5 · Claude Code: slowest high 8.8 s (n 16). Fastest low 5.8 s (n 16). Claude Opus 5.5 · Claude Code: slowest high 10.1 s (n 16). Fastest low 7.5 s (n 16).
Notesn = 16 per row
One line per model and route; default = the effort flag was not passed
Medians only; the per-call ranges are in the total-time chart and they overlap. "default" is placed last because its level is not known: the CLI chose it.
Source: Effort ladder: the hard task set at each effort level
How the setting works
A reasoning model writes a chain of intermediate steps before the final answer. The effort level sets a budget or a bias for that chain. At low effort the model answers with less deliberation; at high effort it explores more, checks more and writes more.
The names differ by vendor. Codex CLI takes an effort value per request. Claude Code takes an effort flag; when the flag is not passed, the CLI chooses, and we call that default. Default is not the same as medium, and its level is not published, so our charts place it after the named levels.
Does more effort buy quality?
Not on our hard set. In the effort ladder study, three models ran 8 hard tasks with strict validators, twice each, at every effort:
- All 11 configurations (Sonnet 5.5 at low, medium, high and default; Opus 5.5 at the same four; GPT-6.1 Sol through the Codex CLI at low, medium and high) passed 16 of 16 calls strictly.
- Each 16/16 has a 95% Wilson interval of 81% to 100%.
- There were 0 format misses and 0 wrong answers in 176 calls.
So pass rate does not separate any effort level here. The set has a ceiling for these models: with 16 calls per cell it cannot rule out a difference of up to about 19 points. "No difference found" is not "no difference".
What more effort did change
More effort changed output and time. Median total time per call:
- Sonnet 5.5: low 5.8 s, medium 7.6 s, high 8.8 s, default 8.0 s.
- Opus 5.5: low 7.5 s, medium 9.7 s, high 10.1 s, default 9.2 s.
- GPT-6.1 Sol (Codex CLI): low 13.6 s, medium 13.1 s, high 18.1 s.
Within each model the fastest and slowest calls of every effort overlap, so these medians describe this run; they are not a tested ranking. Median output tokens for Sonnet went from 667 at low to 770 at medium and 1,192 at high.
What it costs per correct answer
At list price (a calculation; the calls ran on subscriptions), the cost per strict pass went from:
- $0.0122 at low to $0.0167 at high for Sonnet 5.5,
- $0.0212 at low to $0.0337 at high for Opus 5.5,
- $0.0128 at low to $0.0151 at high for GPT-6.1 Sol through the Codex CLI.
Cost does not always rise in order: for GPT-6.1 Sol the medium cell cost more per pass than the high cell in this run. Per-call cost depends on output length, which varies from call to call.
How to choose an effort level
- Start low on tasks your model already passes. If low passes, a higher effort only adds time and cost.
- Raise effort when you see failures that look like reasoning errors, not format errors. Fix the format first; a correct answer in the wrong format is a validator problem.
- Measure on your own hard cases. Our set has a ceiling for these models. Your hardest tasks may not.
- Compare like with like. Change only the effort: same model, same route, same tasks, same validators. Otherwise you measure several things at once.
- Count every call. A higher effort that times out more often must show those failures in its rate.
Frequently asked questions
Is high reasoning effort always better?
No. On our 8 hard tasks, every effort level of Sonnet 5.5, Opus 5.5 and GPT-6.1 Sol passed 16 of 16 calls, so high effort bought no measurable quality there. It did raise median time per call and the list-price cost per pass.
What does "default" effort mean in Claude Code?
It means the effort flag was not passed, so the CLI chose the level. Its level is not published, so we report it as its own configuration and do not treat it as medium.
How much slower is high effort than low?
In our run, Sonnet 5.5 went from a median 5.8 s per call at low to 8.8 s at high, and Opus 5.5 from 7.5 s to 10.1 s. The ranges of single calls overlap, so the gap is not a tested ranking.
How many calls do I need to see an effort difference?
More than 16 per cell when both sides score near 100%. At 16/16 the 95% interval still reaches down to 81%, so a real difference smaller than about 19 points could hide inside it. Harder tasks, where the models fail sometimes, separate effort levels with fewer calls.
Watch the data
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks
All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.
Transcript
- Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
- 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
- All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
- More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
- List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
- 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Start at low effort and measure. Every call, interval and cost online.