Explainer · Reasoning effort

What is reasoning effort?

Definition

Reasoning effort is a setting on a reasoning model that controls how much internal work, usually in the form of hidden thinking tokens, the model does before it gives its answer. Typical levels are low, medium and high. A higher effort can help on hard problems, but it costs more output tokens, more time and more money per call, so the useful question is whether a higher effort changes the result on your tasks.

Agent team · · 4 min read · Every number is from the public studies

Seconds (median)

* default: the effort flag was not passed; its level is not known, so no line joins it.

4 efforts, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol · Codex CLI. Claude Sonnet 5.5 · Claude Code: slowest high 8.8 s (n 16). Fastest low 5.8 s (n 16). Claude Opus 5.5 · Claude Code: slowest high 10.1 s (n 16). Fastest low 7.5 s (n 16).

Notesn = 16 per row

One line per model and route; default = the effort flag was not passed

Medians only; the per-call ranges are in the total-time chart and they overlap. "default" is placed last because its level is not known: the CLI chose it.

Source: Effort ladder: the hard task set at each effort level

How the setting works

A reasoning model writes a chain of intermediate steps before the final answer. The effort level sets a budget or a bias for that chain. At low effort the model answers with less deliberation; at high effort it explores more, checks more and writes more.

The names differ by vendor. Codex CLI takes an effort value per request. Claude Code takes an effort flag; when the flag is not passed, the CLI chooses, and we call that default. Default is not the same as medium, and its level is not published, so our charts place it after the named levels.

Does more effort buy quality?

Not on our hard set. In the effort ladder study, three models ran 8 hard tasks with strict validators, twice each, at every effort:

Every rate is 95% or more
Claude Sonnet 5.5 (low)
Claude Sonnet 5.5 (medium)
Claude Sonnet 5.5 (high)
Claude Sonnet 5.5
Claude Opus 5.5 (low)
Claude Opus 5.5 (medium)
Claude Opus 5.5 (high)
Claude Opus 5.5
GPT-6.1 Sol (low)
GPT-6.1 Sol (medium)
GPT-6.1 Sol (high)

Every interval overlaps every other: this chart does not order these rows.

11 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 16 per row11 of 11 at 100%: this task set cannot separate them.

Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.

Source: Effort ladder: the hard task set at each effort level

  • All 11 configurations (Sonnet 5.5 at low, medium, high and default; Opus 5.5 at the same four; GPT-6.1 Sol through the Codex CLI at low, medium and high) passed 16 of 16 calls strictly.
  • Each 16/16 has a 95% Wilson interval of 81% to 100%.
  • There were 0 format misses and 0 wrong answers in 176 calls.

So pass rate does not separate any effort level here. The set has a ceiling for these models: with 16 calls per cell it cannot rule out a difference of up to about 19 points. "No difference found" is not "no difference".

What more effort did change

More effort changed output and time. Median total time per call:

Entrance: medians race at 13× real timeMotion reduced: press Replay to animateThe slowest median is 18.1 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

11 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 18.1 s (range 11.7 s–92.2 s, n 16). Fastest Claude Sonnet 5.5 (low) · Claude Code 5.8 s (range 2.8 s–20 s, n 16). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 16 per row

Median per configuration; whiskers = fastest and slowest call

Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.

Source: Effort ladder: the hard task set at each effort level

  • Sonnet 5.5: low 5.8 s, medium 7.6 s, high 8.8 s, default 8.0 s.
  • Opus 5.5: low 7.5 s, medium 9.7 s, high 10.1 s, default 9.2 s.
  • GPT-6.1 Sol (Codex CLI): low 13.6 s, medium 13.1 s, high 18.1 s.

Within each model the fastest and slowest calls of every effort overlap, so these medians describe this run; they are not a tested ranking. Median output tokens for Sonnet went from 667 at low to 770 at medium and 1,192 at high.

What it costs per correct answer

Calculation
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (low) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 11 rows. Highest Claude Opus 5.5 (high) · Claude Code $0.034 (n 16). Lowest Claude Sonnet 5.5 (low) · Claude Code $0.012 (n 16).

Notesn = 16 per row

All calls in a configuration divided by its strict passes

Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.

Sources: Effort ladder: the hard task set at each effort level, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

At list price (a calculation; the calls ran on subscriptions), the cost per strict pass went from:

  • $0.0122 at low to $0.0167 at high for Sonnet 5.5,
  • $0.0212 at low to $0.0337 at high for Opus 5.5,
  • $0.0128 at low to $0.0151 at high for GPT-6.1 Sol through the Codex CLI.

Cost does not always rise in order: for GPT-6.1 Sol the medium cell cost more per pass than the high cell in this run. Per-call cost depends on output length, which varies from call to call.

How to choose an effort level

  • Start low on tasks your model already passes. If low passes, a higher effort only adds time and cost.
  • Raise effort when you see failures that look like reasoning errors, not format errors. Fix the format first; a correct answer in the wrong format is a validator problem.
  • Measure on your own hard cases. Our set has a ceiling for these models. Your hardest tasks may not.
  • Compare like with like. Change only the effort: same model, same route, same tasks, same validators. Otherwise you measure several things at once.
  • Count every call. A higher effort that times out more often must show those failures in its rate.

Frequently asked questions

Is high reasoning effort always better?

No. On our 8 hard tasks, every effort level of Sonnet 5.5, Opus 5.5 and GPT-6.1 Sol passed 16 of 16 calls, so high effort bought no measurable quality there. It did raise median time per call and the list-price cost per pass.

What does "default" effort mean in Claude Code?

It means the effort flag was not passed, so the CLI chose the level. Its level is not published, so we report it as its own configuration and do not treat it as medium.

How much slower is high effort than low?

In our run, Sonnet 5.5 went from a median 5.8 s per call at low to 8.8 s at high, and Opus 5.5 from 7.5 s to 10.1 s. The ranges of single calls overlap, so the gap is not a tested ranking.

How many calls do I need to see an effort difference?

More than 16 per cell when both sides score near 100%. At 16/16 the 95% interval still reaches down to 81%, so a real difference smaller than about 19 points could hide inside it. Harder tasks, where the models fail sometimes, separate effort levels with fewer calls.

Watch the data

Live story · 48 sDoes more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.

Transcript
  1. Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
  2. 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  3. All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  4. Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
  5. More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
  6. List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
  7. 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  8. Start at low effort and measure. Every call, interval and cost online.

The data behind this explainer

More explainers

  • LLM router

    What is an LLM router?

    An LLM router picks the model, effort and context for each request. What routers exist, what a decision costs in time and money, and how to judge one.

  • Wilson confidence interval

    Wilson confidence intervals for AI benchmarks

    A Wilson interval shows the range of true pass rates that fit a benchmark result. The formula, a worked example, and why 16 of 16 still means 81% to 100%.

  • Benchmark saturation

    Benchmark saturation: when every model scores 100%

    Benchmark saturation: models reach the top score and the test stops telling them apart. We measured it on our own sets and show which other measures differ.

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.