• Benchmark
  • Effort
  • Reasoning Effort
  • Claude Opus

Opus at low effort or Sonnet at high effort? A bigger model that thinks less, tested

Opus 5.5 at low effort and Sonnet 5.5 at high effort both passed 16/16. Opus low cost 1.3x as much per pass (calculation). Sonnet at low effort cost least.

TL;DR

  • Neither trade bought extra passes. We tested Claude Opus 5.5 at low effort against Claude Sonnet 5.5 at high effort. Each passed 16 of 16 strictly (95% interval 81% to 100%). The set is at its ceiling, so it cannot rank them on quality. Each cell has n = 16 calls on 8 tasks.
  • Time: the median was 7.5 s for Opus low and 8.81 s for Sonnet high. The ranges were 3.3–15.8 s and 2.9–35.8 s (n = 16 each). They overlap, so this is not a ranking.
  • Tokens: Opus low had medians of 594 output tokens and 87 reasoning tokens. Sonnet high had medians of 1,192 and 745 (n = 16 each; ranges below). These are separate medians.
  • Cost per strict pass (list-price calculation): $0.0212 for Opus low and $0.0167 for Sonnet high. Opus low cost about 1.3x as much (calculation).
  • The lowest calculated cost per pass came from Sonnet at low effort: $0.0122 per pass, the lowest of all 11 cells.

Study: /benchmarks/effort-ladder. New to the setting? Read what reasoning effort is.

Live story · 48 sDoes more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.

Transcript
  1. Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
  2. 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  3. All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  4. Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
  5. More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
  6. List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
  7. 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  8. Start at low effort and measure. Every call, interval and cost online.

The question

You can spend on one task in two ways. Pay for a bigger model that thinks little, or a smaller model that thinks a lot. In Claude Code, that is Opus 5.5 with --effort low, or Sonnet 5.5 with --effort high.

We ran both on the same 8 hard tasks: 16 calls per cell (8 tasks × 2 repetitions), strict validators, no retries. Both are new runs from one batch. The dataset has no comparison row for this pair, so we read two cells of one table side by side.

Quality: a tie at the ceiling

Every rate is 95% or more
Claude Sonnet 5.5 (low)
Claude Sonnet 5.5 (medium)
Claude Sonnet 5.5 (high)
Claude Sonnet 5.5
Claude Opus 5.5 (low)
Claude Opus 5.5 (medium)
Claude Opus 5.5 (high)
Claude Opus 5.5
GPT-6.1 Sol (low)
GPT-6.1 Sol (medium)
GPT-6.1 Sol (high)

Every interval overlaps every other: this chart does not order these rows.

11 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 16 per row11 of 11 at 100%: this task set cannot separate them.

Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.

Source: Effort ladder: the hard task set at each effort level

Opus low passed 16 of 16. Sonnet high passed 16 of 16. So did the other 9 cells: 176 of 176 calls passed (95% Wilson interval 97.9% to 100%), with 0 format misses and 0 wrong answers.

A perfect 16/16 has a 95% interval of 81% to 100%. That interval spans about 19 percentage points (a calculation), but it is not an interval for the difference between models. Sonnet at low effort passed all calls on these 8 tasks. This does not show that either trade is unnecessary on other work. The dataset's 16 effort comparisons have no decided row.

Time: not a ranking

Entrance: medians race at 13× real timeMotion reduced: press Replay to animateThe slowest median is 18.1 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

11 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 18.1 s (range 11.7 s–92.2 s, n 16). Fastest Claude Sonnet 5.5 (low) · Claude Code 5.8 s (range 2.8 s–20 s, n 16). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 16 per row

Median per configuration; whiskers = fastest and slowest call

Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.

Source: Effort ladder: the hard task set at each effort level

Each row has n = 16 calls. The ranges show the fastest and slowest call, not a confidence interval.

CellMedian timeFastest to slowest
Sonnet 5.5, low5.82 s2.8 to 20 s
Opus 5.5, low7.5 s3.3 to 15.8 s
Sonnet 5.5, high8.81 s2.9 to 35.8 s
Opus 5.5, high10.11 s3.6 to 63 s

The Opus low median was about 1.3 s below the Sonnet high median (a calculation). The ranges overlap, so neither side is ahead. Sonnet low has the lowest median of all 11 cells, but its range overlaps all the others.

Tokens: Opus low reports fewer reasoning tokens

  • Output tokens
  • of which reasoning tokens (inner bar)
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

11 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 (high) · Claude Code 1,192 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 284 (n 16). Reasoning tokens: highest Claude Sonnet 5.5 (high) · Claude Code 745 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 63 (n 16).

Notesn = 16 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean "not reported". Their content is never captured. More tokens is not better or worse by itself.

Source: Effort ladder: the hard task set at each effort level

Sonnet high had about 2.0x the median output tokens of Opus low (1,192 against 594). It had about 8.6x the median reasoning tokens (745 against 87). Both ratios are calculations on separate medians, with n = 16 per cell.

At low effort, Opus reported fewer median reasoning tokens than Sonnet (87 against 273). Token counts do not measure how much useful reasoning a model did.

Cell (n = 16 each)Output token rangeReported reasoning token range
Opus 5.5, low220 to 1,5690 to 791
Sonnet 5.5, high220 to 4,1870 to 3,610
Sonnet 5.5, low176 to 2,2630 to 1,489

These are per-call ranges, not confidence intervals.

The CLI reports reasoning tokens, and 0 can mean "not reported". More tokens is not better or worse by itself. Here the pass count did not move.

Cost per pass: the bigger model costs more

Calculation
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Opus 5.5 (medium) · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (high) · Codex CLI

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (low) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 11 rows. Highest Claude Opus 5.5 (high) · Claude Code $0.034 (n 16). Lowest Claude Sonnet 5.5 (low) · Claude Code $0.012 (n 16).

Notesn = 16 per row

All calls in a configuration divided by its strict passes

Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.

Sources: Effort ladder: the hard task set at each effort level, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

This is a calculation: list price × reported tokens, divided by strict passes. The calls ran on flat subscriptions. Each cost row uses n = 16 calls and 16 strict passes. These calculations have no intervals; they describe this sample, not an expected bill.

Cache reads use the recorded rates, and cache writes use the one-hour rate. Codex adds a different system prompt and tool schemas, so its cost also reflects the CLI.

CellList-price cost per strict pass
Sonnet 5.5, low$0.0122
GPT-6.1 Sol, low (Codex CLI)$0.0128
Sonnet 5.5, high$0.0167
Opus 5.5, low$0.0212
Opus 5.5, high$0.0337

The study uses these list prices: Opus costs 2.0x Sonnet for plain input and output (a calculation). The rates are $4 against $2 per million input tokens, and $20 against $10 per million output tokens. The mean output token counts were about half as large for Opus low (a calculation). At these prices, the calculated output spend was close.

Our calculation on mean output tokens (665 and 1,284) gives an output spend of $0.0133 and $0.0128 per call. The rest of the call, input and cache, is about $0.008 for Opus low and $0.004 for Sonnet high.

So Opus low costs about 1.3x as much as Sonnet high per pass (a calculation).

Sonnet low had the lowest calculated cost per pass in this sample. Raising Sonnet to high adds $0.0045 per pass (1.4x the total cost). Swapping to Opus at low adds $0.0090 (1.7x the total cost), about 2.0x the added cost (calculations). Both settings still passed 16/16 (95% interval 81% to 100% each).

What a harder set might show

Our tasks are short and tool-free. This study has no result for this exact pair on a harder set. Such a test could show one of four results:

  1. Sonnet high passes more. That would favour this setting on the tested tasks, without proving why.
  2. Opus low passes more. That would favour this setting on the tested tasks, without proving why.
  3. Both fail the same tasks. Neither setting would solve those tasks in that test.
  4. They tie again. Calculated cost could guide a choice, alongside time and the task checks.

A higher pass count alone would not decide the result. We will call a side ahead only when the 95% intervals do not overlap.

What this means for your settings

  1. Start at the cheapest setting that your validators accept. On our set, that was Sonnet at low effort.
  2. When it fails, test both levers on those failures. Keep the one that fixes them at the lower cost.
  3. Pay for the bigger model only when your own checks show a gap.

How we measured

  • Tasks: the 8 hard tasks of the hard head-to-head, each with a deterministic validator. Before any call, 8 of 8 reference answers passed and 26 of 26 wrong answers failed.
  • Isolation: a fresh empty folder, tools off, one turn, a 300 s timeout. We retried nothing, and no call failed.
  • Intervals: Wilson 95% intervals for rates. Ranges, not intervals, for times.
  • Receipts: new effort-ladder calls and hard head-to-head reference calls. The study links its price sources.

Caveats

  • Ceiling. Every cell scored 16/16, so the set cannot separate these trades on quality.
  • Small cells. Each cell has 2 calls per task, so a few slow calls move a median. Repeated calls on 8 tasks do not give 16 independent task samples. The Wilson intervals describe call outcomes, not coverage of harder tasks.
  • One host. CLI times include start-up and the CLI prompt, on one host and network. Cache use and call order can affect the calculated costs.
  • Reference cells. Opus at high effort comes from the hard head-to-head, run in a different hour. Opus low, Sonnet low and Sonnet high ran in one batch.
  • Not a dataset row. We paired two cells by hand. "Low" and "high" are settings, not token budgets.

Test the trade on your own work

Agent records the model, the effort, the tokens and the validation result of each step. Try Agent to measure the trade on your own tasks.

The data behind this post

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.