Opus at low effort or Sonnet at high effort? A bigger model that thinks less, tested
Opus 5.5 at low effort and Sonnet 5.5 at high effort both passed 16/16. Opus low cost 1.3x as much per pass (calculation). Sonnet at low effort cost least.
TL;DR
- Neither trade bought extra passes. We tested Claude Opus 5.5 at low effort against Claude Sonnet 5.5 at high effort. Each passed 16 of 16 strictly (95% interval 81% to 100%). The set is at its ceiling, so it cannot rank them on quality. Each cell has n = 16 calls on 8 tasks.
- Time: the median was 7.5 s for Opus low and 8.81 s for Sonnet high. The ranges were 3.3–15.8 s and 2.9–35.8 s (n = 16 each). They overlap, so this is not a ranking.
- Tokens: Opus low had medians of 594 output tokens and 87 reasoning tokens. Sonnet high had medians of 1,192 and 745 (n = 16 each; ranges below). These are separate medians.
- Cost per strict pass (list-price calculation): $0.0212 for Opus low and $0.0167 for Sonnet high. Opus low cost about 1.3x as much (calculation).
- The lowest calculated cost per pass came from Sonnet at low effort: $0.0122 per pass, the lowest of all 11 cells.
Study: /benchmarks/effort-ladder. New to the setting? Read what reasoning effort is.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks
All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.
Transcript
- Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
- 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
- All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
- More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
- List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
- 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Start at low effort and measure. Every call, interval and cost online.
The question
You can spend on one task in two ways. Pay for a bigger model that thinks little, or a smaller model that thinks a lot. In Claude Code, that is Opus 5.5 with --effort low, or Sonnet 5.5 with --effort high.
We ran both on the same 8 hard tasks: 16 calls per cell (8 tasks × 2 repetitions), strict validators, no retries. Both are new runs from one batch. The dataset has no comparison row for this pair, so we read two cells of one table side by side.
Quality: a tie at the ceiling
Opus low passed 16 of 16. Sonnet high passed 16 of 16. So did the other 9 cells: 176 of 176 calls passed (95% Wilson interval 97.9% to 100%), with 0 format misses and 0 wrong answers.
A perfect 16/16 has a 95% interval of 81% to 100%. That interval spans about 19 percentage points (a calculation), but it is not an interval for the difference between models. Sonnet at low effort passed all calls on these 8 tasks. This does not show that either trade is unnecessary on other work. The dataset's 16 effort comparisons have no decided row.
Time: not a ranking
Each row has n = 16 calls. The ranges show the fastest and slowest call, not a confidence interval.
| Cell | Median time | Fastest to slowest |
|---|---|---|
| Sonnet 5.5, low | 5.82 s | 2.8 to 20 s |
| Opus 5.5, low | 7.5 s | 3.3 to 15.8 s |
| Sonnet 5.5, high | 8.81 s | 2.9 to 35.8 s |
| Opus 5.5, high | 10.11 s | 3.6 to 63 s |
The Opus low median was about 1.3 s below the Sonnet high median (a calculation). The ranges overlap, so neither side is ahead. Sonnet low has the lowest median of all 11 cells, but its range overlaps all the others.
Tokens: Opus low reports fewer reasoning tokens
Sonnet high had about 2.0x the median output tokens of Opus low (1,192 against 594). It had about 8.6x the median reasoning tokens (745 against 87). Both ratios are calculations on separate medians, with n = 16 per cell.
At low effort, Opus reported fewer median reasoning tokens than Sonnet (87 against 273). Token counts do not measure how much useful reasoning a model did.
| Cell (n = 16 each) | Output token range | Reported reasoning token range |
|---|---|---|
| Opus 5.5, low | 220 to 1,569 | 0 to 791 |
| Sonnet 5.5, high | 220 to 4,187 | 0 to 3,610 |
| Sonnet 5.5, low | 176 to 2,263 | 0 to 1,489 |
These are per-call ranges, not confidence intervals.
The CLI reports reasoning tokens, and 0 can mean "not reported". More tokens is not better or worse by itself. Here the pass count did not move.
Cost per pass: the bigger model costs more
This is a calculation: list price × reported tokens, divided by strict passes. The calls ran on flat subscriptions. Each cost row uses n = 16 calls and 16 strict passes. These calculations have no intervals; they describe this sample, not an expected bill.
Cache reads use the recorded rates, and cache writes use the one-hour rate. Codex adds a different system prompt and tool schemas, so its cost also reflects the CLI.
| Cell | List-price cost per strict pass |
|---|---|
| Sonnet 5.5, low | $0.0122 |
| GPT-6.1 Sol, low (Codex CLI) | $0.0128 |
| Sonnet 5.5, high | $0.0167 |
| Opus 5.5, low | $0.0212 |
| Opus 5.5, high | $0.0337 |
The study uses these list prices: Opus costs 2.0x Sonnet for plain input and output (a calculation). The rates are $4 against $2 per million input tokens, and $20 against $10 per million output tokens. The mean output token counts were about half as large for Opus low (a calculation). At these prices, the calculated output spend was close.
Our calculation on mean output tokens (665 and 1,284) gives an output spend of $0.0133 and $0.0128 per call. The rest of the call, input and cache, is about $0.008 for Opus low and $0.004 for Sonnet high.
So Opus low costs about 1.3x as much as Sonnet high per pass (a calculation).
Sonnet low had the lowest calculated cost per pass in this sample. Raising Sonnet to high adds $0.0045 per pass (1.4x the total cost). Swapping to Opus at low adds $0.0090 (1.7x the total cost), about 2.0x the added cost (calculations). Both settings still passed 16/16 (95% interval 81% to 100% each).
What a harder set might show
Our tasks are short and tool-free. This study has no result for this exact pair on a harder set. Such a test could show one of four results:
- Sonnet high passes more. That would favour this setting on the tested tasks, without proving why.
- Opus low passes more. That would favour this setting on the tested tasks, without proving why.
- Both fail the same tasks. Neither setting would solve those tasks in that test.
- They tie again. Calculated cost could guide a choice, alongside time and the task checks.
A higher pass count alone would not decide the result. We will call a side ahead only when the 95% intervals do not overlap.
What this means for your settings
- Start at the cheapest setting that your validators accept. On our set, that was Sonnet at low effort.
- When it fails, test both levers on those failures. Keep the one that fixes them at the lower cost.
- Pay for the bigger model only when your own checks show a gap.
How we measured
- Tasks: the 8 hard tasks of the hard head-to-head, each with a deterministic validator. Before any call, 8 of 8 reference answers passed and 26 of 26 wrong answers failed.
- Isolation: a fresh empty folder, tools off, one turn, a 300 s timeout. We retried nothing, and no call failed.
- Intervals: Wilson 95% intervals for rates. Ranges, not intervals, for times.
- Receipts: new effort-ladder calls and hard head-to-head reference calls. The study links its price sources.
Caveats
- Ceiling. Every cell scored 16/16, so the set cannot separate these trades on quality.
- Small cells. Each cell has 2 calls per task, so a few slow calls move a median. Repeated calls on 8 tasks do not give 16 independent task samples. The Wilson intervals describe call outcomes, not coverage of harder tasks.
- One host. CLI times include start-up and the CLI prompt, on one host and network. Cache use and call order can affect the calculated costs.
- Reference cells. Opus at high effort comes from the hard head-to-head, run in a different hour. Opus low, Sonnet low and Sonnet high ran in one batch.
- Not a dataset row. We paired two cells by hand. "Low" and "high" are settings, not token budgets.
What to read next
- Does reasoning effort buy quality? Claude and Codex, low to high
- Claude Sonnet vs Opus: when is Opus worth the price?
- Pair: Sonnet 5.5 vs Opus 5.5.
Test the trade on your own work
Agent records the model, the effort, the tokens and the validation result of each step. Try Agent to measure the trade on your own tasks.