3 measured metrics · 3 calculated · 2 studies

Claude Sonnet 5.5mediumvshigheffort

“default”: the effort flag was not passed and the CLI chose.

No row separates them: 1 tie, 5 unclear.

The verdict

Claude Sonnet 5.5 at medium effort and Claude Sonnet 5.5 at high effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

4 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Output tokens per call by effort on hard tasks (Output tokens), 1.5x (High effort larger).

Watch it build

A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.

Live story · 48 sDoes more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.

Transcript
  1. Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
  2. 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  3. All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  4. Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
  5. More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
  6. List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
  7. 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  8. Start at low effort and measure. Every call, interval and cost online.

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Medium effort
  • High effort
  • 95% interval
  • fastest–slowest run (not an interval)
  • where the two overlap
  • hollow: list-price calculation

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks

1 tie · 3 unclear
  • Strict pass rate by effort on eight hard tasks: Medium effort 100% (16/16) (n 16, 95% interval 81%–100%); High effort 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Total time per call by effort on hard tasks: Medium effort 7.63 s (n 16, run range 2.7 s–24 s); High effort 8.81 s (n 16, run range 2.9 s–35.8 s). Unclear.
  • Output tokens per call by effort on hard tasks (Output tokens): Medium effort 770 (n 16); High effort 1,192 (n 16). Unclear.
  • List-price cost per strict pass by effort (calculation), calculation: Medium effort $0.014 (n 16); High effort $0.017 (n 16). Unclear.

How much of an AI bill is thinking? Reasoning tokens by model and effort

2 unclear
  • Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass), calculation: Medium effort $0.0059 (n 16); High effort $0.0094 (n 16). Unclear.
  • Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass), calculation: Medium effort $0.014 (n 16); High effort $0.017 (n 16). Unclear.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

6 rows from 2 studies. No row separates them: 1 tie, 5 unclear.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Claude Sonnet 5.5 at medium effort

  • Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass): $0.0059 vs $0.0094. A list-price calculation, not a measured difference. Calculation

When to pick Claude Sonnet 5.5 at high effort

No row in this data puts Claude Sonnet 5.5 at high effort ahead of Claude Sonnet 5.5 at medium effort. Pick on other grounds (price, access, the tasks you run), or measure your own workload.

Side by side

The study charts, showing only Claude Sonnet 5.5 at each effort it ran; the two compared settings are highlighted. Open a study for every configuration.

Every rate is 95% or more
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code

Every interval overlaps every other: this chart does not order these rows.

4 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 16 per row4 of 4 at 100%: this task set cannot separate them.

Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.

Source: Effort ladder: the hard task set at each effort level

Entrance: medians race at 6.3× real timeMotion reduced: press Replay to animateThe slowest median is 8.8 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

4 rows. Slowest Claude Sonnet 5.5 (high) · Claude Code 8.8 s (range 2.9 s–35.8 s, n 16). Fastest Claude Sonnet 5.5 (low) · Claude Code 5.8 s (range 2.8 s–20 s, n 16). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 16 per row

Median per configuration; whiskers = fastest and slowest call

Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.

Source: Effort ladder: the hard task set at each effort level

Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code

4 rows. Highest Claude Sonnet 5.5 (high) · Claude Code 1,192 (n 16). Lowest Claude Sonnet 5.5 (low) · Claude Code 667 (n 16).

Notesn = 16 per row

Median per configuration; reasoning tokens as the CLI reports them

Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean "not reported". Their content is never captured. More tokens is not better or worse by itself.

Source: Effort ladder: the hard task set at each effort level

Calculation
Claude Sonnet 5.5 (low) · Claude Code
Claude Sonnet 5.5 (medium) · Claude Code
Claude Sonnet 5.5 (high) · Claude Code
Claude Sonnet 5.5 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (low) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 4 rows. Highest Claude Sonnet 5.5 (high) · Claude Code $0.017 (n 16). Lowest Claude Sonnet 5.5 (low) · Claude Code $0.012 (n 16).

Notesn = 16 per row

All calls in a configuration divided by its strict passes

Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.

Sources: Effort ladder: the hard task set at each effort level, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Is Claude Sonnet 5.5 better at medium effort or high effort?
Claude Sonnet 5.5 at medium effort and Claude Sonnet 5.5 at high effort share 3 measured metrics and 3 list-price calculations from 2 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 5 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
How were the two effort settings measured?
They share 3 measured metrics and 3 list-price calculations from 2 public studies: Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks; How much of an AI bill is thinking? Reasoning tokens by model and effort. Only the effort setting differs between the two sides of a row.
How does strict pass rate by effort on eight hard tasks change from medium effort to high effort?
Claude Sonnet 5.5 at medium effort: 100% (16/16) (n = 16, 95% interval 81% to 100%). Claude Sonnet 5.5 at high effort: 100% (16/16) (n = 16, 95% interval 81% to 100%). The 95% intervals overlap (Claude Sonnet 5.5 at medium effort 81% to 100%; Claude Sonnet 5.5 at high effort 81% to 100%), so this sample cannot separate them.
How does total time per call by effort on hard tasks change from medium effort to high effort?
Claude Sonnet 5.5 at medium effort: 7.63 s (n = 16, run range 2.7 s to 24 s). Claude Sonnet 5.5 at high effort: 8.81 s (n = 16, run range 2.9 s to 35.8 s). The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 at medium effort 2.71 s to 24.0 s; Claude Sonnet 5.5 at high effort 2.93 s to 35.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.

The studies behind this page

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.