14 measured metrics · 12 calculated · 4 studies

Claude Opus 5.5highvsdefaulteffort

“default”: the effort flag was not passed and the CLI chose.

No row separates them: 4 ties, 22 unclear.

The verdict

Claude Opus 5.5 at high effort and Claude Opus 5.5 at default effort share 14 measured metrics and 12 list-price calculations from 4 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. "default" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 4 ties and 22 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Output tokens per call (Output tokens), 1.2x (High effort larger).

Watch it build

A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.

Live story · 48 sDoes more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.

Transcript
  1. Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
  2. 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  3. All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  4. Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
  5. More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
  6. List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
  7. 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  8. Start at low effort and measure. Every call, interval and cost online.

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • High effort
  • Default effort
  • 95% interval
  • fastest–slowest run (not an interval)
  • where the two overlap
  • hollow: list-price calculation

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

1 tie · 7 unclear
  • Pass rate on five validated tasks: High effort 100% (15/15) (n 15, 95% interval 80%–100%); Default effort 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
  • Total time per call: High effort 2.71 s (n 15, run range 2.5 s–11.8 s); Default effort 2.75 s (n 15, run range 2.5 s–8.9 s). Unclear.
  • Time to first useful output: High effort 2.04 s (n 15, run range 1.4 s–9.9 s); Default effort 1.92 s (n 15, run range 1.6 s–7.2 s). Unclear.
  • Input tokens per call: what the CLI sends (Cache read): High effort 1,463 (n 15); Default effort 1,401 (n 15). Unclear.
  • Input tokens per call: what the CLI sends (Other input): High effort 619 (n 15); Default effort 680 (n 15). Unclear.
  • Output tokens per call (Output tokens): High effort 78 (n 15); Default effort 64 (n 15). Unclear.
  • List-price cost per call (calculation), calculation: High effort $0.0069 (n 15, run range $0.0059–$0.027); Default effort $0.0069 (n 15, run range $0.0059–$0.022). Unclear.
  • List-price cost per passing answer (calculation), calculation: High effort $0.010 (n 15); Default effort $0.010 (n 15). Unclear.

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

2 ties · 4 unclear
  • Pass rate on eight hard tasks (Strict pass): High effort 100% (24/24) (n 24, 95% interval 86%–100%); Default effort 100% (24/24) (n 24, 95% interval 86%–100%). Tie.
  • Pass rate on eight hard tasks (Lenient (format misses counted)): High effort 100% (24/24) (n 24, 95% interval 86%–100%); Default effort 100% (24/24) (n 24, 95% interval 86%–100%). Tie.
  • Total time per call on hard tasks (separate batches): High effort 11.0 s (n 24, run range 3.6 s–63 s); Default effort 9.18 s (n 24, run range 4.2 s–27.2 s). Unclear.
  • Time to first useful output on hard tasks: High effort 7.13 s (n 24, run range 2.2 s–56.2 s); Default effort 6.78 s (n 24, run range 2.4 s–21.8 s). Unclear.
  • Output tokens per call on hard tasks (Output tokens): High effort 1,052 (n 24); Default effort 945 (n 24). Unclear.
  • List-price cost per strict pass on hard tasks (calculation), calculation: High effort $0.033 (n 24); Default effort $0.028 (n 24). Unclear.

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks

1 tie · 3 unclear
  • Strict pass rate by effort on eight hard tasks: High effort 100% (16/16) (n 16, 95% interval 81%–100%); Default effort 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Total time per call by effort on hard tasks: High effort 10.1 s (n 16, run range 3.6 s–63 s); Default effort 9.18 s (n 16, run range 4.2 s–27.2 s). Unclear.
  • Output tokens per call by effort on hard tasks (Output tokens): High effort 1,052 (n 16); Default effort 945 (n 16). Unclear.
  • List-price cost per strict pass by effort (calculation), calculation: High effort $0.034 (n 16); Default effort $0.029 (n 16). Unclear.

How much of an AI bill is thinking? Reasoning tokens by model and effort

8 unclear
  • Reasoning share of output tokens per call on hard tasks (calculation), calculation: High effort 54.4% (n 24, run range 36%–96%); Default effort 54.8% (n 24, run range 30%–96%). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: High effort $0.018 (n 24); Default effort $0.013 (n 24). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: High effort $0.0080 (n 24); Default effort $0.0080 (n 24). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: High effort $0.0074 (n 24); Default effort $0.0077 (n 24). Unclear.
  • Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass), calculation: High effort $0.018 (n 16); Default effort $0.013 (n 16). Unclear.
  • Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass), calculation: High effort $0.034 (n 16); Default effort $0.029 (n 16). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: High effort 54.4% (n 24, run range 36%–96%); Default effort 54.8% (n 24, run range 30%–96%). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: High effort 43.6% (n 15, run range 0%–93%); Default effort 0% (n 15, run range 0%–93%). Unclear.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

26 rows from 4 studies. No row separates them: 4 ties, 22 unclear.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Claude Opus 5.5 at high effort

No row in this data puts Claude Opus 5.5 at high effort ahead of Claude Opus 5.5 at default effort. Pick on other grounds (price, access, the tasks you run), or measure your own workload.

When to pick Claude Opus 5.5 at default effort

No row in this data puts Claude Opus 5.5 at default effort ahead of Claude Opus 5.5 at high effort. Pick on other grounds (price, access, the tasks you run), or measure your own workload.

Side by side

The study charts, showing only Claude Opus 5.5 at each effort it ran; the two compared settings are highlighted. Open a study for every configuration.

Every rate is 95% or more
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code

Every interval overlaps every other: this chart does not order these rows.

3 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 15 per row3 of 3 at 100%: this task set cannot separate them.

Every call counts; failures and timeouts are non-passes

Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 2× real timeMotion reduced: press Replay to animateThe slowest median is 2.8 s. The clock runs at the recorded speed.
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

3 rows. Slowest Claude Opus 5.5 (low) · Claude Code 2.8 s (range 2.4 s–6.6 s, n 15). Fastest Claude Opus 5.5 (high) · Claude Code 2.7 s (range 2.5 s–11.8 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 1.7× real timeMotion reduced: press Replay to animateThe slowest median is 2.4 s. The clock runs at the recorded speed.
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

3 rows. Slowest Claude Opus 5.5 (low) · Claude Code 2.4 s (range 1.5 s–4.9 s, n 15). Fastest Claude Opus 5.5 · Claude Code 1.9 s (range 1.6 s–7.2 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

  • Cache read
  • Other input
Bar length is the total; segments are its parts.
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code

Totals are the sum of the parts shown. Shares are calculated from the same values.

3 rows, 2 series: Cache read, Other input. Cache read: highest Claude Opus 5.5 (high) · Claude Code 1,463 (n 15). Lowest Claude Opus 5.5 · Claude Code 1,401 (n 15). Other input: highest Claude Opus 5.5 · Claude Code 680 (n 15). Lowest Claude Opus 5.5 (low) · Claude Code 618 (n 15).

Notesn = 15 per row

Mean per call, split into prompt-cache reads and other input

The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.

Source: Provider head-to-head: Claude Code models vs Codex efforts

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Is Claude Opus 5.5 better at high effort or default effort (CLI chose)?
Claude Opus 5.5 at high effort and Claude Opus 5.5 at default effort share 14 measured metrics and 12 list-price calculations from 4 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. "default" means the effort flag was not passed, so its level is the CLI’s choice. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 4 ties and 22 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).
How were the two effort settings measured?
They share 14 measured metrics and 12 list-price calculations from 4 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks; How much of an AI bill is thinking? Reasoning tokens by model and effort. Only the effort setting differs between the two sides of a row.
How does pass rate on five validated tasks change from high effort to default effort (CLI chose)?
Claude Opus 5.5 at high effort: 100% (15/15) (n = 15, 95% interval 80% to 100%). Claude Opus 5.5 at default effort: 100% (15/15) (n = 15, 95% interval 80% to 100%). The 95% intervals overlap (Claude Opus 5.5 at high effort 80% to 100%; Claude Opus 5.5 at default effort 80% to 100%), so this sample cannot separate them.
How does total time per call change from high effort to default effort (CLI chose)?
Claude Opus 5.5 at high effort: 2.71 s (n = 15, run range 2.5 s to 11.8 s). Claude Opus 5.5 at default effort: 2.75 s (n = 15, run range 2.5 s to 8.9 s). The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort 2.45 s to 11.8 s; Claude Opus 5.5 at default effort 2.47 s to 8.91 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
How does time to first useful output change from high effort to default effort (CLI chose)?
Claude Opus 5.5 at high effort: 2.04 s (n = 15, run range 1.4 s to 9.9 s). Claude Opus 5.5 at default effort: 1.92 s (n = 15, run range 1.6 s to 7.2 s). The run ranges (fastest to slowest) overlap (Claude Opus 5.5 at high effort 1.40 s to 9.94 s; Claude Opus 5.5 at default effort 1.56 s to 7.23 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
How does pass rate on eight hard tasks (Strict pass) change from high effort to default effort (CLI chose)?
Claude Opus 5.5 at high effort: 100% (24/24) (n = 24, 95% interval 86% to 100%). Claude Opus 5.5 at default effort: 100% (24/24) (n = 24, 95% interval 86% to 100%). The 95% intervals overlap (Claude Opus 5.5 at high effort 86% to 100%; Claude Opus 5.5 at default effort 86% to 100%), so this sample cannot separate them.

The studies behind this page

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.