14 measured metrics · 12 calculated · 4 studies

GPT-6.1 Sol (Codex CLI)mediumvshigheffort

No row separates them: 5 ties, 21 unclear.

The verdict

GPT-6.1 Sol (Codex CLI) at medium effort and GPT-6.1 Sol (Codex CLI) at high effort share 14 measured metrics and 12 list-price calculations from 4 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 5 ties and 21 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Input tokens per call: what the CLI sends (Cache read), 1.3x (High effort larger).

Watch it build

A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.

Live story · 48 sDoes more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.

Transcript
  1. Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
  2. 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  3. All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  4. Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
  5. More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
  6. List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
  7. 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  8. Start at low effort and measure. Every call, interval and cost online.

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Medium effort
  • High effort
  • 95% interval
  • fastest–slowest run (not an interval)
  • where the two overlap
  • hollow: list-price calculation

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

2 ties · 6 unclear
  • Pass rate on five validated tasks: Medium effort 100% (15/15) (n 15, 95% interval 80%–100%); High effort 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
  • Total time per call: Medium effort 5.65 s (n 15, run range 4.1 s–25.5 s); High effort 5.60 s (n 15, run range 4.1 s–19.5 s). Unclear.
  • Time to first useful output: Medium effort 5.05 s (n 15, run range 3.4 s–17.8 s); High effort 5.32 s (n 15, run range 3.6 s–16.4 s). Unclear.
  • Input tokens per call: what the CLI sends (Cache read): Medium effort 5,180 (n 15); High effort 6,716 (n 15). Unclear.
  • Input tokens per call: what the CLI sends (Other input): Medium effort 6,943 (n 15); High effort 5,406 (n 15). Unclear.
  • Output tokens per call (Output tokens): Medium effort 42 (n 15); High effort 42 (n 15). Tie.
  • List-price cost per call (calculation), calculation: Medium effort $0.010 (n 15, run range $0.0054–$0.027); High effort $0.010 (n 15, run range $0.0066–$0.028). Unclear.
  • List-price cost per passing answer (calculation), calculation: Medium effort $0.016 (n 15); High effort $0.013 (n 15). Unclear.

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

2 ties · 4 unclear
  • Pass rate on eight hard tasks (Strict pass): Medium effort 100% (16/16) (n 16, 95% interval 81%–100%); High effort 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Pass rate on eight hard tasks (Lenient (format misses counted)): Medium effort 100% (16/16) (n 16, 95% interval 81%–100%); High effort 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Total time per call on hard tasks (separate batches): Medium effort 13.1 s (n 16, run range 8.5 s–61.6 s); High effort 18.1 s (n 16, run range 11.7 s–92.2 s). Unclear.
  • Time to first useful output on hard tasks: Medium effort 10.2 s (n 16, run range 6.1 s–40.4 s); High effort 12.7 s (n 16, run range 8.9 s–75.9 s). Unclear.
  • Output tokens per call on hard tasks (Output tokens): Medium effort 335 (n 16); High effort 436 (n 16). Unclear.
  • List-price cost per strict pass on hard tasks (calculation), calculation: Medium effort $0.026 (n 16); High effort $0.015 (n 16). Unclear.

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks

1 tie · 3 unclear
  • Strict pass rate by effort on eight hard tasks: Medium effort 100% (16/16) (n 16, 95% interval 81%–100%); High effort 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Total time per call by effort on hard tasks: Medium effort 13.1 s (n 16, run range 8.5 s–61.6 s); High effort 18.1 s (n 16, run range 11.7 s–92.2 s). Unclear.
  • Output tokens per call by effort on hard tasks (Output tokens): Medium effort 335 (n 16); High effort 436 (n 16). Unclear.
  • List-price cost per strict pass by effort (calculation), calculation: Medium effort $0.026 (n 16); High effort $0.015 (n 16). Unclear.

How much of an AI bill is thinking? Reasoning tokens by model and effort

8 unclear
  • Reasoning share of output tokens per call on hard tasks (calculation), calculation: Medium effort 46.3% (n 16, run range 11%–87%); High effort 57% (n 16, run range 29%–91%). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Medium effort $0.0023 (n 16); High effort $0.0041 (n 16). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Medium effort $0.0030 (n 16); High effort $0.0029 (n 16). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Medium effort $0.020 (n 16); High effort $0.0081 (n 16). Unclear.
  • Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass), calculation: Medium effort $0.0023 (n 16); High effort $0.0041 (n 16). Unclear.
  • Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass), calculation: Medium effort $0.026 (n 16); High effort $0.015 (n 16). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Medium effort 46.3% (n 16, run range 11%–87%); High effort 57% (n 16, run range 29%–91%). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Medium effort 41% (n 15, run range 0%–71%); High effort 58.1% (n 15, run range 0%–76%). Unclear.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

26 rows from 4 studies. No row separates them: 5 ties, 21 unclear.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick GPT-6.1 Sol (Codex CLI) at medium effort

  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0023 vs $0.0041. A list-price calculation, not a measured difference. Calculation
  • Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass): $0.0023 vs $0.0041. A list-price calculation, not a measured difference. Calculation

When to pick GPT-6.1 Sol (Codex CLI) at high effort

  • List-price cost per strict pass on hard tasks (calculation): $0.015 vs $0.026. A list-price calculation, not a measured difference. Calculation
  • List-price cost per strict pass by effort (calculation): $0.015 vs $0.026. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)): $0.0081 vs $0.020. A list-price calculation, not a measured difference. Calculation
  • Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass): $0.015 vs $0.026. A list-price calculation, not a measured difference. Calculation

Side by side

The study charts, showing only GPT-6.1 Sol (Codex CLI) at each effort it ran; the two compared settings are highlighted. Open a study for every configuration.

Every rate is 95% or more
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Every interval overlaps every other: this chart does not order these rows.

3 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln 10–15 per row3 of 3 at 100%: this task set cannot separate them.

Every call counts; failures and timeouts are non-passes

Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 4.5× real timeMotion reduced: press Replay to animateThe slowest median is 6.3 s. The clock runs at the recorded speed.
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

3 rows. Slowest GPT-6.1 Sol (low) · Codex CLI 6.3 s (range 4.7 s–10.5 s, n 10). Fastest GPT-6.1 Sol (high) · Codex CLI 5.6 s (range 4.1 s–19.5 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 10–15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 3.8× real timeMotion reduced: press Replay to animateThe slowest median is 5.3 s. The clock runs at the recorded speed.
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

3 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 5.3 s (range 3.6 s–16.4 s, n 15). Fastest GPT-6.1 Sol (medium) · Codex CLI 5.1 s (range 3.4 s–17.8 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 10–15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

  • Cache read
  • Other input
Bar length is the total; segments are its parts.
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Totals are the sum of the parts shown. Shares are calculated from the same values.

3 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (low) · Codex CLI 8,064 (n 10). Lowest GPT-6.1 Sol (medium) · Codex CLI 5,180 (n 15). Other input: highest GPT-6.1 Sol (medium) · Codex CLI 6,943 (n 15). Lowest GPT-6.1 Sol (low) · Codex CLI 4,059 (n 10).

Notesn 10–15 per row

Mean per call, split into prompt-cache reads and other input

The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.

Source: Provider head-to-head: Claude Code models vs Codex efforts

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Is GPT-6.1 Sol (Codex CLI) better at medium effort or high effort?
GPT-6.1 Sol (Codex CLI) at medium effort and GPT-6.1 Sol (Codex CLI) at high effort share 14 measured metrics and 12 list-price calculations from 4 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 5 ties and 21 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).
How were the two effort settings measured?
They share 14 measured metrics and 12 list-price calculations from 4 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks; How much of an AI bill is thinking? Reasoning tokens by model and effort. Only the effort setting differs between the two sides of a row.
How does pass rate on five validated tasks change from medium effort to high effort?
GPT-6.1 Sol (Codex CLI) at medium effort: 100% (15/15) (n = 15, 95% interval 80% to 100%). GPT-6.1 Sol (Codex CLI) at high effort: 100% (15/15) (n = 15, 95% interval 80% to 100%). The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 80% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 80% to 100%), so this sample cannot separate them.
How does total time per call change from medium effort to high effort?
GPT-6.1 Sol (Codex CLI) at medium effort: 5.65 s (n = 15, run range 4.1 s to 25.5 s). GPT-6.1 Sol (Codex CLI) at high effort: 5.60 s (n = 15, run range 4.1 s to 19.5 s). The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 4.10 s to 25.5 s; GPT-6.1 Sol (Codex CLI) at high effort 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
How does time to first useful output change from medium effort to high effort?
GPT-6.1 Sol (Codex CLI) at medium effort: 5.05 s (n = 15, run range 3.4 s to 17.8 s). GPT-6.1 Sol (Codex CLI) at high effort: 5.32 s (n = 15, run range 3.6 s to 16.4 s). The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 3.36 s to 17.8 s; GPT-6.1 Sol (Codex CLI) at high effort 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
How does pass rate on eight hard tasks (Strict pass) change from medium effort to high effort?
GPT-6.1 Sol (Codex CLI) at medium effort: 100% (16/16) (n = 16, 95% interval 81% to 100%). GPT-6.1 Sol (Codex CLI) at high effort: 100% (16/16) (n = 16, 95% interval 81% to 100%). The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 81% to 100%), so this sample cannot separate them.

The studies behind this page

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.