14 measured metrics · 12 calculated · 4 studies
GPT-6.1 Sol (Codex CLI)mediumvshigheffort
No row separates them: 5 ties, 21 unclear.
The verdict
GPT-6.1 Sol (Codex CLI) at medium effort and GPT-6.1 Sol (Codex CLI) at high effort share 14 measured metrics and 12 list-price calculations from 4 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 5 ties and 21 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | GPT-6.1 Sol (Codex CLI) at medium effort | GPT-6.1 Sol (Codex CLI) at high effort | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 100% (15/15)Codex CLI · effort medium · five short validated tasks | 100% (15/15)Codex CLI · effort high · five short validated tasks | 15 | 95% CI: 80%–100% vs 80%–100% | Tie | The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 80% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 5.65 sCodex CLI · effort medium · five short validated tasks | 5.60 sCodex CLI · effort high · five short validated tasks | 15 | range: 4.1 s–25.5 s vs 4.1 s–19.5 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 4.10 s to 25.5 s; GPT-6.1 Sol (Codex CLI) at high effort 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 5.05 sCodex CLI · effort medium · five short validated tasks | 5.32 sCodex CLI · effort high · five short validated tasks | 15 | range: 3.4 s–17.8 s vs 3.6 s–16.4 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 3.36 s to 17.8 s; GPT-6.1 Sol (Codex CLI) at high effort 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 5,180Codex CLI · effort medium · five short validated tasks | 6,716Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 42Codex CLI · effort medium · five short validated tasks | 42Codex CLI · effort high · five short validated tasks | 15 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.016Codex CLI · effort medium · five short validated tasks | $0.013Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.016 vs $0.013) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Input tokens per call: what the CLI sends (Cache read), 1.3x (High effort larger).
Watch it build
A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks
All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.
Transcript
- Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
- 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
- All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
- More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
- List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
- 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Start at low effort and measure. Every call, interval and cost online.
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- Medium effort
- High effort
- 95% interval
- fastest–slowest run (not an interval)
- where the two overlap
- hollow: list-price calculation
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
- Pass rate on five validated tasks100% (15/15)n 15100% (15/15)n 15TiePass rate on five validated tasks: Medium effort 100% (15/15) (n 15, 95% interval 80%–100%); High effort 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
- Total time per call5.65 sn 155.60 sn 15UnclearTotal time per call: Medium effort 5.65 s (n 15, run range 4.1 s–25.5 s); High effort 5.60 s (n 15, run range 4.1 s–19.5 s). Unclear.
- Time to first useful output5.05 sn 155.32 sn 15UnclearTime to first useful output: Medium effort 5.05 s (n 15, run range 3.4 s–17.8 s); High effort 5.32 s (n 15, run range 3.6 s–16.4 s). Unclear.
- Input tokens per call: what the CLI sends (Cache read)5,180n 156,716n 15UnclearInput tokens per call: what the CLI sends (Cache read): Medium effort 5,180 (n 15); High effort 6,716 (n 15). Unclear.
- Input tokens per call: what the CLI sends (Other input)6,943n 155,406n 15UnclearInput tokens per call: what the CLI sends (Other input): Medium effort 6,943 (n 15); High effort 5,406 (n 15). Unclear.
- Output tokens per call (Output tokens)42n 1542n 15TieOutput tokens per call (Output tokens): Medium effort 42 (n 15); High effort 42 (n 15). Tie.
- List-price cost per call (calculation)Calculation$0.010n 15$0.010n 15UnclearList-price cost per call (calculation), calculation: Medium effort $0.010 (n 15, run range $0.0054–$0.027); High effort $0.010 (n 15, run range $0.0066–$0.028). Unclear.
- List-price cost per passing answer (calculation)Calculation$0.016n 15$0.013n 15UnclearList-price cost per passing answer (calculation), calculation: Medium effort $0.016 (n 15); High effort $0.013 (n 15). Unclear.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
- Pass rate on eight hard tasks (Strict pass)100% (16/16)n 16100% (16/16)n 16TiePass rate on eight hard tasks (Strict pass): Medium effort 100% (16/16) (n 16, 95% interval 81%–100%); High effort 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
- Pass rate on eight hard tasks (Lenient (format misses counted))100% (16/16)n 16100% (16/16)n 16TiePass rate on eight hard tasks (Lenient (format misses counted)): Medium effort 100% (16/16) (n 16, 95% interval 81%–100%); High effort 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
- Total time per call on hard tasks (separate batches)13.1 sn 1618.1 sn 16UnclearTotal time per call on hard tasks (separate batches): Medium effort 13.1 s (n 16, run range 8.5 s–61.6 s); High effort 18.1 s (n 16, run range 11.7 s–92.2 s). Unclear.
- Time to first useful output on hard tasks10.2 sn 1612.7 sn 16UnclearTime to first useful output on hard tasks: Medium effort 10.2 s (n 16, run range 6.1 s–40.4 s); High effort 12.7 s (n 16, run range 8.9 s–75.9 s). Unclear.
- Output tokens per call on hard tasks (Output tokens)335n 16436n 16UnclearOutput tokens per call on hard tasks (Output tokens): Medium effort 335 (n 16); High effort 436 (n 16). Unclear.
- List-price cost per strict pass on hard tasks (calculation)Calculation$0.026n 16$0.015n 16UnclearList-price cost per strict pass on hard tasks (calculation), calculation: Medium effort $0.026 (n 16); High effort $0.015 (n 16). Unclear.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
- Strict pass rate by effort on eight hard tasks100% (16/16)n 16100% (16/16)n 16TieStrict pass rate by effort on eight hard tasks: Medium effort 100% (16/16) (n 16, 95% interval 81%–100%); High effort 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
- Total time per call by effort on hard tasks13.1 sn 1618.1 sn 16UnclearTotal time per call by effort on hard tasks: Medium effort 13.1 s (n 16, run range 8.5 s–61.6 s); High effort 18.1 s (n 16, run range 11.7 s–92.2 s). Unclear.
- Output tokens per call by effort on hard tasks (Output tokens)335n 16436n 16UnclearOutput tokens per call by effort on hard tasks (Output tokens): Medium effort 335 (n 16); High effort 436 (n 16). Unclear.
- List-price cost per strict pass by effort (calculation)Calculation$0.026n 16$0.015n 16UnclearList-price cost per strict pass by effort (calculation), calculation: Medium effort $0.026 (n 16); High effort $0.015 (n 16). Unclear.
How much of an AI bill is thinking? Reasoning tokens by model and effort
- Reasoning share of output tokens per call on hard tasks (calculation)Calculation46.3%n 1657%n 16UnclearReasoning share of output tokens per call on hard tasks (calculation), calculation: Medium effort 46.3% (n 16, run range 11%–87%); High effort 57% (n 16, run range 29%–91%). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation$0.0023n 16$0.0041n 16UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Medium effort $0.0023 (n 16); High effort $0.0041 (n 16). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation$0.0030n 16$0.0029n 16UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Medium effort $0.0030 (n 16); High effort $0.0029 (n 16). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation$0.020n 16$0.0081n 16UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Medium effort $0.020 (n 16); High effort $0.0081 (n 16). Unclear.
- Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation$0.0023n 16$0.0041n 16UnclearReasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass), calculation: Medium effort $0.0023 (n 16); High effort $0.0041 (n 16). Unclear.
- Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation$0.026n 16$0.015n 16UnclearReasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass), calculation: Medium effort $0.026 (n 16); High effort $0.015 (n 16). Unclear.
- Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation46.3%n 1657%n 16UnclearReasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Medium effort 46.3% (n 16, run range 11%–87%); High effort 57% (n 16, run range 29%–91%). Unclear.
- Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation41%n 1558.1%n 15UnclearReasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Medium effort 41% (n 15, run range 0%–71%); High effort 58.1% (n 15, run range 0%–76%). Unclear.
| Metric | GPT-6.1 Sol (Codex CLI) at medium effort | GPT-6.1 Sol (Codex CLI) at high effort | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 100% (15/15)Codex CLI · effort medium · five short validated tasks | 100% (15/15)Codex CLI · effort high · five short validated tasks | 15 | 95% CI: 80%–100% vs 80%–100% | Tie | The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 80% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 5.65 sCodex CLI · effort medium · five short validated tasks | 5.60 sCodex CLI · effort high · five short validated tasks | 15 | range: 4.1 s–25.5 s vs 4.1 s–19.5 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 4.10 s to 25.5 s; GPT-6.1 Sol (Codex CLI) at high effort 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 5.05 sCodex CLI · effort medium · five short validated tasks | 5.32 sCodex CLI · effort high · five short validated tasks | 15 | range: 3.4 s–17.8 s vs 3.6 s–16.4 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 3.36 s to 17.8 s; GPT-6.1 Sol (Codex CLI) at high effort 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 5,180Codex CLI · effort medium · five short validated tasks | 6,716Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Other input) | 6,943Codex CLI · effort medium · five short validated tasks | 5,406Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 42Codex CLI · effort medium · five short validated tasks | 42Codex CLI · effort high · five short validated tasks | 15 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per call (calculation)Calculation | $0.010Codex CLI · effort medium · five short validated tasks | $0.010Codex CLI · effort high · five short validated tasks | 15 | range: $0.0054–$0.027 vs $0.0066–$0.028 | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort $0.0054 to $0.027; GPT-6.1 Sol (Codex CLI) at high effort $0.0066 to $0.028); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.016Codex CLI · effort medium · five short validated tasks | $0.013Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.016 vs $0.013) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Pass rate on eight hard tasks (Strict pass) | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks | 100% (16/16)Codex CLI · effort high · eight hard validated tasks | 16 | 95% CI: 81%–100% vs 81%–100% | Tie | The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks | 100% (16/16)Codex CLI · effort high · eight hard validated tasks | 16 | 95% CI: 81%–100% vs 81%–100% | Tie | The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Total time per call on hard tasks (separate batches) | 13.1 sCodex CLI · effort medium · eight hard validated tasks | 18.1 sCodex CLI · effort high · eight hard validated tasks | 16 | range: 8.5 s–61.6 s vs 11.7 s–92.2 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 8.54 s to 61.6 s; GPT-6.1 Sol (Codex CLI) at high effort 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Time to first useful output on hard tasks | 10.2 sCodex CLI · effort medium · eight hard validated tasks | 12.7 sCodex CLI · effort high · eight hard validated tasks | 16 | range: 6.1 s–40.4 s vs 8.9 s–75.9 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 6.09 s to 40.4 s; GPT-6.1 Sol (Codex CLI) at high effort 8.93 s to 75.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call on hard tasks (Output tokens) | 335Codex CLI · effort medium · eight hard validated tasks | 436Codex CLI · effort high · eight hard validated tasks | 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass on hard tasks (calculation)Calculation | $0.026Codex CLI · effort medium · eight hard validated tasks | $0.015Codex CLI · effort high · eight hard validated tasks | 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.026 vs $0.015) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Strict pass rate by effort on eight hard tasks | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks, effort ladder | 100% (16/16)Codex CLI · effort high · eight hard validated tasks, effort ladder | 16 | 95% CI: 81%–100% vs 81%–100% | Tie | The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 81% to 100%), so this sample cannot separate them. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Total time per call by effort on hard tasks | 13.1 sCodex CLI · effort medium · eight hard validated tasks, effort ladder | 18.1 sCodex CLI · effort high · eight hard validated tasks, effort ladder | 16 | range: 8.5 s–61.6 s vs 11.7 s–92.2 s | Unclear | The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 8.54 s to 61.6 s; GPT-6.1 Sol (Codex CLI) at high effort 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call by effort on hard tasks (Output tokens) | 335Codex CLI · effort medium · eight hard validated tasks, effort ladder | 436Codex CLI · effort high · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass by effort (calculation)Calculation | $0.026Codex CLI · effort medium · eight hard validated tasks, effort ladder | $0.015Codex CLI · effort high · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.026 vs $0.015) is not tested against run-to-run variation. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Reasoning share of output tokens per call on hard tasks (calculation)Calculation | 46.3%Codex CLI · effort medium | 57%Codex CLI · effort high | 16 | range: 11%–87% vs 29%–91% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation | $0.0023Codex CLI · effort medium | $0.0041Codex CLI · effort high | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation | $0.0030Codex CLI · effort medium | $0.0029Codex CLI · effort high | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation | $0.020Codex CLI · effort medium | $0.0081Codex CLI · effort high | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation | $0.0023Codex CLI · effort medium | $0.0041Codex CLI · effort high | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation | $0.026Codex CLI · effort medium | $0.015Codex CLI · effort high | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation | 46.3%Codex CLI · effort medium | 57%Codex CLI · effort high | 16 | range: 11%–87% vs 29%–91% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation | 41%Codex CLI · effort medium | 58.1%Codex CLI · effort high | 15 | range: 0%–71% vs 0%–76% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
26 rows from 4 studies. No row separates them: 5 ties, 21 unclear.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick GPT-6.1 Sol (Codex CLI) at medium effort
- List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0023 vs $0.0041. A list-price calculation, not a measured difference. Calculation
- Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass): $0.0023 vs $0.0041. A list-price calculation, not a measured difference. Calculation
When to pick GPT-6.1 Sol (Codex CLI) at high effort
- List-price cost per strict pass on hard tasks (calculation): $0.015 vs $0.026. A list-price calculation, not a measured difference. Calculation
- List-price cost per strict pass by effort (calculation): $0.015 vs $0.026. A list-price calculation, not a measured difference. Calculation
- List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)): $0.0081 vs $0.020. A list-price calculation, not a measured difference. Calculation
- Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass): $0.015 vs $0.026. A list-price calculation, not a measured difference. Calculation
Side by side
The study charts, showing only GPT-6.1 Sol (Codex CLI) at each effort it ran; the two compared settings are highlighted. Open a study for every configuration.
Every interval overlaps every other: this chart does not order these rows.
| Item | Pass rate | 95% interval | n |
|---|---|---|---|
| GPT-6.1 Sol (high) · Codex CLI | 100% | 80%–100% | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | 100% | 80%–100% | 15 |
| GPT-6.1 Sol (low) · Codex CLI | 100% | 72%–100% | 10 |
3 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln 10–15 per row3 of 3 at 100%: this task set cannot separate them.
Every call counts; failures and timeouts are non-passes
Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.
Source: Provider head-to-head: Claude Code models vs Codex efforts
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (high) · Codex CLI | 5.6 s | 4.1 s–19.5 s | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | 5.7 s | 4.1 s–25.5 s | 15 |
| GPT-6.1 Sol (low) · Codex CLI | 6.3 s | 4.7 s–10.5 s | 10 |
3 rows. Slowest GPT-6.1 Sol (low) · Codex CLI 6.3 s (range 4.7 s–10.5 s, n 10). Fastest GPT-6.1 Sol (high) · Codex CLI 5.6 s (range 4.1 s–19.5 s, n 15). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 10–15 per row
Median per configuration; whiskers = fastest and slowest call
One host, one network, one day. Whiskers are a range, not a confidence interval.
Source: Provider head-to-head: Claude Code models vs Codex efforts
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first useful output | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (high) · Codex CLI | 5.3 s | 3.6 s–16.4 s | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | 5.1 s | 3.4 s–17.8 s | 15 |
| GPT-6.1 Sol (low) · Codex CLI | 5.1 s | 4 s–8.5 s | 10 |
3 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 5.3 s (range 3.6 s–16.4 s, n 15). Fastest GPT-6.1 Sol (medium) · Codex CLI 5.1 s (range 3.4 s–17.8 s, n 15). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 10–15 per row
Median per configuration; whiskers = fastest and slowest call
One host, one network, one day. Whiskers are a range, not a confidence interval.
Source: Provider head-to-head: Claude Code models vs Codex efforts
- Cache read
- Other input
Totals are the sum of the parts shown. Shares are calculated from the same values.
| Item | Cache read | Other input | n |
|---|---|---|---|
| GPT-6.1 Sol (high) · Codex CLI | 6,716 | 5,406 | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | 5,180 | 6,943 | 15 |
| GPT-6.1 Sol (low) · Codex CLI | 8,064 | 4,059 | 10 |
3 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (low) · Codex CLI 8,064 (n 10). Lowest GPT-6.1 Sol (medium) · Codex CLI 5,180 (n 15). Other input: highest GPT-6.1 Sol (medium) · Codex CLI 6,943 (n 15). Lowest GPT-6.1 Sol (low) · Codex CLI 4,059 (n 10).
Notesn 10–15 per row
Mean per call, split into prompt-cache reads and other input
The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.
Source: Provider head-to-head: Claude Code models vs Codex efforts
| Item | Output tokens | n |
|---|---|---|
| GPT-6.1 Sol (high) · Codex CLI | 42 | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | 42 | 15 |
| GPT-6.1 Sol (low) · Codex CLI | 42 | 10 |
3 rows. All at 42.
Notesn 10–15 per row
Median per configuration; reasoning tokens where the CLI reports them
Vendors count reasoning differently; compare within a vendor first. A median of 0 can mean the CLI did not report reasoning for most calls.
Source: Provider head-to-head: Claude Code models vs Codex efforts
Every interval overlaps every other: this chart does not order these rows.
| Item | Cost per call | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (high) · Codex CLI | $0.01 | $0.0066–$0.028 | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.01 | $0.0054–$0.027 | 15 |
| GPT-6.1 Sol (low) · Codex CLI | $0.0077 | $0.0074–$0.026 | 10 |
List-price calculation, not a run. 3 rows. Highest GPT-6.1 Sol (high) · Codex CLI $0.01 (range $0.0066–$0.028, n 15). Lowest GPT-6.1 Sol (low) · Codex CLI $0.0077 (range $0.0074–$0.026, n 10). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 10–15 per row
Reported tokens × list price; the calls ran on subscriptions
Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.
Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
Hover or focus a bar for its ratio to GPT-6.1 Sol (low) (the lowest value): a ratio of list-price calculations, not a measurement.
| Item | Cost per pass | n |
|---|---|---|
| GPT-6.1 Sol (low) · Codex CLI | $0.01 | 10 |
| GPT-6.1 Sol (high) · Codex CLI | $0.013 | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.016 | 15 |
List-price calculation, not a run. 3 rows. Highest GPT-6.1 Sol (medium) · Codex CLI $0.016 (n 15). Lowest GPT-6.1 Sol (low) · Codex CLI $0.01 (n 10).
Notesn 10–15 per row
All calls in a configuration, failures included, divided by its passes
Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.
Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Strict pass
- Lenient (format misses counted)
| Item | Strict pass | Lenient (format misses counted) | 95% interval | n |
|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 100% | 100% | Strict pass: 81%–100%; Lenient (format misses counted): 81%–100% | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 100% | 100% | Strict pass: 81%–100%; Lenient (format misses counted): 81%–100% | 16 |
2 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: all at 100%. Lenient (format misses counted): all at 100%.
NotesWhiskers: 95% Wilson intervaln = 16 per row
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call on hard tasks (separate batches) | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 13.1 s | 8.5 s–61.6 s | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 18.1 s | 11.7 s–92.2 s | 16 |
2 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 18.1 s (range 11.7 s–92.2 s, n 16). Fastest GPT-6.1 Sol (medium) · Codex CLI 13.1 s (range 8.5 s–61.6 s, n 16). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 16 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first useful output on hard tasks | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 10.2 s | 6.1 s–40.4 s | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 12.7 s | 8.9 s–75.9 s | 16 |
2 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 12.7 s (range 8.9 s–75.9 s, n 16). Fastest GPT-6.1 Sol (medium) · Codex CLI 10.2 s (range 6.1 s–40.4 s, n 16). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 16 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
| Item | Output tokens | n |
|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 335 | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 436 | 16 |
2 rows. Highest GPT-6.1 Sol (high) · Codex CLI 436 (n 16). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16).
Notesn = 16 per row
Median per configuration; reasoning tokens as the CLI reports them
Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
| Item | Cost per strict pass | n |
|---|---|---|
| GPT-6.1 Sol (high) · Codex CLI | $0.015 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.026 | 16 |
List-price calculation, not a run. 2 rows. Highest GPT-6.1 Sol (medium) · Codex CLI $0.026 (n 16). Lowest GPT-6.1 Sol (high) · Codex CLI $0.015 (n 16).
Notesn = 16 per row
All calls in a configuration, failures and format misses included, divided by its strict passes
Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.
Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
Every interval overlaps every other: this chart does not order these rows.
| Item | Strict pass | 95% interval | n |
|---|---|---|---|
| GPT-6.1 Sol (low) · Codex CLI | 100% | 81%–100% | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | 100% | 81%–100% | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 100% | 81%–100% | 16 |
3 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 16 per row3 of 3 at 100%: this task set cannot separate them.
Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.
Source: Effort ladder: the hard task set at each effort level
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call by effort on hard tasks | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (low) · Codex CLI | 13.6 s | 7.9 s–44.3 s | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | 13.1 s | 8.5 s–61.6 s | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 18.1 s | 11.7 s–92.2 s | 16 |
3 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 18.1 s (range 11.7 s–92.2 s, n 16). Fastest GPT-6.1 Sol (medium) · Codex CLI 13.1 s (range 8.5 s–61.6 s, n 16). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 16 per row
Median per configuration; whiskers = fastest and slowest call
Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.
Source: Effort ladder: the hard task set at each effort level
| Item | Output tokens | n |
|---|---|---|
| GPT-6.1 Sol (low) · Codex CLI | 284 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | 335 | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 436 | 16 |
3 rows. Highest GPT-6.1 Sol (high) · Codex CLI 436 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI 284 (n 16).
Notesn = 16 per row
Median per configuration; reasoning tokens as the CLI reports them
Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean "not reported". Their content is never captured. More tokens is not better or worse by itself.
Source: Effort ladder: the hard task set at each effort level
Hover or focus a bar for its ratio to GPT-6.1 Sol (low) (the lowest value): a ratio of list-price calculations, not a measurement.
| Item | Cost per strict pass | n |
|---|---|---|
| GPT-6.1 Sol (low) · Codex CLI | $0.013 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.026 | 16 |
| GPT-6.1 Sol (high) · Codex CLI | $0.015 | 16 |
List-price calculation, not a run. 3 rows. Highest GPT-6.1 Sol (medium) · Codex CLI $0.026 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI $0.013 (n 16).
Notesn = 16 per row
All calls in a configuration divided by its strict passes
Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.
Sources: Effort ladder: the hard task set at each effort level, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Reasoning (output tokens)
- Remaining output (visible-answer estimate)
- Input (prompt, cache priced)
Totals are the sum of the parts shown. Shares are calculated from the same values.
| Item | Reasoning (output tokens) | Remaining output (visible-answer estimate) | Input (prompt, cache priced) | n |
|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | $0.0023 | $0.003 | $0.02 | 16 |
| GPT-6.1 Sol (high) · Codex CLI | $0.0041 | $0.0029 | $0.0081 | 16 |
List-price calculation, not a run. 2 rows, 3 series: Reasoning (output tokens), Remaining output (visible-answer estimate), Input (prompt, cache priced). Reasoning (output tokens): highest GPT-6.1 Sol (high) · Codex CLI $0.0041 (n 16). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.0023 (n 16). Remaining output (visible-answer estimate): highest GPT-6.1 Sol (medium) · Codex CLI $0.003 (n 16). Lowest GPT-6.1 Sol (high) · Codex CLI $0.0029 (n 16).
Notesn = 16 per row
Mean per call on the hard tasks; the three parts add up to the call
Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
- Reasoning cost per strict pass
- Total cost per strict pass (square)
Gap labels, Total cost per strict pass vs Reasoning cost per strict pass: Total cost per strict pass is x% higher (+) or lower (−) than Reasoning cost per strict pass, calculated from the two values shown (the change counted from Reasoning cost per strict pass’s value).
| Item | Reasoning cost per strict pass | Total cost per strict pass | n |
|---|---|---|---|
| GPT-6.1 Sol (low) · Codex CLI | $0.0012 | $0.013 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.0023 | $0.026 | 16 |
| GPT-6.1 Sol (high) · Codex CLI | $0.0041 | $0.015 | 16 |
List-price calculation, not a run. 3 rows, 2 series: Reasoning cost per strict pass, Total cost per strict pass. Reasoning cost per strict pass: highest GPT-6.1 Sol (high) · Codex CLI $0.0041 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI $0.0012 (n 16). Total cost per strict pass: highest GPT-6.1 Sol (medium) · Codex CLI $0.026 (n 16). Lowest GPT-6.1 Sol (low) · Codex CLI $0.013 (n 16).
Notesn = 16 per row
List price ÷ strict passes; every cell is 8 tasks × 2 repetitions
Calculation, not a bill: reported tokens × list price, divided by the cell's strict passes; the calls ran on flat subscriptions. The effort-ladder cells: new calls plus reference cells reused from the hard head-to-head (Claude repetitions 1-2 only). "Default" means the effort flag was not passed. The total is the same value as the effort-ladder cost-per-pass chart. Effort levels are not the same scale across vendors, and the reference cells ran in a different batch and hour.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
- Eight hard tasks
- Five short tasks (square)
Gap labels, Five short tasks vs Eight hard tasks: Five short tasks is x percentage points higher (+) or lower (−) than Eight hard tasks, calculated from the two values shown; lines are the lowest–highest run (not an interval).
| Item | Eight hard tasks | Five short tasks | Range (lowest–highest run) | n |
|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 46% | 41% | Eight hard tasks: 11%–87%; Five short tasks: 0%–71% | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 57% | 58% | Eight hard tasks: 29%–91%; Five short tasks: 0%–76% | 16 |
List-price calculation, not a run. 2 rows, 2 series: Eight hard tasks, Five short tasks. Eight hard tasks: highest GPT-6.1 Sol (high) · Codex CLI 57% (range 29%–91%, n 16). Lowest GPT-6.1 Sol (medium) · Codex CLI 46% (range 11%–87%, n 16). All run ranges overlap. Five short tasks: highest GPT-6.1 Sol (high) · Codex CLI 58% (range 0%–76%, n 15). Lowest GPT-6.1 Sol (medium) · Codex CLI 41% (range 0%–71%, n 15). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 15–16 per row
Median call per configuration; five short tasks and eight hard tasks
Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Is GPT-6.1 Sol (Codex CLI) better at medium effort or high effort?
- GPT-6.1 Sol (Codex CLI) at medium effort and GPT-6.1 Sol (Codex CLI) at high effort share 14 measured metrics and 12 list-price calculations from 4 studies. Only the effort setting differs between the two sides of a row; the route and the task set are the same. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 5 ties and 21 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 15 at the smallest).
- How were the two effort settings measured?
- They share 14 measured metrics and 12 list-price calculations from 4 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks; How much of an AI bill is thinking? Reasoning tokens by model and effort. Only the effort setting differs between the two sides of a row.
- How does pass rate on five validated tasks change from medium effort to high effort?
- GPT-6.1 Sol (Codex CLI) at medium effort: 100% (15/15) (n = 15, 95% interval 80% to 100%). GPT-6.1 Sol (Codex CLI) at high effort: 100% (15/15) (n = 15, 95% interval 80% to 100%). The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 80% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 80% to 100%), so this sample cannot separate them.
- How does total time per call change from medium effort to high effort?
- GPT-6.1 Sol (Codex CLI) at medium effort: 5.65 s (n = 15, run range 4.1 s to 25.5 s). GPT-6.1 Sol (Codex CLI) at high effort: 5.60 s (n = 15, run range 4.1 s to 19.5 s). The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 4.10 s to 25.5 s; GPT-6.1 Sol (Codex CLI) at high effort 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
- How does time to first useful output change from medium effort to high effort?
- GPT-6.1 Sol (Codex CLI) at medium effort: 5.05 s (n = 15, run range 3.4 s to 17.8 s). GPT-6.1 Sol (Codex CLI) at high effort: 5.32 s (n = 15, run range 3.6 s to 16.4 s). The run ranges (fastest to slowest) overlap (GPT-6.1 Sol (Codex CLI) at medium effort 3.36 s to 17.8 s; GPT-6.1 Sol (Codex CLI) at high effort 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
- How does pass rate on eight hard tasks (Strict pass) change from medium effort to high effort?
- GPT-6.1 Sol (Codex CLI) at medium effort: 100% (16/16) (n = 16, 95% interval 81% to 100%). GPT-6.1 Sol (Codex CLI) at high effort: 100% (16/16) (n = 16, 95% interval 81% to 100%). The 95% intervals overlap (GPT-6.1 Sol (Codex CLI) at medium effort 81% to 100%; GPT-6.1 Sol (Codex CLI) at high effort 81% to 100%), so this sample cannot separate them.
The studies behind this page
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
176 calls on 8 hard tasks at low, medium, high and default effort: Sonnet, Opus and GPT-6.1 Sol. Pass rate with 95% intervals, time, tokens, cost per pass.
How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.