31 measured metrics · 19 calculated · 7 studies
Claude Opus 5.5vsGPT-6.1 Sol (Codex CLI)
No row separates them: 12 ties, 38 unclear.
The verdict
Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) share 31 measured metrics and 19 list-price calculations from 7 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 12 ties and 38 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | Claude Opus 5.5 | GPT-6.1 Sol (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 100% (15/15)Claude Code · effort high · five short validated tasks | 100% (15/15)Codex CLI · effort high · five short validated tasks | 15 | 95% CI: 80%–100% vs 80%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 2.71 sClaude Code · effort high · five short validated tasks | 5.60 sCodex CLI · effort high · five short validated tasks | 15 | range: 2.5 s–11.8 s vs 4.1 s–19.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.45 s to 11.8 s; GPT-6.1 Sol (Codex CLI) 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 2.04 sClaude Code · effort high · five short validated tasks | 5.32 sCodex CLI · effort high · five short validated tasks | 15 | range: 1.4 s–9.9 s vs 3.6 s–16.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.40 s to 9.94 s; GPT-6.1 Sol (Codex CLI) 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 1,463Claude Code · effort high · five short validated tasks | 6,716Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 78Claude Code · effort high · five short validated tasks | 42Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.010Claude Code · effort high · five short validated tasks | $0.013Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.010 vs $0.013) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Input tokens per call: what the CLI sends (Cache read), 4.6x (GPT-6.1 Sol (Codex CLI) larger).
Watch it build
A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.
Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI): what the measurements say
50 comparison rows from 7 studies: 0 rows favour Opus 5.5, 0 favour GPT-6.1 Sol (Codex CLI), 50 are ties or unclear. Cost rows are calculations.
Transcript
- Comparison · 50 rows · 7 studies. Opus 5.5 vs GPT-6.1 Sol (Codex CLI). A winner only where the 95% intervals or run ranges do not overlap.
- 50 comparison rows from 7 studies. None separates them: intervals or run ranges overlap, or none was recorded. Rows where Opus 5.5 is ahead: 0 (of 50). Rows where GPT-6.1 Sol (Codex CLI) is ahead: 0 (of 50). Ties or unclear: 50 (12 ties · 38 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
- Five short tasks: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 8 rows separates them. Table: Five short tasks · 5 of 8 rows · Claude Code vs Codex CLI · n = 15 per side. Source study: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head. Rows shown: Pass rate on five validated tasks; Total time per call; Time to first useful output; Input tokens per call: what the CLI sends (Cache read); Input tokens per call: what the CLI sends (Other input). Recorded settings: Claude Code · effort high · five short validated tasks; Codex CLI · effort high · five short validated tasks. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
- Eight hard tasks: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 6 rows separates them. Table: Eight hard tasks · 5 of 6 rows · Claude Code vs Codex CLI · n = 24 vs 16. Source study: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks. Rows shown: Pass rate on eight hard tasks (Strict pass); Pass rate on eight hard tasks (Lenient (format misses counted)); Total time per call on hard tasks (separate batches); Time to first useful output on hard tasks; Output tokens per call on hard tasks (Output tokens). Recorded settings: Claude Code · effort high · eight hard validated tasks; Codex CLI · effort high · eight hard validated tasks. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
- Coding agents, hidden tests: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 4 rows separates them. Table: Coding agents, hidden tests · Claude Code vs Codex CLI · n = 12 per side. Source study: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks. Rows shown: Coding sessions that passed every hidden check; Time per coding session; Tool calls per coding session; List-price cost per passing coding session (calculation). Recorded settings: Claude Code · six small repository tasks with hidden tests; Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests. Includes a calculation, not a bill or a new run. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
- Effort ladder, default effort: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 4 rows separates them. Table: Effort ladder, default effort · Claude Code vs Codex CLI · n = 16 per side. Source study: Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks. Rows shown: Strict pass rate by effort on eight hard tasks; Total time per call by effort on hard tasks; Output tokens per call by effort on hard tasks (Output tokens); List-price cost per strict pass by effort (calculation). Recorded settings: Claude Code · effort medium · eight hard validated tasks, effort ladder; Codex CLI · effort medium · eight hard validated tasks, effort ladder. Includes a calculation, not a bill or a new run. Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
- No winner where the data shows none. Showing 18 of 50 rows; every row and its reason online.
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- Claude Opus 5.5
- GPT-6.1 Sol (Codex CLI)
- 95% interval
- fastest–slowest run (not an interval)
- where the two overlap
- hollow: list-price calculation
These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
- Pass rate on five validated tasks100% (15/15)n 15100% (15/15)n 15TiePass rate on five validated tasks: Claude Opus 5.5 100% (15/15) (n 15, 95% interval 80%–100%); GPT-6.1 Sol (Codex CLI) 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
- Total time per call2.71 sn 155.60 sn 15UnclearTotal time per call: Claude Opus 5.5 2.71 s (n 15, run range 2.5 s–11.8 s); GPT-6.1 Sol (Codex CLI) 5.60 s (n 15, run range 4.1 s–19.5 s). Unclear.
- Time to first useful output2.04 sn 155.32 sn 15UnclearTime to first useful output: Claude Opus 5.5 2.04 s (n 15, run range 1.4 s–9.9 s); GPT-6.1 Sol (Codex CLI) 5.32 s (n 15, run range 3.6 s–16.4 s). Unclear.
- Input tokens per call: what the CLI sends (Cache read)1,463n 156,716n 15UnclearInput tokens per call: what the CLI sends (Cache read): Claude Opus 5.5 1,463 (n 15); GPT-6.1 Sol (Codex CLI) 6,716 (n 15). Unclear.
- Input tokens per call: what the CLI sends (Other input)619n 155,406n 15UnclearInput tokens per call: what the CLI sends (Other input): Claude Opus 5.5 619 (n 15); GPT-6.1 Sol (Codex CLI) 5,406 (n 15). Unclear.
- Output tokens per call (Output tokens)78n 1542n 15UnclearOutput tokens per call (Output tokens): Claude Opus 5.5 78 (n 15); GPT-6.1 Sol (Codex CLI) 42 (n 15). Unclear.
- List-price cost per call (calculation)Calculation$0.0069n 15$0.010n 15UnclearList-price cost per call (calculation), calculation: Claude Opus 5.5 $0.0069 (n 15, run range $0.0059–$0.027); GPT-6.1 Sol (Codex CLI) $0.010 (n 15, run range $0.0066–$0.028). Unclear.
- List-price cost per passing answer (calculation)Calculation$0.010n 15$0.013n 15UnclearList-price cost per passing answer (calculation), calculation: Claude Opus 5.5 $0.010 (n 15); GPT-6.1 Sol (Codex CLI) $0.013 (n 15). Unclear.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
- Pass rate on eight hard tasks (Strict pass)100% (24/24)n 24100% (16/16)n 16TiePass rate on eight hard tasks (Strict pass): Claude Opus 5.5 100% (24/24) (n 24, 95% interval 86%–100%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
- Pass rate on eight hard tasks (Lenient (format misses counted))100% (24/24)n 24100% (16/16)n 16TiePass rate on eight hard tasks (Lenient (format misses counted)): Claude Opus 5.5 100% (24/24) (n 24, 95% interval 86%–100%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
- Total time per call on hard tasks (separate batches)11.0 sn 2418.1 sn 16UnclearTotal time per call on hard tasks (separate batches): Claude Opus 5.5 11.0 s (n 24, run range 3.6 s–63 s); GPT-6.1 Sol (Codex CLI) 18.1 s (n 16, run range 11.7 s–92.2 s). Unclear.
- Time to first useful output on hard tasks7.13 sn 2412.7 sn 16UnclearTime to first useful output on hard tasks: Claude Opus 5.5 7.13 s (n 24, run range 2.2 s–56.2 s); GPT-6.1 Sol (Codex CLI) 12.7 s (n 16, run range 8.9 s–75.9 s). Unclear.
- Output tokens per call on hard tasks (Output tokens)1,052n 24436n 16UnclearOutput tokens per call on hard tasks (Output tokens): Claude Opus 5.5 1,052 (n 24); GPT-6.1 Sol (Codex CLI) 436 (n 16). Unclear.
- List-price cost per strict pass on hard tasks (calculation)Calculation$0.033n 24$0.015n 16UnclearList-price cost per strict pass on hard tasks (calculation), calculation: Claude Opus 5.5 $0.033 (n 24); GPT-6.1 Sol (Codex CLI) $0.015 (n 16). Unclear.
Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
- Coding sessions that passed every hidden check100% (12/12)n 12100% (12/12)n 12TieCoding sessions that passed every hidden check: Claude Opus 5.5 100% (12/12) (n 12, 95% interval 76%–100%); GPT-6.1 Sol (Codex CLI) 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
- Time per coding session56.9 sn 12113.4 sn 12UnclearTime per coding session: Claude Opus 5.5 56.9 s (n 12, run range 29.8 s–186 s); GPT-6.1 Sol (Codex CLI) 113.4 s (n 12, run range 78.5 s–222 s). Unclear.
- Tool calls per coding session7.5n 1212.5n 12UnclearTool calls per coding session: Claude Opus 5.5 7.5 (n 12, run range 5–14); GPT-6.1 Sol (Codex CLI) 12.5 (n 12, run range 8–18). Unclear.
- List-price cost per passing coding session (calculation)Calculation$0.22n 12$0.098n 12UnclearList-price cost per passing coding session (calculation), calculation: Claude Opus 5.5 $0.22 (n 12); GPT-6.1 Sol (Codex CLI) $0.098 (n 12). Unclear.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
- Strict pass rate by effort on eight hard tasks100% (16/16)n 16100% (16/16)n 16TieStrict pass rate by effort on eight hard tasks: Claude Opus 5.5 100% (16/16) (n 16, 95% interval 81%–100%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
- Total time per call by effort on hard tasks9.72 sn 1613.1 sn 16UnclearTotal time per call by effort on hard tasks: Claude Opus 5.5 9.72 s (n 16, run range 4.8 s–31.4 s); GPT-6.1 Sol (Codex CLI) 13.1 s (n 16, run range 8.5 s–61.6 s). Unclear.
- Output tokens per call by effort on hard tasks (Output tokens)853n 16335n 16UnclearOutput tokens per call by effort on hard tasks (Output tokens): Claude Opus 5.5 853 (n 16); GPT-6.1 Sol (Codex CLI) 335 (n 16). Unclear.
- List-price cost per strict pass by effort (calculation)Calculation$0.029n 16$0.026n 16UnclearList-price cost per strict pass by effort (calculation), calculation: Claude Opus 5.5 $0.029 (n 16); GPT-6.1 Sol (Codex CLI) $0.026 (n 16). Unclear.
How much of an AI bill is thinking? Reasoning tokens by model and effort
- Reasoning share of output tokens per call on hard tasks (calculation)Calculation54.4%n 2457%n 16UnclearReasoning share of output tokens per call on hard tasks (calculation), calculation: Claude Opus 5.5 54.4% (n 24, run range 36%–96%); GPT-6.1 Sol (Codex CLI) 57% (n 16, run range 29%–91%). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation$0.018n 24$0.0041n 16UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Claude Opus 5.5 $0.018 (n 24); GPT-6.1 Sol (Codex CLI) $0.0041 (n 16). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation$0.0080n 24$0.0029n 16UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Claude Opus 5.5 $0.0080 (n 24); GPT-6.1 Sol (Codex CLI) $0.0029 (n 16). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation$0.0074n 24$0.0081n 16UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Claude Opus 5.5 $0.0074 (n 24); GPT-6.1 Sol (Codex CLI) $0.0081 (n 16). Unclear.
- Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation$0.013n 16$0.0023n 16UnclearReasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass), calculation: Claude Opus 5.5 $0.013 (n 16); GPT-6.1 Sol (Codex CLI) $0.0023 (n 16). Unclear.
- Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation$0.029n 16$0.026n 16UnclearReasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass), calculation: Claude Opus 5.5 $0.029 (n 16); GPT-6.1 Sol (Codex CLI) $0.026 (n 16). Unclear.
- Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation54.4%n 2457%n 16UnclearReasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Claude Opus 5.5 54.4% (n 24, run range 36%–96%); GPT-6.1 Sol (Codex CLI) 57% (n 16, run range 29%–91%). Unclear.
- Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation43.6%n 1558.1%n 15UnclearReasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Claude Opus 5.5 43.6% (n 15, run range 0%–93%); GPT-6.1 Sol (Codex CLI) 58.1% (n 15, run range 0%–76%). Unclear.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
- Time to first text: a 250-line answer, six models1.97 sn 43.52 sn 4UnclearTime to first text: a 250-line answer, six models: Claude Opus 5.5 1.97 s (n 4, run range 1.7 s–2.4 s); GPT-6.1 Sol (Codex CLI) 3.52 s (n 4, run range 2.8 s–4.4 s). Unclear.
- Output speed after the first text: visible tokens per second (calculation)Calculation156n 480n 4UnclearOutput speed after the first text: visible tokens per second (calculation), calculation: Claude Opus 5.5 156 (n 4, run range 155–156); GPT-6.1 Sol (Codex CLI) 80 (n 4, run range 72–81). Unclear.
- Output speed in characters per second after the first text (calculation)Calculation347n 4323n 4UnclearOutput speed in characters per second after the first text (calculation), calculation: Claude Opus 5.5 347 (n 4, run range 345–349); GPT-6.1 Sol (Codex CLI) 323 (n 4, run range 291–327). Unclear.
- Time to first text as the prompt grows: 1kCalculation1.51 sn 33.36 sn 3UnclearTime to first text as the prompt grows: 1k, calculation: Claude Opus 5.5 1.51 s (n 3, run range 1.5 s–2 s); GPT-6.1 Sol (Codex CLI) 3.36 s (n 3, run range 3.4 s–4.8 s). Unclear.
- Time to first text as the prompt grows: 16kCalculation1.74 sn 34.02 sn 3UnclearTime to first text as the prompt grows: 16k, calculation: Claude Opus 5.5 1.74 s (n 3, run range 1.7 s–3 s); GPT-6.1 Sol (Codex CLI) 4.02 s (n 3, run range 3.3 s–4.3 s). Unclear.
- Time to first text as the prompt grows: 64kCalculation1.79 sn 33.93 sn 3UnclearTime to first text as the prompt grows: 64k, calculation: Claude Opus 5.5 1.79 s (n 3, run range 1.7 s–3.7 s); GPT-6.1 Sol (Codex CLI) 3.93 s (n 3, run range 3.4 s–4.4 s). Unclear.
- Total time per call by prompt size (1k prompt)1.83 sn 33.43 sn 3UnclearTotal time per call by prompt size (1k prompt): Claude Opus 5.5 1.83 s (n 3, run range 1.8 s–2.4 s); GPT-6.1 Sol (Codex CLI) 3.43 s (n 3, run range 3.4 s–4.9 s). Unclear.
- Total time per call by prompt size (16k prompt)2.36 sn 34.14 sn 3UnclearTotal time per call by prompt size (16k prompt): Claude Opus 5.5 2.36 s (n 3, run range 2.1 s–3.4 s); GPT-6.1 Sol (Codex CLI) 4.14 s (n 3, run range 4 s–4.7 s). Unclear.
- Total time per call by prompt size (64k prompt)2.35 sn 33.96 sn 3UnclearTotal time per call by prompt size (64k prompt): Claude Opus 5.5 2.35 s (n 3, run range 2.3 s–4.3 s); GPT-6.1 Sol (Codex CLI) 3.96 s (n 3, run range 3.5 s–4.4 s). Unclear.
- Exact lookup answers at the 1k, 16k and 64k prompt-size targets56% (5/9)n 9100% (9/9)n 9TieExact lookup answers at the 1k, 16k and 64k prompt-size targets: Claude Opus 5.5 56% (5/9) (n 9, 95% interval 27%–81%); GPT-6.1 Sol (Codex CLI) 100% (9/9) (n 9, 95% interval 70%–100%). Tie.
Calculation: at these rates, about 13 runs per side would separate them.
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
- Pass rate on 4 harder tasks (Strict pass)42% (5/12)n 1269% (11/16)n 16TiePass rate on 4 harder tasks (Strict pass): Claude Opus 5.5 42% (5/12) (n 12, 95% interval 19%–68%); GPT-6.1 Sol (Codex CLI) 69% (11/16) (n 16, 95% interval 44%–86%). Tie.
Calculation: at these rates, about 49 runs per side would separate them.
- Pass rate on 4 harder tasks (Lenient (format misses counted))50% (6/12)n 1269% (11/16)n 16TiePass rate on 4 harder tasks (Lenient (format misses counted)): Claude Opus 5.5 50% (6/12) (n 12, 95% interval 25%–75%); GPT-6.1 Sol (Codex CLI) 69% (11/16) (n 16, 95% interval 44%–86%). Tie.
Calculation: at these rates, about 104 runs per side would separate them.
- Calls that tried a tool although tools were off42% (5/12)n 120% (0/16)n 16UnclearCalls that tried a tool although tools were off: Claude Opus 5.5 42% (5/12) (n 12, 95% interval 19%–68%); GPT-6.1 Sol (Codex CLI) 0% (0/16) (n 16, 95% interval 0%–19%). Unclear.
- Strict pass rate by task: 10x10 nonogram100% (3/3)n 375% (3/4)n 4TieStrict pass rate by task: 10x10 nonogram: Claude Opus 5.5 100% (3/3) (n 3, 95% interval 44%–100%); GPT-6.1 Sol (Codex CLI) 75% (3/4) (n 4, 95% interval 30%–95%). Tie.
Calculation: at these rates, about 27 runs per side would separate them.
- Strict pass rate by task: Sudoku, 22 givens0% (0/3)n 325% (1/4)n 4TieStrict pass rate by task: Sudoku, 22 givens: Claude Opus 5.5 0% (0/3) (n 3, 95% interval 0%–56%); GPT-6.1 Sol (Codex CLI) 25% (1/4) (n 4, 95% interval 4.6%–70%). Tie.
Calculation: at these rates, about 26 runs per side would separate them.
- Strict pass rate by task: 6x6 Skyscrapers33% (1/3)n 3100% (4/4)n 4TieStrict pass rate by task: 6x6 Skyscrapers: Claude Opus 5.5 33% (1/3) (n 3, 95% interval 6.2%–79%); GPT-6.1 Sol (Codex CLI) 100% (4/4) (n 4, 95% interval 51%–100%). Tie.
Calculation: at these rates, about 7 runs per side would separate them.
- Strict pass rate by task: Seeded shuffle output33% (1/3)n 375% (3/4)n 4TieStrict pass rate by task: Seeded shuffle output: Claude Opus 5.5 33% (1/3) (n 3, 95% interval 6.2%–79%); GPT-6.1 Sol (Codex CLI) 75% (3/4) (n 4, 95% interval 30%–95%). Tie.
Calculation: at these rates, about 21 runs per side would separate them.
- Total time per call on harder tasks80.3 sn 9120.2 sn 13UnclearTotal time per call on harder tasks: Claude Opus 5.5 80.3 s (n 9, run range 3.8 s–280 s); GPT-6.1 Sol (Codex CLI) 120.2 s (n 13, run range 46.2 s–273 s). Unclear.
- Output tokens per call on harder tasks (Output tokens)8,420n 94,994n 13UnclearOutput tokens per call on harder tasks (Output tokens): Claude Opus 5.5 8,420 (n 9, run range 279–40,044); GPT-6.1 Sol (Codex CLI) 4,994 (n 13, run range 2,099–13,413). Unclear.
- List-price cost per strict pass on harder tasks (calculation)Calculation$0.59n 12$0.083n 16UnclearList-price cost per strict pass on harder tasks (calculation), calculation: Claude Opus 5.5 $0.59 (n 12); GPT-6.1 Sol (Codex CLI) $0.083 (n 16). Unclear.
| Metric | Claude Opus 5.5 | GPT-6.1 Sol (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 100% (15/15)Claude Code · effort high · five short validated tasks | 100% (15/15)Codex CLI · effort high · five short validated tasks | 15 | 95% CI: 80%–100% vs 80%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 2.71 sClaude Code · effort high · five short validated tasks | 5.60 sCodex CLI · effort high · five short validated tasks | 15 | range: 2.5 s–11.8 s vs 4.1 s–19.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.45 s to 11.8 s; GPT-6.1 Sol (Codex CLI) 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 2.04 sClaude Code · effort high · five short validated tasks | 5.32 sCodex CLI · effort high · five short validated tasks | 15 | range: 1.4 s–9.9 s vs 3.6 s–16.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.40 s to 9.94 s; GPT-6.1 Sol (Codex CLI) 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 1,463Claude Code · effort high · five short validated tasks | 6,716Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Other input) | 619Claude Code · effort high · five short validated tasks | 5,406Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 78Claude Code · effort high · five short validated tasks | 42Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per call (calculation)Calculation | $0.0069Claude Code · effort high · five short validated tasks | $0.010Codex CLI · effort high · five short validated tasks | 15 | range: $0.0059–$0.027 vs $0.0066–$0.028 | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 $0.0059 to $0.027; GPT-6.1 Sol (Codex CLI) $0.0066 to $0.028); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.010Claude Code · effort high · five short validated tasks | $0.013Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.010 vs $0.013) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Pass rate on eight hard tasks (Strict pass) | 100% (24/24)Claude Code · effort high · eight hard validated tasks | 100% (16/16)Codex CLI · effort high · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 100% (24/24)Claude Code · effort high · eight hard validated tasks | 100% (16/16)Codex CLI · effort high · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Total time per call on hard tasks (separate batches) | 11.0 sClaude Code · effort high · eight hard validated tasks | 18.1 sCodex CLI · effort high · eight hard validated tasks | 24 / 16 | range: 3.6 s–63 s vs 11.7 s–92.2 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 3.63 s to 63.0 s; GPT-6.1 Sol (Codex CLI) 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Time to first useful output on hard tasks | 7.13 sClaude Code · effort high · eight hard validated tasks | 12.7 sCodex CLI · effort high · eight hard validated tasks | 24 / 16 | range: 2.2 s–56.2 s vs 8.9 s–75.9 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.15 s to 56.2 s; GPT-6.1 Sol (Codex CLI) 8.93 s to 75.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call on hard tasks (Output tokens) | 1,052Claude Code · effort high · eight hard validated tasks | 436Codex CLI · effort high · eight hard validated tasks | 24 / 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass on hard tasks (calculation)Calculation | $0.033Claude Code · effort high · eight hard validated tasks | $0.015Codex CLI · effort high · eight hard validated tasks | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.033 vs $0.015, 2.2x) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Coding sessions that passed every hidden check | 100% (12/12)Claude Code · six small repository tasks with hidden tests | 100% (12/12)Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Time per coding session | 56.9 sClaude Code · six small repository tasks with hidden tests | 113.4 sCodex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | range: 29.8 s–186 s vs 78.5 s–222 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 29.8 s to 185.8 s; GPT-6.1 Sol (Codex CLI) 78.5 s to 221.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Tool calls per coding session | 7.5Claude Code · six small repository tasks with hidden tests | 12.5Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | range: 5–14 vs 8–18 | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| List-price cost per passing coding session (calculation)Calculation | $0.22Claude Code · six small repository tasks with hidden tests | $0.098Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.22 vs $0.098, 2.3x) is not tested against run-to-run variation. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Strict pass rate by effort on eight hard tasks | 100% (16/16)Claude Code · effort medium · eight hard validated tasks, effort ladder | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks, effort ladder | 16 | 95% CI: 81%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 81% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Total time per call by effort on hard tasks | 9.72 sClaude Code · effort medium · eight hard validated tasks, effort ladder | 13.1 sCodex CLI · effort medium · eight hard validated tasks, effort ladder | 16 | range: 4.8 s–31.4 s vs 8.5 s–61.6 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 4.78 s to 31.4 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call by effort on hard tasks (Output tokens) | 853Claude Code · effort medium · eight hard validated tasks, effort ladder | 335Codex CLI · effort medium · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass by effort (calculation)Calculation | $0.029Claude Code · effort medium · eight hard validated tasks, effort ladder | $0.026Codex CLI · effort medium · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.029 vs $0.026) is not tested against run-to-run variation. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Reasoning share of output tokens per call on hard tasks (calculation)Calculation | 54.4%Claude Code · effort high | 57%Codex CLI · effort high | 24 / 16 | range: 36%–96% vs 29%–91% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation | $0.018Claude Code · effort high | $0.0041Codex CLI · effort high | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation | $0.0080Claude Code · effort high | $0.0029Codex CLI · effort high | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation | $0.0074Claude Code · effort high | $0.0081Codex CLI · effort high | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation | $0.013Claude Code · effort medium | $0.0023Codex CLI · effort medium | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation | $0.029Claude Code · effort medium | $0.026Codex CLI · effort medium | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation | 54.4%Claude Code · effort high | 57%Codex CLI · effort high | 24 / 16 | range: 36%–96% vs 29%–91% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation | 43.6%Claude Code · effort high | 58.1%Codex CLI · effort high | 15 | range: 0%–93% vs 0%–76% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Time to first text: a 250-line answer, six models | 1.97 sClaude Code | 3.52 sCodex CLI · effort low | 4 | range: 1.7 s–2.4 s vs 2.8 s–4.4 s | Unclear | Only 4 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.70 s to 2.35 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s), but 4 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 156Claude Code | 80Codex CLI · effort low | 4 | range: 155–156 vs 72–81 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 347Claude Code | 323Codex CLI · effort low | 4 | range: 345–349 vs 291–327 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 1kCalculation | 1.51 sClaude Code | 3.36 sCodex CLI · effort low | 3 | range: 1.5 s–2 s vs 3.4 s–4.8 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.46 s to 2.01 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 16kCalculation | 1.74 sClaude Code | 4.02 sCodex CLI · effort low | 3 | range: 1.7 s–3 s vs 3.3 s–4.3 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.70 s to 2.97 s; GPT-6.1 Sol (Codex CLI) 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 64kCalculation | 1.79 sClaude Code | 3.93 sCodex CLI · effort low | 3 | range: 1.7 s–3.7 s vs 3.4 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.72 s to 3.72 s; GPT-6.1 Sol (Codex CLI) 3.42 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (1k prompt) | 1.83 sClaude Code | 3.43 sCodex CLI · effort low | 3 | range: 1.8 s–2.4 s vs 3.4 s–4.9 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.82 s to 2.41 s; GPT-6.1 Sol (Codex CLI) 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (16k prompt) | 2.36 sClaude Code | 4.14 sCodex CLI · effort low | 3 | range: 2.1 s–3.4 s vs 4 s–4.7 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 2.11 s to 3.40 s; GPT-6.1 Sol (Codex CLI) 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (64k prompt) | 2.35 sClaude Code | 3.96 sCodex CLI · effort low | 3 | range: 2.3 s–4.3 s vs 3.5 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.26 s to 4.29 s; GPT-6.1 Sol (Codex CLI) 3.47 s to 4.44 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Exact lookup answers at the 1k, 16k and 64k prompt-size targets | 56% (5/9)Claude Code | 100% (9/9)Codex CLI · effort low | 9 | 95% CI: 27%–81% vs 70%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 27% to 81%; GPT-6.1 Sol (Codex CLI) 70% to 100%), so this sample cannot separate them. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Pass rate on 4 harder tasks (Strict pass) | 42% (5/12)Claude Code | 69% (11/16)Codex CLI · effort medium | 12 / 16 | 95% CI: 19%–68% vs 44%–86% | Tie | The 95% intervals overlap (Claude Opus 5.5 19% to 68%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Pass rate on 4 harder tasks (Lenient (format misses counted)) | 50% (6/12)Claude Code | 69% (11/16)Codex CLI · effort medium | 12 / 16 | 95% CI: 25%–75% vs 44%–86% | Tie | The 95% intervals overlap (Claude Opus 5.5 25% to 75%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Calls that tried a tool although tools were off | 42% (5/12)Claude Code | 0% (0/16)Codex CLI · effort medium | 12 / 16 | 95% CI: 19%–68% vs 0%–19% | Unclear | More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 10x10 nonogram | 100% (3/3)Claude Code | 75% (3/4)Codex CLI · effort medium | 3 / 4 | 95% CI: 44%–100% vs 30%–95% | Tie | The 95% intervals overlap (Claude Opus 5.5 44% to 100%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Sudoku, 22 givens | 0% (0/3)Claude Code | 25% (1/4)Codex CLI · effort medium | 3 / 4 | 95% CI: 0%–56% vs 4.6%–70% | Tie | The 95% intervals overlap (Claude Opus 5.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 5% to 70%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 6x6 Skyscrapers | 33% (1/3)Claude Code | 100% (4/4)Codex CLI · effort medium | 3 / 4 | 95% CI: 6.2%–79% vs 51%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 6% to 79%; GPT-6.1 Sol (Codex CLI) 51% to 100%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Seeded shuffle output | 33% (1/3)Claude Code | 75% (3/4)Codex CLI · effort medium | 3 / 4 | 95% CI: 6.2%–79% vs 30%–95% | Tie | The 95% intervals overlap (Claude Opus 5.5 6% to 79%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Total time per call on harder tasks | 80.3 sClaude Code | 120.2 sCodex CLI · effort medium | 9 / 13 | range: 3.8 s–280 s vs 46.2 s–273 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 3.82 s to 279.5 s; GPT-6.1 Sol (Codex CLI) 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Output tokens per call on harder tasks (Output tokens) | 8,420Claude Code | 4,994Codex CLI · effort medium | 9 / 13 | range: 279–40,044 vs 2,099–13,413 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| List-price cost per strict pass on harder tasks (calculation)Calculation | $0.59Claude Code | $0.083Codex CLI · effort medium | 12 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.59 vs $0.083, 7.2x) is not tested against run-to-run variation. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
50 rows from 7 studies. No row separates them: 12 ties, 38 unclear.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick Claude Opus 5.5
- List-price cost per call (calculation): $0.0069 vs $0.010. A list-price calculation, not a measured difference. Calculation
- Time to first text as the prompt grows: 1k: 1.51 s vs 3.36 s. A list-price calculation, not a measured difference. Calculation
- Time to first text as the prompt grows: 16k: 1.74 s vs 4.02 s. A list-price calculation, not a measured difference. Calculation
- Time to first text as the prompt grows: 64k: 1.79 s vs 3.93 s. A list-price calculation, not a measured difference. Calculation
When to pick GPT-6.1 Sol (Codex CLI)
- List-price cost per strict pass on hard tasks (calculation): $0.015 vs $0.033. A list-price calculation, not a measured difference. Calculation
- List-price cost per passing coding session (calculation): $0.098 vs $0.22. A list-price calculation, not a measured difference. Calculation
- List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0041 vs $0.018. A list-price calculation, not a measured difference. Calculation
- List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)): $0.0029 vs $0.0080. A list-price calculation, not a measured difference. Calculation
- Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass): $0.0023 vs $0.013. A list-price calculation, not a measured difference. Calculation
- List-price cost per strict pass on harder tasks (calculation): $0.083 vs $0.59. A list-price calculation, not a measured difference. Calculation
Side by side
The study charts, showing only these two. Open a study for every configuration.
| Item | Pass rate | 95% interval | n |
|---|---|---|---|
| Claude Opus 5.5 (high) · Claude Code | 100% | 80%–100% | 15 |
| GPT-6.1 Sol (high) · Codex CLI | 100% | 80%–100% | 15 |
2 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 15 per row
Every call counts; failures and timeouts are non-passes
Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.
Source: Provider head-to-head: Claude Code models vs Codex efforts
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 (high) · Claude Code | 2.7 s | 2.5 s–11.8 s | 15 |
| GPT-6.1 Sol (high) · Codex CLI | 5.6 s | 4.1 s–19.5 s | 15 |
2 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 5.6 s (range 4.1 s–19.5 s, n 15). Fastest Claude Opus 5.5 (high) · Claude Code 2.7 s (range 2.5 s–11.8 s, n 15). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 15 per row
Median per configuration; whiskers = fastest and slowest call
One host, one network, one day. Whiskers are a range, not a confidence interval.
Source: Provider head-to-head: Claude Code models vs Codex efforts
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first useful output | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 (high) · Claude Code | 2 s | 1.4 s–9.9 s | 15 |
| GPT-6.1 Sol (high) · Codex CLI | 5.3 s | 3.6 s–16.4 s | 15 |
2 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 5.3 s (range 3.6 s–16.4 s, n 15). Fastest Claude Opus 5.5 (high) · Claude Code 2 s (range 1.4 s–9.9 s, n 15). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 15 per row
Median per configuration; whiskers = fastest and slowest call
One host, one network, one day. Whiskers are a range, not a confidence interval.
Source: Provider head-to-head: Claude Code models vs Codex efforts
- Cache read
- Other input
Totals are the sum of the parts shown. Shares are calculated from the same values.
| Item | Cache read | Other input | n |
|---|---|---|---|
| Claude Opus 5.5 (high) · Claude Code | 1,463 | 619 | 15 |
| GPT-6.1 Sol (high) · Codex CLI | 6,716 | 5,406 | 15 |
2 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (high) · Codex CLI 6,716 (n 15). Lowest Claude Opus 5.5 (high) · Claude Code 1,463 (n 15). Other input: highest GPT-6.1 Sol (high) · Codex CLI 5,406 (n 15). Lowest Claude Opus 5.5 (high) · Claude Code 619 (n 15).
Notesn = 15 per row
Mean per call, split into prompt-cache reads and other input
The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.
Source: Provider head-to-head: Claude Code models vs Codex efforts
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Opus 5.5 (high) · Claude Code | 78 | 34 | 15 |
| GPT-6.1 Sol (high) · Codex CLI | 42 | 21 | 15 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Opus 5.5 (high) · Claude Code 78 (n 15). Lowest GPT-6.1 Sol (high) · Codex CLI 42 (n 15). Reasoning tokens: highest Claude Opus 5.5 (high) · Claude Code 34 (n 15). Lowest GPT-6.1 Sol (high) · Codex CLI 21 (n 15).
Notesn = 15 per row
Median per configuration; reasoning tokens where the CLI reports them
Vendors count reasoning differently; compare within a vendor first. A median of 0 can mean the CLI did not report reasoning for most calls.
Source: Provider head-to-head: Claude Code models vs Codex efforts
| Item | Cost per call | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 (high) · Claude Code | $0.0069 | $0.0059–$0.027 | 15 |
| GPT-6.1 Sol (high) · Codex CLI | $0.01 | $0.0066–$0.028 | 15 |
List-price calculation, not a run. 2 rows. Highest GPT-6.1 Sol (high) · Codex CLI $0.01 (range $0.0066–$0.028, n 15). Lowest Claude Opus 5.5 (high) · Claude Code $0.0069 (range $0.0059–$0.027, n 15). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 15 per row
Reported tokens × list price; the calls ran on subscriptions
Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.
Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
| Item | Cost per pass | n |
|---|---|---|
| Claude Opus 5.5 (high) · Claude Code | $0.01 | 15 |
| GPT-6.1 Sol (high) · Codex CLI | $0.013 | 15 |
List-price calculation, not a run. 2 rows. Highest GPT-6.1 Sol (high) · Codex CLI $0.013 (n 15). Lowest Claude Opus 5.5 (high) · Claude Code $0.01 (n 15).
Notesn = 15 per row
All calls in a configuration, failures included, divided by its passes
Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.
Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Strict pass
- Lenient (format misses counted)
| Item | Strict pass | Lenient (format misses counted) | 95% interval | n |
|---|---|---|---|---|
| Claude Opus 5.5 · Claude Code | 100% | 100% | Strict pass: 86%–100%; Lenient (format misses counted): 86%–100% | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 100% | 100% | Strict pass: 81%–100%; Lenient (format misses counted): 81%–100% | 16 |
2 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: all at 100%. Lenient (format misses counted): all at 100%.
NotesWhiskers: 95% Wilson intervaln 16–24 per row
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call on hard tasks (separate batches) | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 (high) · Claude Code | 11 s | 3.6 s–63 s | 24 |
| GPT-6.1 Sol (high) · Codex CLI | 18.1 s | 11.7 s–92.2 s | 16 |
2 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 18.1 s (range 11.7 s–92.2 s, n 16). Fastest Claude Opus 5.5 (high) · Claude Code 11 s (range 3.6 s–63 s, n 24). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first useful output on hard tasks | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 (high) · Claude Code | 7.1 s | 2.2 s–56.2 s | 24 |
| GPT-6.1 Sol (high) · Codex CLI | 12.7 s | 8.9 s–75.9 s | 16 |
2 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 12.7 s (range 8.9 s–75.9 s, n 16). Fastest Claude Opus 5.5 (high) · Claude Code 7.1 s (range 2.2 s–56.2 s, n 24). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Opus 5.5 (high) · Claude Code | 1,052 | 614 | 24 |
| GPT-6.1 Sol (high) · Codex CLI | 436 | 225 | 16 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Opus 5.5 (high) · Claude Code 1,052 (n 24). Lowest GPT-6.1 Sol (high) · Codex CLI 436 (n 16). Reasoning tokens: highest Claude Opus 5.5 (high) · Claude Code 614 (n 24). Lowest GPT-6.1 Sol (high) · Codex CLI 225 (n 16).
Notesn 16–24 per row
Median per configuration; reasoning tokens as the CLI reports them
Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
| Item | Cost per strict pass | n |
|---|---|---|
| GPT-6.1 Sol (high) · Codex CLI | $0.015 | 16 |
| Claude Opus 5.5 (high) · Claude Code | $0.033 | 24 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 (high) · Claude Code $0.033 (n 24). Lowest GPT-6.1 Sol (high) · Codex CLI $0.015 (n 16).
Notesn 16–24 per row
All calls in a configuration, failures and format misses included, divided by its strict passes
Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.
Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
| Item | Passed every hidden check | 95% interval | n |
|---|---|---|---|
| Claude Opus 5.5 · Claude Code | 100% | 76%–100% | 12 |
| GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI | 100% | 76%–100% | 12 |
2 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 12 per row
A pass needs every hidden check · 12 sessions per agent · 95% Wilson intervals
6 tasks × 2 repetitions per agent. All 36 sessions passed, so the chart sits at its ceiling: this task set cannot rank the agents on quality. Before the first session, every base repository failed its hidden checks and every reference patch passed them.
Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Wall time per session | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 · Claude Code | 56.9 s | 29.8 s–186 s | 12 |
| GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI | 113 s | 78.5 s–222 s | 12 |
2 rows. Slowest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI 113 s (range 78.5 s–222 s, n 12). Fastest Claude Opus 5.5 · Claude Code 56.9 s (range 29.8 s–186 s, n 12). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 12 per row
Median wall time; whiskers = fastest and slowest of 12 sessions (not an interval)
CLI process start to exit. One session at a time per lane; the Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts. Codex time includes reading the tester’s notes, writing a work log and trying to commit. A range is not a confidence interval.
Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks
| Item | Tool calls per session | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 · Claude Code | 7.5 | 5–14 | 12 |
| GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI | 12.5 | 8–18 | 12 |
2 rows. Highest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI 12.5 (range 8–18, n 12). Lowest Claude Opus 5.5 · Claude Code 7.5 (range 5–14, n 12). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 12 per row
Median; whiskers = fewest and most of 12 sessions (not an interval)
Claude Code counts its tool calls (Bash, Read, Edit, Write, Glob, Grep). Codex CLI counts shell commands and file changes; it has no separate read tool, so it reads files with shell commands. Turns are not compared: Codex reports one turn per run.
Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks
| Item | List-price cost per pass | n |
|---|---|---|
| Claude Opus 5.5 · Claude Code | $0.22 | 12 |
| GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI | $0.098 | 12 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 · Claude Code $0.22 (n 12). Lowest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI $0.098 (n 12).
Notesn = 12 per row
Reported tokens of all 12 sessions × list price, divided by the passes
Calculation, not a bill: both CLIs ran on flat subscriptions. Claude cache writes are priced at 2× input, as in the other studies (Claude Code’s own estimate gives the same totals); Codex cached input at its cache-read price. Codex input includes its own system prompt and, here, the tester’s AGENTS.md.
Sources: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
| Item | Strict pass | 95% interval | n |
|---|---|---|---|
| Claude Opus 5.5 (low) · Claude Code | 100% | 81%–100% | 16 |
| GPT-6.1 Sol (low) · Codex CLI | 100% | 81%–100% | 16 |
2 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 16 per row
Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.
Source: Effort ladder: the hard task set at each effort level
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call by effort on hard tasks | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 (medium) · Claude Code | 9.7 s | 4.8 s–31.4 s | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | 13.1 s | 8.5 s–61.6 s | 16 |
2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 13.1 s (range 8.5 s–61.6 s, n 16). Fastest Claude Opus 5.5 (medium) · Claude Code 9.7 s (range 4.8 s–31.4 s, n 16). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 16 per row
Median per configuration; whiskers = fastest and slowest call
Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.
Source: Effort ladder: the hard task set at each effort level
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Opus 5.5 (medium) · Claude Code | 853 | 518 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | 335 | 150 | 16 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Opus 5.5 (medium) · Claude Code 853 (n 16). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16). Reasoning tokens: highest Claude Opus 5.5 (medium) · Claude Code 518 (n 16). Lowest GPT-6.1 Sol (medium) · Codex CLI 150 (n 16).
Notesn = 16 per row
Median per configuration; reasoning tokens as the CLI reports them
Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean "not reported". Their content is never captured. More tokens is not better or worse by itself.
Source: Effort ladder: the hard task set at each effort level
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Opus 5.5 (medium) · Claude Code | $0.029 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.026 | 16 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 (medium) · Claude Code $0.029 (n 16). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.026 (n 16).
Notesn = 16 per row
All calls in a configuration divided by its strict passes
Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.
Sources: Effort ladder: the hard task set at each effort level, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Reasoning (output tokens)
- Remaining output (visible-answer estimate)
- Input (prompt, cache priced)
Totals are the sum of the parts shown. Shares are calculated from the same values.
| Item | Reasoning (output tokens) | Remaining output (visible-answer estimate) | Input (prompt, cache priced) | n |
|---|---|---|---|---|
| Claude Opus 5.5 (high) · Claude Code | $0.018 | $0.008 | $0.0074 | 24 |
| GPT-6.1 Sol (high) · Codex CLI | $0.0041 | $0.0029 | $0.0081 | 16 |
List-price calculation, not a run. 2 rows, 3 series: Reasoning (output tokens), Remaining output (visible-answer estimate), Input (prompt, cache priced). Reasoning (output tokens): highest Claude Opus 5.5 (high) · Claude Code $0.018 (n 24). Lowest GPT-6.1 Sol (high) · Codex CLI $0.0041 (n 16). Remaining output (visible-answer estimate): highest Claude Opus 5.5 (high) · Claude Code $0.008 (n 24). Lowest GPT-6.1 Sol (high) · Codex CLI $0.0029 (n 16).
Notesn 16–24 per row
Mean per call on the hard tasks; the three parts add up to the call
Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
- Reasoning cost per strict pass
- Total cost per strict pass (square)
Gap labels, Total cost per strict pass vs Reasoning cost per strict pass: Total cost per strict pass is x% higher (+) or lower (−) than Reasoning cost per strict pass, calculated from the two values shown (the change counted from Reasoning cost per strict pass’s value).
| Item | Reasoning cost per strict pass | Total cost per strict pass | n |
|---|---|---|---|
| Claude Opus 5.5 (medium) · Claude Code | $0.013 | $0.029 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.0023 | $0.026 | 16 |
List-price calculation, not a run. 2 rows, 2 series: Reasoning cost per strict pass, Total cost per strict pass. Reasoning cost per strict pass: highest Claude Opus 5.5 (medium) · Claude Code $0.013 (n 16). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.0023 (n 16). Total cost per strict pass: highest Claude Opus 5.5 (medium) · Claude Code $0.029 (n 16). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.026 (n 16).
Notesn = 16 per row
List price ÷ strict passes; every cell is 8 tasks × 2 repetitions
Calculation, not a bill: reported tokens × list price, divided by the cell's strict passes; the calls ran on flat subscriptions. The effort-ladder cells: new calls plus reference cells reused from the hard head-to-head (Claude repetitions 1-2 only). "Default" means the effort flag was not passed. The total is the same value as the effort-ladder cost-per-pass chart. Effort levels are not the same scale across vendors, and the reference cells ran in a different batch and hour.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
- Eight hard tasks
- Five short tasks (square)
Gap labels, Five short tasks vs Eight hard tasks: Five short tasks is x percentage points higher (+) or lower (−) than Eight hard tasks, calculated from the two values shown; lines are the lowest–highest run (not an interval).
| Item | Eight hard tasks | Five short tasks | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Opus 5.5 (high) · Claude Code | 54% | 44% | Eight hard tasks: 36%–96%; Five short tasks: 0%–93% | 24 |
| GPT-6.1 Sol (high) · Codex CLI | 57% | 58% | Eight hard tasks: 29%–91%; Five short tasks: 0%–76% | 16 |
List-price calculation, not a run. 2 rows, 2 series: Eight hard tasks, Five short tasks. Eight hard tasks: highest GPT-6.1 Sol (high) · Codex CLI 57% (range 29%–91%, n 16). Lowest Claude Opus 5.5 (high) · Claude Code 54% (range 36%–96%, n 24). All run ranges overlap. Five short tasks: highest GPT-6.1 Sol (high) · Codex CLI 58% (range 0%–76%, n 15). Lowest Claude Opus 5.5 (high) · Claude Code 44% (range 0%–93%, n 15). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 15–24 per row
Median call per configuration; five short tasks and eight hard tasks
Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.
| Item | Time to first text | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 · Claude Code | 2 s | 1.7 s–2.4 s | 4 |
| GPT-6.1 Sol (low) · Codex CLI | 3.5 s | 2.8 s–4.4 s | 4 |
2 rows. Slowest GPT-6.1 Sol (low) · Codex CLI 3.5 s (range 2.8 s–4.4 s, n 4). Fastest Claude Opus 5.5 · Claude Code 2 s (range 1.7 s–2.4 s, n 4). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = fastest and slowest call
Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.
Source: LLM speed anatomy
| Item | Visible tokens per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 · Claude Code | 156 | 155–156 | 4 |
| GPT-6.1 Sol (low) · Codex CLI | 80 | 72–81 | 4 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 · Claude Code 156 (range 155–156, n 4). Lowest GPT-6.1 Sol (low) · Codex CLI 80 (range 72–81, n 4). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = slowest and fastest call
Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.
Source: LLM speed anatomy
| Item | Characters per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Opus 5.5 · Claude Code | 347 | 345–349 | 4 |
| GPT-6.1 Sol (low) · Codex CLI | 323 | 291–327 | 4 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 · Claude Code 347 (range 345–349, n 4). Lowest GPT-6.1 Sol (low) · Codex CLI 323 (range 291–327, n 4). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call
Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
- Claude Haiku 4.5 · Claude Code
- Claude Sonnet 5.5 · Claude Code
- Claude Opus 5.5 · Claude Code
- GPT-6.1 Sol (low) · Codex CLI
| Prompt-size target (approximate Haiku tokens; calibration calculation) | Claude Haiku 4.5 · Claude Code | Claude Sonnet 5.5 · Claude Code | Claude Opus 5.5 · Claude Code | GPT-6.1 Sol (low) · Codex CLI | Range (lowest–highest run) | n |
|---|---|---|---|---|---|---|
| 1k | 1.9 s | 1.5 s | 1.5 s | 3.4 s | Claude Haiku 4.5 · Claude Code: 1.9 s–2 s; Claude Sonnet 5.5 · Claude Code: 1.2 s–1.7 s; Claude Opus 5.5 · Claude Code: 1.5 s–2 s; GPT-6.1 Sol (low) · Codex CLI: 3.4 s–4.8 s | 3 |
| 16k | 2.3 s | 1.8 s | 1.7 s | 4 s | Claude Haiku 4.5 · Claude Code: 2.2 s–2.5 s; Claude Sonnet 5.5 · Claude Code: 1.6 s–2.1 s; Claude Opus 5.5 · Claude Code: 1.7 s–3 s; GPT-6.1 Sol (low) · Codex CLI: 3.3 s–4.3 s | 3 |
| 64k | 2.8 s | 3.1 s | 1.8 s | 3.9 s | Claude Haiku 4.5 · Claude Code: 2.5 s–2.9 s; Claude Sonnet 5.5 · Claude Code: 1.4 s–3.6 s; Claude Opus 5.5 · Claude Code: 1.7 s–3.7 s; GPT-6.1 Sol (low) · Codex CLI: 3.4 s–4.4 s | 3 |
List-price calculation, not a run. 3 rows, 4 series: Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (low) · Codex CLI. Claude Haiku 4.5 · Claude Code: slowest 64k 2.8 s (range 2.5 s–2.9 s, n 3). Fastest 1k 1.9 s (range 1.9 s–2 s, n 3). Not all run ranges overlap. Claude Sonnet 5.5 · Claude Code: slowest 64k 3.1 s (range 1.4 s–3.6 s, n 3). Fastest 1k 1.5 s (range 1.2 s–1.7 s, n 3). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Median of 3 calls per size; whiskers = fastest and slowest call
Each call used a new ledger seed. Cache-read counts stayed within the short-prompt baseline (see the cache table). This does not identify which tokens were cached. Sizes name the text we send; each model’s reported input tokens are in the table and include the CLI’s own prefix. The size calibration subtracts estimated prefixes from probe input counts; these are calculations, not measured prefix counts for each call. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
1k prompt
16k prompt
64k prompt
One panel per series, all on the same axis; whiskers are the fastest–slowest run (not an interval).
| Item | 1k prompt | 16k prompt | 64k prompt | Range (lowest–highest run) | n |
|---|---|---|---|---|---|
| Claude Opus 5.5 · Claude Code | 1.8 s | 2.4 s | 2.4 s | 1k prompt: 1.8 s–2.4 s; 16k prompt: 2.1 s–3.4 s; 64k prompt: 2.3 s–4.3 s | 3 |
| GPT-6.1 Sol (low) · Codex CLI | 3.4 s | 4.1 s | 4 s | 1k prompt: 3.4 s–4.9 s; 16k prompt: 4 s–4.7 s; 64k prompt: 3.5 s–4.4 s | 3 |
2 rows, 3 series: 1k prompt, 16k prompt, 64k prompt. 1k prompt: slowest GPT-6.1 Sol (low) · Codex CLI 3.4 s (range 3.4 s–4.9 s, n 3). Fastest Claude Opus 5.5 · Claude Code 1.8 s (range 1.8 s–2.4 s, n 3). Not all run ranges overlap. 16k prompt: slowest GPT-6.1 Sol (low) · Codex CLI 4.1 s (range 4 s–4.7 s, n 3). Fastest Claude Opus 5.5 · Claude Code 2.4 s (range 2.1 s–3.4 s, n 3). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Median of 3 calls per bar; whiskers = fastest and slowest call
Whole call: CLI start-up, first text and the one-line answer. Whiskers are a range of calls, not a confidence interval. Each call used a new ledger.
Source: LLM speed anatomy
| Item | Exact answer | 95% interval | n |
|---|---|---|---|
| Claude Opus 5.5 · Claude Code | 56% | 27%–81% | 9 |
| GPT-6.1 Sol (low) · Codex CLI | 100% | 70%–100% | 9 |
2 rows. Highest GPT-6.1 Sol (low) · Codex CLI 100% (95% interval 70%–100%, n 9). Lowest Claude Opus 5.5 · Claude Code 56% (95% interval 27%–81%, n 9). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 9 per row
All sizes together per model; whiskers = 95% Wilson intervals
Whiskers are 95% Wilson intervals. One lookup question per call; a reply with extra words is a format miss, not a pass. With 9 calls per model, a perfect score still has a wide interval.
Source: LLM speed anatomy
- Strict pass
- Lenient (format misses counted)
| Item | Strict pass | Lenient (format misses counted) | 95% interval | n |
|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 69% | 69% | Strict pass: 44%–86%; Lenient (format misses counted): 44%–86% | 16 |
| Claude Opus 5.5 · Claude Code | 42% | 50% | Strict pass: 19%–68%; Lenient (format misses counted): 25%–75% | 12 |
2 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest GPT-6.1 Sol (medium) · Codex CLI 69% (95% interval 44%–86%, n 16). Lowest Claude Opus 5.5 · Claude Code 42% (95% interval 19%–68%, n 12). All intervals overlap. Lenient (format misses counted): highest GPT-6.1 Sol (medium) · Codex CLI 69% (95% interval 44%–86%, n 16). Lowest Claude Opus 5.5 · Claude Code 50% (95% interval 25%–75%, n 12). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 12–16 per row
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row. Counted calls are new calls.
Source: Harder tasks head-to-head
| Item | Tool attempt | 95% interval | n |
|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 0% | 0%–19% | 16 |
| Claude Opus 5.5 · Claude Code | 42% | 19%–68% | 12 |
2 rows. Highest Claude Opus 5.5 · Claude Code 42% (95% interval 19%–68%, n 12). Lowest GPT-6.1 Sol (medium) · Codex CLI 0% (95% interval 0%–19%, n 16). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 12–16 per row
Share of calls whose reply held tool-call markup or whose tool call the CLI could not parse
Every prompt asked for the answer only. Both CLIs disabled tools. The Codex runner adds a developer instruction not to call tools. The Claude Code runner adds no such instruction. The validator decides the score. Tool-call markup is a separate flag; a flagged call can still pass if its whole reply meets the validator. It is a behaviour, not a quality score: more is not better. No call with a tool attempt passed. Whiskers are 95% Wilson intervals. The flag comes from the reply text, which is not published.
Source: Harder tasks head-to-head
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | GPT-6.1 Sol (medium) · Codex CLI | Claude Opus 5.5 · Claude Code | Claude Sonnet 5.5 · Claude Code | Claude Haiku 4.5 · Claude Code | 95% interval | n |
|---|---|---|---|---|---|---|
| 10x10 nonogram | 75% | 100% | 100% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 30%–95%; Claude Opus 5.5 · Claude Code: 44%–100%; Claude Sonnet 5.5 · Claude Code: 51%–100%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
| Sudoku, 22 givens | 25% | 0% | 0% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 4.6%–70%; Claude Opus 5.5 · Claude Code: 0%–56%; Claude Sonnet 5.5 · Claude Code: 0%–49%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
| 6x6 Skyscrapers | 100% | 33% | 0% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 51%–100%; Claude Opus 5.5 · Claude Code: 6.2%–79%; Claude Sonnet 5.5 · Claude Code: 0%–49%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
3 rows, 4 series: GPT-6.1 Sol (medium) · Codex CLI, Claude Opus 5.5 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Haiku 4.5 · Claude Code. GPT-6.1 Sol (medium) · Codex CLI: highest 6x6 Skyscrapers 100% (95% interval 51%–100%, n 4). Lowest Sudoku, 22 givens 25% (95% interval 4.6%–70%, n 4). All intervals overlap. Claude Opus 5.5 · Claude Code: highest 10x10 nonogram 100% (95% interval 44%–100%, n 3). Lowest Sudoku, 22 givens 0% (95% interval 0%–56%, n 3). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 3–4 per row
One bar per configuration and task; each bar rests on only a few calls, so the 95% intervals are wide
Each bar is 3 to 4 calls, so one call moves a bar by a quarter or a third. Whiskers are 95% Wilson intervals; with this few calls they overlap almost everywhere.
Source: Harder tasks head-to-head
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 120 s | 46.2 s–273 s | 13 |
| Claude Opus 5.5 · Claude Code | 80.3 s | 3.8 s–280 s | 9 |
2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 120 s (range 46.2 s–273 s, n 13). Fastest Claude Opus 5.5 · Claude Code 80.3 s (range 3.8 s–280 s, n 9). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 9–13 per row
Median per configuration; whiskers = fastest and slowest call
Median and range over the calls that completed. Completed calls include wrong answers and format misses. Only timeouts and tool-call parse errors are excluded from this run’s timings. Both count as non-passes in the outcomes chart. One Mac, one network, one session. Arena servers shared the Mac during part of the run. Host load was not controlled, so these times cannot isolate model speed. Whiskers are a range, not a confidence interval. Times include the CLI start-up and the CLI’s own system prompt. Highlighted: configurations that passed every call.
Source: Harder tasks head-to-head
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | Range (lowest–highest run) | n |
|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 4,994 | 4,971 | Output tokens: 2,099–13,413; Reasoning tokens: 2,070–13,372 | 13 |
| Claude Opus 5.5 · Claude Code | 8,420 | 8,352 | Output tokens: 279–40,044; Reasoning tokens: 21–9,897 | 9 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Opus 5.5 · Claude Code 8,420 (range 279–40,044, n 9). Lowest GPT-6.1 Sol (medium) · Codex CLI 4,994 (range 2,099–13,413, n 13). All run ranges overlap. Reasoning tokens: highest Claude Opus 5.5 · Claude Code 8,352 (range 21–9,897, n 9). Lowest GPT-6.1 Sol (medium) · Codex CLI 4,971 (range 2,070–13,372, n 13). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 9–13 per row
Median per configuration; reasoning tokens as the CLI reports them
Medians and minimum-to-maximum token ranges cover completed calls only. Ranges are not confidence intervals. The chart omits unknown reasoning counts. Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. Claude Code used an output-token cap setting of 16,000. Some reported totals exceeded it. Codex CLI had no cap. More tokens is not better or worse by itself.
Source: Harder tasks head-to-head
| Item | Cost per strict pass | n |
|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | $0.083 | 16 |
| Claude Opus 5.5 · Claude Code | $0.59 | 12 |
List-price calculation, not a run. 2 rows. Highest Claude Opus 5.5 · Claude Code $0.59 (n 12). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.083 (n 16).
Notesn 12–16 per row
All calls in a configuration, failures and format misses included, divided by its strict passes
Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Timeout calls report no tokens and are not priced. Unpriced calls: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16. These cells show a lower bound. Assume each unpriced call cost its cell’s median priced call. This sensitivity calculation gives GPT-6.1 Sol (medium) $0.100, Opus 5.5 $0.633 and Sonnet 5.5 $0.303. Opus 5.5 figures are provisional: its cache-read price is under re-check. Highlights mark the observed frontier of these lower-bound costs. Unknown timeout costs can change it; this is not a cost ranking. Claude Haiku 4.5 · Claude Code had no strict pass, so it has no cost per pass.
Sources: Harder tasks head-to-head, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Which is better, Claude Opus 5.5 or GPT-6.1 Sol (Codex CLI)?
- Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) share 31 measured metrics and 19 list-price calculations from 7 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 12 ties and 38 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).
- How were Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) measured?
- They share 31 measured metrics and 19 list-price calculations from 7 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks; Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks; How much of an AI bill is thinking? Reasoning tokens by model and effort; Where the seconds go: first text, output speed and prompt size for 6 LLMs; GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Every row names its configuration, its sample size and its interval or range.
- How do Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) compare on pass rate on five validated tasks?
- Claude Opus 5.5: 100% (15/15) (Claude Code · effort high · five short validated tasks; n = 15; 95% interval 80% to 100%). GPT-6.1 Sol (Codex CLI): 100% (15/15) (Codex CLI · effort high · five short validated tasks; n = 15; 95% interval 80% to 100%). The 95% intervals overlap (Claude Opus 5.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.
- How do Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) compare on total time per call?
- Claude Opus 5.5: 2.71 s (Claude Code · effort high · five short validated tasks; n = 15; run range 2.5 s to 11.8 s). GPT-6.1 Sol (Codex CLI): 5.60 s (Codex CLI · effort high · five short validated tasks; n = 15; run range 4.1 s to 19.5 s). The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.45 s to 11.8 s; GPT-6.1 Sol (Codex CLI) 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
- How do Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) compare on time to first useful output?
- Claude Opus 5.5: 2.04 s (Claude Code · effort high · five short validated tasks; n = 15; run range 1.4 s to 9.9 s). GPT-6.1 Sol (Codex CLI): 5.32 s (Codex CLI · effort high · five short validated tasks; n = 15; run range 3.6 s to 16.4 s). The run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.40 s to 9.94 s; GPT-6.1 Sol (Codex CLI) 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.
- How do Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) compare on pass rate on eight hard tasks (Strict pass)?
- Claude Opus 5.5: 100% (24/24) (Claude Code · effort high · eight hard validated tasks; n = 24; 95% interval 86% to 100%). GPT-6.1 Sol (Codex CLI): 100% (16/16) (Codex CLI · effort high · eight hard validated tasks; n = 16; 95% interval 81% to 100%). The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.
- How do Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) compare on pass rate on eight hard tasks (Lenient (format misses counted))?
- Claude Opus 5.5: 100% (24/24) (Claude Code · effort high · eight hard validated tasks; n = 24; 95% interval 86% to 100%). GPT-6.1 Sol (Codex CLI): 100% (16/16) (Codex CLI · effort high · eight hard validated tasks; n = 16; 95% interval 81% to 100%). The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.
The studies behind this page
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.
Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
36 graded sessions: Claude Code with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol. All passed every hidden test; time, tool calls and diffs differ.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
176 calls on 8 hard tasks at low, medium, high and default effort: Sonnet, Opus and GPT-6.1 Sol. Pass rate with 95% intervals, time, tokens, cost per pass.
How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.