- Benchmarks
- Compare
- Claude vs GPT
16 measured pairs · updated October 7, 2026
Claude vs GPT, on the tasks we measured.
Read every recorded Anthropic and OpenAI head-to-head in this dataset. A model, its route and a coding CLI are distinct systems. These runs do not measure the ChatGPT app.
Each row keeps its study, task set, configuration and sample size. Intervals and ranges stay labelled; ties and unclear results stay unchanged. List-price calculations are not vendor bills. There is no combined rank across studies.
Browse all model, CLI and router comparisons
Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI): what the measurements say
70 comparison rows from 10 studies: 4 rows favour Sonnet 5.5, 1 favour GPT-6.1 Sol (Codex CLI), 65 are ties or unclear. Cost rows are calculations.
Transcript
- Comparison · 70 rows · 10 studies. Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI). A winner only where the 95% intervals or run ranges do not overlap.
- 70 comparison rows from 10 studies: Sonnet 5.5 ahead on 4, GPT-6.1 Sol (Codex CLI) ahead on 1. The rest do not separate them. Rows where Sonnet 5.5 is ahead: 4 (of 70). Rows where GPT-6.1 Sol (Codex CLI) is ahead: 1 (of 70). Ties or unclear: 65 (22 ties · 43 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
- Coding agents, hidden tests: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 4 rows separates them. Table: Coding agents, hidden tests · Claude Code vs Codex CLI · n = 12 per side. Source study: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks. Rows shown: Coding sessions that passed every hidden check; Time per coding session; Tool calls per coding session; List-price cost per passing coding session (calculation). Recorded settings: Claude Code · six small repository tasks with hidden tests; Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests. Includes a calculation, not a bill or a new run. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
- Caching sessions: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 2 of 9 rows separate them. Table: Caching sessions · 5 of 9 rows · Claude Code vs Codex CLI · n = 10 per side. Source study: Prompt caching and run-to-run consistency in Claude Code and Codex CLI. Rows shown: Same prompt, 10 times: strict pass rate (Exact number); Same prompt, 10 times: strict pass rate (JSON object); Same prompt, 10 times: strict pass rate (Code fix); Same prompt, 10 times: time per call (Exact number); Same prompt, 10 times: time per call (Code fix). Recorded settings: Claude Code · same prompt repeated 10 times; Codex CLI · effort medium · same prompt repeated 10 times. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- Instructions vs JSON schema: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 8 rows separates them. Table: Instructions vs JSON schema · 5 of 8 rows · Claude Code vs Codex CLI · n = 12 per side. Source study: Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI. Rows shown: Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON); What each call produced: strict pass, format miss, wrong values or error (Strict pass); What each call produced: strict pass, format miss, wrong values or error (Format miss); What each call produced: strict pass, format miss, wrong values or error (Wrong values); Time per call, instructions vs schema mode. Recorded settings: Claude Code · instructions; Codex CLI · effort low · instructions. Includes a calculation, not a bill or a new run. Caveat: The calls repeat only three fixed prompts. Wilson intervals describe call outcomes under a binomial assumption; they do not measure accuracy across unseen tasks. The paired p-values also assume independent pairs and do not remove this limit.
- Four harder tasks: pass rate 38% vs 69%, tie: 95% intervals overlap. 1 of 10 rows separates them. Table: Four harder tasks · 5 of 10 rows · Claude Code vs Codex CLI · n = 4–16 per side. Source study: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Rows shown: Pass rate on 4 harder tasks (Strict pass); Pass rate on 4 harder tasks (Lenient (format misses counted)); Calls that tried a tool although tools were off; Strict pass rate by task: 10x10 nonogram; Strict pass rate by task: 6x6 Skyscrapers. Recorded settings: Claude Code; Codex CLI · effort medium. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.
- No winner where the data shows none. Showing 19 of 70 rows; every row and its reason online.
Measured pairs
16 of 16 measured pairs
Claude Code vs Codex CLI
Claude Code and Codex CLI share 66 measured metrics and 22 list-price calculations from 12 studies. Claude Code leads on 7 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 4 more. Codex CLI leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 31 ties and 49 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Each run pairs a CLI with a model, so these rows cannot separate the CLI from the model; the contexts name both. Some rows rest on small samples (n = 2 at the smallest).
Read all 88 metric rows for Claude Code vs Codex CLI
| Metric | Claude Code | Codex CLI | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 80% (12/15)Claude Sonnet 5.5 · five short validated tasks | 100% (15/15)GPT-6.1 Sol · effort medium · five short validated tasks | 15 | 95% CI: 55%–93% vs 80%–100% | Tie | The 95% intervals overlap (Claude Code 55% to 93%; Codex CLI 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 2.31 sClaude Sonnet 5.5 · five short validated tasks | 5.65 sGPT-6.1 Sol · effort medium · five short validated tasks | 15 | range: 2.2 s–7.7 s vs 4.1 s–25.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 2.17 s to 7.73 s; Codex CLI 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 1.56 sClaude Sonnet 5.5 · five short validated tasks | 5.05 sGPT-6.1 Sol · effort medium · five short validated tasks | 15 | range: 1 s–6.4 s vs 3.4 s–17.8 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 0.99 s to 6.39 s; Codex CLI 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 1,401Claude Sonnet 5.5 · five short validated tasks | 5,180GPT-6.1 Sol · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Other input) | 685Claude Sonnet 5.5 · five short validated tasks | 6,943GPT-6.1 Sol · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 107Claude Sonnet 5.5 · five short validated tasks | 42GPT-6.1 Sol · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per call (calculation)Calculation | $0.0036Claude Sonnet 5.5 · five short validated tasks | $0.010GPT-6.1 Sol · effort medium · five short validated tasks | 15 | range: $0.0034–$0.01 vs $0.0054–$0.027 | Unclear | The run ranges (fastest to slowest) overlap (Claude Code $0.0034 to $0.010; Codex CLI $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.0062Claude Sonnet 5.5 · five short validated tasks | $0.016GPT-6.1 Sol · effort medium · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.0062 vs $0.016, 2.5x) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Pass rate on eight hard tasks (Strict pass) | 100% (24/24)Claude Sonnet 5.5 · eight hard validated tasks | 100% (16/16)GPT-6.1 Sol · effort medium · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Code 86% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 100% (24/24)Claude Sonnet 5.5 · eight hard validated tasks | 100% (16/16)GPT-6.1 Sol · effort medium · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Code 86% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Total time per call on hard tasks (separate batches) | 7.75 sClaude Sonnet 5.5 · eight hard validated tasks | 13.1 sGPT-6.1 Sol · effort medium · eight hard validated tasks | 24 / 16 | range: 2.3 s–34.8 s vs 8.5 s–61.6 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 2.26 s to 34.8 s; Codex CLI 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Time to first useful output on hard tasks | 5.95 sClaude Sonnet 5.5 · eight hard validated tasks | 10.2 sGPT-6.1 Sol · effort medium · eight hard validated tasks | 24 / 16 | range: 0.9 s–30.6 s vs 6.1 s–40.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 0.86 s to 30.6 s; Codex CLI 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call on hard tasks (Output tokens) | 1,050Claude Sonnet 5.5 · eight hard validated tasks | 335GPT-6.1 Sol · effort medium · eight hard validated tasks | 24 / 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass on hard tasks (calculation)Calculation | $0.014Claude Sonnet 5.5 · eight hard validated tasks | $0.026GPT-6.1 Sol · effort medium · eight hard validated tasks | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Coding sessions that passed every hidden check | 100% (12/12)Claude Sonnet 5.5 · six small repository tasks with hidden tests | 100% (12/12)GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Time per coding session | 23.1 sClaude Sonnet 5.5 · six small repository tasks with hidden tests | 113.4 sGPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | range: 18.7 s–44.5 s vs 78.5 s–222 s | Claude Code ahead | The run ranges (fastest to slowest) do not overlap (Claude Code 18.7 s to 44.5 s; Codex CLI 78.5 s to 221.9 s). A range is not a confidence interval. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Tool calls per coding session | 7.5Claude Sonnet 5.5 · six small repository tasks with hidden tests | 12.5GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | range: 3–14 vs 8–18 | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| List-price cost per passing coding session (calculation)Calculation | $0.085Claude Sonnet 5.5 · six small repository tasks with hidden tests | $0.098GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.085 vs $0.098) is not tested against run-to-run variation. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Strict pass rate by effort on eight hard tasks | 100% (16/16)Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder | 100% (16/16)GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder | 16 | 95% CI: 81%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Code 81% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Total time per call by effort on hard tasks | 7.63 sClaude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder | 13.1 sGPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder | 16 | range: 2.7 s–24 s vs 8.5 s–61.6 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 2.71 s to 24.0 s; Codex CLI 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call by effort on hard tasks (Output tokens) | 770Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder | 335GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass by effort (calculation)Calculation | $0.014Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder | $0.026GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Same prompt, 10 times: strict pass rate (Exact number) | 100% (10/10)Claude Sonnet 5.5 · same prompt repeated 10 times | 100% (10/10)GPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | 95% CI: 72%–100% vs 72%–100% | Tie | The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: strict pass rate (JSON object) | 100% (10/10)Claude Sonnet 5.5 · same prompt repeated 10 times | 100% (10/10)GPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | 95% CI: 72%–100% vs 72%–100% | Tie | The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: strict pass rate (Code fix) | 100% (10/10)Claude Sonnet 5.5 · same prompt repeated 10 times | 100% (10/10)GPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | 95% CI: 72%–100% vs 72%–100% | Tie | The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (Exact number) | 1Claude Sonnet 5.5 · same prompt repeated 10 times | 1GPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (JSON object) | 1Claude Sonnet 5.5 · same prompt repeated 10 times | 1GPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (Code fix) | 3Claude Sonnet 5.5 · same prompt repeated 10 times | 6GPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | none recorded | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (Exact number) | 6.89 sClaude Sonnet 5.5 · same prompt repeated 10 times | 13.4 sGPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | range: 5.8 s–7.8 s vs 12.3 s–18 s | Claude Code ahead | The run ranges (fastest to slowest) do not overlap (Claude Code 5.81 s to 7.81 s; Codex CLI 12.3 s to 18.0 s). A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (JSON object) | 2.89 sClaude Sonnet 5.5 · same prompt repeated 10 times | 6.42 sGPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | range: 2.7 s–5.3 s vs 5.3 s–8.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 2.68 s to 5.30 s; Codex CLI 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (Code fix) | 2.67 sClaude Sonnet 5.5 · same prompt repeated 10 times | 11.3 sGPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | range: 2.3 s–4.3 s vs 9.1 s–14.9 s | Claude Code ahead | The run ranges (fastest to slowest) do not overlap (Claude Code 2.32 s to 4.34 s; Codex CLI 9.08 s to 14.8 s). A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| CLI start-up tax on a one-word answer (First output event) | 563 msClaude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs | 489 msdefault model · CLI start-up, one-word prompt, 5 runs | 5 | range: 519 ms–726 ms vs 354 ms–1.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 519 ms to 726 ms; Codex CLI 354 ms to 1,304 ms); the medians alone do not show a reliable difference. A range is not a confidence interval. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| CLI start-up tax on a one-word answer (First model output) | 1,461 msClaude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs | 5,059 msdefault model · CLI start-up, one-word prompt, 5 runs | 5 | range: 1.21 s–2.31 s vs 4.39 s–5.48 s | Claude Code ahead | The run ranges (fastest to slowest) do not overlap (Claude Code 1,206 ms to 2,308 ms; Codex CLI 4,391 ms to 5,478 ms). A range is not a confidence interval. Samples are small (5 runs per side). | Routing overhead: deterministic policy vs LLM routers vs Jev |
| CLI start-up tax on a one-word answer (Total wall time) | 2,529 msClaude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs | 5,999 msdefault model · CLI start-up, one-word prompt, 5 runs | 5 | range: 2.27 s–3.38 s vs 5.37 s–6.51 s | Claude Code ahead | The run ranges (fastest to slowest) do not overlap (Claude Code 2,273 ms to 3,382 ms; Codex CLI 5,367 ms to 6,506 ms). A range is not a confidence interval. Samples are small (5 runs per side). | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Input tokens a CLI sends for a one-word answer | 6,761Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs | 17,051default model · CLI start-up, one-word prompt, 5 runs | 5 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Repairing a scheduler: Claude Code vs Codex vs API (Total time) | 15.0 sSonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs | 61.2 sGPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs | 3 | range: 13.9 s–15.9 s vs 59.9 s–69.5 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 13.9 s to 15.9 s; Codex CLI 59.9 s to 69.5 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Repairing a scheduler: Claude Code vs Codex vs API (First useful output) | 7.55 sSonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs | 15.6 sGPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs | 3 | range: 6.8 s–7.6 s vs 13.7 s–23 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 6.77 s to 7.63 s; Codex CLI 13.7 s to 23.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Output tokens to repair the scheduler (Output tokens) | 2,227Sonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs | 1,181GPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs | 3 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Strict pass rate: single call vs agent loop on eight hard tasks | 100% (24/24)Claude Sonnet 5.5 · single call | 63% (10/16)GPT-6 Luna · single call | 24 / 16 | 95% CI: 86%–100% vs 39%–82% | Claude Code ahead | The 95% intervals do not overlap (Claude Code 86% to 100%; Codex CLI 39% to 82%). | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Interval merge fix | 100% (3/3)Claude Sonnet 5.5 · single call | 100% (2/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: DST day-length fix | 100% (3/3)Claude Sonnet 5.5 · single call | 100% (2/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: CSV parser | 100% (3/3)Claude Sonnet 5.5 · single call | 100% (2/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Event-loop order | 100% (3/3)Claude Sonnet 5.5 · single call | 0% (0/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 0%–66% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 0% to 66%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Room schedule | 100% (3/3)Claude Sonnet 5.5 · single call | 50% (1/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 9.4%–91% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 9% to 91%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SemVer regex | 100% (3/3)Claude Sonnet 5.5 · single call | 50% (1/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 9.4%–91% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 9% to 91%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Money refactor | 100% (3/3)Claude Sonnet 5.5 · single call | 0% (0/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 0%–66% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 0% to 66%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SQL report | 100% (3/3)Claude Sonnet 5.5 · single call | 100% (2/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Total time per attempt: single call vs agent loop | 7.75 sClaude Sonnet 5.5 · single call | 5.16 sGPT-6 Luna · single call | 24 / 16 | range: 2.3 s–34.8 s vs 3.6 s–11.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 2.26 s to 34.8 s; Codex CLI 3.59 s to 11.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Input tokens (cache reads included)) | 2,281Claude Sonnet 5.5 · single call | 11,582GPT-6 Luna · single call | 24 / 16 | range: 2,234–2,669 vs 11,526–11,818 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Output tokens) | 1,050Claude Sonnet 5.5 · single call | 345GPT-6 Luna · single call | 24 / 16 | range: 176–3,895 vs 36–634 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tool calls per agent-loop attempt | 0Claude Sonnet 5.5 · agent loop | 0GPT-6 Luna · agent loop | 16 / 14 | range: 0–3 vs 0–1 | Tie | Same value. More or fewer is not better by itself for this metric. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| List-price cost per strict pass: single call vs agent loop (calculation)Calculation | $0.014Claude Sonnet 5.5 · single call | $0.0012GPT-6 Luna · single call | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.014 vs $0.0012, 12x) is not tested against run-to-run variation. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)Calculation | 100% (12/12)Claude Sonnet 5.5 · instructions | 100% (12/12)GPT-6.1 Sol · effort low · instructions | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))Calculation | 100% (12/12)Claude Sonnet 5.5 · instructions | 100% (12/12)GPT-6.1 Sol · effort low · instructions | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Strict pass) | 12Claude Sonnet 5.5 · instructions | 12GPT-6.1 Sol · effort low · instructions | 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Format miss) | 0Claude Sonnet 5.5 · instructions | 0GPT-6.1 Sol · effort low · instructions | 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Wrong values) | 0Claude Sonnet 5.5 · instructions | 0GPT-6.1 Sol · effort low · instructions | 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Error) | 0Claude Sonnet 5.5 · instructions | 0GPT-6.1 Sol · effort low · instructions | 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Time per call, instructions vs schema mode | 3.52 sClaude Sonnet 5.5 · instructions | 6.21 sGPT-6.1 Sol · effort low · instructions | 12 | range: 2.7 s–4.1 s vs 4.2 s–12.3 s | Claude Code ahead | The run ranges (fastest to slowest) do not overlap (Claude Code 2.67 s to 4.12 s; Codex CLI 4.20 s to 12.3 s). A range is not a confidence interval. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call) | 368Claude Sonnet 5.5 · instructions | 117GPT-6.1 Sol · effort low · instructions | 12 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Reasoning share of output tokens per call on hard tasks (calculation)Calculation | 54.5%Claude Sonnet 5.5 | 46.3%GPT-6.1 Sol · effort medium | 24 / 16 | range: 0%–96% vs 11%–87% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation | $0.0067Claude Sonnet 5.5 | $0.0023GPT-6.1 Sol · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation | $0.0037Claude Sonnet 5.5 | $0.0030GPT-6.1 Sol · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation | $0.0040Claude Sonnet 5.5 | $0.020GPT-6.1 Sol · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation | $0.0059Claude Sonnet 5.5 · effort medium | $0.0023GPT-6.1 Sol · effort medium | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation | $0.014Claude Sonnet 5.5 · effort medium | $0.026GPT-6.1 Sol · effort medium | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation | 54.5%Claude Sonnet 5.5 | 46.3%GPT-6.1 Sol · effort medium | 24 / 16 | range: 0%–96% vs 11%–87% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation | 0%Claude Sonnet 5.5 | 41%GPT-6.1 Sol · effort medium | 15 | range: 0%–73% vs 0%–71% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Time to first text: a 250-line answer, six models | 1.96 sClaude Sonnet 5.5 | 3.52 sGPT-6.1 Sol · effort low | 4 | range: 0.9 s–4.1 s vs 2.8 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 0.88 s to 4.09 s; Codex CLI 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 232Claude Sonnet 5.5 | 80GPT-6.1 Sol · effort low | 4 | range: 230–233 vs 72–81 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 517Claude Sonnet 5.5 | 323GPT-6.1 Sol · effort low | 4 | range: 513–519 vs 291–327 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 1kCalculation | 1.45 sClaude Sonnet 5.5 | 3.36 sGPT-6.1 Sol · effort low | 3 | range: 1.2 s–1.7 s vs 3.4 s–4.8 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.23 s to 1.72 s; Codex CLI 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 16kCalculation | 1.78 sClaude Sonnet 5.5 | 4.02 sGPT-6.1 Sol · effort low | 3 | range: 1.6 s–2.1 s vs 3.3 s–4.3 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.64 s to 2.11 s; Codex CLI 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 64kCalculation | 3.07 sClaude Sonnet 5.5 | 3.93 sGPT-6.1 Sol · effort low | 3 | range: 1.4 s–3.6 s vs 3.4 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 1.38 s to 3.61 s; Codex CLI 3.42 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (1k prompt) | 1.78 sClaude Sonnet 5.5 | 3.43 sGPT-6.1 Sol · effort low | 3 | range: 1.6 s–2.1 s vs 3.4 s–4.9 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.57 s to 2.12 s; Codex CLI 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (16k prompt) | 2.10 sClaude Sonnet 5.5 | 4.14 sGPT-6.1 Sol · effort low | 3 | range: 2 s–2.5 s vs 4 s–4.7 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.98 s to 2.48 s; Codex CLI 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (64k prompt) | 3.44 sClaude Sonnet 5.5 | 3.96 sGPT-6.1 Sol · effort low | 3 | range: 1.7 s–4.4 s vs 3.5 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 1.74 s to 4.38 s; Codex CLI 3.47 s to 4.44 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Exact lookup answers at the 1k, 16k and 64k prompt-size targets | 100% (9/9)Claude Sonnet 5.5 | 100% (9/9)GPT-6.1 Sol · effort low | 9 | 95% CI: 70%–100% vs 70%–100% | Tie | The 95% intervals overlap (Claude Code 70% to 100%; Codex CLI 70% to 100%), so this sample cannot separate them. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Pass rate on 4 harder tasks (Strict pass) | 38% (6/16)Claude Sonnet 5.5 | 69% (11/16)GPT-6.1 Sol · effort medium | 16 | 95% CI: 18%–61% vs 44%–86% | Tie | The 95% intervals overlap (Claude Code 18% to 61%; Codex CLI 44% to 86%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Pass rate on 4 harder tasks (Lenient (format misses counted)) | 38% (6/16)Claude Sonnet 5.5 | 69% (11/16)GPT-6.1 Sol · effort medium | 16 | 95% CI: 18%–61% vs 44%–86% | Tie | The 95% intervals overlap (Claude Code 18% to 61%; Codex CLI 44% to 86%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Calls that tried a tool although tools were off | 31% (5/16)Claude Sonnet 5.5 | 0% (0/16)GPT-6.1 Sol · effort medium | 16 | 95% CI: 14%–56% vs 0%–19% | Unclear | More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 10x10 nonogram | 100% (4/4)Claude Sonnet 5.5 | 75% (3/4)GPT-6.1 Sol · effort medium | 4 | 95% CI: 51%–100% vs 30%–95% | Tie | The 95% intervals overlap (Claude Code 51% to 100%; Codex CLI 30% to 95%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Sudoku, 22 givens | 0% (0/4)Claude Sonnet 5.5 | 25% (1/4)GPT-6.1 Sol · effort medium | 4 | 95% CI: 0%–49% vs 4.6%–70% | Tie | The 95% intervals overlap (Claude Code 0% to 49%; Codex CLI 5% to 70%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 6x6 Skyscrapers | 0% (0/4)Claude Sonnet 5.5 | 100% (4/4)GPT-6.1 Sol · effort medium | 4 | 95% CI: 0%–49% vs 51%–100% | Codex CLI ahead | The 95% intervals do not overlap (Claude Code 0% to 49%; Codex CLI 51% to 100%). | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Seeded shuffle output | 50% (2/4)Claude Sonnet 5.5 | 75% (3/4)GPT-6.1 Sol · effort medium | 4 | 95% CI: 15%–85% vs 30%–95% | Tie | The 95% intervals overlap (Claude Code 15% to 85%; Codex CLI 30% to 95%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Total time per call on harder tasks | 70.4 sClaude Sonnet 5.5 | 120.2 sGPT-6.1 Sol · effort medium | 12 / 13 | range: 4.3 s–210 s vs 46.2 s–273 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 4.32 s to 210.1 s; Codex CLI 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Output tokens per call on harder tasks (Output tokens) | 9,287Claude Sonnet 5.5 | 4,994GPT-6.1 Sol · effort medium | 12 / 13 | range: 407–27,921 vs 2,099–13,413 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| List-price cost per strict pass on harder tasks (calculation)Calculation | $0.24Claude Sonnet 5.5 | $0.083GPT-6.1 Sol · effort medium | 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.24 vs $0.083, 2.9x) is not tested against run-to-run variation. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)
Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) share 49 measured metrics and 21 list-price calculations from 10 studies. Claude Sonnet 5.5 leads on 4 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 1 more. GPT-6.1 Sol (Codex CLI) leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 22 ties and 43 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).
Read all 70 metric rows for Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)
| Metric | Claude Sonnet 5.5 | GPT-6.1 Sol (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 80% (12/15)Claude Code · five short validated tasks | 100% (15/15)Codex CLI · effort medium · five short validated tasks | 15 | 95% CI: 55%–93% vs 80%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 55% to 93%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 2.31 sClaude Code · five short validated tasks | 5.65 sCodex CLI · effort medium · five short validated tasks | 15 | range: 2.2 s–7.7 s vs 4.1 s–25.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.17 s to 7.73 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 1.56 sClaude Code · five short validated tasks | 5.05 sCodex CLI · effort medium · five short validated tasks | 15 | range: 1 s–6.4 s vs 3.4 s–17.8 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.99 s to 6.39 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 1,401Claude Code · five short validated tasks | 5,180Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Other input) | 685Claude Code · five short validated tasks | 6,943Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 107Claude Code · five short validated tasks | 42Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per call (calculation)Calculation | $0.0036Claude Code · five short validated tasks | $0.010Codex CLI · effort medium · five short validated tasks | 15 | range: $0.0034–$0.01 vs $0.0054–$0.027 | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 $0.0034 to $0.010; GPT-6.1 Sol (Codex CLI) $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.0062Claude Code · five short validated tasks | $0.016Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.0062 vs $0.016, 2.5x) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Pass rate on eight hard tasks (Strict pass) | 100% (24/24)Claude Code · eight hard validated tasks | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 100% (24/24)Claude Code · eight hard validated tasks | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Total time per call on hard tasks (separate batches) | 7.75 sClaude Code · eight hard validated tasks | 13.1 sCodex CLI · effort medium · eight hard validated tasks | 24 / 16 | range: 2.3 s–34.8 s vs 8.5 s–61.6 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Time to first useful output on hard tasks | 5.95 sClaude Code · eight hard validated tasks | 10.2 sCodex CLI · effort medium · eight hard validated tasks | 24 / 16 | range: 0.9 s–30.6 s vs 6.1 s–40.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.86 s to 30.6 s; GPT-6.1 Sol (Codex CLI) 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call on hard tasks (Output tokens) | 1,050Claude Code · eight hard validated tasks | 335Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass on hard tasks (calculation)Calculation | $0.014Claude Code · eight hard validated tasks | $0.026Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Coding sessions that passed every hidden check | 100% (12/12)Claude Code · six small repository tasks with hidden tests | 100% (12/12)Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Time per coding session | 23.1 sClaude Code · six small repository tasks with hidden tests | 113.4 sCodex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | range: 18.7 s–44.5 s vs 78.5 s–222 s | Claude Sonnet 5.5 ahead | The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 18.7 s to 44.5 s; GPT-6.1 Sol (Codex CLI) 78.5 s to 221.9 s). A range is not a confidence interval. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Tool calls per coding session | 7.5Claude Code · six small repository tasks with hidden tests | 12.5Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | range: 3–14 vs 8–18 | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| List-price cost per passing coding session (calculation)Calculation | $0.085Claude Code · six small repository tasks with hidden tests | $0.098Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.085 vs $0.098) is not tested against run-to-run variation. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Strict pass rate by effort on eight hard tasks | 100% (16/16)Claude Code · effort medium · eight hard validated tasks, effort ladder | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks, effort ladder | 16 | 95% CI: 81%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 81% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Total time per call by effort on hard tasks | 7.63 sClaude Code · effort medium · eight hard validated tasks, effort ladder | 13.1 sCodex CLI · effort medium · eight hard validated tasks, effort ladder | 16 | range: 2.7 s–24 s vs 8.5 s–61.6 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.71 s to 24.0 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call by effort on hard tasks (Output tokens) | 770Claude Code · effort medium · eight hard validated tasks, effort ladder | 335Codex CLI · effort medium · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass by effort (calculation)Calculation | $0.014Claude Code · effort medium · eight hard validated tasks, effort ladder | $0.026Codex CLI · effort medium · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Same prompt, 10 times: strict pass rate (Exact number) | 100% (10/10)Claude Code · same prompt repeated 10 times | 100% (10/10)Codex CLI · effort medium · same prompt repeated 10 times | 10 | 95% CI: 72%–100% vs 72%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: strict pass rate (JSON object) | 100% (10/10)Claude Code · same prompt repeated 10 times | 100% (10/10)Codex CLI · effort medium · same prompt repeated 10 times | 10 | 95% CI: 72%–100% vs 72%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: strict pass rate (Code fix) | 100% (10/10)Claude Code · same prompt repeated 10 times | 100% (10/10)Codex CLI · effort medium · same prompt repeated 10 times | 10 | 95% CI: 72%–100% vs 72%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (Exact number) | 1Claude Code · same prompt repeated 10 times | 1Codex CLI · effort medium · same prompt repeated 10 times | 10 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (JSON object) | 1Claude Code · same prompt repeated 10 times | 1Codex CLI · effort medium · same prompt repeated 10 times | 10 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (Code fix) | 3Claude Code · same prompt repeated 10 times | 6Codex CLI · effort medium · same prompt repeated 10 times | 10 | none recorded | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (Exact number) | 6.89 sClaude Code · same prompt repeated 10 times | 13.4 sCodex CLI · effort medium · same prompt repeated 10 times | 10 | range: 5.8 s–7.8 s vs 12.3 s–18 s | Claude Sonnet 5.5 ahead | The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 5.81 s to 7.81 s; GPT-6.1 Sol (Codex CLI) 12.3 s to 18.0 s). A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (JSON object) | 2.89 sClaude Code · same prompt repeated 10 times | 6.42 sCodex CLI · effort medium · same prompt repeated 10 times | 10 | range: 2.7 s–5.3 s vs 5.3 s–8.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.68 s to 5.30 s; GPT-6.1 Sol (Codex CLI) 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (Code fix) | 2.67 sClaude Code · same prompt repeated 10 times | 11.3 sCodex CLI · effort medium · same prompt repeated 10 times | 10 | range: 2.3 s–4.3 s vs 9.1 s–14.9 s | Claude Sonnet 5.5 ahead | The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 2.32 s to 4.34 s; GPT-6.1 Sol (Codex CLI) 9.08 s to 14.8 s). A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Repairing a scheduler: Claude Code vs Codex vs API (Total time) | 15.0 sClaude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs | 61.2 sCodex CLI · effort medium · scheduler repair, 296 checks, 3 runs | 3 | range: 13.9 s–15.9 s vs 59.9 s–69.5 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 13.9 s to 15.9 s; GPT-6.1 Sol (Codex CLI) 59.9 s to 69.5 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Repairing a scheduler: Claude Code vs Codex vs API (First useful output) | 7.55 sClaude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs | 15.6 sCodex CLI · effort medium · scheduler repair, 296 checks, 3 runs | 3 | range: 6.8 s–7.6 s vs 13.7 s–23 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 6.77 s to 7.63 s; GPT-6.1 Sol (Codex CLI) 13.7 s to 23.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Output tokens to repair the scheduler (Output tokens) | 2,227Claude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs | 1,181Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs | 3 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)Calculation | 100% (12/12)Claude Code · instructions | 100% (12/12)Codex CLI · effort low · instructions | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))Calculation | 100% (12/12)Claude Code · instructions | 100% (12/12)Codex CLI · effort low · instructions | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Strict pass) | 12Claude Code · instructions | 12Codex CLI · effort low · instructions | 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Format miss) | 0Claude Code · instructions | 0Codex CLI · effort low · instructions | 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Wrong values) | 0Claude Code · instructions | 0Codex CLI · effort low · instructions | 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Error) | 0Claude Code · instructions | 0Codex CLI · effort low · instructions | 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Time per call, instructions vs schema mode | 3.52 sClaude Code · instructions | 6.21 sCodex CLI · effort low · instructions | 12 | range: 2.7 s–4.1 s vs 4.2 s–12.3 s | Claude Sonnet 5.5 ahead | The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 2.67 s to 4.12 s; GPT-6.1 Sol (Codex CLI) 4.20 s to 12.3 s). A range is not a confidence interval. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call) | 368Claude Code · instructions | 117Codex CLI · effort low · instructions | 12 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Reasoning share of output tokens per call on hard tasks (calculation)Calculation | 54.5%Claude Code | 46.3%Codex CLI · effort medium | 24 / 16 | range: 0%–96% vs 11%–87% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation | $0.0067Claude Code | $0.0023Codex CLI · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation | $0.0037Claude Code | $0.0030Codex CLI · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation | $0.0040Claude Code | $0.020Codex CLI · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation | $0.0059Claude Code · effort medium | $0.0023Codex CLI · effort medium | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation | $0.014Claude Code · effort medium | $0.026Codex CLI · effort medium | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation | 54.5%Claude Code | 46.3%Codex CLI · effort medium | 24 / 16 | range: 0%–96% vs 11%–87% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation | 0%Claude Code | 41%Codex CLI · effort medium | 15 | range: 0%–73% vs 0%–71% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Time to first text: a 250-line answer, six models | 1.96 sClaude Code | 3.52 sCodex CLI · effort low | 4 | range: 0.9 s–4.1 s vs 2.8 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.88 s to 4.09 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 232Claude Code | 80Codex CLI · effort low | 4 | range: 230–233 vs 72–81 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 517Claude Code | 323Codex CLI · effort low | 4 | range: 513–519 vs 291–327 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 1kCalculation | 1.45 sClaude Code | 3.36 sCodex CLI · effort low | 3 | range: 1.2 s–1.7 s vs 3.4 s–4.8 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 1.23 s to 1.72 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 16kCalculation | 1.78 sClaude Code | 4.02 sCodex CLI · effort low | 3 | range: 1.6 s–2.1 s vs 3.3 s–4.3 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 1.64 s to 2.11 s; GPT-6.1 Sol (Codex CLI) 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 64kCalculation | 3.07 sClaude Code | 3.93 sCodex CLI · effort low | 3 | range: 1.4 s–3.6 s vs 3.4 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.38 s to 3.61 s; GPT-6.1 Sol (Codex CLI) 3.42 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (1k prompt) | 1.78 sClaude Code | 3.43 sCodex CLI · effort low | 3 | range: 1.6 s–2.1 s vs 3.4 s–4.9 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 1.57 s to 2.12 s; GPT-6.1 Sol (Codex CLI) 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (16k prompt) | 2.10 sClaude Code | 4.14 sCodex CLI · effort low | 3 | range: 2 s–2.5 s vs 4 s–4.7 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 1.98 s to 2.48 s; GPT-6.1 Sol (Codex CLI) 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (64k prompt) | 3.44 sClaude Code | 3.96 sCodex CLI · effort low | 3 | range: 1.7 s–4.4 s vs 3.5 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.74 s to 4.38 s; GPT-6.1 Sol (Codex CLI) 3.47 s to 4.44 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Exact lookup answers at the 1k, 16k and 64k prompt-size targets | 100% (9/9)Claude Code | 100% (9/9)Codex CLI · effort low | 9 | 95% CI: 70%–100% vs 70%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 70% to 100%; GPT-6.1 Sol (Codex CLI) 70% to 100%), so this sample cannot separate them. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Pass rate on 4 harder tasks (Strict pass) | 38% (6/16)Claude Code | 69% (11/16)Codex CLI · effort medium | 16 | 95% CI: 18%–61% vs 44%–86% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 18% to 61%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Pass rate on 4 harder tasks (Lenient (format misses counted)) | 38% (6/16)Claude Code | 69% (11/16)Codex CLI · effort medium | 16 | 95% CI: 18%–61% vs 44%–86% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 18% to 61%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Calls that tried a tool although tools were off | 31% (5/16)Claude Code | 0% (0/16)Codex CLI · effort medium | 16 | 95% CI: 14%–56% vs 0%–19% | Unclear | More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 10x10 nonogram | 100% (4/4)Claude Code | 75% (3/4)Codex CLI · effort medium | 4 | 95% CI: 51%–100% vs 30%–95% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 51% to 100%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Sudoku, 22 givens | 0% (0/4)Claude Code | 25% (1/4)Codex CLI · effort medium | 4 | 95% CI: 0%–49% vs 4.6%–70% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 0% to 49%; GPT-6.1 Sol (Codex CLI) 5% to 70%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 6x6 Skyscrapers | 0% (0/4)Claude Code | 100% (4/4)Codex CLI · effort medium | 4 | 95% CI: 0%–49% vs 51%–100% | GPT-6.1 Sol (Codex CLI) ahead | The 95% intervals do not overlap (Claude Sonnet 5.5 0% to 49%; GPT-6.1 Sol (Codex CLI) 51% to 100%). | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Seeded shuffle output | 50% (2/4)Claude Code | 75% (3/4)Codex CLI · effort medium | 4 | 95% CI: 15%–85% vs 30%–95% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 15% to 85%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Total time per call on harder tasks | 70.4 sClaude Code | 120.2 sCodex CLI · effort medium | 12 / 13 | range: 4.3 s–210 s vs 46.2 s–273 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 4.32 s to 210.1 s; GPT-6.1 Sol (Codex CLI) 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Output tokens per call on harder tasks (Output tokens) | 9,287Claude Code | 4,994Codex CLI · effort medium | 12 / 13 | range: 407–27,921 vs 2,099–13,413 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| List-price cost per strict pass on harder tasks (calculation)Calculation | $0.24Claude Code | $0.083Codex CLI · effort medium | 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.24 vs $0.083, 2.9x) is not tested against run-to-run variation. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI)
Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) share 31 measured metrics and 19 list-price calculations from 7 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 12 ties and 38 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).
Read all 50 metric rows for Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI)
| Metric | Claude Opus 5.5 | GPT-6.1 Sol (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 100% (15/15)Claude Code · effort high · five short validated tasks | 100% (15/15)Codex CLI · effort high · five short validated tasks | 15 | 95% CI: 80%–100% vs 80%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 2.71 sClaude Code · effort high · five short validated tasks | 5.60 sCodex CLI · effort high · five short validated tasks | 15 | range: 2.5 s–11.8 s vs 4.1 s–19.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.45 s to 11.8 s; GPT-6.1 Sol (Codex CLI) 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 2.04 sClaude Code · effort high · five short validated tasks | 5.32 sCodex CLI · effort high · five short validated tasks | 15 | range: 1.4 s–9.9 s vs 3.6 s–16.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.40 s to 9.94 s; GPT-6.1 Sol (Codex CLI) 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 1,463Claude Code · effort high · five short validated tasks | 6,716Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Other input) | 619Claude Code · effort high · five short validated tasks | 5,406Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 78Claude Code · effort high · five short validated tasks | 42Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per call (calculation)Calculation | $0.0069Claude Code · effort high · five short validated tasks | $0.010Codex CLI · effort high · five short validated tasks | 15 | range: $0.0059–$0.027 vs $0.0066–$0.028 | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 $0.0059 to $0.027; GPT-6.1 Sol (Codex CLI) $0.0066 to $0.028); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.010Claude Code · effort high · five short validated tasks | $0.013Codex CLI · effort high · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.010 vs $0.013) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Pass rate on eight hard tasks (Strict pass) | 100% (24/24)Claude Code · effort high · eight hard validated tasks | 100% (16/16)Codex CLI · effort high · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 100% (24/24)Claude Code · effort high · eight hard validated tasks | 100% (16/16)Codex CLI · effort high · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Total time per call on hard tasks (separate batches) | 11.0 sClaude Code · effort high · eight hard validated tasks | 18.1 sCodex CLI · effort high · eight hard validated tasks | 24 / 16 | range: 3.6 s–63 s vs 11.7 s–92.2 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 3.63 s to 63.0 s; GPT-6.1 Sol (Codex CLI) 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Time to first useful output on hard tasks | 7.13 sClaude Code · effort high · eight hard validated tasks | 12.7 sCodex CLI · effort high · eight hard validated tasks | 24 / 16 | range: 2.2 s–56.2 s vs 8.9 s–75.9 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.15 s to 56.2 s; GPT-6.1 Sol (Codex CLI) 8.93 s to 75.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call on hard tasks (Output tokens) | 1,052Claude Code · effort high · eight hard validated tasks | 436Codex CLI · effort high · eight hard validated tasks | 24 / 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass on hard tasks (calculation)Calculation | $0.033Claude Code · effort high · eight hard validated tasks | $0.015Codex CLI · effort high · eight hard validated tasks | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.033 vs $0.015, 2.2x) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Coding sessions that passed every hidden check | 100% (12/12)Claude Code · six small repository tasks with hidden tests | 100% (12/12)Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Time per coding session | 56.9 sClaude Code · six small repository tasks with hidden tests | 113.4 sCodex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | range: 29.8 s–186 s vs 78.5 s–222 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 29.8 s to 185.8 s; GPT-6.1 Sol (Codex CLI) 78.5 s to 221.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Tool calls per coding session | 7.5Claude Code · six small repository tasks with hidden tests | 12.5Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | range: 5–14 vs 8–18 | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| List-price cost per passing coding session (calculation)Calculation | $0.22Claude Code · six small repository tasks with hidden tests | $0.098Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.22 vs $0.098, 2.3x) is not tested against run-to-run variation. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Strict pass rate by effort on eight hard tasks | 100% (16/16)Claude Code · effort medium · eight hard validated tasks, effort ladder | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks, effort ladder | 16 | 95% CI: 81%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 81% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Total time per call by effort on hard tasks | 9.72 sClaude Code · effort medium · eight hard validated tasks, effort ladder | 13.1 sCodex CLI · effort medium · eight hard validated tasks, effort ladder | 16 | range: 4.8 s–31.4 s vs 8.5 s–61.6 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 4.78 s to 31.4 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call by effort on hard tasks (Output tokens) | 853Claude Code · effort medium · eight hard validated tasks, effort ladder | 335Codex CLI · effort medium · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass by effort (calculation)Calculation | $0.029Claude Code · effort medium · eight hard validated tasks, effort ladder | $0.026Codex CLI · effort medium · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.029 vs $0.026) is not tested against run-to-run variation. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Reasoning share of output tokens per call on hard tasks (calculation)Calculation | 54.4%Claude Code · effort high | 57%Codex CLI · effort high | 24 / 16 | range: 36%–96% vs 29%–91% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation | $0.018Claude Code · effort high | $0.0041Codex CLI · effort high | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation | $0.0080Claude Code · effort high | $0.0029Codex CLI · effort high | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation | $0.0074Claude Code · effort high | $0.0081Codex CLI · effort high | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation | $0.013Claude Code · effort medium | $0.0023Codex CLI · effort medium | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation | $0.029Claude Code · effort medium | $0.026Codex CLI · effort medium | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation | 54.4%Claude Code · effort high | 57%Codex CLI · effort high | 24 / 16 | range: 36%–96% vs 29%–91% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation | 43.6%Claude Code · effort high | 58.1%Codex CLI · effort high | 15 | range: 0%–93% vs 0%–76% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Time to first text: a 250-line answer, six models | 1.97 sClaude Code | 3.52 sCodex CLI · effort low | 4 | range: 1.7 s–2.4 s vs 2.8 s–4.4 s | Unclear | Only 4 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.70 s to 2.35 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s), but 4 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 156Claude Code | 80Codex CLI · effort low | 4 | range: 155–156 vs 72–81 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 347Claude Code | 323Codex CLI · effort low | 4 | range: 345–349 vs 291–327 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 1kCalculation | 1.51 sClaude Code | 3.36 sCodex CLI · effort low | 3 | range: 1.5 s–2 s vs 3.4 s–4.8 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.46 s to 2.01 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 16kCalculation | 1.74 sClaude Code | 4.02 sCodex CLI · effort low | 3 | range: 1.7 s–3 s vs 3.3 s–4.3 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.70 s to 2.97 s; GPT-6.1 Sol (Codex CLI) 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 64kCalculation | 1.79 sClaude Code | 3.93 sCodex CLI · effort low | 3 | range: 1.7 s–3.7 s vs 3.4 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.72 s to 3.72 s; GPT-6.1 Sol (Codex CLI) 3.42 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (1k prompt) | 1.83 sClaude Code | 3.43 sCodex CLI · effort low | 3 | range: 1.8 s–2.4 s vs 3.4 s–4.9 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.82 s to 2.41 s; GPT-6.1 Sol (Codex CLI) 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (16k prompt) | 2.36 sClaude Code | 4.14 sCodex CLI · effort low | 3 | range: 2.1 s–3.4 s vs 4 s–4.7 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 2.11 s to 3.40 s; GPT-6.1 Sol (Codex CLI) 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (64k prompt) | 2.35 sClaude Code | 3.96 sCodex CLI · effort low | 3 | range: 2.3 s–4.3 s vs 3.5 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.26 s to 4.29 s; GPT-6.1 Sol (Codex CLI) 3.47 s to 4.44 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Exact lookup answers at the 1k, 16k and 64k prompt-size targets | 56% (5/9)Claude Code | 100% (9/9)Codex CLI · effort low | 9 | 95% CI: 27%–81% vs 70%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 27% to 81%; GPT-6.1 Sol (Codex CLI) 70% to 100%), so this sample cannot separate them. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Pass rate on 4 harder tasks (Strict pass) | 42% (5/12)Claude Code | 69% (11/16)Codex CLI · effort medium | 12 / 16 | 95% CI: 19%–68% vs 44%–86% | Tie | The 95% intervals overlap (Claude Opus 5.5 19% to 68%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Pass rate on 4 harder tasks (Lenient (format misses counted)) | 50% (6/12)Claude Code | 69% (11/16)Codex CLI · effort medium | 12 / 16 | 95% CI: 25%–75% vs 44%–86% | Tie | The 95% intervals overlap (Claude Opus 5.5 25% to 75%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Calls that tried a tool although tools were off | 42% (5/12)Claude Code | 0% (0/16)Codex CLI · effort medium | 12 / 16 | 95% CI: 19%–68% vs 0%–19% | Unclear | More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 10x10 nonogram | 100% (3/3)Claude Code | 75% (3/4)Codex CLI · effort medium | 3 / 4 | 95% CI: 44%–100% vs 30%–95% | Tie | The 95% intervals overlap (Claude Opus 5.5 44% to 100%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Sudoku, 22 givens | 0% (0/3)Claude Code | 25% (1/4)Codex CLI · effort medium | 3 / 4 | 95% CI: 0%–56% vs 4.6%–70% | Tie | The 95% intervals overlap (Claude Opus 5.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 5% to 70%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 6x6 Skyscrapers | 33% (1/3)Claude Code | 100% (4/4)Codex CLI · effort medium | 3 / 4 | 95% CI: 6.2%–79% vs 51%–100% | Tie | The 95% intervals overlap (Claude Opus 5.5 6% to 79%; GPT-6.1 Sol (Codex CLI) 51% to 100%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Seeded shuffle output | 33% (1/3)Claude Code | 75% (3/4)Codex CLI · effort medium | 3 / 4 | 95% CI: 6.2%–79% vs 30%–95% | Tie | The 95% intervals overlap (Claude Opus 5.5 6% to 79%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Total time per call on harder tasks | 80.3 sClaude Code | 120.2 sCodex CLI · effort medium | 9 / 13 | range: 3.8 s–280 s vs 46.2 s–273 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Opus 5.5 3.82 s to 279.5 s; GPT-6.1 Sol (Codex CLI) 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Output tokens per call on harder tasks (Output tokens) | 8,420Claude Code | 4,994Codex CLI · effort medium | 9 / 13 | range: 279–40,044 vs 2,099–13,413 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| List-price cost per strict pass on harder tasks (calculation)Calculation | $0.59Claude Code | $0.083Codex CLI · effort medium | 12 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.59 vs $0.083, 7.2x) is not tested against run-to-run variation. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)
Claude Haiku 4.5 and GPT-6.1 Sol (Codex CLI) share 40 measured metrics and 16 list-price calculations from 7 studies. Claude Haiku 4.5 leads on 2 rows: Same prompt, 10 times: time per call (Exact number), 5.06 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 5.95 s vs 11.3 s. GPT-6.1 Sol (Codex CLI) leads on 6 rows: Pass rate on eight hard tasks (Strict pass), 100% (16/16) vs 46% (11/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); Same prompt, 10 times: strict pass rate (JSON object), 100% (10/10) vs 10% (1/10); and 3 more. On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 13 ties and 35 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).
Read all 56 metric rows for Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)
| Metric | Claude Haiku 4.5 | GPT-6.1 Sol (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 100% (15/15)Claude Code · five short validated tasks | 100% (15/15)Codex CLI · effort medium · five short validated tasks | 15 | 95% CI: 80%–100% vs 80%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 4.43 sClaude Code · five short validated tasks | 5.65 sCodex CLI · effort medium · five short validated tasks | 15 | range: 3.2 s–23.6 s vs 4.1 s–25.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 3.63 sClaude Code · five short validated tasks | 5.05 sCodex CLI · effort medium · five short validated tasks | 15 | range: 2.8 s–22.3 s vs 3.4 s–17.8 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.78 s to 22.3 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 0Claude Code · five short validated tasks | 5,180Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Other input) | 3,790Claude Code · five short validated tasks | 6,943Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 367Claude Code · five short validated tasks | 42Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per call (calculation)Calculation | $0.0057Claude Code · five short validated tasks | $0.010Codex CLI · effort medium · five short validated tasks | 15 | range: $0.0051–$0.018 vs $0.0054–$0.027 | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 $0.0051 to $0.018; GPT-6.1 Sol (Codex CLI) $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.0084Claude Code · five short validated tasks | $0.016Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.0084 vs $0.016) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Pass rate on eight hard tasks (Strict pass) | 46% (11/24)Claude Code · eight hard validated tasks | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | 95% CI: 28%–65% vs 81%–100% | GPT-6.1 Sol (Codex CLI) ahead | The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; GPT-6.1 Sol (Codex CLI) 81% to 100%). | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 67% (16/24)Claude Code · eight hard validated tasks | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | 95% CI: 47%–82% vs 81%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 47% to 82%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Total time per call on hard tasks (separate batches) | 39.0 sClaude Code · eight hard validated tasks | 13.1 sCodex CLI · effort medium · eight hard validated tasks | 24 / 16 | range: 15.3 s–75.1 s vs 8.5 s–61.6 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Time to first useful output on hard tasks | 35.5 sClaude Code · eight hard validated tasks | 10.2 sCodex CLI · effort medium · eight hard validated tasks | 24 / 16 | range: 12.9 s–70.3 s vs 6.1 s–40.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 12.9 s to 70.3 s; GPT-6.1 Sol (Codex CLI) 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call on hard tasks (Output tokens) | 5,064Claude Code · eight hard validated tasks | 335Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass on hard tasks (calculation)Calculation | $0.067Claude Code · eight hard validated tasks | $0.026Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.067 vs $0.026, 2.6x) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Same prompt, 10 times: strict pass rate (Exact number) | 0% (0/10)Claude Code · same prompt repeated 10 times | 100% (10/10)Codex CLI · effort medium · same prompt repeated 10 times | 10 | 95% CI: 0%–28% vs 72%–100% | GPT-6.1 Sol (Codex CLI) ahead | The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; GPT-6.1 Sol (Codex CLI) 72% to 100%). | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: strict pass rate (JSON object) | 10% (1/10)Claude Code · same prompt repeated 10 times | 100% (10/10)Codex CLI · effort medium · same prompt repeated 10 times | 10 | 95% CI: 1.8%–40% vs 72%–100% | GPT-6.1 Sol (Codex CLI) ahead | The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; GPT-6.1 Sol (Codex CLI) 72% to 100%). | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: strict pass rate (Code fix) | 100% (10/10)Claude Code · same prompt repeated 10 times | 100% (10/10)Codex CLI · effort medium · same prompt repeated 10 times | 10 | 95% CI: 72%–100% vs 72%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (Exact number) | 1Claude Code · same prompt repeated 10 times | 1Codex CLI · effort medium · same prompt repeated 10 times | 10 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (JSON object) | 1Claude Code · same prompt repeated 10 times | 1Codex CLI · effort medium · same prompt repeated 10 times | 10 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (Code fix) | 6Claude Code · same prompt repeated 10 times | 6Codex CLI · effort medium · same prompt repeated 10 times | 10 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (Exact number) | 5.06 sClaude Code · same prompt repeated 10 times | 13.4 sCodex CLI · effort medium · same prompt repeated 10 times | 10 | range: 4.4 s–6.2 s vs 12.3 s–18 s | Claude Haiku 4.5 ahead | The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.42 s to 6.20 s; GPT-6.1 Sol (Codex CLI) 12.3 s to 18.0 s). A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (JSON object) | 7.03 sClaude Code · same prompt repeated 10 times | 6.42 sCodex CLI · effort medium · same prompt repeated 10 times | 10 | range: 5.3 s–12.3 s vs 5.3 s–8.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 5.28 s to 12.3 s; GPT-6.1 Sol (Codex CLI) 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (Code fix) | 5.95 sClaude Code · same prompt repeated 10 times | 11.3 sCodex CLI · effort medium · same prompt repeated 10 times | 10 | range: 4.9 s–7.3 s vs 9.1 s–14.9 s | Claude Haiku 4.5 ahead | The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; GPT-6.1 Sol (Codex CLI) 9.08 s to 14.8 s). A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)Calculation | 0% (0/24)Claude Code · instructions | 100% (12/12)Codex CLI · effort low · instructions | 24 / 12 | 95% CI: 0%–14% vs 76%–100% | GPT-6.1 Sol (Codex CLI) ahead | The 95% intervals do not overlap (Claude Haiku 4.5 0% to 14%; GPT-6.1 Sol (Codex CLI) 76% to 100%). | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))Calculation | 71% (17/24)Claude Code · instructions | 100% (12/12)Codex CLI · effort low · instructions | 24 / 12 | 95% CI: 51%–85% vs 76%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 51% to 85%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Strict pass) | 0Claude Code · instructions | 12Codex CLI · effort low · instructions | 24 / 12 | none recorded | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Format miss) | 17Claude Code · instructions | 0Codex CLI · effort low · instructions | 24 / 12 | none recorded | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Wrong values) | 7Claude Code · instructions | 0Codex CLI · effort low · instructions | 24 / 12 | none recorded | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Error) | 0Claude Code · instructions | 0Codex CLI · effort low · instructions | 24 / 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Time per call, instructions vs schema mode | 9.52 sClaude Code · instructions | 6.21 sCodex CLI · effort low · instructions | 24 / 12 | range: 5.7 s–17 s vs 4.2 s–12.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 5.67 s to 17.0 s; GPT-6.1 Sol (Codex CLI) 4.20 s to 12.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call) | 1,128Claude Code · instructions | 117Codex CLI · effort low · instructions | 24 / 12 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Reasoning share of output tokens per call on hard tasks (calculation)Calculation | 91.7%Claude Code | 46.3%Codex CLI · effort medium | 24 / 16 | range: 76%–99% vs 11%–87% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation | $0.024Claude Code | $0.0023Codex CLI · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation | $0.0018Claude Code | $0.0030Codex CLI · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation | $0.0045Claude Code | $0.020Codex CLI · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation | 91.7%Claude Code | 46.3%Codex CLI · effort medium | 24 / 16 | range: 76%–99% vs 11%–87% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation | 90.2%Claude Code | 41%Codex CLI · effort medium | 15 | range: 73%–98% vs 0%–71% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Time to first text: a 250-line answer, six models | 4.00 sClaude Code | 3.52 sCodex CLI · effort low | 4 | range: 2.8 s–6.4 s vs 2.8 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 153Claude Code | 80Codex CLI · effort low | 4 | range: 153–216 vs 72–81 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 547Claude Code | 323Codex CLI · effort low | 3 / 4 | range: 546–548 vs 291–327 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 1kCalculation | 1.93 sClaude Code | 3.36 sCodex CLI · effort low | 3 | range: 1.9 s–2 s vs 3.4 s–4.8 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 1.85 s to 2.04 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 16kCalculation | 2.27 sClaude Code | 4.02 sCodex CLI · effort low | 3 | range: 2.2 s–2.5 s vs 3.3 s–4.3 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.47 s; GPT-6.1 Sol (Codex CLI) 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 64kCalculation | 2.78 sClaude Code | 3.93 sCodex CLI · effort low | 3 | range: 2.5 s–2.9 s vs 3.4 s–4.4 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.45 s to 2.89 s; GPT-6.1 Sol (Codex CLI) 3.42 s to 4.38 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (1k prompt) | 2.34 sClaude Code | 3.43 sCodex CLI · effort low | 3 | range: 2.2 s–2.5 s vs 3.4 s–4.9 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.46 s; GPT-6.1 Sol (Codex CLI) 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (16k prompt) | 2.79 sClaude Code | 4.14 sCodex CLI · effort low | 3 | range: 2.6 s–2.8 s vs 4 s–4.7 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.58 s to 2.84 s; GPT-6.1 Sol (Codex CLI) 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (64k prompt) | 3.13 sClaude Code | 3.96 sCodex CLI · effort low | 3 | range: 2.8 s–3.3 s vs 3.5 s–4.4 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.84 s to 3.28 s; GPT-6.1 Sol (Codex CLI) 3.47 s to 4.44 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Exact lookup answers at the 1k, 16k and 64k prompt-size targets | 100% (9/9)Claude Code | 100% (9/9)Codex CLI · effort low | 9 | 95% CI: 70%–100% vs 70%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 70% to 100%; GPT-6.1 Sol (Codex CLI) 70% to 100%), so this sample cannot separate them. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Pass rate on 4 harder tasks (Strict pass) | 0% (0/12)Claude Code | 69% (11/16)Codex CLI · effort medium | 12 / 16 | 95% CI: 0%–24% vs 44%–86% | GPT-6.1 Sol (Codex CLI) ahead | The 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; GPT-6.1 Sol (Codex CLI) 44% to 86%). | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Pass rate on 4 harder tasks (Lenient (format misses counted)) | 0% (0/12)Claude Code | 69% (11/16)Codex CLI · effort medium | 12 / 16 | 95% CI: 0%–24% vs 44%–86% | GPT-6.1 Sol (Codex CLI) ahead | The 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; GPT-6.1 Sol (Codex CLI) 44% to 86%). | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Calls that tried a tool although tools were off | 8% (1/12)Claude Code | 0% (0/16)Codex CLI · effort medium | 12 / 16 | 95% CI: 1.5%–35% vs 0%–19% | Unclear | More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 10x10 nonogram | 0% (0/3)Claude Code | 75% (3/4)Codex CLI · effort medium | 3 / 4 | 95% CI: 0%–56% vs 30%–95% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Sudoku, 22 givens | 0% (0/3)Claude Code | 25% (1/4)Codex CLI · effort medium | 3 / 4 | 95% CI: 0%–56% vs 4.6%–70% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 5% to 70%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 6x6 Skyscrapers | 0% (0/3)Claude Code | 100% (4/4)Codex CLI · effort medium | 3 / 4 | 95% CI: 0%–56% vs 51%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 51% to 100%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Seeded shuffle output | 0% (0/3)Claude Code | 75% (3/4)Codex CLI · effort medium | 3 / 4 | 95% CI: 0%–56% vs 30%–95% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Total time per call on harder tasks | 109.0 sClaude Code | 120.2 sCodex CLI · effort medium | 10 / 13 | range: 25.7 s–224 s vs 46.2 s–273 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 25.7 s to 223.9 s; GPT-6.1 Sol (Codex CLI) 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Output tokens per call on harder tasks (Output tokens) | 12,508Claude Code | 4,994Codex CLI · effort medium | 10 / 13 | range: 2,965–26,532 vs 2,099–13,413 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI)
Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) share 12 measured metrics and 11 list-price calculations from 4 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 20 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 4 at the smallest).
Read all 23 metric rows for Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI)
| Metric | Claude Fable 5.1 | GPT-6.1 Sol (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 100% (15/15)Claude Code · five short validated tasks | 100% (15/15)Codex CLI · effort medium · five short validated tasks | 15 | 95% CI: 80%–100% vs 80%–100% | Tie | The 95% intervals overlap (Claude Fable 5.1 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 1.94 sClaude Code · five short validated tasks | 5.65 sCodex CLI · effort medium · five short validated tasks | 15 | range: 1.4 s–9.8 s vs 4.1 s–25.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 1.41 s to 9.83 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 1.20 sClaude Code · five short validated tasks | 5.05 sCodex CLI · effort medium · five short validated tasks | 15 | range: 1 s–7.9 s vs 3.4 s–17.8 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 0.95 s to 7.90 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 2,760Claude Code · five short validated tasks | 5,180Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Other input) | 473Claude Code · five short validated tasks | 6,943Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 64Claude Code · five short validated tasks | 42Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per call (calculation)Calculation | $0.0099Claude Code · five short validated tasks | $0.010Codex CLI · effort medium · five short validated tasks | 15 | range: $0.0049–$0.058 vs $0.0054–$0.027 | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 $0.0049 to $0.058; GPT-6.1 Sol (Codex CLI) $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.021Claude Code · five short validated tasks | $0.016Codex CLI · effort medium · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.021 vs $0.016) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Pass rate on eight hard tasks (Strict pass) | 100% (24/24)Claude Code · eight hard validated tasks | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 100% (24/24)Claude Code · eight hard validated tasks | 100% (16/16)Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Total time per call on hard tasks (separate batches) | 16.1 sClaude Code · eight hard validated tasks | 13.1 sCodex CLI · effort medium · eight hard validated tasks | 24 / 16 | range: 4.5 s–90 s vs 8.5 s–61.6 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 4.46 s to 90.0 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Time to first useful output on hard tasks | 11.6 sClaude Code · eight hard validated tasks | 10.2 sCodex CLI · effort medium · eight hard validated tasks | 24 / 16 | range: 2 s–85.3 s vs 6.1 s–40.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 2.00 s to 85.3 s; GPT-6.1 Sol (Codex CLI) 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call on hard tasks (Output tokens) | 1,366Claude Code · eight hard validated tasks | 335Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass on hard tasks (calculation)Calculation | $0.093Claude Code · eight hard validated tasks | $0.026Codex CLI · effort medium · eight hard validated tasks | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.093 vs $0.026, 3.6x) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Reasoning share of output tokens per call on hard tasks (calculation)Calculation | 64.2%Claude Code | 46.3%Codex CLI · effort medium | 24 / 16 | range: 23%–97% vs 11%–87% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation | $0.054Claude Code | $0.0023Codex CLI · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation | $0.019Claude Code | $0.0030Codex CLI · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation | $0.021Claude Code | $0.020Codex CLI · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation | 64.2%Claude Code | 46.3%Codex CLI · effort medium | 24 / 16 | range: 23%–97% vs 11%–87% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation | 0%Claude Code | 41%Codex CLI · effort medium | 15 | range: 0%–74% vs 0%–71% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Time to first text: a 250-line answer, six models | 4.43 sClaude Code | 3.52 sCodex CLI · effort low | 4 | range: 2.3 s–4.6 s vs 2.8 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Fable 5.1 2.27 s to 4.64 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 123Claude Code | 80Codex CLI · effort low | 4 | range: 121–131 vs 72–81 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 273Claude Code | 323Codex CLI · effort low | 4 | range: 270–293 vs 291–327 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
Claude Haiku 4.5 vs GPT-6 Luna (Codex CLI)
Claude Haiku 4.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. GPT-6 Luna (Codex CLI) leads on 1 row: Total time per attempt: single call vs agent loop, 5.16 s vs 39.0 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).
Read all 17 metric rows for Claude Haiku 4.5 vs GPT-6 Luna (Codex CLI)
| Metric | Claude Haiku 4.5 | GPT-6 Luna (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Strict pass rate: single call vs agent loop on eight hard tasks | 46% (11/24)Claude Code · single call | 63% (10/16)Codex CLI · single call | 24 / 16 | 95% CI: 28%–65% vs 39%–82% | Tie | The 95% intervals overlap (Claude Haiku 4.5 28% to 65%; GPT-6 Luna (Codex CLI) 39% to 82%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Interval merge fix | 100% (3/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: DST day-length fix | 33% (1/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 6.2%–79% vs 34%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 6% to 79%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: CSV parser | 67% (2/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 21%–94% vs 34%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Event-loop order | 0% (0/3)Claude Code · single call | 0% (0/2)Codex CLI · single call | 3 / 2 | 95% CI: 0%–56% vs 0%–66% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Room schedule | 0% (0/3)Claude Code · single call | 50% (1/2)Codex CLI · single call | 3 / 2 | 95% CI: 0%–56% vs 9.4%–91% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SemVer regex | 100% (3/3)Claude Code · single call | 50% (1/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 9.4%–91% | Tie | The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Money refactor | 67% (2/3)Claude Code · single call | 0% (0/2)Codex CLI · single call | 3 / 2 | 95% CI: 21%–94% vs 0%–66% | Tie | The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SQL report | 0% (0/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 0%–56% vs 34%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Total time per attempt: single call vs agent loop | 39.0 sClaude Code · single call | 5.16 sCodex CLI · single call | 24 / 16 | range: 15.3 s–75.1 s vs 3.6 s–11.3 s | GPT-6 Luna (Codex CLI) ahead | The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 15.3 s to 75.1 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s). A range is not a confidence interval. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Input tokens (cache reads included)) | 3,941Claude Code · single call | 11,582Codex CLI · single call | 24 / 16 | range: 3,879–4,221 vs 11,526–11,818 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Output tokens) | 5,064Claude Code · single call | 345Codex CLI · single call | 24 / 16 | range: 1,899–9,321 vs 36–634 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tool calls per agent-loop attempt | 3Claude Code · agent loop | 0Codex CLI · agent loop | 24 / 14 | range: 2–18 vs 0–1 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| List-price cost per strict pass: single call vs agent loop (calculation)Calculation | $0.067Claude Code · single call | $0.0012Codex CLI · single call | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.067 vs $0.0012, 58x) is not tested against run-to-run variation. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Time to first text: a 250-line answer, six models | 4.00 sClaude Code | 3.30 sCodex CLI · effort low | 4 | range: 2.8 s–6.4 s vs 3.2 s–3.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 153Claude Code | 129Codex CLI · effort low | 4 | range: 153–216 vs 56–259 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 547Claude Code | 524Codex CLI · effort low | 3 / 4 | range: 546–548 vs 225–1,052 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
Claude Sonnet 5.5 vs GPT-6 Luna (Codex CLI)
Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. Claude Sonnet 5.5 leads on 1 row: Strict pass rate: single call vs agent loop on eight hard tasks, 100% (24/24) vs 63% (10/16). On those rows the 95% intervals do not overlap. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).
Read all 17 metric rows for Claude Sonnet 5.5 vs GPT-6 Luna (Codex CLI)
| Metric | Claude Sonnet 5.5 | GPT-6 Luna (Codex CLI) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Strict pass rate: single call vs agent loop on eight hard tasks | 100% (24/24)Claude Code · single call | 63% (10/16)Codex CLI · single call | 24 / 16 | 95% CI: 86%–100% vs 39%–82% | Claude Sonnet 5.5 ahead | The 95% intervals do not overlap (Claude Sonnet 5.5 86% to 100%; GPT-6 Luna (Codex CLI) 39% to 82%). | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Interval merge fix | 100% (3/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: DST day-length fix | 100% (3/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: CSV parser | 100% (3/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Event-loop order | 100% (3/3)Claude Code · single call | 0% (0/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 0%–66% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Room schedule | 100% (3/3)Claude Code · single call | 50% (1/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 9.4%–91% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SemVer regex | 100% (3/3)Claude Code · single call | 50% (1/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 9.4%–91% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Money refactor | 100% (3/3)Claude Code · single call | 0% (0/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 0%–66% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SQL report | 100% (3/3)Claude Code · single call | 100% (2/2)Codex CLI · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Total time per attempt: single call vs agent loop | 7.75 sClaude Code · single call | 5.16 sCodex CLI · single call | 24 / 16 | range: 2.3 s–34.8 s vs 3.6 s–11.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Input tokens (cache reads included)) | 2,281Claude Code · single call | 11,582Codex CLI · single call | 24 / 16 | range: 2,234–2,669 vs 11,526–11,818 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Output tokens) | 1,050Claude Code · single call | 345Codex CLI · single call | 24 / 16 | range: 176–3,895 vs 36–634 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tool calls per agent-loop attempt | 0Claude Code · agent loop | 0Codex CLI · agent loop | 16 / 14 | range: 0–3 vs 0–1 | Tie | Same value. More or fewer is not better by itself for this metric. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| List-price cost per strict pass: single call vs agent loop (calculation)Calculation | $0.014Claude Code · single call | $0.0012Codex CLI · single call | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.014 vs $0.0012, 12x) is not tested against run-to-run variation. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Time to first text: a 250-line answer, six models | 1.96 sClaude Code | 3.30 sCodex CLI · effort low | 4 | range: 0.9 s–4.1 s vs 3.2 s–3.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.88 s to 4.09 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 232Claude Code | 129Codex CLI · effort low | 4 | range: 230–233 vs 56–259 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 517Claude Code | 524Codex CLI · effort low | 4 | range: 513–519 vs 225–1,052 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
Claude Haiku 4.5 vs GPT 5.2
Claude Haiku 4.5 and GPT 5.2 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.
Read all 3 metric rows for Claude Haiku 4.5 vs GPT 5.2
| Metric | Claude Haiku 4.5 | GPT 5.2 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 76% (25/33)effort high · public mini-SWE-agent v2 run, same instances | 85% (28/33)effort high · public mini-SWE-agent v2 run, same instances | 33 | 95% CI: 59%–87% vs 69%–93% | Tie | The 95% intervals overlap (Claude Haiku 4.5 59% to 87%; GPT 5.2 69% to 93%), so this sample cannot separate them. | Agent on SWE-bench Verified vs 11 public models |
| Model calls per instance | 68.5effort high · public mini-SWE-agent v2 run, same instances | 35.6effort high · public mini-SWE-agent v2 run, same instances | 33 | none recorded | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Agent on SWE-bench Verified vs 11 public models |
| Recorded cost per resolved instance: Agent vs the public panel | $0.48effort high · public mini-SWE-agent v2 run, same instances | $0.63effort high · public mini-SWE-agent v2 run, same instances | 25 / 28 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.48 vs $0.63) is not tested against run-to-run variation. | What if every call ran on Opus? Repricing real agent tokens |
Claude Haiku 4.5 vs GPT 5 mini
Claude Haiku 4.5 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.
Read all 3 metric rows for Claude Haiku 4.5 vs GPT 5 mini
| Metric | Claude Haiku 4.5 | GPT 5 mini | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 76% (25/33)effort high · public mini-SWE-agent v2 run, same instances | 64% (21/33)public mini-SWE-agent v2 run, same instances | 33 | 95% CI: 59%–87% vs 47%–78% | Tie | The 95% intervals overlap (Claude Haiku 4.5 59% to 87%; GPT 5 mini 47% to 78%), so this sample cannot separate them. | Agent on SWE-bench Verified vs 11 public models |
| Model calls per instance | 68.5effort high · public mini-SWE-agent v2 run, same instances | 20.8public mini-SWE-agent v2 run, same instances | 33 | none recorded | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Agent on SWE-bench Verified vs 11 public models |
| Recorded cost per resolved instance: Agent vs the public panel | $0.48effort high · public mini-SWE-agent v2 run, same instances | $0.080public mini-SWE-agent v2 run, same instances | 25 / 21 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.48 vs $0.080, 6.0x) is not tested against run-to-run variation. | What if every call ran on Opus? Repricing real agent tokens |
Claude Opus 4.5 vs GPT 5 mini
Claude Opus 4.5 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.
Read all 3 metric rows for Claude Opus 4.5 vs GPT 5 mini
| Metric | Claude Opus 4.5 | GPT 5 mini | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 73% (24/33)effort high · public mini-SWE-agent v2 run, same instances | 64% (21/33)public mini-SWE-agent v2 run, same instances | 33 | 95% CI: 56%–85% vs 47%–78% | Tie | The 95% intervals overlap (Claude Opus 4.5 56% to 85%; GPT 5 mini 47% to 78%), so this sample cannot separate them. | Agent on SWE-bench Verified vs 11 public models |
| Model calls per instance | 35.9effort high · public mini-SWE-agent v2 run, same instances | 20.8public mini-SWE-agent v2 run, same instances | 33 | none recorded | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Agent on SWE-bench Verified vs 11 public models |
| Recorded cost per resolved instance: Agent vs the public panel | $1.18effort high · public mini-SWE-agent v2 run, same instances | $0.080public mini-SWE-agent v2 run, same instances | 24 / 21 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($1.18 vs $0.080, 15x) is not tested against run-to-run variation. | What if every call ran on Opus? Repricing real agent tokens |
Claude Opus 4.6 vs GPT 5 mini
Claude Opus 4.6 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.
Read all 3 metric rows for Claude Opus 4.6 vs GPT 5 mini
| Metric | Claude Opus 4.6 | GPT 5 mini | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 70% (23/33)public mini-SWE-agent v2 run, same instances | 64% (21/33)public mini-SWE-agent v2 run, same instances | 33 | 95% CI: 53%–83% vs 47%–78% | Tie | The 95% intervals overlap (Claude Opus 4.6 53% to 83%; GPT 5 mini 47% to 78%), so this sample cannot separate them. | Agent on SWE-bench Verified vs 11 public models |
| Model calls per instance | 28.9public mini-SWE-agent v2 run, same instances | 20.8public mini-SWE-agent v2 run, same instances | 33 | none recorded | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Agent on SWE-bench Verified vs 11 public models |
| Recorded cost per resolved instance: Agent vs the public panel | $0.88public mini-SWE-agent v2 run, same instances | $0.080public mini-SWE-agent v2 run, same instances | 23 / 21 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.88 vs $0.080, 11x) is not tested against run-to-run variation. | What if every call ran on Opus? Repricing real agent tokens |
Claude Sonnet 4.5 vs GPT 5 mini
Claude Sonnet 4.5 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.
Read all 3 metric rows for Claude Sonnet 4.5 vs GPT 5 mini
| Metric | Claude Sonnet 4.5 | GPT 5 mini | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 76% (25/33)effort high · public mini-SWE-agent v2 run, same instances | 64% (21/33)public mini-SWE-agent v2 run, same instances | 33 | 95% CI: 59%–87% vs 47%–78% | Tie | The 95% intervals overlap (Claude Sonnet 4.5 59% to 87%; GPT 5 mini 47% to 78%), so this sample cannot separate them. | Agent on SWE-bench Verified vs 11 public models |
| Model calls per instance | 51effort high · public mini-SWE-agent v2 run, same instances | 20.8public mini-SWE-agent v2 run, same instances | 33 | none recorded | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Agent on SWE-bench Verified vs 11 public models |
| Recorded cost per resolved instance: Agent vs the public panel | $0.91effort high · public mini-SWE-agent v2 run, same instances | $0.080public mini-SWE-agent v2 run, same instances | 25 / 21 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.91 vs $0.080, 11x) is not tested against run-to-run variation. | What if every call ran on Opus? Repricing real agent tokens |
Claude Sonnet 5.5 vs GPT-6.1 Sol (OpenAI API)
Claude Sonnet 5.5 and GPT-6.1 Sol (OpenAI API) share 3 measured metrics from one study. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 unclear; each row says why. Every row ran the two sides through different routes (for example Claude Code vs OpenAI API), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).
Read all 3 metric rows for Claude Sonnet 5.5 vs GPT-6.1 Sol (OpenAI API)
| Metric | Claude Sonnet 5.5 | GPT-6.1 Sol (OpenAI API) | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Repairing a scheduler: Claude Code vs Codex vs API (Total time) | 15.0 sClaude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs | 17.3 sOpenAI API · effort medium · scheduler repair, 296 checks, 3 runs | 3 | range: 13.9 s–15.9 s vs 16.3 s–18.6 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 13.9 s to 15.9 s; GPT-6.1 Sol (OpenAI API) 16.3 s to 18.6 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Repairing a scheduler: Claude Code vs Codex vs API (First useful output) | 7.55 sClaude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs | 7.46 sOpenAI API · effort medium · scheduler repair, 296 checks, 3 runs | 3 | range: 6.8 s–7.6 s vs 6.7 s–9.1 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 6.77 s to 7.63 s; GPT-6.1 Sol (OpenAI API) 6.68 s to 9.05 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Output tokens to repair the scheduler (Output tokens) | 2,227Claude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs | 1,313OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs | 3 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
GPT 5.2 vs Claude Opus 4.5
GPT 5.2 and Claude Opus 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.
Read all 3 metric rows for GPT 5.2 vs Claude Opus 4.5
| Metric | GPT 5.2 | Claude Opus 4.5 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 85% (28/33)effort high · public mini-SWE-agent v2 run, same instances | 73% (24/33)effort high · public mini-SWE-agent v2 run, same instances | 33 | 95% CI: 69%–93% vs 56%–85% | Tie | The 95% intervals overlap (GPT 5.2 69% to 93%; Claude Opus 4.5 56% to 85%), so this sample cannot separate them. | Agent on SWE-bench Verified vs 11 public models |
| Model calls per instance | 35.6effort high · public mini-SWE-agent v2 run, same instances | 35.9effort high · public mini-SWE-agent v2 run, same instances | 33 | none recorded | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Agent on SWE-bench Verified vs 11 public models |
| Recorded cost per resolved instance: Agent vs the public panel | $0.63effort high · public mini-SWE-agent v2 run, same instances | $1.18effort high · public mini-SWE-agent v2 run, same instances | 28 / 24 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.63 vs $1.18) is not tested against run-to-run variation. | What if every call ran on Opus? Repricing real agent tokens |
GPT 5.2 vs Claude Opus 4.6
GPT 5.2 and Claude Opus 4.6 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.
Read all 3 metric rows for GPT 5.2 vs Claude Opus 4.6
| Metric | GPT 5.2 | Claude Opus 4.6 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 85% (28/33)effort high · public mini-SWE-agent v2 run, same instances | 70% (23/33)public mini-SWE-agent v2 run, same instances | 33 | 95% CI: 69%–93% vs 53%–83% | Tie | The 95% intervals overlap (GPT 5.2 69% to 93%; Claude Opus 4.6 53% to 83%), so this sample cannot separate them. | Agent on SWE-bench Verified vs 11 public models |
| Model calls per instance | 35.6effort high · public mini-SWE-agent v2 run, same instances | 28.9public mini-SWE-agent v2 run, same instances | 33 | none recorded | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Agent on SWE-bench Verified vs 11 public models |
| Recorded cost per resolved instance: Agent vs the public panel | $0.63effort high · public mini-SWE-agent v2 run, same instances | $0.88public mini-SWE-agent v2 run, same instances | 28 / 23 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.63 vs $0.88) is not tested against run-to-run variation. | What if every call ran on Opus? Repricing real agent tokens |
GPT 5.2 vs Claude Sonnet 4.5
GPT 5.2 and Claude Sonnet 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.
Read all 3 metric rows for GPT 5.2 vs Claude Sonnet 4.5
| Metric | GPT 5.2 | Claude Sonnet 4.5 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 85% (28/33)effort high · public mini-SWE-agent v2 run, same instances | 76% (25/33)effort high · public mini-SWE-agent v2 run, same instances | 33 | 95% CI: 69%–93% vs 59%–87% | Tie | The 95% intervals overlap (GPT 5.2 69% to 93%; Claude Sonnet 4.5 59% to 87%), so this sample cannot separate them. | Agent on SWE-bench Verified vs 11 public models |
| Model calls per instance | 35.6effort high · public mini-SWE-agent v2 run, same instances | 51effort high · public mini-SWE-agent v2 run, same instances | 33 | none recorded | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Agent on SWE-bench Verified vs 11 public models |
| Recorded cost per resolved instance: Agent vs the public panel | $0.63effort high · public mini-SWE-agent v2 run, same instances | $0.91effort high · public mini-SWE-agent v2 run, same instances | 28 / 25 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.63 vs $0.91) is not tested against run-to-run variation. | What if every call ran on Opus? Repricing real agent tokens |