16 measured pairs · updated October 7, 2026

Claude vs GPT, on the tasks we measured.

Read every recorded Anthropic and OpenAI head-to-head in this dataset. A model, its route and a coding CLI are distinct systems. These runs do not measure the ChatGPT app.

Each row keeps its study, task set, configuration and sample size. Intervals and ranges stay labelled; ties and unclear results stay unchanged. List-price calculations are not vendor bills. There is no combined rank across studies.

Browse all model, CLI and router comparisons

Live story · 42 sClaude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI): what the measurements say

Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI): what the measurements say

70 comparison rows from 10 studies: 4 rows favour Sonnet 5.5, 1 favour GPT-6.1 Sol (Codex CLI), 65 are ties or unclear. Cost rows are calculations.

Transcript
  1. Comparison · 70 rows · 10 studies. Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI). A winner only where the 95% intervals or run ranges do not overlap.
  2. 70 comparison rows from 10 studies: Sonnet 5.5 ahead on 4, GPT-6.1 Sol (Codex CLI) ahead on 1. The rest do not separate them. Rows where Sonnet 5.5 is ahead: 4 (of 70). Rows where GPT-6.1 Sol (Codex CLI) is ahead: 1 (of 70). Ties or unclear: 65 (22 ties · 43 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
  3. Coding agents, hidden tests: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 4 rows separates them. Table: Coding agents, hidden tests · Claude Code vs Codex CLI · n = 12 per side. Source study: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks. Rows shown: Coding sessions that passed every hidden check; Time per coding session; Tool calls per coding session; List-price cost per passing coding session (calculation). Recorded settings: Claude Code · six small repository tasks with hidden tests; Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests. Includes a calculation, not a bill or a new run. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
  4. Caching sessions: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 2 of 9 rows separate them. Table: Caching sessions · 5 of 9 rows · Claude Code vs Codex CLI · n = 10 per side. Source study: Prompt caching and run-to-run consistency in Claude Code and Codex CLI. Rows shown: Same prompt, 10 times: strict pass rate (Exact number); Same prompt, 10 times: strict pass rate (JSON object); Same prompt, 10 times: strict pass rate (Code fix); Same prompt, 10 times: time per call (Exact number); Same prompt, 10 times: time per call (Code fix). Recorded settings: Claude Code · same prompt repeated 10 times; Codex CLI · effort medium · same prompt repeated 10 times. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  5. Instructions vs JSON schema: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 8 rows separates them. Table: Instructions vs JSON schema · 5 of 8 rows · Claude Code vs Codex CLI · n = 12 per side. Source study: Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI. Rows shown: Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON); What each call produced: strict pass, format miss, wrong values or error (Strict pass); What each call produced: strict pass, format miss, wrong values or error (Format miss); What each call produced: strict pass, format miss, wrong values or error (Wrong values); Time per call, instructions vs schema mode. Recorded settings: Claude Code · instructions; Codex CLI · effort low · instructions. Includes a calculation, not a bill or a new run. Caveat: The calls repeat only three fixed prompts. Wilson intervals describe call outcomes under a binomial assumption; they do not measure accuracy across unseen tasks. The paired p-values also assume independent pairs and do not remove this limit.
  6. Four harder tasks: pass rate 38% vs 69%, tie: 95% intervals overlap. 1 of 10 rows separates them. Table: Four harder tasks · 5 of 10 rows · Claude Code vs Codex CLI · n = 4–16 per side. Source study: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Rows shown: Pass rate on 4 harder tasks (Strict pass); Pass rate on 4 harder tasks (Lenient (format misses counted)); Calls that tried a tool although tools were off; Strict pass rate by task: 10x10 nonogram; Strict pass rate by task: 6x6 Skyscrapers. Recorded settings: Claude Code; Codex CLI · effort medium. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.
  7. No winner where the data shows none. Showing 19 of 70 rows; every row and its reason online.

Measured pairs

16 of 16 measured pairs

Coding CLI vs coding cli

Claude Code vs Codex CLI

Claude Code and Codex CLI share 66 measured metrics and 22 list-price calculations from 12 studies. Claude Code leads on 7 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 4 more. Codex CLI leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 31 ties and 49 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Each run pairs a CLI with a model, so these rows cannot separate the CLI from the model; the contexts name both. Some rows rest on small samples (n = 2 at the smallest).

66 measured rows · 22 labelled calculations

Read all 88 metric rows for Claude Code vs Codex CLI
MetricClaude CodeCodex CLInInterval or rangeOutcomeBasisStudy
Pass rate on five validated tasks80% (12/15)Claude Sonnet 5.5 · five short validated tasks100% (15/15)GPT-6.1 Sol · effort medium · five short validated tasks1595% CI: 55%–93% vs 80%–100%TieThe 95% intervals overlap (Claude Code 55% to 93%; Codex CLI 80% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Total time per call2.31 sClaude Sonnet 5.5 · five short validated tasks5.65 sGPT-6.1 Sol · effort medium · five short validated tasks15range: 2.2 s–7.7 s vs 4.1 s–25.5 sUnclearThe run ranges (fastest to slowest) overlap (Claude Code 2.17 s to 7.73 s; Codex CLI 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Time to first useful output1.56 sClaude Sonnet 5.5 · five short validated tasks5.05 sGPT-6.1 Sol · effort medium · five short validated tasks15range: 1 s–6.4 s vs 3.4 s–17.8 sUnclearThe run ranges (fastest to slowest) overlap (Claude Code 0.99 s to 6.39 s; Codex CLI 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Input tokens per call: what the CLI sends (Cache read)1,401Claude Sonnet 5.5 · five short validated tasks5,180GPT-6.1 Sol · effort medium · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Input tokens per call: what the CLI sends (Other input)685Claude Sonnet 5.5 · five short validated tasks6,943GPT-6.1 Sol · effort medium · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Output tokens per call (Output tokens)107Claude Sonnet 5.5 · five short validated tasks42GPT-6.1 Sol · effort medium · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
List-price cost per call (calculation)Calculation$0.0036Claude Sonnet 5.5 · five short validated tasks$0.010GPT-6.1 Sol · effort medium · five short validated tasks15range: $0.0034–$0.01 vs $0.0054–$0.027UnclearThe run ranges (fastest to slowest) overlap (Claude Code $0.0034 to $0.010; Codex CLI $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
List-price cost per passing answer (calculation)Calculation$0.0062Claude Sonnet 5.5 · five short validated tasks$0.016GPT-6.1 Sol · effort medium · five short validated tasks15none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.0062 vs $0.016, 2.5x) is not tested against run-to-run variation.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Pass rate on eight hard tasks (Strict pass)100% (24/24)Claude Sonnet 5.5 · eight hard validated tasks100% (16/16)GPT-6.1 Sol · effort medium · eight hard validated tasks24 / 1695% CI: 86%–100% vs 81%–100%TieThe 95% intervals overlap (Claude Code 86% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Pass rate on eight hard tasks (Lenient (format misses counted))100% (24/24)Claude Sonnet 5.5 · eight hard validated tasks100% (16/16)GPT-6.1 Sol · effort medium · eight hard validated tasks24 / 1695% CI: 86%–100% vs 81%–100%TieThe 95% intervals overlap (Claude Code 86% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Total time per call on hard tasks (separate batches)7.75 sClaude Sonnet 5.5 · eight hard validated tasks13.1 sGPT-6.1 Sol · effort medium · eight hard validated tasks24 / 16range: 2.3 s–34.8 s vs 8.5 s–61.6 sUnclearThe run ranges (fastest to slowest) overlap (Claude Code 2.26 s to 34.8 s; Codex CLI 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Time to first useful output on hard tasks5.95 sClaude Sonnet 5.5 · eight hard validated tasks10.2 sGPT-6.1 Sol · effort medium · eight hard validated tasks24 / 16range: 0.9 s–30.6 s vs 6.1 s–40.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Code 0.86 s to 30.6 s; Codex CLI 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Output tokens per call on hard tasks (Output tokens)1,050Claude Sonnet 5.5 · eight hard validated tasks335GPT-6.1 Sol · effort medium · eight hard validated tasks24 / 16none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
List-price cost per strict pass on hard tasks (calculation)Calculation$0.014Claude Sonnet 5.5 · eight hard validated tasks$0.026GPT-6.1 Sol · effort medium · eight hard validated tasks24 / 16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Coding sessions that passed every hidden check100% (12/12)Claude Sonnet 5.5 · six small repository tasks with hidden tests100% (12/12)GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests1295% CI: 76%–100% vs 76%–100%TieThe 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them.Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
Time per coding session23.1 sClaude Sonnet 5.5 · six small repository tasks with hidden tests113.4 sGPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests12range: 18.7 s–44.5 s vs 78.5 s–222 sClaude Code aheadThe run ranges (fastest to slowest) do not overlap (Claude Code 18.7 s to 44.5 s; Codex CLI 78.5 s to 221.9 s). A range is not a confidence interval.Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
Tool calls per coding session7.5Claude Sonnet 5.5 · six small repository tasks with hidden tests12.5GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests12range: 3–14 vs 8–18UnclearMore or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
List-price cost per passing coding session (calculation)Calculation$0.085Claude Sonnet 5.5 · six small repository tasks with hidden tests$0.098GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests12none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.085 vs $0.098) is not tested against run-to-run variation.Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
Strict pass rate by effort on eight hard tasks100% (16/16)Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder100% (16/16)GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder1695% CI: 81%–100% vs 81%–100%TieThe 95% intervals overlap (Claude Code 81% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them.Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
Total time per call by effort on hard tasks7.63 sClaude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder13.1 sGPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder16range: 2.7 s–24 s vs 8.5 s–61.6 sUnclearThe run ranges (fastest to slowest) overlap (Claude Code 2.71 s to 24.0 s; Codex CLI 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
Output tokens per call by effort on hard tasks (Output tokens)770Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder335GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder16none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
List-price cost per strict pass by effort (calculation)Calculation$0.014Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder$0.026GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
Same prompt, 10 times: strict pass rate (Exact number)100% (10/10)Claude Sonnet 5.5 · same prompt repeated 10 times100% (10/10)GPT-6.1 Sol · effort medium · same prompt repeated 10 times1095% CI: 72%–100% vs 72%–100%TieThe 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: strict pass rate (JSON object)100% (10/10)Claude Sonnet 5.5 · same prompt repeated 10 times100% (10/10)GPT-6.1 Sol · effort medium · same prompt repeated 10 times1095% CI: 72%–100% vs 72%–100%TieThe 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: strict pass rate (Code fix)100% (10/10)Claude Sonnet 5.5 · same prompt repeated 10 times100% (10/10)GPT-6.1 Sol · effort medium · same prompt repeated 10 times1095% CI: 72%–100% vs 72%–100%TieThe 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: how many different answers (Exact number)1Claude Sonnet 5.5 · same prompt repeated 10 times1GPT-6.1 Sol · effort medium · same prompt repeated 10 times10none recordedTieSame value. More or fewer is not better by itself for this metric.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: how many different answers (JSON object)1Claude Sonnet 5.5 · same prompt repeated 10 times1GPT-6.1 Sol · effort medium · same prompt repeated 10 times10none recordedTieSame value. More or fewer is not better by itself for this metric.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: how many different answers (Code fix)3Claude Sonnet 5.5 · same prompt repeated 10 times6GPT-6.1 Sol · effort medium · same prompt repeated 10 times10none recordedUnclearMore or fewer count is not better or worse by itself; this row describes behaviour, not a winner.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: time per call (Exact number)6.89 sClaude Sonnet 5.5 · same prompt repeated 10 times13.4 sGPT-6.1 Sol · effort medium · same prompt repeated 10 times10range: 5.8 s–7.8 s vs 12.3 s–18 sClaude Code aheadThe run ranges (fastest to slowest) do not overlap (Claude Code 5.81 s to 7.81 s; Codex CLI 12.3 s to 18.0 s). A range is not a confidence interval.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: time per call (JSON object)2.89 sClaude Sonnet 5.5 · same prompt repeated 10 times6.42 sGPT-6.1 Sol · effort medium · same prompt repeated 10 times10range: 2.7 s–5.3 s vs 5.3 s–8.3 sUnclearThe run ranges (fastest to slowest) overlap (Claude Code 2.68 s to 5.30 s; Codex CLI 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: time per call (Code fix)2.67 sClaude Sonnet 5.5 · same prompt repeated 10 times11.3 sGPT-6.1 Sol · effort medium · same prompt repeated 10 times10range: 2.3 s–4.3 s vs 9.1 s–14.9 sClaude Code aheadThe run ranges (fastest to slowest) do not overlap (Claude Code 2.32 s to 4.34 s; Codex CLI 9.08 s to 14.8 s). A range is not a confidence interval.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
CLI start-up tax on a one-word answer (First output event)563 msClaude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs489 msdefault model · CLI start-up, one-word prompt, 5 runs5range: 519 ms–726 ms vs 354 ms–1.3 sUnclearThe run ranges (fastest to slowest) overlap (Claude Code 519 ms to 726 ms; Codex CLI 354 ms to 1,304 ms); the medians alone do not show a reliable difference. A range is not a confidence interval.Routing overhead: deterministic policy vs LLM routers vs Jev
CLI start-up tax on a one-word answer (First model output)1,461 msClaude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs5,059 msdefault model · CLI start-up, one-word prompt, 5 runs5range: 1.21 s–2.31 s vs 4.39 s–5.48 sClaude Code aheadThe run ranges (fastest to slowest) do not overlap (Claude Code 1,206 ms to 2,308 ms; Codex CLI 4,391 ms to 5,478 ms). A range is not a confidence interval. Samples are small (5 runs per side).Routing overhead: deterministic policy vs LLM routers vs Jev
CLI start-up tax on a one-word answer (Total wall time)2,529 msClaude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs5,999 msdefault model · CLI start-up, one-word prompt, 5 runs5range: 2.27 s–3.38 s vs 5.37 s–6.51 sClaude Code aheadThe run ranges (fastest to slowest) do not overlap (Claude Code 2,273 ms to 3,382 ms; Codex CLI 5,367 ms to 6,506 ms). A range is not a confidence interval. Samples are small (5 runs per side).Routing overhead: deterministic policy vs LLM routers vs Jev
Input tokens a CLI sends for a one-word answer6,761Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs17,051default model · CLI start-up, one-word prompt, 5 runs5none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Routing overhead: deterministic policy vs LLM routers vs Jev
Repairing a scheduler: Claude Code vs Codex vs API (Total time)15.0 sSonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs61.2 sGPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs3range: 13.9 s–15.9 s vs 59.9 s–69.5 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 13.9 s to 15.9 s; Codex CLI 59.9 s to 69.5 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Claude Code CLI vs Codex CLI vs the API: latency and tokens
Repairing a scheduler: Claude Code vs Codex vs API (First useful output)7.55 sSonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs15.6 sGPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs3range: 6.8 s–7.6 s vs 13.7 s–23 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 6.77 s to 7.63 s; Codex CLI 13.7 s to 23.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Claude Code CLI vs Codex CLI vs the API: latency and tokens
Output tokens to repair the scheduler (Output tokens)2,227Sonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs1,181GPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs3none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Claude Code CLI vs Codex CLI vs the API: latency and tokens
Strict pass rate: single call vs agent loop on eight hard tasks100% (24/24)Claude Sonnet 5.5 · single call63% (10/16)GPT-6 Luna · single call24 / 1695% CI: 86%–100% vs 39%–82%Claude Code aheadThe 95% intervals do not overlap (Claude Code 86% to 100%; Codex CLI 39% to 82%).Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: Interval merge fix100% (3/3)Claude Sonnet 5.5 · single call100% (2/2)GPT-6 Luna · single call3 / 295% CI: 44%–100% vs 34%–100%TieThe 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: DST day-length fix100% (3/3)Claude Sonnet 5.5 · single call100% (2/2)GPT-6 Luna · single call3 / 295% CI: 44%–100% vs 34%–100%TieThe 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: CSV parser100% (3/3)Claude Sonnet 5.5 · single call100% (2/2)GPT-6 Luna · single call3 / 295% CI: 44%–100% vs 34%–100%TieThe 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: Event-loop order100% (3/3)Claude Sonnet 5.5 · single call0% (0/2)GPT-6 Luna · single call3 / 295% CI: 44%–100% vs 0%–66%TieThe 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 0% to 66%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: Room schedule100% (3/3)Claude Sonnet 5.5 · single call50% (1/2)GPT-6 Luna · single call3 / 295% CI: 44%–100% vs 9.4%–91%TieThe 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 9% to 91%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: SemVer regex100% (3/3)Claude Sonnet 5.5 · single call50% (1/2)GPT-6 Luna · single call3 / 295% CI: 44%–100% vs 9.4%–91%TieThe 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 9% to 91%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: Money refactor100% (3/3)Claude Sonnet 5.5 · single call0% (0/2)GPT-6 Luna · single call3 / 295% CI: 44%–100% vs 0%–66%TieThe 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 0% to 66%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: SQL report100% (3/3)Claude Sonnet 5.5 · single call100% (2/2)GPT-6 Luna · single call3 / 295% CI: 44%–100% vs 34%–100%TieThe 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Total time per attempt: single call vs agent loop7.75 sClaude Sonnet 5.5 · single call5.16 sGPT-6 Luna · single call24 / 16range: 2.3 s–34.8 s vs 3.6 s–11.3 sUnclearThe run ranges (fastest to slowest) overlap (Claude Code 2.26 s to 34.8 s; Codex CLI 3.59 s to 11.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))2,281Claude Sonnet 5.5 · single call11,582GPT-6 Luna · single call24 / 16range: 2,234–2,669 vs 11,526–11,818UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Tokens per attempt: single call vs agent loop (Output tokens)1,050Claude Sonnet 5.5 · single call345GPT-6 Luna · single call24 / 16range: 176–3,895 vs 36–634UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Tool calls per agent-loop attempt0Claude Sonnet 5.5 · agent loop0GPT-6 Luna · agent loop16 / 14range: 0–3 vs 0–1TieSame value. More or fewer is not better by itself for this metric.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
List-price cost per strict pass: single call vs agent loop (calculation)Calculation$0.014Claude Sonnet 5.5 · single call$0.0012GPT-6 Luna · single call24 / 16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.014 vs $0.0012, 12x) is not tested against run-to-run variation.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)Calculation100% (12/12)Claude Sonnet 5.5 · instructions100% (12/12)GPT-6.1 Sol · effort low · instructions1295% CI: 76%–100% vs 76%–100%TieThe 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))Calculation100% (12/12)Claude Sonnet 5.5 · instructions100% (12/12)GPT-6.1 Sol · effort low · instructions1295% CI: 76%–100% vs 76%–100%TieThe 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
What each call produced: strict pass, format miss, wrong values or error (Strict pass)12Claude Sonnet 5.5 · instructions12GPT-6.1 Sol · effort low · instructions12none recordedTieSame value. More or fewer is not better by itself for this metric.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
What each call produced: strict pass, format miss, wrong values or error (Format miss)0Claude Sonnet 5.5 · instructions0GPT-6.1 Sol · effort low · instructions12none recordedTieSame value. More or fewer is not better by itself for this metric.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
What each call produced: strict pass, format miss, wrong values or error (Wrong values)0Claude Sonnet 5.5 · instructions0GPT-6.1 Sol · effort low · instructions12none recordedTieSame value. More or fewer is not better by itself for this metric.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
What each call produced: strict pass, format miss, wrong values or error (Error)0Claude Sonnet 5.5 · instructions0GPT-6.1 Sol · effort low · instructions12none recordedTieSame value. More or fewer is not better by itself for this metric.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
Time per call, instructions vs schema mode3.52 sClaude Sonnet 5.5 · instructions6.21 sGPT-6.1 Sol · effort low · instructions12range: 2.7 s–4.1 s vs 4.2 s–12.3 sClaude Code aheadThe run ranges (fastest to slowest) do not overlap (Claude Code 2.67 s to 4.12 s; Codex CLI 4.20 s to 12.3 s). A range is not a confidence interval.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call)368Claude Sonnet 5.5 · instructions117GPT-6.1 Sol · effort low · instructions12none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
Reasoning share of output tokens per call on hard tasks (calculation)Calculation54.5%Claude Sonnet 5.546.3%GPT-6.1 Sol · effort medium24 / 16range: 0%–96% vs 11%–87%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation$0.0067Claude Sonnet 5.5$0.0023GPT-6.1 Sol · effort medium24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation$0.0037Claude Sonnet 5.5$0.0030GPT-6.1 Sol · effort medium24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation$0.0040Claude Sonnet 5.5$0.020GPT-6.1 Sol · effort medium24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation$0.0059Claude Sonnet 5.5 · effort medium$0.0023GPT-6.1 Sol · effort medium16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation$0.014Claude Sonnet 5.5 · effort medium$0.026GPT-6.1 Sol · effort medium16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation54.5%Claude Sonnet 5.546.3%GPT-6.1 Sol · effort medium24 / 16range: 0%–96% vs 11%–87%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation0%Claude Sonnet 5.541%GPT-6.1 Sol · effort medium15range: 0%–73% vs 0%–71%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Time to first text: a 250-line answer, six models1.96 sClaude Sonnet 5.53.52 sGPT-6.1 Sol · effort low4range: 0.9 s–4.1 s vs 2.8 s–4.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Code 0.88 s to 4.09 s; Codex CLI 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed after the first text: visible tokens per second (calculation)Calculation232Claude Sonnet 5.580GPT-6.1 Sol · effort low4range: 230–233 vs 72–81UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed in characters per second after the first text (calculation)Calculation517Claude Sonnet 5.5323GPT-6.1 Sol · effort low4range: 513–519 vs 291–327UnclearMore or fewer count is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Time to first text as the prompt grows: 1kCalculation1.45 sClaude Sonnet 5.53.36 sGPT-6.1 Sol · effort low3range: 1.2 s–1.7 s vs 3.4 s–4.8 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.23 s to 1.72 s; Codex CLI 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Time to first text as the prompt grows: 16kCalculation1.78 sClaude Sonnet 5.54.02 sGPT-6.1 Sol · effort low3range: 1.6 s–2.1 s vs 3.3 s–4.3 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.64 s to 2.11 s; Codex CLI 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Time to first text as the prompt grows: 64kCalculation3.07 sClaude Sonnet 5.53.93 sGPT-6.1 Sol · effort low3range: 1.4 s–3.6 s vs 3.4 s–4.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Code 1.38 s to 3.61 s; Codex CLI 3.42 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Total time per call by prompt size (1k prompt)1.78 sClaude Sonnet 5.53.43 sGPT-6.1 Sol · effort low3range: 1.6 s–2.1 s vs 3.4 s–4.9 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.57 s to 2.12 s; Codex CLI 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Total time per call by prompt size (16k prompt)2.10 sClaude Sonnet 5.54.14 sGPT-6.1 Sol · effort low3range: 2 s–2.5 s vs 4 s–4.7 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.98 s to 2.48 s; Codex CLI 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Total time per call by prompt size (64k prompt)3.44 sClaude Sonnet 5.53.96 sGPT-6.1 Sol · effort low3range: 1.7 s–4.4 s vs 3.5 s–4.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Code 1.74 s to 4.38 s; Codex CLI 3.47 s to 4.44 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Exact lookup answers at the 1k, 16k and 64k prompt-size targets100% (9/9)Claude Sonnet 5.5100% (9/9)GPT-6.1 Sol · effort low995% CI: 70%–100% vs 70%–100%TieThe 95% intervals overlap (Claude Code 70% to 100%; Codex CLI 70% to 100%), so this sample cannot separate them.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Pass rate on 4 harder tasks (Strict pass)38% (6/16)Claude Sonnet 5.569% (11/16)GPT-6.1 Sol · effort medium1695% CI: 18%–61% vs 44%–86%TieThe 95% intervals overlap (Claude Code 18% to 61%; Codex CLI 44% to 86%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Pass rate on 4 harder tasks (Lenient (format misses counted))38% (6/16)Claude Sonnet 5.569% (11/16)GPT-6.1 Sol · effort medium1695% CI: 18%–61% vs 44%–86%TieThe 95% intervals overlap (Claude Code 18% to 61%; Codex CLI 44% to 86%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Calls that tried a tool although tools were off31% (5/16)Claude Sonnet 5.50% (0/16)GPT-6.1 Sol · effort medium1695% CI: 14%–56% vs 0%–19%UnclearMore or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: 10x10 nonogram100% (4/4)Claude Sonnet 5.575% (3/4)GPT-6.1 Sol · effort medium495% CI: 51%–100% vs 30%–95%TieThe 95% intervals overlap (Claude Code 51% to 100%; Codex CLI 30% to 95%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: Sudoku, 22 givens0% (0/4)Claude Sonnet 5.525% (1/4)GPT-6.1 Sol · effort medium495% CI: 0%–49% vs 4.6%–70%TieThe 95% intervals overlap (Claude Code 0% to 49%; Codex CLI 5% to 70%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: 6x6 Skyscrapers0% (0/4)Claude Sonnet 5.5100% (4/4)GPT-6.1 Sol · effort medium495% CI: 0%–49% vs 51%–100%Codex CLI aheadThe 95% intervals do not overlap (Claude Code 0% to 49%; Codex CLI 51% to 100%).GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: Seeded shuffle output50% (2/4)Claude Sonnet 5.575% (3/4)GPT-6.1 Sol · effort medium495% CI: 15%–85% vs 30%–95%TieThe 95% intervals overlap (Claude Code 15% to 85%; Codex CLI 30% to 95%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Total time per call on harder tasks70.4 sClaude Sonnet 5.5120.2 sGPT-6.1 Sol · effort medium12 / 13range: 4.3 s–210 s vs 46.2 s–273 sUnclearThe run ranges (fastest to slowest) overlap (Claude Code 4.32 s to 210.1 s; Codex CLI 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Output tokens per call on harder tasks (Output tokens)9,287Claude Sonnet 5.54,994GPT-6.1 Sol · effort medium12 / 13range: 407–27,921 vs 2,099–13,413UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
List-price cost per strict pass on harder tasks (calculation)Calculation$0.24Claude Sonnet 5.5$0.083GPT-6.1 Sol · effort medium16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.24 vs $0.083, 2.9x) is not tested against run-to-run variation.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

Model vs model

Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)

Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) share 49 measured metrics and 21 list-price calculations from 10 studies. Claude Sonnet 5.5 leads on 4 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 1 more. GPT-6.1 Sol (Codex CLI) leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 22 ties and 43 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).

49 measured rows · 21 labelled calculations

Read all 70 metric rows for Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)
MetricClaude Sonnet 5.5GPT-6.1 Sol (Codex CLI)nInterval or rangeOutcomeBasisStudy
Pass rate on five validated tasks80% (12/15)Claude Code · five short validated tasks100% (15/15)Codex CLI · effort medium · five short validated tasks1595% CI: 55%–93% vs 80%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 55% to 93%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Total time per call2.31 sClaude Code · five short validated tasks5.65 sCodex CLI · effort medium · five short validated tasks15range: 2.2 s–7.7 s vs 4.1 s–25.5 sUnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.17 s to 7.73 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Time to first useful output1.56 sClaude Code · five short validated tasks5.05 sCodex CLI · effort medium · five short validated tasks15range: 1 s–6.4 s vs 3.4 s–17.8 sUnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.99 s to 6.39 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Input tokens per call: what the CLI sends (Cache read)1,401Claude Code · five short validated tasks5,180Codex CLI · effort medium · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Input tokens per call: what the CLI sends (Other input)685Claude Code · five short validated tasks6,943Codex CLI · effort medium · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Output tokens per call (Output tokens)107Claude Code · five short validated tasks42Codex CLI · effort medium · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
List-price cost per call (calculation)Calculation$0.0036Claude Code · five short validated tasks$0.010Codex CLI · effort medium · five short validated tasks15range: $0.0034–$0.01 vs $0.0054–$0.027UnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 $0.0034 to $0.010; GPT-6.1 Sol (Codex CLI) $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
List-price cost per passing answer (calculation)Calculation$0.0062Claude Code · five short validated tasks$0.016Codex CLI · effort medium · five short validated tasks15none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.0062 vs $0.016, 2.5x) is not tested against run-to-run variation.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Pass rate on eight hard tasks (Strict pass)100% (24/24)Claude Code · eight hard validated tasks100% (16/16)Codex CLI · effort medium · eight hard validated tasks24 / 1695% CI: 86%–100% vs 81%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Pass rate on eight hard tasks (Lenient (format misses counted))100% (24/24)Claude Code · eight hard validated tasks100% (16/16)Codex CLI · effort medium · eight hard validated tasks24 / 1695% CI: 86%–100% vs 81%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Total time per call on hard tasks (separate batches)7.75 sClaude Code · eight hard validated tasks13.1 sCodex CLI · effort medium · eight hard validated tasks24 / 16range: 2.3 s–34.8 s vs 8.5 s–61.6 sUnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Time to first useful output on hard tasks5.95 sClaude Code · eight hard validated tasks10.2 sCodex CLI · effort medium · eight hard validated tasks24 / 16range: 0.9 s–30.6 s vs 6.1 s–40.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.86 s to 30.6 s; GPT-6.1 Sol (Codex CLI) 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Output tokens per call on hard tasks (Output tokens)1,050Claude Code · eight hard validated tasks335Codex CLI · effort medium · eight hard validated tasks24 / 16none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
List-price cost per strict pass on hard tasks (calculation)Calculation$0.014Claude Code · eight hard validated tasks$0.026Codex CLI · effort medium · eight hard validated tasks24 / 16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Coding sessions that passed every hidden check100% (12/12)Claude Code · six small repository tasks with hidden tests100% (12/12)Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests1295% CI: 76%–100% vs 76%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
Time per coding session23.1 sClaude Code · six small repository tasks with hidden tests113.4 sCodex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests12range: 18.7 s–44.5 s vs 78.5 s–222 sClaude Sonnet 5.5 aheadThe run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 18.7 s to 44.5 s; GPT-6.1 Sol (Codex CLI) 78.5 s to 221.9 s). A range is not a confidence interval.Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
Tool calls per coding session7.5Claude Code · six small repository tasks with hidden tests12.5Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests12range: 3–14 vs 8–18UnclearMore or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
List-price cost per passing coding session (calculation)Calculation$0.085Claude Code · six small repository tasks with hidden tests$0.098Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests12none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.085 vs $0.098) is not tested against run-to-run variation.Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
Strict pass rate by effort on eight hard tasks100% (16/16)Claude Code · effort medium · eight hard validated tasks, effort ladder100% (16/16)Codex CLI · effort medium · eight hard validated tasks, effort ladder1695% CI: 81%–100% vs 81%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 81% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
Total time per call by effort on hard tasks7.63 sClaude Code · effort medium · eight hard validated tasks, effort ladder13.1 sCodex CLI · effort medium · eight hard validated tasks, effort ladder16range: 2.7 s–24 s vs 8.5 s–61.6 sUnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.71 s to 24.0 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
Output tokens per call by effort on hard tasks (Output tokens)770Claude Code · effort medium · eight hard validated tasks, effort ladder335Codex CLI · effort medium · eight hard validated tasks, effort ladder16none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
List-price cost per strict pass by effort (calculation)Calculation$0.014Claude Code · effort medium · eight hard validated tasks, effort ladder$0.026Codex CLI · effort medium · eight hard validated tasks, effort ladder16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation.Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
Same prompt, 10 times: strict pass rate (Exact number)100% (10/10)Claude Code · same prompt repeated 10 times100% (10/10)Codex CLI · effort medium · same prompt repeated 10 times1095% CI: 72%–100% vs 72%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: strict pass rate (JSON object)100% (10/10)Claude Code · same prompt repeated 10 times100% (10/10)Codex CLI · effort medium · same prompt repeated 10 times1095% CI: 72%–100% vs 72%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: strict pass rate (Code fix)100% (10/10)Claude Code · same prompt repeated 10 times100% (10/10)Codex CLI · effort medium · same prompt repeated 10 times1095% CI: 72%–100% vs 72%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: how many different answers (Exact number)1Claude Code · same prompt repeated 10 times1Codex CLI · effort medium · same prompt repeated 10 times10none recordedTieSame value. More or fewer is not better by itself for this metric.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: how many different answers (JSON object)1Claude Code · same prompt repeated 10 times1Codex CLI · effort medium · same prompt repeated 10 times10none recordedTieSame value. More or fewer is not better by itself for this metric.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: how many different answers (Code fix)3Claude Code · same prompt repeated 10 times6Codex CLI · effort medium · same prompt repeated 10 times10none recordedUnclearMore or fewer count is not better or worse by itself; this row describes behaviour, not a winner.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: time per call (Exact number)6.89 sClaude Code · same prompt repeated 10 times13.4 sCodex CLI · effort medium · same prompt repeated 10 times10range: 5.8 s–7.8 s vs 12.3 s–18 sClaude Sonnet 5.5 aheadThe run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 5.81 s to 7.81 s; GPT-6.1 Sol (Codex CLI) 12.3 s to 18.0 s). A range is not a confidence interval.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: time per call (JSON object)2.89 sClaude Code · same prompt repeated 10 times6.42 sCodex CLI · effort medium · same prompt repeated 10 times10range: 2.7 s–5.3 s vs 5.3 s–8.3 sUnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.68 s to 5.30 s; GPT-6.1 Sol (Codex CLI) 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: time per call (Code fix)2.67 sClaude Code · same prompt repeated 10 times11.3 sCodex CLI · effort medium · same prompt repeated 10 times10range: 2.3 s–4.3 s vs 9.1 s–14.9 sClaude Sonnet 5.5 aheadThe run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 2.32 s to 4.34 s; GPT-6.1 Sol (Codex CLI) 9.08 s to 14.8 s). A range is not a confidence interval.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Repairing a scheduler: Claude Code vs Codex vs API (Total time)15.0 sClaude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs61.2 sCodex CLI · effort medium · scheduler repair, 296 checks, 3 runs3range: 13.9 s–15.9 s vs 59.9 s–69.5 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 13.9 s to 15.9 s; GPT-6.1 Sol (Codex CLI) 59.9 s to 69.5 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Claude Code CLI vs Codex CLI vs the API: latency and tokens
Repairing a scheduler: Claude Code vs Codex vs API (First useful output)7.55 sClaude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs15.6 sCodex CLI · effort medium · scheduler repair, 296 checks, 3 runs3range: 6.8 s–7.6 s vs 13.7 s–23 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 6.77 s to 7.63 s; GPT-6.1 Sol (Codex CLI) 13.7 s to 23.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Claude Code CLI vs Codex CLI vs the API: latency and tokens
Output tokens to repair the scheduler (Output tokens)2,227Claude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs1,181Codex CLI · effort medium · scheduler repair, 296 checks, 3 runs3none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Claude Code CLI vs Codex CLI vs the API: latency and tokens
Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)Calculation100% (12/12)Claude Code · instructions100% (12/12)Codex CLI · effort low · instructions1295% CI: 76%–100% vs 76%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))Calculation100% (12/12)Claude Code · instructions100% (12/12)Codex CLI · effort low · instructions1295% CI: 76%–100% vs 76%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
What each call produced: strict pass, format miss, wrong values or error (Strict pass)12Claude Code · instructions12Codex CLI · effort low · instructions12none recordedTieSame value. More or fewer is not better by itself for this metric.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
What each call produced: strict pass, format miss, wrong values or error (Format miss)0Claude Code · instructions0Codex CLI · effort low · instructions12none recordedTieSame value. More or fewer is not better by itself for this metric.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
What each call produced: strict pass, format miss, wrong values or error (Wrong values)0Claude Code · instructions0Codex CLI · effort low · instructions12none recordedTieSame value. More or fewer is not better by itself for this metric.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
What each call produced: strict pass, format miss, wrong values or error (Error)0Claude Code · instructions0Codex CLI · effort low · instructions12none recordedTieSame value. More or fewer is not better by itself for this metric.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
Time per call, instructions vs schema mode3.52 sClaude Code · instructions6.21 sCodex CLI · effort low · instructions12range: 2.7 s–4.1 s vs 4.2 s–12.3 sClaude Sonnet 5.5 aheadThe run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 2.67 s to 4.12 s; GPT-6.1 Sol (Codex CLI) 4.20 s to 12.3 s). A range is not a confidence interval.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call)368Claude Code · instructions117Codex CLI · effort low · instructions12none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
Reasoning share of output tokens per call on hard tasks (calculation)Calculation54.5%Claude Code46.3%Codex CLI · effort medium24 / 16range: 0%–96% vs 11%–87%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation$0.0067Claude Code$0.0023Codex CLI · effort medium24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation$0.0037Claude Code$0.0030Codex CLI · effort medium24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation$0.0040Claude Code$0.020Codex CLI · effort medium24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation$0.0059Claude Code · effort medium$0.0023Codex CLI · effort medium16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation$0.014Claude Code · effort medium$0.026Codex CLI · effort medium16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation54.5%Claude Code46.3%Codex CLI · effort medium24 / 16range: 0%–96% vs 11%–87%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation0%Claude Code41%Codex CLI · effort medium15range: 0%–73% vs 0%–71%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Time to first text: a 250-line answer, six models1.96 sClaude Code3.52 sCodex CLI · effort low4range: 0.9 s–4.1 s vs 2.8 s–4.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.88 s to 4.09 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed after the first text: visible tokens per second (calculation)Calculation232Claude Code80Codex CLI · effort low4range: 230–233 vs 72–81UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed in characters per second after the first text (calculation)Calculation517Claude Code323Codex CLI · effort low4range: 513–519 vs 291–327UnclearMore or fewer count is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Time to first text as the prompt grows: 1kCalculation1.45 sClaude Code3.36 sCodex CLI · effort low3range: 1.2 s–1.7 s vs 3.4 s–4.8 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 1.23 s to 1.72 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Time to first text as the prompt grows: 16kCalculation1.78 sClaude Code4.02 sCodex CLI · effort low3range: 1.6 s–2.1 s vs 3.3 s–4.3 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 1.64 s to 2.11 s; GPT-6.1 Sol (Codex CLI) 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Time to first text as the prompt grows: 64kCalculation3.07 sClaude Code3.93 sCodex CLI · effort low3range: 1.4 s–3.6 s vs 3.4 s–4.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.38 s to 3.61 s; GPT-6.1 Sol (Codex CLI) 3.42 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Total time per call by prompt size (1k prompt)1.78 sClaude Code3.43 sCodex CLI · effort low3range: 1.6 s–2.1 s vs 3.4 s–4.9 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 1.57 s to 2.12 s; GPT-6.1 Sol (Codex CLI) 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Total time per call by prompt size (16k prompt)2.10 sClaude Code4.14 sCodex CLI · effort low3range: 2 s–2.5 s vs 4 s–4.7 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 1.98 s to 2.48 s; GPT-6.1 Sol (Codex CLI) 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Total time per call by prompt size (64k prompt)3.44 sClaude Code3.96 sCodex CLI · effort low3range: 1.7 s–4.4 s vs 3.5 s–4.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 1.74 s to 4.38 s; GPT-6.1 Sol (Codex CLI) 3.47 s to 4.44 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Exact lookup answers at the 1k, 16k and 64k prompt-size targets100% (9/9)Claude Code100% (9/9)Codex CLI · effort low995% CI: 70%–100% vs 70%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 70% to 100%; GPT-6.1 Sol (Codex CLI) 70% to 100%), so this sample cannot separate them.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Pass rate on 4 harder tasks (Strict pass)38% (6/16)Claude Code69% (11/16)Codex CLI · effort medium1695% CI: 18%–61% vs 44%–86%TieThe 95% intervals overlap (Claude Sonnet 5.5 18% to 61%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Pass rate on 4 harder tasks (Lenient (format misses counted))38% (6/16)Claude Code69% (11/16)Codex CLI · effort medium1695% CI: 18%–61% vs 44%–86%TieThe 95% intervals overlap (Claude Sonnet 5.5 18% to 61%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Calls that tried a tool although tools were off31% (5/16)Claude Code0% (0/16)Codex CLI · effort medium1695% CI: 14%–56% vs 0%–19%UnclearMore or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: 10x10 nonogram100% (4/4)Claude Code75% (3/4)Codex CLI · effort medium495% CI: 51%–100% vs 30%–95%TieThe 95% intervals overlap (Claude Sonnet 5.5 51% to 100%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: Sudoku, 22 givens0% (0/4)Claude Code25% (1/4)Codex CLI · effort medium495% CI: 0%–49% vs 4.6%–70%TieThe 95% intervals overlap (Claude Sonnet 5.5 0% to 49%; GPT-6.1 Sol (Codex CLI) 5% to 70%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: 6x6 Skyscrapers0% (0/4)Claude Code100% (4/4)Codex CLI · effort medium495% CI: 0%–49% vs 51%–100%GPT-6.1 Sol (Codex CLI) aheadThe 95% intervals do not overlap (Claude Sonnet 5.5 0% to 49%; GPT-6.1 Sol (Codex CLI) 51% to 100%).GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: Seeded shuffle output50% (2/4)Claude Code75% (3/4)Codex CLI · effort medium495% CI: 15%–85% vs 30%–95%TieThe 95% intervals overlap (Claude Sonnet 5.5 15% to 85%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Total time per call on harder tasks70.4 sClaude Code120.2 sCodex CLI · effort medium12 / 13range: 4.3 s–210 s vs 46.2 s–273 sUnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 4.32 s to 210.1 s; GPT-6.1 Sol (Codex CLI) 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Output tokens per call on harder tasks (Output tokens)9,287Claude Code4,994Codex CLI · effort medium12 / 13range: 407–27,921 vs 2,099–13,413UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
List-price cost per strict pass on harder tasks (calculation)Calculation$0.24Claude Code$0.083Codex CLI · effort medium16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.24 vs $0.083, 2.9x) is not tested against run-to-run variation.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

Model vs model

Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI)

Claude Opus 5.5 and GPT-6.1 Sol (Codex CLI) share 31 measured metrics and 19 list-price calculations from 7 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 12 ties and 38 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).

31 measured rows · 19 labelled calculations

Read all 50 metric rows for Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI)
MetricClaude Opus 5.5GPT-6.1 Sol (Codex CLI)nInterval or rangeOutcomeBasisStudy
Pass rate on five validated tasks100% (15/15)Claude Code · effort high · five short validated tasks100% (15/15)Codex CLI · effort high · five short validated tasks1595% CI: 80%–100% vs 80%–100%TieThe 95% intervals overlap (Claude Opus 5.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Total time per call2.71 sClaude Code · effort high · five short validated tasks5.60 sCodex CLI · effort high · five short validated tasks15range: 2.5 s–11.8 s vs 4.1 s–19.5 sUnclearThe run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.45 s to 11.8 s; GPT-6.1 Sol (Codex CLI) 4.05 s to 19.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Time to first useful output2.04 sClaude Code · effort high · five short validated tasks5.32 sCodex CLI · effort high · five short validated tasks15range: 1.4 s–9.9 s vs 3.6 s–16.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.40 s to 9.94 s; GPT-6.1 Sol (Codex CLI) 3.64 s to 16.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Input tokens per call: what the CLI sends (Cache read)1,463Claude Code · effort high · five short validated tasks6,716Codex CLI · effort high · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Input tokens per call: what the CLI sends (Other input)619Claude Code · effort high · five short validated tasks5,406Codex CLI · effort high · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Output tokens per call (Output tokens)78Claude Code · effort high · five short validated tasks42Codex CLI · effort high · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
List-price cost per call (calculation)Calculation$0.0069Claude Code · effort high · five short validated tasks$0.010Codex CLI · effort high · five short validated tasks15range: $0.0059–$0.027 vs $0.0066–$0.028UnclearThe run ranges (fastest to slowest) overlap (Claude Opus 5.5 $0.0059 to $0.027; GPT-6.1 Sol (Codex CLI) $0.0066 to $0.028); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
List-price cost per passing answer (calculation)Calculation$0.010Claude Code · effort high · five short validated tasks$0.013Codex CLI · effort high · five short validated tasks15none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.010 vs $0.013) is not tested against run-to-run variation.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Pass rate on eight hard tasks (Strict pass)100% (24/24)Claude Code · effort high · eight hard validated tasks100% (16/16)Codex CLI · effort high · eight hard validated tasks24 / 1695% CI: 86%–100% vs 81%–100%TieThe 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Pass rate on eight hard tasks (Lenient (format misses counted))100% (24/24)Claude Code · effort high · eight hard validated tasks100% (16/16)Codex CLI · effort high · eight hard validated tasks24 / 1695% CI: 86%–100% vs 81%–100%TieThe 95% intervals overlap (Claude Opus 5.5 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Total time per call on hard tasks (separate batches)11.0 sClaude Code · effort high · eight hard validated tasks18.1 sCodex CLI · effort high · eight hard validated tasks24 / 16range: 3.6 s–63 s vs 11.7 s–92.2 sUnclearThe run ranges (fastest to slowest) overlap (Claude Opus 5.5 3.63 s to 63.0 s; GPT-6.1 Sol (Codex CLI) 11.7 s to 92.2 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Time to first useful output on hard tasks7.13 sClaude Code · effort high · eight hard validated tasks12.7 sCodex CLI · effort high · eight hard validated tasks24 / 16range: 2.2 s–56.2 s vs 8.9 s–75.9 sUnclearThe run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.15 s to 56.2 s; GPT-6.1 Sol (Codex CLI) 8.93 s to 75.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Output tokens per call on hard tasks (Output tokens)1,052Claude Code · effort high · eight hard validated tasks436Codex CLI · effort high · eight hard validated tasks24 / 16none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
List-price cost per strict pass on hard tasks (calculation)Calculation$0.033Claude Code · effort high · eight hard validated tasks$0.015Codex CLI · effort high · eight hard validated tasks24 / 16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.033 vs $0.015, 2.2x) is not tested against run-to-run variation.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Coding sessions that passed every hidden check100% (12/12)Claude Code · six small repository tasks with hidden tests100% (12/12)Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests1295% CI: 76%–100% vs 76%–100%TieThe 95% intervals overlap (Claude Opus 5.5 76% to 100%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
Time per coding session56.9 sClaude Code · six small repository tasks with hidden tests113.4 sCodex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests12range: 29.8 s–186 s vs 78.5 s–222 sUnclearThe run ranges (fastest to slowest) overlap (Claude Opus 5.5 29.8 s to 185.8 s; GPT-6.1 Sol (Codex CLI) 78.5 s to 221.9 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
Tool calls per coding session7.5Claude Code · six small repository tasks with hidden tests12.5Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests12range: 5–14 vs 8–18UnclearMore or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
List-price cost per passing coding session (calculation)Calculation$0.22Claude Code · six small repository tasks with hidden tests$0.098Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests12none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.22 vs $0.098, 2.3x) is not tested against run-to-run variation.Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
Strict pass rate by effort on eight hard tasks100% (16/16)Claude Code · effort medium · eight hard validated tasks, effort ladder100% (16/16)Codex CLI · effort medium · eight hard validated tasks, effort ladder1695% CI: 81%–100% vs 81%–100%TieThe 95% intervals overlap (Claude Opus 5.5 81% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
Total time per call by effort on hard tasks9.72 sClaude Code · effort medium · eight hard validated tasks, effort ladder13.1 sCodex CLI · effort medium · eight hard validated tasks, effort ladder16range: 4.8 s–31.4 s vs 8.5 s–61.6 sUnclearThe run ranges (fastest to slowest) overlap (Claude Opus 5.5 4.78 s to 31.4 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
Output tokens per call by effort on hard tasks (Output tokens)853Claude Code · effort medium · eight hard validated tasks, effort ladder335Codex CLI · effort medium · eight hard validated tasks, effort ladder16none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
List-price cost per strict pass by effort (calculation)Calculation$0.029Claude Code · effort medium · eight hard validated tasks, effort ladder$0.026Codex CLI · effort medium · eight hard validated tasks, effort ladder16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.029 vs $0.026) is not tested against run-to-run variation.Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
Reasoning share of output tokens per call on hard tasks (calculation)Calculation54.4%Claude Code · effort high57%Codex CLI · effort high24 / 16range: 36%–96% vs 29%–91%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation$0.018Claude Code · effort high$0.0041Codex CLI · effort high24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation$0.0080Claude Code · effort high$0.0029Codex CLI · effort high24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation$0.0074Claude Code · effort high$0.0081Codex CLI · effort high24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation$0.013Claude Code · effort medium$0.0023Codex CLI · effort medium16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation$0.029Claude Code · effort medium$0.026Codex CLI · effort medium16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation54.4%Claude Code · effort high57%Codex CLI · effort high24 / 16range: 36%–96% vs 29%–91%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation43.6%Claude Code · effort high58.1%Codex CLI · effort high15range: 0%–93% vs 0%–76%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Time to first text: a 250-line answer, six models1.97 sClaude Code3.52 sCodex CLI · effort low4range: 1.7 s–2.4 s vs 2.8 s–4.4 sUnclearOnly 4 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.70 s to 2.35 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s), but 4 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed after the first text: visible tokens per second (calculation)Calculation156Claude Code80Codex CLI · effort low4range: 155–156 vs 72–81UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed in characters per second after the first text (calculation)Calculation347Claude Code323Codex CLI · effort low4range: 345–349 vs 291–327UnclearMore or fewer count is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Time to first text as the prompt grows: 1kCalculation1.51 sClaude Code3.36 sCodex CLI · effort low3range: 1.5 s–2 s vs 3.4 s–4.8 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.46 s to 2.01 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Time to first text as the prompt grows: 16kCalculation1.74 sClaude Code4.02 sCodex CLI · effort low3range: 1.7 s–3 s vs 3.3 s–4.3 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.70 s to 2.97 s; GPT-6.1 Sol (Codex CLI) 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Time to first text as the prompt grows: 64kCalculation1.79 sClaude Code3.93 sCodex CLI · effort low3range: 1.7 s–3.7 s vs 3.4 s–4.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Opus 5.5 1.72 s to 3.72 s; GPT-6.1 Sol (Codex CLI) 3.42 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Total time per call by prompt size (1k prompt)1.83 sClaude Code3.43 sCodex CLI · effort low3range: 1.8 s–2.4 s vs 3.4 s–4.9 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 1.82 s to 2.41 s; GPT-6.1 Sol (Codex CLI) 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Total time per call by prompt size (16k prompt)2.36 sClaude Code4.14 sCodex CLI · effort low3range: 2.1 s–3.4 s vs 4 s–4.7 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Opus 5.5 2.11 s to 3.40 s; GPT-6.1 Sol (Codex CLI) 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Total time per call by prompt size (64k prompt)2.35 sClaude Code3.96 sCodex CLI · effort low3range: 2.3 s–4.3 s vs 3.5 s–4.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Opus 5.5 2.26 s to 4.29 s; GPT-6.1 Sol (Codex CLI) 3.47 s to 4.44 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Exact lookup answers at the 1k, 16k and 64k prompt-size targets56% (5/9)Claude Code100% (9/9)Codex CLI · effort low995% CI: 27%–81% vs 70%–100%TieThe 95% intervals overlap (Claude Opus 5.5 27% to 81%; GPT-6.1 Sol (Codex CLI) 70% to 100%), so this sample cannot separate them.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Pass rate on 4 harder tasks (Strict pass)42% (5/12)Claude Code69% (11/16)Codex CLI · effort medium12 / 1695% CI: 19%–68% vs 44%–86%TieThe 95% intervals overlap (Claude Opus 5.5 19% to 68%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Pass rate on 4 harder tasks (Lenient (format misses counted))50% (6/12)Claude Code69% (11/16)Codex CLI · effort medium12 / 1695% CI: 25%–75% vs 44%–86%TieThe 95% intervals overlap (Claude Opus 5.5 25% to 75%; GPT-6.1 Sol (Codex CLI) 44% to 86%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Calls that tried a tool although tools were off42% (5/12)Claude Code0% (0/16)Codex CLI · effort medium12 / 1695% CI: 19%–68% vs 0%–19%UnclearMore or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: 10x10 nonogram100% (3/3)Claude Code75% (3/4)Codex CLI · effort medium3 / 495% CI: 44%–100% vs 30%–95%TieThe 95% intervals overlap (Claude Opus 5.5 44% to 100%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: Sudoku, 22 givens0% (0/3)Claude Code25% (1/4)Codex CLI · effort medium3 / 495% CI: 0%–56% vs 4.6%–70%TieThe 95% intervals overlap (Claude Opus 5.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 5% to 70%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: 6x6 Skyscrapers33% (1/3)Claude Code100% (4/4)Codex CLI · effort medium3 / 495% CI: 6.2%–79% vs 51%–100%TieThe 95% intervals overlap (Claude Opus 5.5 6% to 79%; GPT-6.1 Sol (Codex CLI) 51% to 100%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: Seeded shuffle output33% (1/3)Claude Code75% (3/4)Codex CLI · effort medium3 / 495% CI: 6.2%–79% vs 30%–95%TieThe 95% intervals overlap (Claude Opus 5.5 6% to 79%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Total time per call on harder tasks80.3 sClaude Code120.2 sCodex CLI · effort medium9 / 13range: 3.8 s–280 s vs 46.2 s–273 sUnclearThe run ranges (fastest to slowest) overlap (Claude Opus 5.5 3.82 s to 279.5 s; GPT-6.1 Sol (Codex CLI) 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Output tokens per call on harder tasks (Output tokens)8,420Claude Code4,994Codex CLI · effort medium9 / 13range: 279–40,044 vs 2,099–13,413UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
List-price cost per strict pass on harder tasks (calculation)Calculation$0.59Claude Code$0.083Codex CLI · effort medium12 / 16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.59 vs $0.083, 7.2x) is not tested against run-to-run variation.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

Model vs model

Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)

Claude Haiku 4.5 and GPT-6.1 Sol (Codex CLI) share 40 measured metrics and 16 list-price calculations from 7 studies. Claude Haiku 4.5 leads on 2 rows: Same prompt, 10 times: time per call (Exact number), 5.06 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 5.95 s vs 11.3 s. GPT-6.1 Sol (Codex CLI) leads on 6 rows: Pass rate on eight hard tasks (Strict pass), 100% (16/16) vs 46% (11/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); Same prompt, 10 times: strict pass rate (JSON object), 100% (10/10) vs 10% (1/10); and 3 more. On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 13 ties and 35 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).

40 measured rows · 16 labelled calculations

Read all 56 metric rows for Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)
MetricClaude Haiku 4.5GPT-6.1 Sol (Codex CLI)nInterval or rangeOutcomeBasisStudy
Pass rate on five validated tasks100% (15/15)Claude Code · five short validated tasks100% (15/15)Codex CLI · effort medium · five short validated tasks1595% CI: 80%–100% vs 80%–100%TieThe 95% intervals overlap (Claude Haiku 4.5 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Total time per call4.43 sClaude Code · five short validated tasks5.65 sCodex CLI · effort medium · five short validated tasks15range: 3.2 s–23.6 s vs 4.1 s–25.5 sUnclearThe run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Time to first useful output3.63 sClaude Code · five short validated tasks5.05 sCodex CLI · effort medium · five short validated tasks15range: 2.8 s–22.3 s vs 3.4 s–17.8 sUnclearThe run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.78 s to 22.3 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Input tokens per call: what the CLI sends (Cache read)0Claude Code · five short validated tasks5,180Codex CLI · effort medium · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Input tokens per call: what the CLI sends (Other input)3,790Claude Code · five short validated tasks6,943Codex CLI · effort medium · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Output tokens per call (Output tokens)367Claude Code · five short validated tasks42Codex CLI · effort medium · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
List-price cost per call (calculation)Calculation$0.0057Claude Code · five short validated tasks$0.010Codex CLI · effort medium · five short validated tasks15range: $0.0051–$0.018 vs $0.0054–$0.027UnclearThe run ranges (fastest to slowest) overlap (Claude Haiku 4.5 $0.0051 to $0.018; GPT-6.1 Sol (Codex CLI) $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
List-price cost per passing answer (calculation)Calculation$0.0084Claude Code · five short validated tasks$0.016Codex CLI · effort medium · five short validated tasks15none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.0084 vs $0.016) is not tested against run-to-run variation.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Pass rate on eight hard tasks (Strict pass)46% (11/24)Claude Code · eight hard validated tasks100% (16/16)Codex CLI · effort medium · eight hard validated tasks24 / 1695% CI: 28%–65% vs 81%–100%GPT-6.1 Sol (Codex CLI) aheadThe 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; GPT-6.1 Sol (Codex CLI) 81% to 100%).Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Pass rate on eight hard tasks (Lenient (format misses counted))67% (16/24)Claude Code · eight hard validated tasks100% (16/16)Codex CLI · effort medium · eight hard validated tasks24 / 1695% CI: 47%–82% vs 81%–100%TieThe 95% intervals overlap (Claude Haiku 4.5 47% to 82%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Total time per call on hard tasks (separate batches)39.0 sClaude Code · eight hard validated tasks13.1 sCodex CLI · effort medium · eight hard validated tasks24 / 16range: 15.3 s–75.1 s vs 8.5 s–61.6 sUnclearThe run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Time to first useful output on hard tasks35.5 sClaude Code · eight hard validated tasks10.2 sCodex CLI · effort medium · eight hard validated tasks24 / 16range: 12.9 s–70.3 s vs 6.1 s–40.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Haiku 4.5 12.9 s to 70.3 s; GPT-6.1 Sol (Codex CLI) 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Output tokens per call on hard tasks (Output tokens)5,064Claude Code · eight hard validated tasks335Codex CLI · effort medium · eight hard validated tasks24 / 16none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
List-price cost per strict pass on hard tasks (calculation)Calculation$0.067Claude Code · eight hard validated tasks$0.026Codex CLI · effort medium · eight hard validated tasks24 / 16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.067 vs $0.026, 2.6x) is not tested against run-to-run variation.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Same prompt, 10 times: strict pass rate (Exact number)0% (0/10)Claude Code · same prompt repeated 10 times100% (10/10)Codex CLI · effort medium · same prompt repeated 10 times1095% CI: 0%–28% vs 72%–100%GPT-6.1 Sol (Codex CLI) aheadThe 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; GPT-6.1 Sol (Codex CLI) 72% to 100%).Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: strict pass rate (JSON object)10% (1/10)Claude Code · same prompt repeated 10 times100% (10/10)Codex CLI · effort medium · same prompt repeated 10 times1095% CI: 1.8%–40% vs 72%–100%GPT-6.1 Sol (Codex CLI) aheadThe 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; GPT-6.1 Sol (Codex CLI) 72% to 100%).Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: strict pass rate (Code fix)100% (10/10)Claude Code · same prompt repeated 10 times100% (10/10)Codex CLI · effort medium · same prompt repeated 10 times1095% CI: 72%–100% vs 72%–100%TieThe 95% intervals overlap (Claude Haiku 4.5 72% to 100%; GPT-6.1 Sol (Codex CLI) 72% to 100%), so this sample cannot separate them.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: how many different answers (Exact number)1Claude Code · same prompt repeated 10 times1Codex CLI · effort medium · same prompt repeated 10 times10none recordedTieSame value. More or fewer is not better by itself for this metric.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: how many different answers (JSON object)1Claude Code · same prompt repeated 10 times1Codex CLI · effort medium · same prompt repeated 10 times10none recordedTieSame value. More or fewer is not better by itself for this metric.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: how many different answers (Code fix)6Claude Code · same prompt repeated 10 times6Codex CLI · effort medium · same prompt repeated 10 times10none recordedTieSame value. More or fewer is not better by itself for this metric.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: time per call (Exact number)5.06 sClaude Code · same prompt repeated 10 times13.4 sCodex CLI · effort medium · same prompt repeated 10 times10range: 4.4 s–6.2 s vs 12.3 s–18 sClaude Haiku 4.5 aheadThe run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.42 s to 6.20 s; GPT-6.1 Sol (Codex CLI) 12.3 s to 18.0 s). A range is not a confidence interval.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: time per call (JSON object)7.03 sClaude Code · same prompt repeated 10 times6.42 sCodex CLI · effort medium · same prompt repeated 10 times10range: 5.3 s–12.3 s vs 5.3 s–8.3 sUnclearThe run ranges (fastest to slowest) overlap (Claude Haiku 4.5 5.28 s to 12.3 s; GPT-6.1 Sol (Codex CLI) 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Same prompt, 10 times: time per call (Code fix)5.95 sClaude Code · same prompt repeated 10 times11.3 sCodex CLI · effort medium · same prompt repeated 10 times10range: 4.9 s–7.3 s vs 9.1 s–14.9 sClaude Haiku 4.5 aheadThe run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; GPT-6.1 Sol (Codex CLI) 9.08 s to 14.8 s). A range is not a confidence interval.Prompt caching and run-to-run consistency in Claude Code and Codex CLI
Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)Calculation0% (0/24)Claude Code · instructions100% (12/12)Codex CLI · effort low · instructions24 / 1295% CI: 0%–14% vs 76%–100%GPT-6.1 Sol (Codex CLI) aheadThe 95% intervals do not overlap (Claude Haiku 4.5 0% to 14%; GPT-6.1 Sol (Codex CLI) 76% to 100%).Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))Calculation71% (17/24)Claude Code · instructions100% (12/12)Codex CLI · effort low · instructions24 / 1295% CI: 51%–85% vs 76%–100%TieThe 95% intervals overlap (Claude Haiku 4.5 51% to 85%; GPT-6.1 Sol (Codex CLI) 76% to 100%), so this sample cannot separate them.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
What each call produced: strict pass, format miss, wrong values or error (Strict pass)0Claude Code · instructions12Codex CLI · effort low · instructions24 / 12none recordedUnclearMore or fewer count is not better or worse by itself; this row describes behaviour, not a winner.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
What each call produced: strict pass, format miss, wrong values or error (Format miss)17Claude Code · instructions0Codex CLI · effort low · instructions24 / 12none recordedUnclearMore or fewer count is not better or worse by itself; this row describes behaviour, not a winner.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
What each call produced: strict pass, format miss, wrong values or error (Wrong values)7Claude Code · instructions0Codex CLI · effort low · instructions24 / 12none recordedUnclearMore or fewer count is not better or worse by itself; this row describes behaviour, not a winner.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
What each call produced: strict pass, format miss, wrong values or error (Error)0Claude Code · instructions0Codex CLI · effort low · instructions24 / 12none recordedTieSame value. More or fewer is not better by itself for this metric.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
Time per call, instructions vs schema mode9.52 sClaude Code · instructions6.21 sCodex CLI · effort low · instructions24 / 12range: 5.7 s–17 s vs 4.2 s–12.3 sUnclearThe run ranges (fastest to slowest) overlap (Claude Haiku 4.5 5.67 s to 17.0 s; GPT-6.1 Sol (Codex CLI) 4.20 s to 12.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call)1,128Claude Code · instructions117Codex CLI · effort low · instructions24 / 12none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
Reasoning share of output tokens per call on hard tasks (calculation)Calculation91.7%Claude Code46.3%Codex CLI · effort medium24 / 16range: 76%–99% vs 11%–87%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation$0.024Claude Code$0.0023Codex CLI · effort medium24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation$0.0018Claude Code$0.0030Codex CLI · effort medium24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation$0.0045Claude Code$0.020Codex CLI · effort medium24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation91.7%Claude Code46.3%Codex CLI · effort medium24 / 16range: 76%–99% vs 11%–87%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation90.2%Claude Code41%Codex CLI · effort medium15range: 73%–98% vs 0%–71%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Time to first text: a 250-line answer, six models4.00 sClaude Code3.52 sCodex CLI · effort low4range: 2.8 s–6.4 s vs 2.8 s–4.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed after the first text: visible tokens per second (calculation)Calculation153Claude Code80Codex CLI · effort low4range: 153–216 vs 72–81UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed in characters per second after the first text (calculation)Calculation547Claude Code323Codex CLI · effort low3 / 4range: 546–548 vs 291–327UnclearMore or fewer count is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Time to first text as the prompt grows: 1kCalculation1.93 sClaude Code3.36 sCodex CLI · effort low3range: 1.9 s–2 s vs 3.4 s–4.8 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 1.85 s to 2.04 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Time to first text as the prompt grows: 16kCalculation2.27 sClaude Code4.02 sCodex CLI · effort low3range: 2.2 s–2.5 s vs 3.3 s–4.3 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.47 s; GPT-6.1 Sol (Codex CLI) 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Time to first text as the prompt grows: 64kCalculation2.78 sClaude Code3.93 sCodex CLI · effort low3range: 2.5 s–2.9 s vs 3.4 s–4.4 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.45 s to 2.89 s; GPT-6.1 Sol (Codex CLI) 3.42 s to 4.38 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Total time per call by prompt size (1k prompt)2.34 sClaude Code3.43 sCodex CLI · effort low3range: 2.2 s–2.5 s vs 3.4 s–4.9 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.46 s; GPT-6.1 Sol (Codex CLI) 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Total time per call by prompt size (16k prompt)2.79 sClaude Code4.14 sCodex CLI · effort low3range: 2.6 s–2.8 s vs 4 s–4.7 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.58 s to 2.84 s; GPT-6.1 Sol (Codex CLI) 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Total time per call by prompt size (64k prompt)3.13 sClaude Code3.96 sCodex CLI · effort low3range: 2.8 s–3.3 s vs 3.5 s–4.4 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.84 s to 3.28 s; GPT-6.1 Sol (Codex CLI) 3.47 s to 4.44 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Exact lookup answers at the 1k, 16k and 64k prompt-size targets100% (9/9)Claude Code100% (9/9)Codex CLI · effort low995% CI: 70%–100% vs 70%–100%TieThe 95% intervals overlap (Claude Haiku 4.5 70% to 100%; GPT-6.1 Sol (Codex CLI) 70% to 100%), so this sample cannot separate them.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Pass rate on 4 harder tasks (Strict pass)0% (0/12)Claude Code69% (11/16)Codex CLI · effort medium12 / 1695% CI: 0%–24% vs 44%–86%GPT-6.1 Sol (Codex CLI) aheadThe 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; GPT-6.1 Sol (Codex CLI) 44% to 86%).GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Pass rate on 4 harder tasks (Lenient (format misses counted))0% (0/12)Claude Code69% (11/16)Codex CLI · effort medium12 / 1695% CI: 0%–24% vs 44%–86%GPT-6.1 Sol (Codex CLI) aheadThe 95% intervals do not overlap (Claude Haiku 4.5 0% to 24%; GPT-6.1 Sol (Codex CLI) 44% to 86%).GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Calls that tried a tool although tools were off8% (1/12)Claude Code0% (0/16)Codex CLI · effort medium12 / 1695% CI: 1.5%–35% vs 0%–19%UnclearMore or fewer rate is not better or worse by itself; this row describes behaviour, not a winner.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: 10x10 nonogram0% (0/3)Claude Code75% (3/4)Codex CLI · effort medium3 / 495% CI: 0%–56% vs 30%–95%TieThe 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: Sudoku, 22 givens0% (0/3)Claude Code25% (1/4)Codex CLI · effort medium3 / 495% CI: 0%–56% vs 4.6%–70%TieThe 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 5% to 70%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: 6x6 Skyscrapers0% (0/3)Claude Code100% (4/4)Codex CLI · effort medium3 / 495% CI: 0%–56% vs 51%–100%TieThe 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 51% to 100%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Strict pass rate by task: Seeded shuffle output0% (0/3)Claude Code75% (3/4)Codex CLI · effort medium3 / 495% CI: 0%–56% vs 30%–95%TieThe 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6.1 Sol (Codex CLI) 30% to 95%), so this sample cannot separate them.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Total time per call on harder tasks109.0 sClaude Code120.2 sCodex CLI · effort medium10 / 13range: 25.7 s–224 s vs 46.2 s–273 sUnclearThe run ranges (fastest to slowest) overlap (Claude Haiku 4.5 25.7 s to 223.9 s; GPT-6.1 Sol (Codex CLI) 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
Output tokens per call on harder tasks (Output tokens)12,508Claude Code4,994Codex CLI · effort medium10 / 13range: 2,965–26,532 vs 2,099–13,413UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

Model vs model

Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI)

Claude Fable 5.1 and GPT-6.1 Sol (Codex CLI) share 12 measured metrics and 11 list-price calculations from 4 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 ties and 20 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 4 at the smallest).

12 measured rows · 11 labelled calculations

Read all 23 metric rows for Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI)
MetricClaude Fable 5.1GPT-6.1 Sol (Codex CLI)nInterval or rangeOutcomeBasisStudy
Pass rate on five validated tasks100% (15/15)Claude Code · five short validated tasks100% (15/15)Codex CLI · effort medium · five short validated tasks1595% CI: 80%–100% vs 80%–100%TieThe 95% intervals overlap (Claude Fable 5.1 80% to 100%; GPT-6.1 Sol (Codex CLI) 80% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Total time per call1.94 sClaude Code · five short validated tasks5.65 sCodex CLI · effort medium · five short validated tasks15range: 1.4 s–9.8 s vs 4.1 s–25.5 sUnclearThe run ranges (fastest to slowest) overlap (Claude Fable 5.1 1.41 s to 9.83 s; GPT-6.1 Sol (Codex CLI) 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Time to first useful output1.20 sClaude Code · five short validated tasks5.05 sCodex CLI · effort medium · five short validated tasks15range: 1 s–7.9 s vs 3.4 s–17.8 sUnclearThe run ranges (fastest to slowest) overlap (Claude Fable 5.1 0.95 s to 7.90 s; GPT-6.1 Sol (Codex CLI) 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Input tokens per call: what the CLI sends (Cache read)2,760Claude Code · five short validated tasks5,180Codex CLI · effort medium · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Input tokens per call: what the CLI sends (Other input)473Claude Code · five short validated tasks6,943Codex CLI · effort medium · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Output tokens per call (Output tokens)64Claude Code · five short validated tasks42Codex CLI · effort medium · five short validated tasks15none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
List-price cost per call (calculation)Calculation$0.0099Claude Code · five short validated tasks$0.010Codex CLI · effort medium · five short validated tasks15range: $0.0049–$0.058 vs $0.0054–$0.027UnclearThe run ranges (fastest to slowest) overlap (Claude Fable 5.1 $0.0049 to $0.058; GPT-6.1 Sol (Codex CLI) $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
List-price cost per passing answer (calculation)Calculation$0.021Claude Code · five short validated tasks$0.016Codex CLI · effort medium · five short validated tasks15none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.021 vs $0.016) is not tested against run-to-run variation.Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Pass rate on eight hard tasks (Strict pass)100% (24/24)Claude Code · eight hard validated tasks100% (16/16)Codex CLI · effort medium · eight hard validated tasks24 / 1695% CI: 86%–100% vs 81%–100%TieThe 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Pass rate on eight hard tasks (Lenient (format misses counted))100% (24/24)Claude Code · eight hard validated tasks100% (16/16)Codex CLI · effort medium · eight hard validated tasks24 / 1695% CI: 86%–100% vs 81%–100%TieThe 95% intervals overlap (Claude Fable 5.1 86% to 100%; GPT-6.1 Sol (Codex CLI) 81% to 100%), so this sample cannot separate them.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Total time per call on hard tasks (separate batches)16.1 sClaude Code · eight hard validated tasks13.1 sCodex CLI · effort medium · eight hard validated tasks24 / 16range: 4.5 s–90 s vs 8.5 s–61.6 sUnclearThe run ranges (fastest to slowest) overlap (Claude Fable 5.1 4.46 s to 90.0 s; GPT-6.1 Sol (Codex CLI) 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Time to first useful output on hard tasks11.6 sClaude Code · eight hard validated tasks10.2 sCodex CLI · effort medium · eight hard validated tasks24 / 16range: 2 s–85.3 s vs 6.1 s–40.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Fable 5.1 2.00 s to 85.3 s; GPT-6.1 Sol (Codex CLI) 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Output tokens per call on hard tasks (Output tokens)1,366Claude Code · eight hard validated tasks335Codex CLI · effort medium · eight hard validated tasks24 / 16none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
List-price cost per strict pass on hard tasks (calculation)Calculation$0.093Claude Code · eight hard validated tasks$0.026Codex CLI · effort medium · eight hard validated tasks24 / 16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.093 vs $0.026, 3.6x) is not tested against run-to-run variation.Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Reasoning share of output tokens per call on hard tasks (calculation)Calculation64.2%Claude Code46.3%Codex CLI · effort medium24 / 16range: 23%–97% vs 11%–87%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation$0.054Claude Code$0.0023Codex CLI · effort medium24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation$0.019Claude Code$0.0030Codex CLI · effort medium24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation$0.021Claude Code$0.020Codex CLI · effort medium24 / 16none recordedUnclearMore or fewer usd is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation64.2%Claude Code46.3%Codex CLI · effort medium24 / 16range: 23%–97% vs 11%–87%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation0%Claude Code41%Codex CLI · effort medium15range: 0%–74% vs 0%–71%UnclearMore or fewer percent is not better or worse by itself; this row describes behaviour, not a winner.How much of an AI bill is thinking? Reasoning tokens by model and effort
Time to first text: a 250-line answer, six models4.43 sClaude Code3.52 sCodex CLI · effort low4range: 2.3 s–4.6 s vs 2.8 s–4.4 sUnclearThe run ranges (fastest to slowest) overlap (Claude Fable 5.1 2.27 s to 4.64 s; GPT-6.1 Sol (Codex CLI) 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed after the first text: visible tokens per second (calculation)Calculation123Claude Code80Codex CLI · effort low4range: 121–131 vs 72–81UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed in characters per second after the first text (calculation)Calculation273Claude Code323Codex CLI · effort low4range: 270–293 vs 291–327UnclearMore or fewer count is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs

Model vs model

Claude Haiku 4.5 vs GPT-6 Luna (Codex CLI)

Claude Haiku 4.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. GPT-6 Luna (Codex CLI) leads on 1 row: Total time per attempt: single call vs agent loop, 5.16 s vs 39.0 s. On those rows the run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).

14 measured rows · 3 labelled calculations

Read all 17 metric rows for Claude Haiku 4.5 vs GPT-6 Luna (Codex CLI)
MetricClaude Haiku 4.5GPT-6 Luna (Codex CLI)nInterval or rangeOutcomeBasisStudy
Strict pass rate: single call vs agent loop on eight hard tasks46% (11/24)Claude Code · single call63% (10/16)Codex CLI · single call24 / 1695% CI: 28%–65% vs 39%–82%TieThe 95% intervals overlap (Claude Haiku 4.5 28% to 65%; GPT-6 Luna (Codex CLI) 39% to 82%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: Interval merge fix100% (3/3)Claude Code · single call100% (2/2)Codex CLI · single call3 / 295% CI: 44%–100% vs 34%–100%TieThe 95% intervals overlap (Claude Haiku 4.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: DST day-length fix33% (1/3)Claude Code · single call100% (2/2)Codex CLI · single call3 / 295% CI: 6.2%–79% vs 34%–100%TieThe 95% intervals overlap (Claude Haiku 4.5 6% to 79%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: CSV parser67% (2/3)Claude Code · single call100% (2/2)Codex CLI · single call3 / 295% CI: 21%–94% vs 34%–100%TieThe 95% intervals overlap (Claude Haiku 4.5 21% to 94%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: Event-loop order0% (0/3)Claude Code · single call0% (0/2)Codex CLI · single call3 / 295% CI: 0%–56% vs 0%–66%TieThe 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: Room schedule0% (0/3)Claude Code · single call50% (1/2)Codex CLI · single call3 / 295% CI: 0%–56% vs 9.4%–91%TieThe 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: SemVer regex100% (3/3)Claude Code · single call50% (1/2)Codex CLI · single call3 / 295% CI: 44%–100% vs 9.4%–91%TieThe 95% intervals overlap (Claude Haiku 4.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: Money refactor67% (2/3)Claude Code · single call0% (0/2)Codex CLI · single call3 / 295% CI: 21%–94% vs 0%–66%TieThe 95% intervals overlap (Claude Haiku 4.5 21% to 94%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: SQL report0% (0/3)Claude Code · single call100% (2/2)Codex CLI · single call3 / 295% CI: 0%–56% vs 34%–100%TieThe 95% intervals overlap (Claude Haiku 4.5 0% to 56%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Total time per attempt: single call vs agent loop39.0 sClaude Code · single call5.16 sCodex CLI · single call24 / 16range: 15.3 s–75.1 s vs 3.6 s–11.3 sGPT-6 Luna (Codex CLI) aheadThe run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 15.3 s to 75.1 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s). A range is not a confidence interval.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))3,941Claude Code · single call11,582Codex CLI · single call24 / 16range: 3,879–4,221 vs 11,526–11,818UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Tokens per attempt: single call vs agent loop (Output tokens)5,064Claude Code · single call345Codex CLI · single call24 / 16range: 1,899–9,321 vs 36–634UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Tool calls per agent-loop attempt3Claude Code · agent loop0Codex CLI · agent loop24 / 14range: 2–18 vs 0–1UnclearMore or fewer count is not better or worse by itself; this row describes behaviour, not a winner.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
List-price cost per strict pass: single call vs agent loop (calculation)Calculation$0.067Claude Code · single call$0.0012Codex CLI · single call24 / 16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.067 vs $0.0012, 58x) is not tested against run-to-run variation.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Time to first text: a 250-line answer, six models4.00 sClaude Code3.30 sCodex CLI · effort low4range: 2.8 s–6.4 s vs 3.2 s–3.5 sUnclearThe run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed after the first text: visible tokens per second (calculation)Calculation153Claude Code129Codex CLI · effort low4range: 153–216 vs 56–259UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed in characters per second after the first text (calculation)Calculation547Claude Code524Codex CLI · effort low3 / 4range: 546–548 vs 225–1,052UnclearMore or fewer count is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs

Model vs model

Claude Sonnet 5.5 vs GPT-6 Luna (Codex CLI)

Claude Sonnet 5.5 and GPT-6 Luna (Codex CLI) share 14 measured metrics and 3 list-price calculations from 2 studies. Claude Sonnet 5.5 leads on 1 row: Strict pass rate: single call vs agent loop on eight hard tasks, 100% (24/24) vs 63% (10/16). On those rows the 95% intervals do not overlap. The other rows are 9 ties and 7 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 2 at the smallest).

14 measured rows · 3 labelled calculations

Read all 17 metric rows for Claude Sonnet 5.5 vs GPT-6 Luna (Codex CLI)
MetricClaude Sonnet 5.5GPT-6 Luna (Codex CLI)nInterval or rangeOutcomeBasisStudy
Strict pass rate: single call vs agent loop on eight hard tasks100% (24/24)Claude Code · single call63% (10/16)Codex CLI · single call24 / 1695% CI: 86%–100% vs 39%–82%Claude Sonnet 5.5 aheadThe 95% intervals do not overlap (Claude Sonnet 5.5 86% to 100%; GPT-6 Luna (Codex CLI) 39% to 82%).Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: Interval merge fix100% (3/3)Claude Code · single call100% (2/2)Codex CLI · single call3 / 295% CI: 44%–100% vs 34%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: DST day-length fix100% (3/3)Claude Code · single call100% (2/2)Codex CLI · single call3 / 295% CI: 44%–100% vs 34%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: CSV parser100% (3/3)Claude Code · single call100% (2/2)Codex CLI · single call3 / 295% CI: 44%–100% vs 34%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: Event-loop order100% (3/3)Claude Code · single call0% (0/2)Codex CLI · single call3 / 295% CI: 44%–100% vs 0%–66%TieThe 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: Room schedule100% (3/3)Claude Code · single call50% (1/2)Codex CLI · single call3 / 295% CI: 44%–100% vs 9.4%–91%TieThe 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: SemVer regex100% (3/3)Claude Code · single call50% (1/2)Codex CLI · single call3 / 295% CI: 44%–100% vs 9.4%–91%TieThe 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 9% to 91%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: Money refactor100% (3/3)Claude Code · single call0% (0/2)Codex CLI · single call3 / 295% CI: 44%–100% vs 0%–66%TieThe 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 0% to 66%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Strict passes per task: single call vs agent loop: SQL report100% (3/3)Claude Code · single call100% (2/2)Codex CLI · single call3 / 295% CI: 44%–100% vs 34%–100%TieThe 95% intervals overlap (Claude Sonnet 5.5 44% to 100%; GPT-6 Luna (Codex CLI) 34% to 100%), so this sample cannot separate them.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Total time per attempt: single call vs agent loop7.75 sClaude Code · single call5.16 sCodex CLI · single call24 / 16range: 2.3 s–34.8 s vs 3.6 s–11.3 sUnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 2.26 s to 34.8 s; GPT-6 Luna (Codex CLI) 3.59 s to 11.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))2,281Claude Code · single call11,582Codex CLI · single call24 / 16range: 2,234–2,669 vs 11,526–11,818UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Tokens per attempt: single call vs agent loop (Output tokens)1,050Claude Code · single call345Codex CLI · single call24 / 16range: 176–3,895 vs 36–634UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Tool calls per agent-loop attempt0Claude Code · agent loop0Codex CLI · agent loop16 / 14range: 0–3 vs 0–1TieSame value. More or fewer is not better by itself for this metric.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
List-price cost per strict pass: single call vs agent loop (calculation)Calculation$0.014Claude Code · single call$0.0012Codex CLI · single call24 / 16none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.014 vs $0.0012, 12x) is not tested against run-to-run variation.Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
Time to first text: a 250-line answer, six models1.96 sClaude Code3.30 sCodex CLI · effort low4range: 0.9 s–4.1 s vs 3.2 s–3.5 sUnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 0.88 s to 4.09 s; GPT-6 Luna (Codex CLI) 3.19 s to 3.47 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed after the first text: visible tokens per second (calculation)Calculation232Claude Code129Codex CLI · effort low4range: 230–233 vs 56–259UnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs
Output speed in characters per second after the first text (calculation)Calculation517Claude Code524Codex CLI · effort low4range: 513–519 vs 225–1,052UnclearMore or fewer count is not better or worse by itself; this row describes behaviour, not a winner.Where the seconds go: first text, output speed and prompt size for 6 LLMs

Model vs model

Claude Haiku 4.5 vs GPT 5.2

Claude Haiku 4.5 and GPT 5.2 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.

3 measured rows · 0 labelled calculations

Read all 3 metric rows for Claude Haiku 4.5 vs GPT 5.2
MetricClaude Haiku 4.5GPT 5.2nInterval or rangeOutcomeBasisStudy
Resolved rate on the same 33 SWE-bench Verified instances76% (25/33)effort high · public mini-SWE-agent v2 run, same instances85% (28/33)effort high · public mini-SWE-agent v2 run, same instances3395% CI: 59%–87% vs 69%–93%TieThe 95% intervals overlap (Claude Haiku 4.5 59% to 87%; GPT 5.2 69% to 93%), so this sample cannot separate them.Agent on SWE-bench Verified vs 11 public models
Model calls per instance68.5effort high · public mini-SWE-agent v2 run, same instances35.6effort high · public mini-SWE-agent v2 run, same instances33none recordedUnclearMore or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.Agent on SWE-bench Verified vs 11 public models
Recorded cost per resolved instance: Agent vs the public panel$0.48effort high · public mini-SWE-agent v2 run, same instances$0.63effort high · public mini-SWE-agent v2 run, same instances25 / 28none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.48 vs $0.63) is not tested against run-to-run variation.What if every call ran on Opus? Repricing real agent tokens

Model vs model

Claude Haiku 4.5 vs GPT 5 mini

Claude Haiku 4.5 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.

3 measured rows · 0 labelled calculations

Read all 3 metric rows for Claude Haiku 4.5 vs GPT 5 mini
MetricClaude Haiku 4.5GPT 5 mininInterval or rangeOutcomeBasisStudy
Resolved rate on the same 33 SWE-bench Verified instances76% (25/33)effort high · public mini-SWE-agent v2 run, same instances64% (21/33)public mini-SWE-agent v2 run, same instances3395% CI: 59%–87% vs 47%–78%TieThe 95% intervals overlap (Claude Haiku 4.5 59% to 87%; GPT 5 mini 47% to 78%), so this sample cannot separate them.Agent on SWE-bench Verified vs 11 public models
Model calls per instance68.5effort high · public mini-SWE-agent v2 run, same instances20.8public mini-SWE-agent v2 run, same instances33none recordedUnclearMore or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.Agent on SWE-bench Verified vs 11 public models
Recorded cost per resolved instance: Agent vs the public panel$0.48effort high · public mini-SWE-agent v2 run, same instances$0.080public mini-SWE-agent v2 run, same instances25 / 21none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.48 vs $0.080, 6.0x) is not tested against run-to-run variation.What if every call ran on Opus? Repricing real agent tokens

Model vs model

Claude Opus 4.5 vs GPT 5 mini

Claude Opus 4.5 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.

3 measured rows · 0 labelled calculations

Read all 3 metric rows for Claude Opus 4.5 vs GPT 5 mini
MetricClaude Opus 4.5GPT 5 mininInterval or rangeOutcomeBasisStudy
Resolved rate on the same 33 SWE-bench Verified instances73% (24/33)effort high · public mini-SWE-agent v2 run, same instances64% (21/33)public mini-SWE-agent v2 run, same instances3395% CI: 56%–85% vs 47%–78%TieThe 95% intervals overlap (Claude Opus 4.5 56% to 85%; GPT 5 mini 47% to 78%), so this sample cannot separate them.Agent on SWE-bench Verified vs 11 public models
Model calls per instance35.9effort high · public mini-SWE-agent v2 run, same instances20.8public mini-SWE-agent v2 run, same instances33none recordedUnclearMore or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.Agent on SWE-bench Verified vs 11 public models
Recorded cost per resolved instance: Agent vs the public panel$1.18effort high · public mini-SWE-agent v2 run, same instances$0.080public mini-SWE-agent v2 run, same instances24 / 21none recordedUnclearNo interval or range was recorded for either side, so the gap ($1.18 vs $0.080, 15x) is not tested against run-to-run variation.What if every call ran on Opus? Repricing real agent tokens

Model vs model

Claude Opus 4.6 vs GPT 5 mini

Claude Opus 4.6 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.

3 measured rows · 0 labelled calculations

Read all 3 metric rows for Claude Opus 4.6 vs GPT 5 mini
MetricClaude Opus 4.6GPT 5 mininInterval or rangeOutcomeBasisStudy
Resolved rate on the same 33 SWE-bench Verified instances70% (23/33)public mini-SWE-agent v2 run, same instances64% (21/33)public mini-SWE-agent v2 run, same instances3395% CI: 53%–83% vs 47%–78%TieThe 95% intervals overlap (Claude Opus 4.6 53% to 83%; GPT 5 mini 47% to 78%), so this sample cannot separate them.Agent on SWE-bench Verified vs 11 public models
Model calls per instance28.9public mini-SWE-agent v2 run, same instances20.8public mini-SWE-agent v2 run, same instances33none recordedUnclearMore or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.Agent on SWE-bench Verified vs 11 public models
Recorded cost per resolved instance: Agent vs the public panel$0.88public mini-SWE-agent v2 run, same instances$0.080public mini-SWE-agent v2 run, same instances23 / 21none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.88 vs $0.080, 11x) is not tested against run-to-run variation.What if every call ran on Opus? Repricing real agent tokens

Model vs model

Claude Sonnet 4.5 vs GPT 5 mini

Claude Sonnet 4.5 and GPT 5 mini share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.

3 measured rows · 0 labelled calculations

Read all 3 metric rows for Claude Sonnet 4.5 vs GPT 5 mini
MetricClaude Sonnet 4.5GPT 5 mininInterval or rangeOutcomeBasisStudy
Resolved rate on the same 33 SWE-bench Verified instances76% (25/33)effort high · public mini-SWE-agent v2 run, same instances64% (21/33)public mini-SWE-agent v2 run, same instances3395% CI: 59%–87% vs 47%–78%TieThe 95% intervals overlap (Claude Sonnet 4.5 59% to 87%; GPT 5 mini 47% to 78%), so this sample cannot separate them.Agent on SWE-bench Verified vs 11 public models
Model calls per instance51effort high · public mini-SWE-agent v2 run, same instances20.8public mini-SWE-agent v2 run, same instances33none recordedUnclearMore or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.Agent on SWE-bench Verified vs 11 public models
Recorded cost per resolved instance: Agent vs the public panel$0.91effort high · public mini-SWE-agent v2 run, same instances$0.080public mini-SWE-agent v2 run, same instances25 / 21none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.91 vs $0.080, 11x) is not tested against run-to-run variation.What if every call ran on Opus? Repricing real agent tokens

Model vs model

Claude Sonnet 5.5 vs GPT-6.1 Sol (OpenAI API)

Claude Sonnet 5.5 and GPT-6.1 Sol (OpenAI API) share 3 measured metrics from one study. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 3 unclear; each row says why. Every row ran the two sides through different routes (for example Claude Code vs OpenAI API), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).

3 measured rows · 0 labelled calculations

Read all 3 metric rows for Claude Sonnet 5.5 vs GPT-6.1 Sol (OpenAI API)
MetricClaude Sonnet 5.5GPT-6.1 Sol (OpenAI API)nInterval or rangeOutcomeBasisStudy
Repairing a scheduler: Claude Code vs Codex vs API (Total time)15.0 sClaude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs17.3 sOpenAI API · effort medium · scheduler repair, 296 checks, 3 runs3range: 13.9 s–15.9 s vs 16.3 s–18.6 sUnclearOnly 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 13.9 s to 15.9 s; GPT-6.1 Sol (OpenAI API) 16.3 s to 18.6 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval.Claude Code CLI vs Codex CLI vs the API: latency and tokens
Repairing a scheduler: Claude Code vs Codex vs API (First useful output)7.55 sClaude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs7.46 sOpenAI API · effort medium · scheduler repair, 296 checks, 3 runs3range: 6.8 s–7.6 s vs 6.7 s–9.1 sUnclearThe run ranges (fastest to slowest) overlap (Claude Sonnet 5.5 6.77 s to 7.63 s; GPT-6.1 Sol (OpenAI API) 6.68 s to 9.05 s); the medians alone do not show a reliable difference. A range is not a confidence interval.Claude Code CLI vs Codex CLI vs the API: latency and tokens
Output tokens to repair the scheduler (Output tokens)2,227Claude Code CLI · effort medium · scheduler repair, 296 checks, 3 runs1,313OpenAI API · effort medium · scheduler repair, 296 checks, 3 runs3none recordedUnclearMore or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner.Claude Code CLI vs Codex CLI vs the API: latency and tokens

Model vs model

GPT 5.2 vs Claude Opus 4.5

GPT 5.2 and Claude Opus 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.

3 measured rows · 0 labelled calculations

Read all 3 metric rows for GPT 5.2 vs Claude Opus 4.5
MetricGPT 5.2Claude Opus 4.5nInterval or rangeOutcomeBasisStudy
Resolved rate on the same 33 SWE-bench Verified instances85% (28/33)effort high · public mini-SWE-agent v2 run, same instances73% (24/33)effort high · public mini-SWE-agent v2 run, same instances3395% CI: 69%–93% vs 56%–85%TieThe 95% intervals overlap (GPT 5.2 69% to 93%; Claude Opus 4.5 56% to 85%), so this sample cannot separate them.Agent on SWE-bench Verified vs 11 public models
Model calls per instance35.6effort high · public mini-SWE-agent v2 run, same instances35.9effort high · public mini-SWE-agent v2 run, same instances33none recordedUnclearMore or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.Agent on SWE-bench Verified vs 11 public models
Recorded cost per resolved instance: Agent vs the public panel$0.63effort high · public mini-SWE-agent v2 run, same instances$1.18effort high · public mini-SWE-agent v2 run, same instances28 / 24none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.63 vs $1.18) is not tested against run-to-run variation.What if every call ran on Opus? Repricing real agent tokens

Model vs model

GPT 5.2 vs Claude Opus 4.6

GPT 5.2 and Claude Opus 4.6 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.

3 measured rows · 0 labelled calculations

Read all 3 metric rows for GPT 5.2 vs Claude Opus 4.6
MetricGPT 5.2Claude Opus 4.6nInterval or rangeOutcomeBasisStudy
Resolved rate on the same 33 SWE-bench Verified instances85% (28/33)effort high · public mini-SWE-agent v2 run, same instances70% (23/33)public mini-SWE-agent v2 run, same instances3395% CI: 69%–93% vs 53%–83%TieThe 95% intervals overlap (GPT 5.2 69% to 93%; Claude Opus 4.6 53% to 83%), so this sample cannot separate them.Agent on SWE-bench Verified vs 11 public models
Model calls per instance35.6effort high · public mini-SWE-agent v2 run, same instances28.9public mini-SWE-agent v2 run, same instances33none recordedUnclearMore or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.Agent on SWE-bench Verified vs 11 public models
Recorded cost per resolved instance: Agent vs the public panel$0.63effort high · public mini-SWE-agent v2 run, same instances$0.88public mini-SWE-agent v2 run, same instances28 / 23none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.63 vs $0.88) is not tested against run-to-run variation.What if every call ran on Opus? Repricing real agent tokens

Model vs model

GPT 5.2 vs Claude Sonnet 4.5

GPT 5.2 and Claude Sonnet 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.

3 measured rows · 0 labelled calculations

Read all 3 metric rows for GPT 5.2 vs Claude Sonnet 4.5
MetricGPT 5.2Claude Sonnet 4.5nInterval or rangeOutcomeBasisStudy
Resolved rate on the same 33 SWE-bench Verified instances85% (28/33)effort high · public mini-SWE-agent v2 run, same instances76% (25/33)effort high · public mini-SWE-agent v2 run, same instances3395% CI: 69%–93% vs 59%–87%TieThe 95% intervals overlap (GPT 5.2 69% to 93%; Claude Sonnet 4.5 59% to 87%), so this sample cannot separate them.Agent on SWE-bench Verified vs 11 public models
Model calls per instance35.6effort high · public mini-SWE-agent v2 run, same instances51effort high · public mini-SWE-agent v2 run, same instances33none recordedUnclearMore or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.Agent on SWE-bench Verified vs 11 public models
Recorded cost per resolved instance: Agent vs the public panel$0.63effort high · public mini-SWE-agent v2 run, same instances$0.91effort high · public mini-SWE-agent v2 run, same instances28 / 25none recordedUnclearNo interval or range was recorded for either side, so the gap ($0.63 vs $0.91) is not tested against run-to-run variation.What if every call ran on Opus? Repricing real agent tokens

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.