66 measured metrics · 22 calculated · 12 studies
Claude CodevsCodex CLI
Claude Code ahead on 7, Codex CLI ahead on 1; 31 ties, 49 unclear. A side is ahead only where the intervals or ranges do not overlap.
The verdict
Claude Code and Codex CLI share 66 measured metrics and 22 list-price calculations from 12 studies. Claude Code leads on 7 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 4 more. Codex CLI leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 31 ties and 49 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Each run pairs a CLI with a model, so these rows cannot separate the CLI from the model; the contexts name both. Some rows rest on small samples (n = 2 at the smallest).
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | Claude Code | Codex CLI | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 80% (12/15)Claude Sonnet 5.5 · five short validated tasks | 100% (15/15)GPT-6.1 Sol · effort medium · five short validated tasks | 15 | 95% CI: 55%–93% vs 80%–100% | Tie | The 95% intervals overlap (Claude Code 55% to 93%; Codex CLI 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 2.31 sClaude Sonnet 5.5 · five short validated tasks | 5.65 sGPT-6.1 Sol · effort medium · five short validated tasks | 15 | range: 2.2 s–7.7 s vs 4.1 s–25.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 2.17 s to 7.73 s; Codex CLI 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 1.56 sClaude Sonnet 5.5 · five short validated tasks | 5.05 sGPT-6.1 Sol · effort medium · five short validated tasks | 15 | range: 1 s–6.4 s vs 3.4 s–17.8 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 0.99 s to 6.39 s; Codex CLI 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 1,401Claude Sonnet 5.5 · five short validated tasks | 5,180GPT-6.1 Sol · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 107Claude Sonnet 5.5 · five short validated tasks | 42GPT-6.1 Sol · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.0062Claude Sonnet 5.5 · five short validated tasks | $0.016GPT-6.1 Sol · effort medium · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.0062 vs $0.016, 2.5x) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Input tokens per call: what the CLI sends (Cache read), 3.7x (Codex CLI larger).
Watch it build
A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.
Claude Code vs Codex CLI: what the measurements say
88 comparison rows from 12 studies: 7 rows favour Claude Code, 1 favour Codex CLI, 80 are ties or unclear. Cost rows are calculations.
Transcript
- Comparison · 88 rows · 12 studies. Claude Code vs Codex CLI. A winner only where the 95% intervals or run ranges do not overlap.
- 88 comparison rows from 12 studies: Claude Code ahead on 7, Codex CLI ahead on 1. The rest do not separate them. Rows where Claude Code is ahead: 7 (of 88). Rows where Codex CLI is ahead: 1 (of 88). Ties or unclear: 80 (31 ties · 49 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
- Coding agents, hidden tests: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 4 rows separates them. Table: Coding agents, hidden tests · Claude Sonnet 5.5 vs GPT-6.1 Sol · n = 12 per side. Source study: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks. Rows shown: Coding sessions that passed every hidden check; Time per coding session; Tool calls per coding session; List-price cost per passing coding session (calculation). Recorded settings: Claude Sonnet 5.5 · six small repository tasks with hidden tests; GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests. Includes a calculation, not a bill or a new run. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
- Caching sessions: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 2 of 9 rows separate them. Table: Caching sessions · 5 of 9 rows · Claude Sonnet 5.5 vs GPT-6.1 Sol · n = 10 per side. Source study: Prompt caching and run-to-run consistency in Claude Code and Codex CLI. Rows shown: Same prompt, 10 times: strict pass rate (Exact number); Same prompt, 10 times: strict pass rate (JSON object); Same prompt, 10 times: strict pass rate (Code fix); Same prompt, 10 times: time per call (Exact number); Same prompt, 10 times: time per call (Code fix). Recorded settings: Claude Sonnet 5.5 · same prompt repeated 10 times; GPT-6.1 Sol · effort medium · same prompt repeated 10 times. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- Routing overhead: pass rate not measured. 2 of 4 rows separate them. Table: Routing overhead · Claude Haiku 4.5 vs default model · n = 5 per side. Source study: Routing overhead: deterministic policy vs LLM routers vs Jev. Rows shown: CLI start-up tax on a one-word answer (First output event); CLI start-up tax on a one-word answer (First model output); CLI start-up tax on a one-word answer (Total wall time); Input tokens a CLI sends for a one-word answer. Recorded settings: Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs; default model · CLI start-up, one-word prompt, 5 runs. Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
- Single call vs agent loop: pass rate 100% vs 63%, Claude Code ahead: 95% intervals separate. 1 of 14 rows separates them. Table: Single call vs agent loop · 5 of 14 rows · Claude Sonnet 5.5 vs GPT-6 Luna · n = 3–24 vs 2–16. Source study: Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks. Rows shown: Strict pass rate: single call vs agent loop on eight hard tasks; Strict passes per task: single call vs agent loop: Interval merge fix; Strict passes per task: single call vs agent loop: DST day-length fix; Strict passes per task: single call vs agent loop: CSV parser; Strict passes per task: single call vs agent loop: Event-loop order. Recorded settings: Claude Sonnet 5.5 · single call; GPT-6 Luna · single call. Caveat: Each row is a CLI + model pair. Claude Code and Codex CLI add their own system prompts and tool schemas, and Codex CLI also loads the account’s user-level instruction file. A gap between Claude and GPT-6 Luna rows is partly the CLI.
- No winner where the data shows none. Showing 18 of 88 rows; every row and its reason online.
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- Claude Code
- Codex CLI
- 95% interval
- fastest–slowest run (not an interval)
- where the two overlap
- hollow: list-price calculation
These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
- Pass rate on five validated tasks80% (12/15)n 15100% (15/15)n 15TiePass rate on five validated tasks: Claude Code 80% (12/15) (n 15, 95% interval 55%–93%); Codex CLI 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
Calculation: at these rates, about 33 runs per side would separate them.
- Total time per call2.31 sn 155.65 sn 15UnclearTotal time per call: Claude Code 2.31 s (n 15, run range 2.2 s–7.7 s); Codex CLI 5.65 s (n 15, run range 4.1 s–25.5 s). Unclear.
- Time to first useful output1.56 sn 155.05 sn 15UnclearTime to first useful output: Claude Code 1.56 s (n 15, run range 1 s–6.4 s); Codex CLI 5.05 s (n 15, run range 3.4 s–17.8 s). Unclear.
- Input tokens per call: what the CLI sends (Cache read)1,401n 155,180n 15UnclearInput tokens per call: what the CLI sends (Cache read): Claude Code 1,401 (n 15); Codex CLI 5,180 (n 15). Unclear.
- Input tokens per call: what the CLI sends (Other input)685n 156,943n 15UnclearInput tokens per call: what the CLI sends (Other input): Claude Code 685 (n 15); Codex CLI 6,943 (n 15). Unclear.
- Output tokens per call (Output tokens)107n 1542n 15UnclearOutput tokens per call (Output tokens): Claude Code 107 (n 15); Codex CLI 42 (n 15). Unclear.
- List-price cost per call (calculation)Calculation$0.0036n 15$0.010n 15UnclearList-price cost per call (calculation), calculation: Claude Code $0.0036 (n 15, run range $0.0034–$0.01); Codex CLI $0.010 (n 15, run range $0.0054–$0.027). Unclear.
- List-price cost per passing answer (calculation)Calculation$0.0062n 15$0.016n 15UnclearList-price cost per passing answer (calculation), calculation: Claude Code $0.0062 (n 15); Codex CLI $0.016 (n 15). Unclear.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
- Pass rate on eight hard tasks (Strict pass)100% (24/24)n 24100% (16/16)n 16TiePass rate on eight hard tasks (Strict pass): Claude Code 100% (24/24) (n 24, 95% interval 86%–100%); Codex CLI 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
- Pass rate on eight hard tasks (Lenient (format misses counted))100% (24/24)n 24100% (16/16)n 16TiePass rate on eight hard tasks (Lenient (format misses counted)): Claude Code 100% (24/24) (n 24, 95% interval 86%–100%); Codex CLI 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
- Total time per call on hard tasks (separate batches)7.75 sn 2413.1 sn 16UnclearTotal time per call on hard tasks (separate batches): Claude Code 7.75 s (n 24, run range 2.3 s–34.8 s); Codex CLI 13.1 s (n 16, run range 8.5 s–61.6 s). Unclear.
- Time to first useful output on hard tasks5.95 sn 2410.2 sn 16UnclearTime to first useful output on hard tasks: Claude Code 5.95 s (n 24, run range 0.9 s–30.6 s); Codex CLI 10.2 s (n 16, run range 6.1 s–40.4 s). Unclear.
- Output tokens per call on hard tasks (Output tokens)1,050n 24335n 16UnclearOutput tokens per call on hard tasks (Output tokens): Claude Code 1,050 (n 24); Codex CLI 335 (n 16). Unclear.
- List-price cost per strict pass on hard tasks (calculation)Calculation$0.014n 24$0.026n 16UnclearList-price cost per strict pass on hard tasks (calculation), calculation: Claude Code $0.014 (n 24); Codex CLI $0.026 (n 16). Unclear.
Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
- Coding sessions that passed every hidden check100% (12/12)n 12100% (12/12)n 12TieCoding sessions that passed every hidden check: Claude Code 100% (12/12) (n 12, 95% interval 76%–100%); Codex CLI 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
- Time per coding session23.1 sn 12113.4 sn 12Claude Code aheadTime per coding session: Claude Code 23.1 s (n 12, run range 18.7 s–44.5 s); Codex CLI 113.4 s (n 12, run range 78.5 s–222 s). Claude Code ahead.
- Tool calls per coding session7.5n 1212.5n 12UnclearTool calls per coding session: Claude Code 7.5 (n 12, run range 3–14); Codex CLI 12.5 (n 12, run range 8–18). Unclear.
- List-price cost per passing coding session (calculation)Calculation$0.085n 12$0.098n 12UnclearList-price cost per passing coding session (calculation), calculation: Claude Code $0.085 (n 12); Codex CLI $0.098 (n 12). Unclear.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
- Strict pass rate by effort on eight hard tasks100% (16/16)n 16100% (16/16)n 16TieStrict pass rate by effort on eight hard tasks: Claude Code 100% (16/16) (n 16, 95% interval 81%–100%); Codex CLI 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
- Total time per call by effort on hard tasks7.63 sn 1613.1 sn 16UnclearTotal time per call by effort on hard tasks: Claude Code 7.63 s (n 16, run range 2.7 s–24 s); Codex CLI 13.1 s (n 16, run range 8.5 s–61.6 s). Unclear.
- Output tokens per call by effort on hard tasks (Output tokens)770n 16335n 16UnclearOutput tokens per call by effort on hard tasks (Output tokens): Claude Code 770 (n 16); Codex CLI 335 (n 16). Unclear.
- List-price cost per strict pass by effort (calculation)Calculation$0.014n 16$0.026n 16UnclearList-price cost per strict pass by effort (calculation), calculation: Claude Code $0.014 (n 16); Codex CLI $0.026 (n 16). Unclear.
Prompt caching and run-to-run consistency in Claude Code and Codex CLI
- Same prompt, 10 times: strict pass rate (Exact number)100% (10/10)n 10100% (10/10)n 10TieSame prompt, 10 times: strict pass rate (Exact number): Claude Code 100% (10/10) (n 10, 95% interval 72%–100%); Codex CLI 100% (10/10) (n 10, 95% interval 72%–100%). Tie.
- Same prompt, 10 times: strict pass rate (JSON object)100% (10/10)n 10100% (10/10)n 10TieSame prompt, 10 times: strict pass rate (JSON object): Claude Code 100% (10/10) (n 10, 95% interval 72%–100%); Codex CLI 100% (10/10) (n 10, 95% interval 72%–100%). Tie.
- Same prompt, 10 times: strict pass rate (Code fix)100% (10/10)n 10100% (10/10)n 10TieSame prompt, 10 times: strict pass rate (Code fix): Claude Code 100% (10/10) (n 10, 95% interval 72%–100%); Codex CLI 100% (10/10) (n 10, 95% interval 72%–100%). Tie.
- Same prompt, 10 times: how many different answers (Exact number)1n 101n 10TieSame prompt, 10 times: how many different answers (Exact number): Claude Code 1 (n 10); Codex CLI 1 (n 10). Tie.
- Same prompt, 10 times: how many different answers (JSON object)1n 101n 10TieSame prompt, 10 times: how many different answers (JSON object): Claude Code 1 (n 10); Codex CLI 1 (n 10). Tie.
- Same prompt, 10 times: how many different answers (Code fix)3n 106n 10UnclearSame prompt, 10 times: how many different answers (Code fix): Claude Code 3 (n 10); Codex CLI 6 (n 10). Unclear.
- Same prompt, 10 times: time per call (Exact number)6.89 sn 1013.4 sn 10Claude Code aheadSame prompt, 10 times: time per call (Exact number): Claude Code 6.89 s (n 10, run range 5.8 s–7.8 s); Codex CLI 13.4 s (n 10, run range 12.3 s–18 s). Claude Code ahead.
- Same prompt, 10 times: time per call (JSON object)2.89 sn 106.42 sn 10UnclearSame prompt, 10 times: time per call (JSON object): Claude Code 2.89 s (n 10, run range 2.7 s–5.3 s); Codex CLI 6.42 s (n 10, run range 5.3 s–8.3 s). Unclear.
- Same prompt, 10 times: time per call (Code fix)2.67 sn 1011.3 sn 10Claude Code aheadSame prompt, 10 times: time per call (Code fix): Claude Code 2.67 s (n 10, run range 2.3 s–4.3 s); Codex CLI 11.3 s (n 10, run range 9.1 s–14.9 s). Claude Code ahead.
Routing overhead: deterministic policy vs LLM routers vs Jev
- CLI start-up tax on a one-word answer (First output event)563 msn 5489 msn 5UnclearCLI start-up tax on a one-word answer (First output event): Claude Code 563 ms (n 5, run range 519 ms–726 ms); Codex CLI 489 ms (n 5, run range 354 ms–1.3 s). Unclear.
- CLI start-up tax on a one-word answer (First model output)1,461 msn 55,059 msn 5Claude Code aheadCLI start-up tax on a one-word answer (First model output): Claude Code 1,461 ms (n 5, run range 1.21 s–2.31 s); Codex CLI 5,059 ms (n 5, run range 4.39 s–5.48 s). Claude Code ahead.
- CLI start-up tax on a one-word answer (Total wall time)2,529 msn 55,999 msn 5Claude Code aheadCLI start-up tax on a one-word answer (Total wall time): Claude Code 2,529 ms (n 5, run range 2.27 s–3.38 s); Codex CLI 5,999 ms (n 5, run range 5.37 s–6.51 s). Claude Code ahead.
- Input tokens a CLI sends for a one-word answer6,761n 517,051n 5UnclearInput tokens a CLI sends for a one-word answer: Claude Code 6,761 (n 5); Codex CLI 17,051 (n 5). Unclear.
Claude Code CLI vs Codex CLI vs the API: latency and tokens
- Repairing a scheduler: Claude Code vs Codex vs API (Total time)15.0 sn 361.2 sn 3UnclearRepairing a scheduler: Claude Code vs Codex vs API (Total time): Claude Code 15.0 s (n 3, run range 13.9 s–15.9 s); Codex CLI 61.2 s (n 3, run range 59.9 s–69.5 s). Unclear.
- Repairing a scheduler: Claude Code vs Codex vs API (First useful output)7.55 sn 315.6 sn 3UnclearRepairing a scheduler: Claude Code vs Codex vs API (First useful output): Claude Code 7.55 s (n 3, run range 6.8 s–7.6 s); Codex CLI 15.6 s (n 3, run range 13.7 s–23 s). Unclear.
- Output tokens to repair the scheduler (Output tokens)2,227n 31,181n 3UnclearOutput tokens to repair the scheduler (Output tokens): Claude Code 2,227 (n 3); Codex CLI 1,181 (n 3). Unclear.
Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
- Strict pass rate: single call vs agent loop on eight hard tasks100% (24/24)n 2463% (10/16)n 16Claude Code aheadStrict pass rate: single call vs agent loop on eight hard tasks: Claude Code 100% (24/24) (n 24, 95% interval 86%–100%); Codex CLI 63% (10/16) (n 16, 95% interval 39%–82%). Claude Code ahead.
- Strict passes per task: single call vs agent loop: Interval merge fix100% (3/3)n 3100% (2/2)n 2TieStrict passes per task: single call vs agent loop: Interval merge fix: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
- Strict passes per task: single call vs agent loop: DST day-length fix100% (3/3)n 3100% (2/2)n 2TieStrict passes per task: single call vs agent loop: DST day-length fix: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
- Strict passes per task: single call vs agent loop: CSV parser100% (3/3)n 3100% (2/2)n 2TieStrict passes per task: single call vs agent loop: CSV parser: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
- Strict passes per task: single call vs agent loop: Event-loop order100% (3/3)n 30% (0/2)n 2TieStrict passes per task: single call vs agent loop: Event-loop order: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 0% (0/2) (n 2, 95% interval 0%–66%). Tie.
Calculation: at these rates, about 4 runs per side would separate them.
- Strict passes per task: single call vs agent loop: Room schedule100% (3/3)n 350% (1/2)n 2TieStrict passes per task: single call vs agent loop: Room schedule: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 50% (1/2) (n 2, 95% interval 9.4%–91%). Tie.
Calculation: at these rates, about 12 runs per side would separate them.
- Strict passes per task: single call vs agent loop: SemVer regex100% (3/3)n 350% (1/2)n 2TieStrict passes per task: single call vs agent loop: SemVer regex: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 50% (1/2) (n 2, 95% interval 9.4%–91%). Tie.
Calculation: at these rates, about 12 runs per side would separate them.
- Strict passes per task: single call vs agent loop: Money refactor100% (3/3)n 30% (0/2)n 2TieStrict passes per task: single call vs agent loop: Money refactor: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 0% (0/2) (n 2, 95% interval 0%–66%). Tie.
Calculation: at these rates, about 4 runs per side would separate them.
- Strict passes per task: single call vs agent loop: SQL report100% (3/3)n 3100% (2/2)n 2TieStrict passes per task: single call vs agent loop: SQL report: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
- Total time per attempt: single call vs agent loop7.75 sn 245.16 sn 16UnclearTotal time per attempt: single call vs agent loop: Claude Code 7.75 s (n 24, run range 2.3 s–34.8 s); Codex CLI 5.16 s (n 16, run range 3.6 s–11.3 s). Unclear.
- Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))2,281n 2411,582n 16UnclearTokens per attempt: single call vs agent loop (Input tokens (cache reads included)): Claude Code 2,281 (n 24, run range 2,234–2,669); Codex CLI 11,582 (n 16, run range 11,526–11,818). Unclear.
- Tokens per attempt: single call vs agent loop (Output tokens)1,050n 24345n 16UnclearTokens per attempt: single call vs agent loop (Output tokens): Claude Code 1,050 (n 24, run range 176–3,895); Codex CLI 345 (n 16, run range 36–634). Unclear.
- Tool calls per agent-loop attempt0n 160n 14TieTool calls per agent-loop attempt: Claude Code 0 (n 16, run range 0–3); Codex CLI 0 (n 14, run range 0–1). Tie.
- List-price cost per strict pass: single call vs agent loop (calculation)Calculation$0.014n 24$0.0012n 16UnclearList-price cost per strict pass: single call vs agent loop (calculation), calculation: Claude Code $0.014 (n 24); Codex CLI $0.0012 (n 16). Unclear.
Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
- Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)Calculation100% (12/12)n 12100% (12/12)n 12TieDoes a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON), calculation: Claude Code 100% (12/12) (n 12, 95% interval 76%–100%); Codex CLI 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
- Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))Calculation100% (12/12)n 12100% (12/12)n 12TieDoes a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss)), calculation: Claude Code 100% (12/12) (n 12, 95% interval 76%–100%); Codex CLI 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
- What each call produced: strict pass, format miss, wrong values or error (Strict pass)12n 1212n 12TieWhat each call produced: strict pass, format miss, wrong values or error (Strict pass): Claude Code 12 (n 12); Codex CLI 12 (n 12). Tie.
- What each call produced: strict pass, format miss, wrong values or error (Format miss)0n 120n 12TieWhat each call produced: strict pass, format miss, wrong values or error (Format miss): Claude Code 0 (n 12); Codex CLI 0 (n 12). Tie.
- What each call produced: strict pass, format miss, wrong values or error (Wrong values)0n 120n 12TieWhat each call produced: strict pass, format miss, wrong values or error (Wrong values): Claude Code 0 (n 12); Codex CLI 0 (n 12). Tie.
- What each call produced: strict pass, format miss, wrong values or error (Error)0n 120n 12TieWhat each call produced: strict pass, format miss, wrong values or error (Error): Claude Code 0 (n 12); Codex CLI 0 (n 12). Tie.
- Time per call, instructions vs schema mode3.52 sn 126.21 sn 12Claude Code aheadTime per call, instructions vs schema mode: Claude Code 3.52 s (n 12, run range 2.7 s–4.1 s); Codex CLI 6.21 s (n 12, run range 4.2 s–12.3 s). Claude Code ahead.
- Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call)368n 12117n 12UnclearOutput and reasoning tokens per call, instructions vs schema mode (Median output tokens per call): Claude Code 368 (n 12); Codex CLI 117 (n 12). Unclear.
How much of an AI bill is thinking? Reasoning tokens by model and effort
- Reasoning share of output tokens per call on hard tasks (calculation)Calculation54.5%n 2446.3%n 16UnclearReasoning share of output tokens per call on hard tasks (calculation), calculation: Claude Code 54.5% (n 24, run range 0%–96%); Codex CLI 46.3% (n 16, run range 11%–87%). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation$0.0067n 24$0.0023n 16UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Claude Code $0.0067 (n 24); Codex CLI $0.0023 (n 16). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation$0.0037n 24$0.0030n 16UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Claude Code $0.0037 (n 24); Codex CLI $0.0030 (n 16). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation$0.0040n 24$0.020n 16UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Claude Code $0.0040 (n 24); Codex CLI $0.020 (n 16). Unclear.
- Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation$0.0059n 16$0.0023n 16UnclearReasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass), calculation: Claude Code $0.0059 (n 16); Codex CLI $0.0023 (n 16). Unclear.
- Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation$0.014n 16$0.026n 16UnclearReasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass), calculation: Claude Code $0.014 (n 16); Codex CLI $0.026 (n 16). Unclear.
- Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation54.5%n 2446.3%n 16UnclearReasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Claude Code 54.5% (n 24, run range 0%–96%); Codex CLI 46.3% (n 16, run range 11%–87%). Unclear.
- Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation0%n 1541%n 15UnclearReasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Claude Code 0% (n 15, run range 0%–73%); Codex CLI 41% (n 15, run range 0%–71%). Unclear.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
- Time to first text: a 250-line answer, six models1.96 sn 43.52 sn 4UnclearTime to first text: a 250-line answer, six models: Claude Code 1.96 s (n 4, run range 0.9 s–4.1 s); Codex CLI 3.52 s (n 4, run range 2.8 s–4.4 s). Unclear.
- Output speed after the first text: visible tokens per second (calculation)Calculation232n 480n 4UnclearOutput speed after the first text: visible tokens per second (calculation), calculation: Claude Code 232 (n 4, run range 230–233); Codex CLI 80 (n 4, run range 72–81). Unclear.
- Output speed in characters per second after the first text (calculation)Calculation517n 4323n 4UnclearOutput speed in characters per second after the first text (calculation), calculation: Claude Code 517 (n 4, run range 513–519); Codex CLI 323 (n 4, run range 291–327). Unclear.
- Time to first text as the prompt grows: 1kCalculation1.45 sn 33.36 sn 3UnclearTime to first text as the prompt grows: 1k, calculation: Claude Code 1.45 s (n 3, run range 1.2 s–1.7 s); Codex CLI 3.36 s (n 3, run range 3.4 s–4.8 s). Unclear.
- Time to first text as the prompt grows: 16kCalculation1.78 sn 34.02 sn 3UnclearTime to first text as the prompt grows: 16k, calculation: Claude Code 1.78 s (n 3, run range 1.6 s–2.1 s); Codex CLI 4.02 s (n 3, run range 3.3 s–4.3 s). Unclear.
- Time to first text as the prompt grows: 64kCalculation3.07 sn 33.93 sn 3UnclearTime to first text as the prompt grows: 64k, calculation: Claude Code 3.07 s (n 3, run range 1.4 s–3.6 s); Codex CLI 3.93 s (n 3, run range 3.4 s–4.4 s). Unclear.
- Total time per call by prompt size (1k prompt)1.78 sn 33.43 sn 3UnclearTotal time per call by prompt size (1k prompt): Claude Code 1.78 s (n 3, run range 1.6 s–2.1 s); Codex CLI 3.43 s (n 3, run range 3.4 s–4.9 s). Unclear.
- Total time per call by prompt size (16k prompt)2.10 sn 34.14 sn 3UnclearTotal time per call by prompt size (16k prompt): Claude Code 2.10 s (n 3, run range 2 s–2.5 s); Codex CLI 4.14 s (n 3, run range 4 s–4.7 s). Unclear.
- Total time per call by prompt size (64k prompt)3.44 sn 33.96 sn 3UnclearTotal time per call by prompt size (64k prompt): Claude Code 3.44 s (n 3, run range 1.7 s–4.4 s); Codex CLI 3.96 s (n 3, run range 3.5 s–4.4 s). Unclear.
- Exact lookup answers at the 1k, 16k and 64k prompt-size targets100% (9/9)n 9100% (9/9)n 9TieExact lookup answers at the 1k, 16k and 64k prompt-size targets: Claude Code 100% (9/9) (n 9, 95% interval 70%–100%); Codex CLI 100% (9/9) (n 9, 95% interval 70%–100%). Tie.
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
- Pass rate on 4 harder tasks (Strict pass)38% (6/16)n 1669% (11/16)n 16TiePass rate on 4 harder tasks (Strict pass): Claude Code 38% (6/16) (n 16, 95% interval 18%–61%); Codex CLI 69% (11/16) (n 16, 95% interval 44%–86%). Tie.
Calculation: at these rates, about 40 runs per side would separate them.
- Pass rate on 4 harder tasks (Lenient (format misses counted))38% (6/16)n 1669% (11/16)n 16TiePass rate on 4 harder tasks (Lenient (format misses counted)): Claude Code 38% (6/16) (n 16, 95% interval 18%–61%); Codex CLI 69% (11/16) (n 16, 95% interval 44%–86%). Tie.
Calculation: at these rates, about 40 runs per side would separate them.
- Calls that tried a tool although tools were off31% (5/16)n 160% (0/16)n 16UnclearCalls that tried a tool although tools were off: Claude Code 31% (5/16) (n 16, 95% interval 14%–56%); Codex CLI 0% (0/16) (n 16, 95% interval 0%–19%). Unclear.
- Strict pass rate by task: 10x10 nonogram100% (4/4)n 475% (3/4)n 4TieStrict pass rate by task: 10x10 nonogram: Claude Code 100% (4/4) (n 4, 95% interval 51%–100%); Codex CLI 75% (3/4) (n 4, 95% interval 30%–95%). Tie.
Calculation: at these rates, about 27 runs per side would separate them.
- Strict pass rate by task: Sudoku, 22 givens0% (0/4)n 425% (1/4)n 4TieStrict pass rate by task: Sudoku, 22 givens: Claude Code 0% (0/4) (n 4, 95% interval 0%–49%); Codex CLI 25% (1/4) (n 4, 95% interval 4.6%–70%). Tie.
Calculation: at these rates, about 26 runs per side would separate them.
- Strict pass rate by task: 6x6 Skyscrapers0% (0/4)n 4100% (4/4)n 4Codex CLI aheadStrict pass rate by task: 6x6 Skyscrapers: Claude Code 0% (0/4) (n 4, 95% interval 0%–49%); Codex CLI 100% (4/4) (n 4, 95% interval 51%–100%). Codex CLI ahead.
- Strict pass rate by task: Seeded shuffle output50% (2/4)n 475% (3/4)n 4TieStrict pass rate by task: Seeded shuffle output: Claude Code 50% (2/4) (n 4, 95% interval 15%–85%); Codex CLI 75% (3/4) (n 4, 95% interval 30%–95%). Tie.
Calculation: at these rates, about 54 runs per side would separate them.
- Total time per call on harder tasks70.4 sn 12120.2 sn 13UnclearTotal time per call on harder tasks: Claude Code 70.4 s (n 12, run range 4.3 s–210 s); Codex CLI 120.2 s (n 13, run range 46.2 s–273 s). Unclear.
- Output tokens per call on harder tasks (Output tokens)9,287n 124,994n 13UnclearOutput tokens per call on harder tasks (Output tokens): Claude Code 9,287 (n 12, run range 407–27,921); Codex CLI 4,994 (n 13, run range 2,099–13,413). Unclear.
- List-price cost per strict pass on harder tasks (calculation)Calculation$0.24n 16$0.083n 16UnclearList-price cost per strict pass on harder tasks (calculation), calculation: Claude Code $0.24 (n 16); Codex CLI $0.083 (n 16). Unclear.
| Metric | Claude Code | Codex CLI | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 80% (12/15)Claude Sonnet 5.5 · five short validated tasks | 100% (15/15)GPT-6.1 Sol · effort medium · five short validated tasks | 15 | 95% CI: 55%–93% vs 80%–100% | Tie | The 95% intervals overlap (Claude Code 55% to 93%; Codex CLI 80% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 2.31 sClaude Sonnet 5.5 · five short validated tasks | 5.65 sGPT-6.1 Sol · effort medium · five short validated tasks | 15 | range: 2.2 s–7.7 s vs 4.1 s–25.5 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 2.17 s to 7.73 s; Codex CLI 4.10 s to 25.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 1.56 sClaude Sonnet 5.5 · five short validated tasks | 5.05 sGPT-6.1 Sol · effort medium · five short validated tasks | 15 | range: 1 s–6.4 s vs 3.4 s–17.8 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 0.99 s to 6.39 s; Codex CLI 3.36 s to 17.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 1,401Claude Sonnet 5.5 · five short validated tasks | 5,180GPT-6.1 Sol · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Other input) | 685Claude Sonnet 5.5 · five short validated tasks | 6,943GPT-6.1 Sol · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 107Claude Sonnet 5.5 · five short validated tasks | 42GPT-6.1 Sol · effort medium · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per call (calculation)Calculation | $0.0036Claude Sonnet 5.5 · five short validated tasks | $0.010GPT-6.1 Sol · effort medium · five short validated tasks | 15 | range: $0.0034–$0.01 vs $0.0054–$0.027 | Unclear | The run ranges (fastest to slowest) overlap (Claude Code $0.0034 to $0.010; Codex CLI $0.0054 to $0.027); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.0062Claude Sonnet 5.5 · five short validated tasks | $0.016GPT-6.1 Sol · effort medium · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.0062 vs $0.016, 2.5x) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Pass rate on eight hard tasks (Strict pass) | 100% (24/24)Claude Sonnet 5.5 · eight hard validated tasks | 100% (16/16)GPT-6.1 Sol · effort medium · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Code 86% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 100% (24/24)Claude Sonnet 5.5 · eight hard validated tasks | 100% (16/16)GPT-6.1 Sol · effort medium · eight hard validated tasks | 24 / 16 | 95% CI: 86%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Code 86% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Total time per call on hard tasks (separate batches) | 7.75 sClaude Sonnet 5.5 · eight hard validated tasks | 13.1 sGPT-6.1 Sol · effort medium · eight hard validated tasks | 24 / 16 | range: 2.3 s–34.8 s vs 8.5 s–61.6 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 2.26 s to 34.8 s; Codex CLI 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Time to first useful output on hard tasks | 5.95 sClaude Sonnet 5.5 · eight hard validated tasks | 10.2 sGPT-6.1 Sol · effort medium · eight hard validated tasks | 24 / 16 | range: 0.9 s–30.6 s vs 6.1 s–40.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 0.86 s to 30.6 s; Codex CLI 6.09 s to 40.4 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call on hard tasks (Output tokens) | 1,050Claude Sonnet 5.5 · eight hard validated tasks | 335GPT-6.1 Sol · effort medium · eight hard validated tasks | 24 / 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass on hard tasks (calculation)Calculation | $0.014Claude Sonnet 5.5 · eight hard validated tasks | $0.026GPT-6.1 Sol · effort medium · eight hard validated tasks | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Coding sessions that passed every hidden check | 100% (12/12)Claude Sonnet 5.5 · six small repository tasks with hidden tests | 100% (12/12)GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Time per coding session | 23.1 sClaude Sonnet 5.5 · six small repository tasks with hidden tests | 113.4 sGPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | range: 18.7 s–44.5 s vs 78.5 s–222 s | Claude Code ahead | The run ranges (fastest to slowest) do not overlap (Claude Code 18.7 s to 44.5 s; Codex CLI 78.5 s to 221.9 s). A range is not a confidence interval. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Tool calls per coding session | 7.5Claude Sonnet 5.5 · six small repository tasks with hidden tests | 12.5GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | range: 3–14 vs 8–18 | Unclear | More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| List-price cost per passing coding session (calculation)Calculation | $0.085Claude Sonnet 5.5 · six small repository tasks with hidden tests | $0.098GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests | 12 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.085 vs $0.098) is not tested against run-to-run variation. | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks |
| Strict pass rate by effort on eight hard tasks | 100% (16/16)Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder | 100% (16/16)GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder | 16 | 95% CI: 81%–100% vs 81%–100% | Tie | The 95% intervals overlap (Claude Code 81% to 100%; Codex CLI 81% to 100%), so this sample cannot separate them. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Total time per call by effort on hard tasks | 7.63 sClaude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder | 13.1 sGPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder | 16 | range: 2.7 s–24 s vs 8.5 s–61.6 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 2.71 s to 24.0 s; Codex CLI 8.54 s to 61.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call by effort on hard tasks (Output tokens) | 770Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder | 335GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass by effort (calculation)Calculation | $0.014Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder | $0.026GPT-6.1 Sol · effort medium · eight hard validated tasks, effort ladder | 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.014 vs $0.026) is not tested against run-to-run variation. | Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks |
| Same prompt, 10 times: strict pass rate (Exact number) | 100% (10/10)Claude Sonnet 5.5 · same prompt repeated 10 times | 100% (10/10)GPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | 95% CI: 72%–100% vs 72%–100% | Tie | The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: strict pass rate (JSON object) | 100% (10/10)Claude Sonnet 5.5 · same prompt repeated 10 times | 100% (10/10)GPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | 95% CI: 72%–100% vs 72%–100% | Tie | The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: strict pass rate (Code fix) | 100% (10/10)Claude Sonnet 5.5 · same prompt repeated 10 times | 100% (10/10)GPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | 95% CI: 72%–100% vs 72%–100% | Tie | The 95% intervals overlap (Claude Code 72% to 100%; Codex CLI 72% to 100%), so this sample cannot separate them. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (Exact number) | 1Claude Sonnet 5.5 · same prompt repeated 10 times | 1GPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (JSON object) | 1Claude Sonnet 5.5 · same prompt repeated 10 times | 1GPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (Code fix) | 3Claude Sonnet 5.5 · same prompt repeated 10 times | 6GPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | none recorded | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (Exact number) | 6.89 sClaude Sonnet 5.5 · same prompt repeated 10 times | 13.4 sGPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | range: 5.8 s–7.8 s vs 12.3 s–18 s | Claude Code ahead | The run ranges (fastest to slowest) do not overlap (Claude Code 5.81 s to 7.81 s; Codex CLI 12.3 s to 18.0 s). A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (JSON object) | 2.89 sClaude Sonnet 5.5 · same prompt repeated 10 times | 6.42 sGPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | range: 2.7 s–5.3 s vs 5.3 s–8.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 2.68 s to 5.30 s; Codex CLI 5.25 s to 8.26 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (Code fix) | 2.67 sClaude Sonnet 5.5 · same prompt repeated 10 times | 11.3 sGPT-6.1 Sol · effort medium · same prompt repeated 10 times | 10 | range: 2.3 s–4.3 s vs 9.1 s–14.9 s | Claude Code ahead | The run ranges (fastest to slowest) do not overlap (Claude Code 2.32 s to 4.34 s; Codex CLI 9.08 s to 14.8 s). A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| CLI start-up tax on a one-word answer (First output event) | 563 msClaude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs | 489 msdefault model · CLI start-up, one-word prompt, 5 runs | 5 | range: 519 ms–726 ms vs 354 ms–1.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 519 ms to 726 ms; Codex CLI 354 ms to 1,304 ms); the medians alone do not show a reliable difference. A range is not a confidence interval. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| CLI start-up tax on a one-word answer (First model output) | 1,461 msClaude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs | 5,059 msdefault model · CLI start-up, one-word prompt, 5 runs | 5 | range: 1.21 s–2.31 s vs 4.39 s–5.48 s | Claude Code ahead | The run ranges (fastest to slowest) do not overlap (Claude Code 1,206 ms to 2,308 ms; Codex CLI 4,391 ms to 5,478 ms). A range is not a confidence interval. Samples are small (5 runs per side). | Routing overhead: deterministic policy vs LLM routers vs Jev |
| CLI start-up tax on a one-word answer (Total wall time) | 2,529 msClaude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs | 5,999 msdefault model · CLI start-up, one-word prompt, 5 runs | 5 | range: 2.27 s–3.38 s vs 5.37 s–6.51 s | Claude Code ahead | The run ranges (fastest to slowest) do not overlap (Claude Code 2,273 ms to 3,382 ms; Codex CLI 5,367 ms to 6,506 ms). A range is not a confidence interval. Samples are small (5 runs per side). | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Input tokens a CLI sends for a one-word answer | 6,761Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs | 17,051default model · CLI start-up, one-word prompt, 5 runs | 5 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Repairing a scheduler: Claude Code vs Codex vs API (Total time) | 15.0 sSonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs | 61.2 sGPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs | 3 | range: 13.9 s–15.9 s vs 59.9 s–69.5 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 13.9 s to 15.9 s; Codex CLI 59.9 s to 69.5 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Repairing a scheduler: Claude Code vs Codex vs API (First useful output) | 7.55 sSonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs | 15.6 sGPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs | 3 | range: 6.8 s–7.6 s vs 13.7 s–23 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 6.77 s to 7.63 s; Codex CLI 13.7 s to 23.0 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Output tokens to repair the scheduler (Output tokens) | 2,227Sonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs | 1,181GPT-6.1 Sol · effort medium · scheduler repair, 296 checks, 3 runs | 3 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Claude Code CLI vs Codex CLI vs the API: latency and tokens |
| Strict pass rate: single call vs agent loop on eight hard tasks | 100% (24/24)Claude Sonnet 5.5 · single call | 63% (10/16)GPT-6 Luna · single call | 24 / 16 | 95% CI: 86%–100% vs 39%–82% | Claude Code ahead | The 95% intervals do not overlap (Claude Code 86% to 100%; Codex CLI 39% to 82%). | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Interval merge fix | 100% (3/3)Claude Sonnet 5.5 · single call | 100% (2/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: DST day-length fix | 100% (3/3)Claude Sonnet 5.5 · single call | 100% (2/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: CSV parser | 100% (3/3)Claude Sonnet 5.5 · single call | 100% (2/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Event-loop order | 100% (3/3)Claude Sonnet 5.5 · single call | 0% (0/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 0%–66% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 0% to 66%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Room schedule | 100% (3/3)Claude Sonnet 5.5 · single call | 50% (1/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 9.4%–91% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 9% to 91%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SemVer regex | 100% (3/3)Claude Sonnet 5.5 · single call | 50% (1/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 9.4%–91% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 9% to 91%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Money refactor | 100% (3/3)Claude Sonnet 5.5 · single call | 0% (0/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 0%–66% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 0% to 66%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SQL report | 100% (3/3)Claude Sonnet 5.5 · single call | 100% (2/2)GPT-6 Luna · single call | 3 / 2 | 95% CI: 44%–100% vs 34%–100% | Tie | The 95% intervals overlap (Claude Code 44% to 100%; Codex CLI 34% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Total time per attempt: single call vs agent loop | 7.75 sClaude Sonnet 5.5 · single call | 5.16 sGPT-6 Luna · single call | 24 / 16 | range: 2.3 s–34.8 s vs 3.6 s–11.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 2.26 s to 34.8 s; Codex CLI 3.59 s to 11.3 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Input tokens (cache reads included)) | 2,281Claude Sonnet 5.5 · single call | 11,582GPT-6 Luna · single call | 24 / 16 | range: 2,234–2,669 vs 11,526–11,818 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Output tokens) | 1,050Claude Sonnet 5.5 · single call | 345GPT-6 Luna · single call | 24 / 16 | range: 176–3,895 vs 36–634 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tool calls per agent-loop attempt | 0Claude Sonnet 5.5 · agent loop | 0GPT-6 Luna · agent loop | 16 / 14 | range: 0–3 vs 0–1 | Tie | Same value. More or fewer is not better by itself for this metric. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| List-price cost per strict pass: single call vs agent loop (calculation)Calculation | $0.014Claude Sonnet 5.5 · single call | $0.0012GPT-6 Luna · single call | 24 / 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.014 vs $0.0012, 12x) is not tested against run-to-run variation. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)Calculation | 100% (12/12)Claude Sonnet 5.5 · instructions | 100% (12/12)GPT-6.1 Sol · effort low · instructions | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))Calculation | 100% (12/12)Claude Sonnet 5.5 · instructions | 100% (12/12)GPT-6.1 Sol · effort low · instructions | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Claude Code 76% to 100%; Codex CLI 76% to 100%), so this sample cannot separate them. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Strict pass) | 12Claude Sonnet 5.5 · instructions | 12GPT-6.1 Sol · effort low · instructions | 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Format miss) | 0Claude Sonnet 5.5 · instructions | 0GPT-6.1 Sol · effort low · instructions | 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Wrong values) | 0Claude Sonnet 5.5 · instructions | 0GPT-6.1 Sol · effort low · instructions | 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Error) | 0Claude Sonnet 5.5 · instructions | 0GPT-6.1 Sol · effort low · instructions | 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Time per call, instructions vs schema mode | 3.52 sClaude Sonnet 5.5 · instructions | 6.21 sGPT-6.1 Sol · effort low · instructions | 12 | range: 2.7 s–4.1 s vs 4.2 s–12.3 s | Claude Code ahead | The run ranges (fastest to slowest) do not overlap (Claude Code 2.67 s to 4.12 s; Codex CLI 4.20 s to 12.3 s). A range is not a confidence interval. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call) | 368Claude Sonnet 5.5 · instructions | 117GPT-6.1 Sol · effort low · instructions | 12 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Reasoning share of output tokens per call on hard tasks (calculation)Calculation | 54.5%Claude Sonnet 5.5 | 46.3%GPT-6.1 Sol · effort medium | 24 / 16 | range: 0%–96% vs 11%–87% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation | $0.0067Claude Sonnet 5.5 | $0.0023GPT-6.1 Sol · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation | $0.0037Claude Sonnet 5.5 | $0.0030GPT-6.1 Sol · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation | $0.0040Claude Sonnet 5.5 | $0.020GPT-6.1 Sol · effort medium | 24 / 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass)Calculation | $0.0059Claude Sonnet 5.5 · effort medium | $0.0023GPT-6.1 Sol · effort medium | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass)Calculation | $0.014Claude Sonnet 5.5 · effort medium | $0.026GPT-6.1 Sol · effort medium | 16 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation | 54.5%Claude Sonnet 5.5 | 46.3%GPT-6.1 Sol · effort medium | 24 / 16 | range: 0%–96% vs 11%–87% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation | 0%Claude Sonnet 5.5 | 41%GPT-6.1 Sol · effort medium | 15 | range: 0%–73% vs 0%–71% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Time to first text: a 250-line answer, six models | 1.96 sClaude Sonnet 5.5 | 3.52 sGPT-6.1 Sol · effort low | 4 | range: 0.9 s–4.1 s vs 2.8 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 0.88 s to 4.09 s; Codex CLI 2.75 s to 4.42 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 232Claude Sonnet 5.5 | 80GPT-6.1 Sol · effort low | 4 | range: 230–233 vs 72–81 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 517Claude Sonnet 5.5 | 323GPT-6.1 Sol · effort low | 4 | range: 513–519 vs 291–327 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 1kCalculation | 1.45 sClaude Sonnet 5.5 | 3.36 sGPT-6.1 Sol · effort low | 3 | range: 1.2 s–1.7 s vs 3.4 s–4.8 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.23 s to 1.72 s; Codex CLI 3.36 s to 4.75 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 16kCalculation | 1.78 sClaude Sonnet 5.5 | 4.02 sGPT-6.1 Sol · effort low | 3 | range: 1.6 s–2.1 s vs 3.3 s–4.3 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.64 s to 2.11 s; Codex CLI 3.30 s to 4.28 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 64kCalculation | 3.07 sClaude Sonnet 5.5 | 3.93 sGPT-6.1 Sol · effort low | 3 | range: 1.4 s–3.6 s vs 3.4 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 1.38 s to 3.61 s; Codex CLI 3.42 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (1k prompt) | 1.78 sClaude Sonnet 5.5 | 3.43 sGPT-6.1 Sol · effort low | 3 | range: 1.6 s–2.1 s vs 3.4 s–4.9 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.57 s to 2.12 s; Codex CLI 3.43 s to 4.92 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (16k prompt) | 2.10 sClaude Sonnet 5.5 | 4.14 sGPT-6.1 Sol · effort low | 3 | range: 2 s–2.5 s vs 4 s–4.7 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Code 1.98 s to 2.48 s; Codex CLI 3.96 s to 4.68 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (64k prompt) | 3.44 sClaude Sonnet 5.5 | 3.96 sGPT-6.1 Sol · effort low | 3 | range: 1.7 s–4.4 s vs 3.5 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 1.74 s to 4.38 s; Codex CLI 3.47 s to 4.44 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Exact lookup answers at the 1k, 16k and 64k prompt-size targets | 100% (9/9)Claude Sonnet 5.5 | 100% (9/9)GPT-6.1 Sol · effort low | 9 | 95% CI: 70%–100% vs 70%–100% | Tie | The 95% intervals overlap (Claude Code 70% to 100%; Codex CLI 70% to 100%), so this sample cannot separate them. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Pass rate on 4 harder tasks (Strict pass) | 38% (6/16)Claude Sonnet 5.5 | 69% (11/16)GPT-6.1 Sol · effort medium | 16 | 95% CI: 18%–61% vs 44%–86% | Tie | The 95% intervals overlap (Claude Code 18% to 61%; Codex CLI 44% to 86%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Pass rate on 4 harder tasks (Lenient (format misses counted)) | 38% (6/16)Claude Sonnet 5.5 | 69% (11/16)GPT-6.1 Sol · effort medium | 16 | 95% CI: 18%–61% vs 44%–86% | Tie | The 95% intervals overlap (Claude Code 18% to 61%; Codex CLI 44% to 86%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Calls that tried a tool although tools were off | 31% (5/16)Claude Sonnet 5.5 | 0% (0/16)GPT-6.1 Sol · effort medium | 16 | 95% CI: 14%–56% vs 0%–19% | Unclear | More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 10x10 nonogram | 100% (4/4)Claude Sonnet 5.5 | 75% (3/4)GPT-6.1 Sol · effort medium | 4 | 95% CI: 51%–100% vs 30%–95% | Tie | The 95% intervals overlap (Claude Code 51% to 100%; Codex CLI 30% to 95%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Sudoku, 22 givens | 0% (0/4)Claude Sonnet 5.5 | 25% (1/4)GPT-6.1 Sol · effort medium | 4 | 95% CI: 0%–49% vs 4.6%–70% | Tie | The 95% intervals overlap (Claude Code 0% to 49%; Codex CLI 5% to 70%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 6x6 Skyscrapers | 0% (0/4)Claude Sonnet 5.5 | 100% (4/4)GPT-6.1 Sol · effort medium | 4 | 95% CI: 0%–49% vs 51%–100% | Codex CLI ahead | The 95% intervals do not overlap (Claude Code 0% to 49%; Codex CLI 51% to 100%). | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Seeded shuffle output | 50% (2/4)Claude Sonnet 5.5 | 75% (3/4)GPT-6.1 Sol · effort medium | 4 | 95% CI: 15%–85% vs 30%–95% | Tie | The 95% intervals overlap (Claude Code 15% to 85%; Codex CLI 30% to 95%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Total time per call on harder tasks | 70.4 sClaude Sonnet 5.5 | 120.2 sGPT-6.1 Sol · effort medium | 12 / 13 | range: 4.3 s–210 s vs 46.2 s–273 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Code 4.32 s to 210.1 s; Codex CLI 46.2 s to 273.5 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Output tokens per call on harder tasks (Output tokens) | 9,287Claude Sonnet 5.5 | 4,994GPT-6.1 Sol · effort medium | 12 / 13 | range: 407–27,921 vs 2,099–13,413 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| List-price cost per strict pass on harder tasks (calculation)Calculation | $0.24Claude Sonnet 5.5 | $0.083GPT-6.1 Sol · effort medium | 16 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.24 vs $0.083, 2.9x) is not tested against run-to-run variation. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
88 rows from 12 studies. Claude Code ahead on 7, Codex CLI ahead on 1; 31 ties, 49 unclear. A side is ahead only where the intervals or ranges do not overlap.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick Claude Code
- List-price cost per call (calculation): $0.0036 vs $0.010. A list-price calculation, not a measured difference. Calculation
- List-price cost per passing answer (calculation): $0.0062 vs $0.016. A list-price calculation, not a measured difference. Calculation
- List-price cost per strict pass on hard tasks (calculation): $0.014 vs $0.026. A list-price calculation, not a measured difference. Calculation
- Time per coding session: 23.1 s vs 113.4 s. The run ranges (fastest to slowest) do not overlap (Claude Code 18.7 s to 44.5 s; Codex CLI 78.5 s to 221.9 s). A range is not a confidence interval.
- List-price cost per strict pass by effort (calculation): $0.014 vs $0.026. A list-price calculation, not a measured difference. Calculation
- Same prompt, 10 times: time per call (Exact number): 6.89 s vs 13.4 s. The run ranges (fastest to slowest) do not overlap (Claude Code 5.81 s to 7.81 s; Codex CLI 12.3 s to 18.0 s). A range is not a confidence interval.
- Same prompt, 10 times: time per call (Code fix): 2.67 s vs 11.3 s. The run ranges (fastest to slowest) do not overlap (Claude Code 2.32 s to 4.34 s; Codex CLI 9.08 s to 14.8 s). A range is not a confidence interval.
- CLI start-up tax on a one-word answer (First model output): 1,461 ms vs 5,059 ms. The run ranges (fastest to slowest) do not overlap (Claude Code 1,206 ms to 2,308 ms; Codex CLI 4,391 ms to 5,478 ms). A range is not a confidence interval. Samples are small (5 runs per side).
- CLI start-up tax on a one-word answer (Total wall time): 2,529 ms vs 5,999 ms. The run ranges (fastest to slowest) do not overlap (Claude Code 2,273 ms to 3,382 ms; Codex CLI 5,367 ms to 6,506 ms). A range is not a confidence interval. Samples are small (5 runs per side).
- Strict pass rate: single call vs agent loop on eight hard tasks: 100% (24/24) vs 63% (10/16). The 95% intervals do not overlap (Claude Code 86% to 100%; Codex CLI 39% to 82%).
- Time per call, instructions vs schema mode: 3.52 s vs 6.21 s. The run ranges (fastest to slowest) do not overlap (Claude Code 2.67 s to 4.12 s; Codex CLI 4.20 s to 12.3 s). A range is not a confidence interval.
- List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)): $0.0040 vs $0.020. A list-price calculation, not a measured difference. Calculation
- Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass): $0.014 vs $0.026. A list-price calculation, not a measured difference. Calculation
- Time to first text as the prompt grows: 1k: 1.45 s vs 3.36 s. A list-price calculation, not a measured difference. Calculation
- Time to first text as the prompt grows: 16k: 1.78 s vs 4.02 s. A list-price calculation, not a measured difference. Calculation
When to pick Codex CLI
- List-price cost per strict pass: single call vs agent loop (calculation): $0.0012 vs $0.014. A list-price calculation, not a measured difference. Calculation
- List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0023 vs $0.0067. A list-price calculation, not a measured difference. Calculation
- Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass): $0.0023 vs $0.0059. A list-price calculation, not a measured difference. Calculation
- Strict pass rate by task: 6x6 Skyscrapers: 100% (4/4) vs 0% (0/4). The 95% intervals do not overlap (Claude Code 0% to 49%; Codex CLI 51% to 100%).
- List-price cost per strict pass on harder tasks (calculation): $0.083 vs $0.24. A list-price calculation, not a measured difference. Calculation
Side by side
The study charts, showing only these two. Open a study for every configuration.
| Item | Pass rate | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 80% | 55%–93% | 15 |
| GPT-6.1 Sol (high) · Codex CLI | 100% | 80%–100% | 15 |
2 rows. Highest GPT-6.1 Sol (high) · Codex CLI 100% (95% interval 80%–100%, n 15). Lowest Claude Sonnet 5.5 · Claude Code 80% (95% interval 55%–93%, n 15). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 15 per row
Every call counts; failures and timeouts are non-passes
Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.
Source: Provider head-to-head: Claude Code models vs Codex efforts
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 2.3 s | 2.2 s–7.7 s | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | 5.7 s | 4.1 s–25.5 s | 15 |
2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 5.7 s (range 4.1 s–25.5 s, n 15). Fastest Claude Sonnet 5.5 · Claude Code 2.3 s (range 2.2 s–7.7 s, n 15). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 15 per row
Median per configuration; whiskers = fastest and slowest call
One host, one network, one day. Whiskers are a range, not a confidence interval.
Source: Provider head-to-head: Claude Code models vs Codex efforts
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first useful output | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1.6 s | 1 s–6.4 s | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | 5.1 s | 3.4 s–17.8 s | 15 |
2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 5.1 s (range 3.4 s–17.8 s, n 15). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1 s–6.4 s, n 15). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 15 per row
Median per configuration; whiskers = fastest and slowest call
One host, one network, one day. Whiskers are a range, not a confidence interval.
Source: Provider head-to-head: Claude Code models vs Codex efforts
- Cache read
- Other input
Totals are the sum of the parts shown. Shares are calculated from the same values.
| Item | Cache read | Other input | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1,401 | 685 | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | 5,180 | 6,943 | 15 |
2 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (medium) · Codex CLI 5,180 (n 15). Lowest Claude Sonnet 5.5 · Claude Code 1,401 (n 15). Other input: highest GPT-6.1 Sol (medium) · Codex CLI 6,943 (n 15). Lowest Claude Sonnet 5.5 · Claude Code 685 (n 15).
Notesn = 15 per row
Mean per call, split into prompt-cache reads and other input
The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.
Source: Provider head-to-head: Claude Code models vs Codex efforts
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 107 | 0 | 15 |
| GPT-6.1 Sol (high) · Codex CLI | 42 | 21 | 15 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 · Claude Code 107 (n 15). Lowest GPT-6.1 Sol (high) · Codex CLI 42 (n 15). Reasoning tokens: highest GPT-6.1 Sol (high) · Codex CLI 21 (n 15). Lowest Claude Sonnet 5.5 · Claude Code 0 (n 15).
Notesn = 15 per row
Median per configuration; reasoning tokens where the CLI reports them
Vendors count reasoning differently; compare within a vendor first. A median of 0 can mean the CLI did not report reasoning for most calls.
Source: Provider head-to-head: Claude Code models vs Codex efforts
| Item | Cost per call | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.0036 | $0.0034–$0.01 | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.01 | $0.0054–$0.027 | 15 |
List-price calculation, not a run. 2 rows. Highest GPT-6.1 Sol (medium) · Codex CLI $0.01 (range $0.0054–$0.027, n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.0036 (range $0.0034–$0.01, n 15). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 15 per row
Reported tokens × list price; the calls ran on subscriptions
Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.
Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
| Item | Cost per pass | n |
|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.0062 | 15 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.016 | 15 |
List-price calculation, not a run. 2 rows. Highest GPT-6.1 Sol (medium) · Codex CLI $0.016 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.0062 (n 15).
Notesn = 15 per row
All calls in a configuration, failures included, divided by its passes
Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.
Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Strict pass
- Lenient (format misses counted)
| Item | Strict pass | Lenient (format misses counted) | 95% interval | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 100% | 100% | Strict pass: 86%–100%; Lenient (format misses counted): 86%–100% | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 100% | 100% | Strict pass: 81%–100%; Lenient (format misses counted): 81%–100% | 16 |
2 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: all at 100%. Lenient (format misses counted): all at 100%.
NotesWhiskers: 95% Wilson intervaln 16–24 per row
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call on hard tasks (separate batches) | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 7.8 s | 2.3 s–34.8 s | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 13.1 s | 8.5 s–61.6 s | 16 |
2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 13.1 s (range 8.5 s–61.6 s, n 16). Fastest Claude Sonnet 5.5 · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first useful output on hard tasks | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 6 s | 0.9 s–30.6 s | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 10.2 s | 6.1 s–40.4 s | 16 |
2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 10.2 s (range 6.1 s–40.4 s, n 16). Fastest Claude Sonnet 5.5 · Claude Code 6 s (range 0.9 s–30.6 s, n 24). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1,050 | 585 | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 335 | 150 | 16 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 · Claude Code 1,050 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16). Reasoning tokens: highest Claude Sonnet 5.5 · Claude Code 585 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 150 (n 16).
Notesn 16–24 per row
Median per configuration; reasoning tokens as the CLI reports them
Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.014 | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.026 | 16 |
List-price calculation, not a run. 2 rows. Highest GPT-6.1 Sol (medium) · Codex CLI $0.026 (n 16). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).
Notesn 16–24 per row
All calls in a configuration, failures and format misses included, divided by its strict passes
Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.
Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
| Item | Passed every hidden check | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 100% | 76%–100% | 12 |
| GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI | 100% | 76%–100% | 12 |
2 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 12 per row
A pass needs every hidden check · 12 sessions per agent · 95% Wilson intervals
6 tasks × 2 repetitions per agent. All 36 sessions passed, so the chart sits at its ceiling: this task set cannot rank the agents on quality. Before the first session, every base repository failed its hidden checks and every reference patch passed them.
Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks
1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.
| Item | Wall time per session | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 23.1 s | 18.7 s–44.5 s | 12 |
| GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI | 113 s | 78.5 s–222 s | 12 |
2 rows. Slowest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI 113 s (range 78.5 s–222 s, n 12). Fastest Claude Sonnet 5.5 · Claude Code 23.1 s (range 18.7 s–44.5 s, n 12). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 12 per row
Median wall time; whiskers = fastest and slowest of 12 sessions (not an interval)
CLI process start to exit. One session at a time per lane; the Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts. Codex time includes reading the tester’s notes, writing a work log and trying to commit. A range is not a confidence interval.
Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks
| Item | Tool calls per session | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 7.5 | 3–14 | 12 |
| GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI | 12.5 | 8–18 | 12 |
2 rows. Highest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI 12.5 (range 8–18, n 12). Lowest Claude Sonnet 5.5 · Claude Code 7.5 (range 3–14, n 12). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 12 per row
Median; whiskers = fewest and most of 12 sessions (not an interval)
Claude Code counts its tool calls (Bash, Read, Edit, Write, Glob, Grep). Codex CLI counts shell commands and file changes; it has no separate read tool, so it reads files with shell commands. Turns are not compared: Codex reports one turn per run.
Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks
| Item | List-price cost per pass | n |
|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.085 | 12 |
| GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI | $0.098 | 12 |
List-price calculation, not a run. 2 rows. Highest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI $0.098 (n 12). Lowest Claude Sonnet 5.5 · Claude Code $0.085 (n 12).
Notesn = 12 per row
Reported tokens of all 12 sessions × list price, divided by the passes
Calculation, not a bill: both CLIs ran on flat subscriptions. Claude cache writes are priced at 2× input, as in the other studies (Claude Code’s own estimate gives the same totals); Codex cached input at its cache-read price. Codex input includes its own system prompt and, here, the tester’s AGENTS.md.
Sources: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
| Item | Strict pass | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 (low) · Claude Code | 100% | 81%–100% | 16 |
| GPT-6.1 Sol (low) · Codex CLI | 100% | 81%–100% | 16 |
2 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 16 per row
Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.
Source: Effort ladder: the hard task set at each effort level
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call by effort on hard tasks | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 (medium) · Claude Code | 7.6 s | 2.7 s–24 s | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | 13.1 s | 8.5 s–61.6 s | 16 |
2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 13.1 s (range 8.5 s–61.6 s, n 16). Fastest Claude Sonnet 5.5 (medium) · Claude Code 7.6 s (range 2.7 s–24 s, n 16). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 16 per row
Median per configuration; whiskers = fastest and slowest call
Whiskers are a range (fastest and slowest call), not a confidence interval. One host, one network. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions. The reference cells ran in a different hour.
Source: Effort ladder: the hard task set at each effort level
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Sonnet 5.5 (medium) · Claude Code | 770 | 422 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | 335 | 150 | 16 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 (medium) · Claude Code 770 (n 16). Lowest GPT-6.1 Sol (medium) · Codex CLI 335 (n 16). Reasoning tokens: highest Claude Sonnet 5.5 (medium) · Claude Code 422 (n 16). Lowest GPT-6.1 Sol (medium) · Codex CLI 150 (n 16).
Notesn = 16 per row
Median per configuration; reasoning tokens as the CLI reports them
Reasoning tokens are part of the output tokens where the CLI reports them; 0 can mean "not reported". Their content is never captured. More tokens is not better or worse by itself.
Source: Effort ladder: the hard task set at each effort level
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Sonnet 5.5 (medium) · Claude Code | $0.014 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.026 | 16 |
List-price calculation, not a run. 2 rows. Highest GPT-6.1 Sol (medium) · Codex CLI $0.026 (n 16). Lowest Claude Sonnet 5.5 (medium) · Claude Code $0.014 (n 16).
Notesn = 16 per row
All calls in a configuration divided by its strict passes
Calculation, not a bill: reported tokens × list price, cache reads and writes priced as in the hard head-to-head; the calls ran on flat subscriptions. Codex input includes its own system prompt and tool schemas, so a Claude-vs-Codex gap is partly the CLI.
Sources: Effort ladder: the hard task set at each effort level, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Exact number
- JSON object
- Code fix
| Item | Exact number | JSON object | Code fix | 95% interval | n |
|---|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 100% | 100% | 100% | Exact number: 72%–100%; JSON object: 72%–100%; Code fix: 72%–100% | 10 |
| GPT-6.1 Sol (medium) · Codex CLI | 100% | 100% | 100% | Exact number: 72%–100%; JSON object: 72%–100%; Code fix: 72%–100% | 10 |
2 rows, 3 series: Exact number, JSON object, Code fix. Exact number: all at 100%. JSON object: all at 100%.
NotesWhiskers: 95% Wilson intervaln = 10 per row
One series per prompt; whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.
Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)
Exact number
JSON object
Code fix
One panel per series, all on the same axis.
| Item | Exact number | JSON object | Code fix | n |
|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 1 | 1 | 6 | 10 |
| Claude Sonnet 5.5 · Claude Code | 1 | 1 | 3 | 10 |
| GPT-6.1 Sol (medium) · Codex CLI | 1 | 1 | 6 | 10 |
3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: all at 1. JSON object: all at 1.
Notesn = 10 per row
Distinct normalized answers over 10 repetitions (1 = the same answer every time)
Normalized: the last number for the exact-number prompt; key-sorted JSON for the JSON prompt; for the code fix, the code without fences, comments, whitespace, quote style and line-end semicolons. One distinct answer is not the same as a correct answer: an answer can be the same and wrong every time. Different code can be equally correct.
Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)
- Exact number
- JSON object
- Code fix
| Item | Exact number | JSON object | Code fix | Range (lowest–highest run) | n |
|---|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 6.9 s | 2.9 s | 2.7 s | Exact number: 5.8 s–7.8 s; JSON object: 2.7 s–5.3 s; Code fix: 2.3 s–4.3 s | 10 |
| GPT-6.1 Sol (medium) · Codex CLI | 13.4 s | 6.4 s | 11.3 s | Exact number: 12.3 s–18 s; JSON object: 5.3 s–8.3 s; Code fix: 9.1 s–14.9 s | 10 |
2 rows, 3 series: Exact number, JSON object, Code fix. Exact number: slowest GPT-6.1 Sol (medium) · Codex CLI 13.4 s (range 12.3 s–18 s, n 10). Fastest Claude Sonnet 5.5 · Claude Code 6.9 s (range 5.8 s–7.8 s, n 10). Not all run ranges overlap. JSON object: slowest GPT-6.1 Sol (medium) · Codex CLI 6.4 s (range 5.3 s–8.3 s, n 10). Fastest Claude Sonnet 5.5 · Claude Code 2.9 s (range 2.7 s–5.3 s, n 10). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 10 per row
Median; whiskers = fastest and slowest of 10 calls
Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.
Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)
- First output event
- First model output
- Total wall time
| Item | First output event | First model output | Total wall time | Range (lowest–highest run) | n |
|---|---|---|---|---|---|
| Claude Code · Claude Haiku 4.5 | 563 ms | 1.46 s | 2.53 s | First output event: 519 ms–726 ms; First model output: 1.21 s–2.31 s; Total wall time: 2.27 s–3.38 s | 5 |
| Codex CLI (default model) | 489 ms | 5.06 s | 6 s | First output event: 354 ms–1.3 s; First model output: 4.39 s–5.48 s; Total wall time: 5.37 s–6.51 s | 5 |
2 rows, 3 series: First output event, First model output, Total wall time. First output event: slowest Claude Code · Claude Haiku 4.5 563 ms (range 519 ms–726 ms, n 5). Fastest Codex CLI (default model) 489 ms (range 354 ms–1.3 s, n 5). All run ranges overlap. First model output: slowest Codex CLI (default model) 5.06 s (range 4.39 s–5.48 s, n 5). Fastest Claude Code · Claude Haiku 4.5 1.46 s (range 1.21 s–2.31 s, n 5). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 5 per row
Median of 5 runs; whiskers = fastest and slowest run
Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.
Source: Routing overhead runs: policy microbenchmark and CLI start-up
| Item | Input tokens per call | n |
|---|---|---|
| Claude Code · Claude Haiku 4.5 | 6,761 | 5 |
| Codex CLI (default model) | 17,051 | 5 |
2 rows. Highest Codex CLI (default model) 17,051 (n 5). Lowest Claude Code · Claude Haiku 4.5 6,761 (n 5).
Notesn = 5 per row
Per call, mostly the CLI’s own system prompt and tool definitions
Claude Code sums its disjoint input, cache-read and cache-write fields. Codex CLI reports 17,051 input tokens, 13,184 of them read from the cache. The prompt itself is a few tokens.
Source: Routing overhead runs: policy microbenchmark and CLI start-up
- Total time
- First useful output
| Item | Total time | First useful output | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Code CLI · Sonnet 5.5 · medium | 15 s | 7.6 s | Total time: 13.9 s–15.9 s; First useful output: 6.8 s–7.6 s | 3 |
| Codex CLI · GPT-6.1 Sol · medium | 61.2 s | 15.6 s | Total time: 59.9 s–69.5 s; First useful output: 13.7 s–23 s | 3 |
2 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · medium 61.2 s (range 59.9 s–69.5 s, n 3). Fastest Claude Code CLI · Sonnet 5.5 · medium 15 s (range 13.9 s–15.9 s, n 3). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · medium 15.6 s (range 13.7 s–23 s, n 3). Fastest Claude Code CLI · Sonnet 5.5 · medium 7.6 s (range 6.8 s–7.6 s, n 3). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Same prompt, medium effort, 296 behavioral checks, 3 runs each
All 9 runs passed all 296 checks. Dot = median, whiskers = range. Different models (Sonnet 5.5 vs GPT-6.1 Sol), so this compares route + model pairs, not routes alone.
Source: Provider explorer receipts: CLI vs API
- Output tokens
- of which reasoning tokens (reported) (inner bar)
| Item | Output tokens | Reasoning tokens (reported) | n |
|---|---|---|---|
| Claude Code CLI · Sonnet 5.5 · medium | 2,227 | 0 | 3 |
| Codex CLI · GPT-6.1 Sol · medium | 1,181 | 156 | 3 |
2 rows, 2 series: Output tokens, Reasoning tokens (reported). Output tokens: highest Claude Code CLI · Sonnet 5.5 · medium 2,227 (n 3). Lowest Codex CLI · GPT-6.1 Sol · medium 1,181 (n 3). Reasoning tokens (reported): highest Codex CLI · GPT-6.1 Sol · medium 156 (n 3). Lowest Claude Code CLI · Sonnet 5.5 · medium 0 (n 0).
Notesn 0–3 per row
Median per run; reasoning tokens shown separately where reported
The Claude CLI does not report reasoning tokens separately; 0 there means "not reported", not "none".
Source: Provider explorer receipts: CLI vs API
| Item | Strict pass | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 (single call) · Claude Code | 100% | 86%–100% | 24 |
| GPT-6 Luna (single call) · Codex CLI | 63% | 39%–82% | 16 |
2 rows. Highest Claude Sonnet 5.5 (single call) · Claude Code 100% (95% interval 86%–100%, n 24). Lowest GPT-6 Luna (single call) · Codex CLI 63% (95% interval 39%–82%, n 16). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 16–24 per row
Same tasks and validators. Whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
Claude Haiku 4.5 (single call) · Claude Code
Claude Haiku 4.5 (agent loop) · Claude Code
Claude Sonnet 5.5 (single call) · Claude Code
Claude Sonnet 5.5 (agent loop) · Claude Code
GPT-6 Luna (single call) · Codex CLI
GPT-6 Luna (agent loop) · Codex CLI
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | Claude Haiku 4.5 (single call) · Claude Code | Claude Haiku 4.5 (agent loop) · Claude Code | Claude Sonnet 5.5 (single call) · Claude Code | Claude Sonnet 5.5 (agent loop) · Claude Code | GPT-6 Luna (single call) · Codex CLI | GPT-6 Luna (agent loop) · Codex CLI | 95% interval | n |
|---|---|---|---|---|---|---|---|---|
| Interval merge fix | 100% | 67% | 100% | 100% | 100% | 100% | Claude Haiku 4.5 (single call) · Claude Code: 44%–100%; Claude Haiku 4.5 (agent loop) · Claude Code: 21%–94%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 34%–100%; GPT-6 Luna (agent loop) · Codex CLI: 34%–100% | 3 |
| Event-loop order | 0% | 100% | 100% | 100% | 0% | 100% | Claude Haiku 4.5 (single call) · Claude Code: 0%–56%; Claude Haiku 4.5 (agent loop) · Claude Code: 44%–100%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 0%–66%; GPT-6 Luna (agent loop) · Codex CLI: 34%–100% | 3 |
| Room schedule | 0% | 67% | 100% | 100% | 50% | 100% | Claude Haiku 4.5 (single call) · Claude Code: 0%–56%; Claude Haiku 4.5 (agent loop) · Claude Code: 21%–94%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 9.4%–91%; GPT-6 Luna (agent loop) · Codex CLI: 34%–100% | 3 |
3 rows, 6 series: Claude Haiku 4.5 (single call) · Claude Code, Claude Haiku 4.5 (agent loop) · Claude Code, Claude Sonnet 5.5 (single call) · Claude Code, Claude Sonnet 5.5 (agent loop) · Claude Code, GPT-6 Luna (single call) · Codex CLI, GPT-6 Luna (agent loop) · Codex CLI. Claude Haiku 4.5 (single call) · Claude Code: highest Interval merge fix 100% (95% interval 44%–100%, n 3). Lowest Room schedule 0% (95% interval 0%–56%, n 3). All intervals overlap. Claude Haiku 4.5 (agent loop) · Claude Code: highest Event-loop order 100% (95% interval 44%–100%, n 3). Lowest Room schedule 67% (95% interval 21%–94%, n 3). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 2–3 per row
Share of attempts per task that passed strictly; 2 or 3 attempts per task and configuration
Whiskers are 95% Wilson intervals on 2 or 3 attempts, so they are very wide: read this chart for where a configuration failed, not for a ranking. Contaminated attempts (a read outside the work folder) are left out.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per attempt | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 (single call) · Claude Code | 7.8 s | 2.3 s–34.8 s | 24 |
| GPT-6 Luna (single call) · Codex CLI | 5.2 s | 3.6 s–11.3 s | 16 |
2 rows. Slowest Claude Sonnet 5.5 (single call) · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). Fastest GPT-6 Luna (single call) · Codex CLI 5.2 s (range 3.6 s–11.3 s, n 16). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fastest and slowest attempt
Whiskers are a range (fastest and slowest attempt), not a confidence interval. Agent loop: wall time of the whole session, failures and time-outs included. One host, one network. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun. The reference cells ran in a different hour.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
- Input tokens (cache reads included)
- Output tokens
| Item | Input tokens (cache reads included) | Output tokens | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 (single call) · Claude Code | 2,281 | 1,050 | Input tokens (cache reads included): 2,234–2,669; Output tokens: 176–3,895 | 24 |
| GPT-6 Luna (single call) · Codex CLI | 11,582 | 345 | Input tokens (cache reads included): 11,526–11,818; Output tokens: 36–634 | 16 |
2 rows, 2 series: Input tokens (cache reads included), Output tokens. Input tokens (cache reads included): highest GPT-6 Luna (single call) · Codex CLI 11,582 (range 11,526–11,818, n 16). Lowest Claude Sonnet 5.5 (single call) · Claude Code 2,281 (range 2,234–2,669, n 24). Not all run ranges overlap. Output tokens: highest Claude Sonnet 5.5 (single call) · Claude Code 1,050 (range 176–3,895, n 24). Lowest GPT-6 Luna (single call) · Codex CLI 345 (range 36–634, n 16). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fewest and most
Whiskers are a range (fewest and most), not a confidence interval. Input counts the whole prompt of every model request in the attempt, cache reads and writes included; an agent loop re-sends its growing context each turn. Codex input includes its own system prompt and tool schemas. More tokens is not better or worse by itself.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
| Item | Tool calls per attempt | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 (agent loop) · Claude Code | 0 | 0–3 | 16 |
| GPT-6 Luna (agent loop) · Codex CLI | 0 | 0–1 | 14 |
2 rows. All at 0.
NotesLines: lowest–highest run (not an interval)n 14–16 per row
Median per configuration; whiskers = fewest and most. A single call makes none
Whiskers are a range (fewest and most), not a confidence interval. Claude Code tools: shell, read, edit, write, glob, grep. Codex CLI: shell commands and file changes. The model chose whether to test its answer; the prompt allowed it but did not require it.
Source: Single call vs agent loop
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Sonnet 5.5 (single call) · Claude Code | $0.014 | 24 |
| GPT-6 Luna (single call) · Codex CLI | $0.0012 | 16 |
List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 (single call) · Claude Code $0.014 (n 24). Lowest GPT-6 Luna (single call) · Codex CLI $0.0012 (n 16).
Notesn 16–24 per row
All attempts in a configuration divided by its strict passes
Calculation, not a bill: reported tokens × list price for every attempt (per model when a session used several), cache reads at the cache-read price and cache writes at the one-hour write price, divided by the strict passes. The calls ran on flat subscriptions.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Strict pass: the whole reply is the right JSON
- Right answer in any format (strict pass or format miss)
| Item | Strict pass: the whole reply is the right JSON | Right answer in any format (strict pass or format miss) | 95% interval | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 (instructions) · Claude Code | 100% | 100% | Strict pass: the whole reply is the right JSON: 76%–100%; Right answer in any format (strict pass or format miss): 76%–100% | 12 |
| GPT-6.1 Sol (low, instructions) · Codex CLI | 100% | 100% | Strict pass: the whole reply is the right JSON: 76%–100%; Right answer in any format (strict pass or format miss): 76%–100% | 12 |
List-price calculation, not a run. 2 rows, 2 series: Strict pass: the whole reply is the right JSON, Right answer in any format (strict pass or format miss). Strict pass: the whole reply is the right JSON: all at 100%. Right answer in any format (strict pass or format miss): all at 100%.
NotesWhiskers: 95% Wilson intervaln = 12 per row
Three extraction prompts pooled; whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals (a calculation) over 24 calls and 12 calls per configuration; every error counts as a fail. Strict: the whole reply parses as JSON and matches the expected answer exactly. A format miss is a right answer inside a code fence or prose, so it is never a strict pass.
Source: JSON schema vs instructions
- Strict pass
- Format miss
- Wrong values
- Error
One square per call; counts at the right are exact and in legend order.
| Item | Strict pass | Format miss | Wrong values | Error | n |
|---|---|---|---|---|---|
| Claude Haiku 4.5 (instructions) · Claude Code | 0 | 17 | 7 | 0 | 24 |
| Claude Sonnet 5.5 (instructions) · Claude Code | 12 | 0 | 0 | 0 | 12 |
| GPT-6.1 Sol (low, instructions) · Codex CLI | 12 | 0 | 0 | 0 | 12 |
3 rows, 4 series: Strict pass, Format miss, Wrong values, Error. Strict pass: highest Claude Sonnet 5.5 (instructions) · Claude Code 12 (n 12). Lowest Claude Haiku 4.5 (instructions) · Claude Code 0 (n 24). Format miss: highest Claude Haiku 4.5 (instructions) · Claude Code 17 (n 24). Lowest GPT-6.1 Sol (low, instructions) · Codex CLI 0 (n 12).
Notesn 12–24 per row
Counts of calls per configuration; the three prompts pooled
Counts of calls, not rates; the pass-rate chart carries the same results with 95% intervals. Format miss: the right answer inside a code fence or prose. Wrong values: any other completed reply, with a wrong value, key or type (a reply that sits in a code fence and also has a wrong value is counted here). Error: the call did not complete.
Source: JSON schema vs instructions
1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.
| Item | Median time per call (the three prompts pooled) | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 (instructions) · Claude Code | 3.5 s | 2.7 s–4.1 s | 12 |
| GPT-6.1 Sol (low, instructions) · Codex CLI | 6.2 s | 4.2 s–12.3 s | 12 |
2 rows. Slowest GPT-6.1 Sol (low, instructions) · Codex CLI 6.2 s (range 4.2 s–12.3 s, n 12). Fastest Claude Sonnet 5.5 (instructions) · Claude Code 3.5 s (range 2.7 s–4.1 s, n 12). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 12 per row
Median; whiskers = fastest and slowest completed call
Whiskers are a range (fastest and slowest call), not a confidence interval. Wall time from process start to exit, so it includes CLI start-up; Codex CLI timings include its larger system prompt. The three prompts differ in length, which widens every range.
Source: JSON schema vs instructions
- Median output tokens per call
- of which median reasoning tokens per call (thinking) (inner bar)
| Item | Median output tokens per call | Median reasoning tokens per call (thinking) | n |
|---|---|---|---|
| Claude Sonnet 5.5 (instructions) · Claude Code | 368 | 182 | 12 |
| GPT-6.1 Sol (low, instructions) · Codex CLI | 117 | 21 | 12 |
2 rows, 2 series: Median output tokens per call, Median reasoning tokens per call (thinking). Median output tokens per call: highest Claude Sonnet 5.5 (instructions) · Claude Code 368 (n 12). Lowest GPT-6.1 Sol (low, instructions) · Codex CLI 117 (n 12). Median reasoning tokens per call (thinking): highest Claude Sonnet 5.5 (instructions) · Claude Code 182 (n 12). Lowest GPT-6.1 Sol (low, instructions) · Codex CLI 21 (n 12).
Notesn = 12 per row
Median per call; ranges and sample sizes are in the note
Output tokens include reasoning tokens. Claude schema-mode output includes the CLI’s structured-output tool call. Input counts include CLI context and are not compared. Ranges (not intervals): Claude Haiku 4.5 (instructions) · Claude Code: output 698 to 2059 (n = 24); reasoning 559 to 1830 (n = 24); Claude Haiku 4.5 (JSON schema) · Claude Code: output 716 to 1462 (n = 24); reasoning 522 to 1187 (n = 24); Claude Sonnet 5.5 (instructions) · Claude Code: output 154 to 456 (n = 12); reasoning 54 to 278 (n = 12); Claude Sonnet 5.5 (JSON schema) · Claude Code: output 273 to 541 (n = 12); reasoning 0 to 260 (n = 12); GPT-6.1 Sol (low, instructions) · Codex CLI: output 69 to 259 (n = 12); reasoning 0 to 60 (n = 12); GPT-6.1 Sol (low, JSON schema) · Codex CLI: output 98 to 181 (n = 12); reasoning 0 to 49 (n = 12).
Source: JSON schema vs instructions
- Reasoning (output tokens)
- Remaining output (visible-answer estimate)
- Input (prompt, cache priced)
Totals are the sum of the parts shown. Shares are calculated from the same values.
| Item | Reasoning (output tokens) | Remaining output (visible-answer estimate) | Input (prompt, cache priced) | n |
|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | $0.0023 | $0.003 | $0.02 | 16 |
| Claude Sonnet 5.5 · Claude Code | $0.0067 | $0.0037 | $0.004 | 24 |
List-price calculation, not a run. 2 rows, 3 series: Reasoning (output tokens), Remaining output (visible-answer estimate), Input (prompt, cache priced). Reasoning (output tokens): highest Claude Sonnet 5.5 · Claude Code $0.0067 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.0023 (n 16). Remaining output (visible-answer estimate): highest Claude Sonnet 5.5 · Claude Code $0.0037 (n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.003 (n 16).
Notesn 16–24 per row
Mean per call on the hard tasks; the three parts add up to the call
Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
- Reasoning cost per strict pass
- Total cost per strict pass (square)
Gap labels, Total cost per strict pass vs Reasoning cost per strict pass: Total cost per strict pass is x% higher (+) or lower (−) than Reasoning cost per strict pass, calculated from the two values shown (the change counted from Reasoning cost per strict pass’s value).
| Item | Reasoning cost per strict pass | Total cost per strict pass | n |
|---|---|---|---|
| Claude Sonnet 5.5 (medium) · Claude Code | $0.0059 | $0.014 | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | $0.0023 | $0.026 | 16 |
List-price calculation, not a run. 2 rows, 2 series: Reasoning cost per strict pass, Total cost per strict pass. Reasoning cost per strict pass: highest Claude Sonnet 5.5 (medium) · Claude Code $0.0059 (n 16). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.0023 (n 16). Total cost per strict pass: highest GPT-6.1 Sol (medium) · Codex CLI $0.026 (n 16). Lowest Claude Sonnet 5.5 (medium) · Claude Code $0.014 (n 16).
Notesn = 16 per row
List price ÷ strict passes; every cell is 8 tasks × 2 repetitions
Calculation, not a bill: reported tokens × list price, divided by the cell's strict passes; the calls ran on flat subscriptions. The effort-ladder cells: new calls plus reference cells reused from the hard head-to-head (Claude repetitions 1-2 only). "Default" means the effort flag was not passed. The total is the same value as the effort-ladder cost-per-pass chart. Effort levels are not the same scale across vendors, and the reference cells ran in a different batch and hour.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
- Eight hard tasks
- Five short tasks (square)
Gap labels, Five short tasks vs Eight hard tasks: Five short tasks is x percentage points higher (+) or lower (−) than Eight hard tasks, calculated from the two values shown; lines are the lowest–highest run (not an interval).
| Item | Eight hard tasks | Five short tasks | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 55% | 0% | Eight hard tasks: 0%–96%; Five short tasks: 0%–73% | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | 46% | 41% | Eight hard tasks: 11%–87%; Five short tasks: 0%–71% | 16 |
List-price calculation, not a run. 2 rows, 2 series: Eight hard tasks, Five short tasks. Eight hard tasks: highest Claude Sonnet 5.5 · Claude Code 55% (range 0%–96%, n 24). Lowest GPT-6.1 Sol (medium) · Codex CLI 46% (range 11%–87%, n 16). All run ranges overlap. Five short tasks: highest GPT-6.1 Sol (medium) · Codex CLI 41% (range 0%–71%, n 15). Lowest Claude Sonnet 5.5 · Claude Code 0% (range 0%–73%, n 15). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 15–24 per row
Median call per configuration; five short tasks and eight hard tasks
Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first text | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 2 s | 0.9 s–4.1 s | 4 |
| GPT-6.1 Sol (low) · Codex CLI | 3.5 s | 2.8 s–4.4 s | 4 |
2 rows. Slowest GPT-6.1 Sol (low) · Codex CLI 3.5 s (range 2.8 s–4.4 s, n 4). Fastest Claude Sonnet 5.5 · Claude Code 2 s (range 0.9 s–4.1 s, n 4). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = fastest and slowest call
Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.
Source: LLM speed anatomy
| Item | Visible tokens per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 232 | 230–233 | 4 |
| GPT-6.1 Sol (low) · Codex CLI | 80 | 72–81 | 4 |
List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 · Claude Code 232 (range 230–233, n 4). Lowest GPT-6.1 Sol (low) · Codex CLI 80 (range 72–81, n 4). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = slowest and fastest call
Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.
Source: LLM speed anatomy
| Item | Characters per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 517 | 513–519 | 4 |
| GPT-6.1 Sol (low) · Codex CLI | 323 | 291–327 | 4 |
List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 · Claude Code 517 (range 513–519, n 4). Lowest GPT-6.1 Sol (low) · Codex CLI 323 (range 291–327, n 4). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call
Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
- Claude Haiku 4.5 · Claude Code
- Claude Sonnet 5.5 · Claude Code
- Claude Opus 5.5 · Claude Code
- GPT-6.1 Sol (low) · Codex CLI
| Prompt-size target (approximate Haiku tokens; calibration calculation) | Claude Haiku 4.5 · Claude Code | Claude Sonnet 5.5 · Claude Code | Claude Opus 5.5 · Claude Code | GPT-6.1 Sol (low) · Codex CLI | Range (lowest–highest run) | n |
|---|---|---|---|---|---|---|
| 1k | 1.9 s | 1.5 s | 1.5 s | 3.4 s | Claude Haiku 4.5 · Claude Code: 1.9 s–2 s; Claude Sonnet 5.5 · Claude Code: 1.2 s–1.7 s; Claude Opus 5.5 · Claude Code: 1.5 s–2 s; GPT-6.1 Sol (low) · Codex CLI: 3.4 s–4.8 s | 3 |
| 16k | 2.3 s | 1.8 s | 1.7 s | 4 s | Claude Haiku 4.5 · Claude Code: 2.2 s–2.5 s; Claude Sonnet 5.5 · Claude Code: 1.6 s–2.1 s; Claude Opus 5.5 · Claude Code: 1.7 s–3 s; GPT-6.1 Sol (low) · Codex CLI: 3.3 s–4.3 s | 3 |
| 64k | 2.8 s | 3.1 s | 1.8 s | 3.9 s | Claude Haiku 4.5 · Claude Code: 2.5 s–2.9 s; Claude Sonnet 5.5 · Claude Code: 1.4 s–3.6 s; Claude Opus 5.5 · Claude Code: 1.7 s–3.7 s; GPT-6.1 Sol (low) · Codex CLI: 3.4 s–4.4 s | 3 |
List-price calculation, not a run. 3 rows, 4 series: Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (low) · Codex CLI. Claude Haiku 4.5 · Claude Code: slowest 64k 2.8 s (range 2.5 s–2.9 s, n 3). Fastest 1k 1.9 s (range 1.9 s–2 s, n 3). Not all run ranges overlap. Claude Sonnet 5.5 · Claude Code: slowest 64k 3.1 s (range 1.4 s–3.6 s, n 3). Fastest 1k 1.5 s (range 1.2 s–1.7 s, n 3). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Median of 3 calls per size; whiskers = fastest and slowest call
Each call used a new ledger seed. Cache-read counts stayed within the short-prompt baseline (see the cache table). This does not identify which tokens were cached. Sizes name the text we send; each model’s reported input tokens are in the table and include the CLI’s own prefix. The size calibration subtracts estimated prefixes from probe input counts; these are calculations, not measured prefix counts for each call. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
1k prompt
16k prompt
64k prompt
One panel per series, all on the same axis; whiskers are the fastest–slowest run (not an interval).
| Item | 1k prompt | 16k prompt | 64k prompt | Range (lowest–highest run) | n |
|---|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1.8 s | 2.1 s | 3.4 s | 1k prompt: 1.6 s–2.1 s; 16k prompt: 2 s–2.5 s; 64k prompt: 1.7 s–4.4 s | 3 |
| GPT-6.1 Sol (low) · Codex CLI | 3.4 s | 4.1 s | 4 s | 1k prompt: 3.4 s–4.9 s; 16k prompt: 4 s–4.7 s; 64k prompt: 3.5 s–4.4 s | 3 |
2 rows, 3 series: 1k prompt, 16k prompt, 64k prompt. 1k prompt: slowest GPT-6.1 Sol (low) · Codex CLI 3.4 s (range 3.4 s–4.9 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 1.8 s (range 1.6 s–2.1 s, n 3). Not all run ranges overlap. 16k prompt: slowest GPT-6.1 Sol (low) · Codex CLI 4.1 s (range 4 s–4.7 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 2.1 s (range 2 s–2.5 s, n 3). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Median of 3 calls per bar; whiskers = fastest and slowest call
Whole call: CLI start-up, first text and the one-line answer. Whiskers are a range of calls, not a confidence interval. Each call used a new ledger.
Source: LLM speed anatomy
| Item | Exact answer | 95% interval | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 100% | 70%–100% | 9 |
| GPT-6.1 Sol (low) · Codex CLI | 100% | 70%–100% | 9 |
2 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 9 per row
All sizes together per model; whiskers = 95% Wilson intervals
Whiskers are 95% Wilson intervals. One lookup question per call; a reply with extra words is a format miss, not a pass. With 9 calls per model, a perfect score still has a wide interval.
Source: LLM speed anatomy
- Strict pass
- Lenient (format misses counted)
| Item | Strict pass | Lenient (format misses counted) | 95% interval | n |
|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 69% | 69% | Strict pass: 44%–86%; Lenient (format misses counted): 44%–86% | 16 |
| Claude Sonnet 5.5 · Claude Code | 38% | 38% | Strict pass: 18%–61%; Lenient (format misses counted): 18%–61% | 16 |
2 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest GPT-6.1 Sol (medium) · Codex CLI 69% (95% interval 44%–86%, n 16). Lowest Claude Sonnet 5.5 · Claude Code 38% (95% interval 18%–61%, n 16). All intervals overlap. Lenient (format misses counted): highest GPT-6.1 Sol (medium) · Codex CLI 69% (95% interval 44%–86%, n 16). Lowest Claude Sonnet 5.5 · Claude Code 38% (95% interval 18%–61%, n 16). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 16 per row
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row. Counted calls are new calls.
Source: Harder tasks head-to-head
| Item | Tool attempt | 95% interval | n |
|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 0% | 0%–19% | 16 |
| Claude Sonnet 5.5 · Claude Code | 31% | 14%–56% | 16 |
2 rows. Highest Claude Sonnet 5.5 · Claude Code 31% (95% interval 14%–56%, n 16). Lowest GPT-6.1 Sol (medium) · Codex CLI 0% (95% interval 0%–19%, n 16). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 16 per row
Share of calls whose reply held tool-call markup or whose tool call the CLI could not parse
Every prompt asked for the answer only. Both CLIs disabled tools. The Codex runner adds a developer instruction not to call tools. The Claude Code runner adds no such instruction. The validator decides the score. Tool-call markup is a separate flag; a flagged call can still pass if its whole reply meets the validator. It is a behaviour, not a quality score: more is not better. No call with a tool attempt passed. Whiskers are 95% Wilson intervals. The flag comes from the reply text, which is not published.
Source: Harder tasks head-to-head
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | GPT-6.1 Sol (medium) · Codex CLI | Claude Opus 5.5 · Claude Code | Claude Sonnet 5.5 · Claude Code | Claude Haiku 4.5 · Claude Code | 95% interval | n |
|---|---|---|---|---|---|---|
| 10x10 nonogram | 75% | 100% | 100% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 30%–95%; Claude Opus 5.5 · Claude Code: 44%–100%; Claude Sonnet 5.5 · Claude Code: 51%–100%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
| Sudoku, 22 givens | 25% | 0% | 0% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 4.6%–70%; Claude Opus 5.5 · Claude Code: 0%–56%; Claude Sonnet 5.5 · Claude Code: 0%–49%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
| 6x6 Skyscrapers | 100% | 33% | 0% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 51%–100%; Claude Opus 5.5 · Claude Code: 6.2%–79%; Claude Sonnet 5.5 · Claude Code: 0%–49%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
| Seeded shuffle output | 75% | 33% | 50% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 30%–95%; Claude Opus 5.5 · Claude Code: 6.2%–79%; Claude Sonnet 5.5 · Claude Code: 15%–85%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
4 rows, 4 series: GPT-6.1 Sol (medium) · Codex CLI, Claude Opus 5.5 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Haiku 4.5 · Claude Code. GPT-6.1 Sol (medium) · Codex CLI: highest 6x6 Skyscrapers 100% (95% interval 51%–100%, n 4). Lowest Sudoku, 22 givens 25% (95% interval 4.6%–70%, n 4). All intervals overlap. Claude Opus 5.5 · Claude Code: highest 10x10 nonogram 100% (95% interval 44%–100%, n 3). Lowest Sudoku, 22 givens 0% (95% interval 0%–56%, n 3). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 3–4 per row
One bar per configuration and task; each bar rests on only a few calls, so the 95% intervals are wide
Each bar is 3 to 4 calls, so one call moves a bar by a quarter or a third. Whiskers are 95% Wilson intervals; with this few calls they overlap almost everywhere.
Source: Harder tasks head-to-head
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call | Range (lowest–highest run) | n |
|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 120 s | 46.2 s–273 s | 13 |
| Claude Sonnet 5.5 · Claude Code | 70.4 s | 4.3 s–210 s | 12 |
2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 120 s (range 46.2 s–273 s, n 13). Fastest Claude Sonnet 5.5 · Claude Code 70.4 s (range 4.3 s–210 s, n 12). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 12–13 per row
Median per configuration; whiskers = fastest and slowest call
Median and range over the calls that completed. Completed calls include wrong answers and format misses. Only timeouts and tool-call parse errors are excluded from this run’s timings. Both count as non-passes in the outcomes chart. One Mac, one network, one session. Arena servers shared the Mac during part of the run. Host load was not controlled, so these times cannot isolate model speed. Whiskers are a range, not a confidence interval. Times include the CLI start-up and the CLI’s own system prompt. Highlighted: configurations that passed every call.
Source: Harder tasks head-to-head
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | Range (lowest–highest run) | n |
|---|---|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | 4,994 | 4,971 | Output tokens: 2,099–13,413; Reasoning tokens: 2,070–13,372 | 13 |
| Claude Sonnet 5.5 · Claude Code | 9,287 | 6,557 | Output tokens: 407–27,921; Reasoning tokens: 63–27,902 | 12 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Sonnet 5.5 · Claude Code 9,287 (range 407–27,921, n 12). Lowest GPT-6.1 Sol (medium) · Codex CLI 4,994 (range 2,099–13,413, n 13). All run ranges overlap. Reasoning tokens: highest Claude Sonnet 5.5 · Claude Code 6,557 (range 63–27,902, n 12). Lowest GPT-6.1 Sol (medium) · Codex CLI 4,971 (range 2,070–13,372, n 13). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 12–13 per row
Median per configuration; reasoning tokens as the CLI reports them
Medians and minimum-to-maximum token ranges cover completed calls only. Ranges are not confidence intervals. The chart omits unknown reasoning counts. Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. Claude Code used an output-token cap setting of 16,000. Some reported totals exceeded it. Codex CLI had no cap. More tokens is not better or worse by itself.
Source: Harder tasks head-to-head
| Item | Cost per strict pass | n |
|---|---|---|
| GPT-6.1 Sol (medium) · Codex CLI | $0.083 | 16 |
| Claude Sonnet 5.5 · Claude Code | $0.24 | 16 |
List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 · Claude Code $0.24 (n 16). Lowest GPT-6.1 Sol (medium) · Codex CLI $0.083 (n 16).
Notesn = 16 per row
All calls in a configuration, failures and format misses included, divided by its strict passes
Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Timeout calls report no tokens and are not priced. Unpriced calls: GPT-6.1 Sol (medium) 3 of 16, Opus 5.5 1 of 12 and Sonnet 5.5 4 of 16. These cells show a lower bound. Assume each unpriced call cost its cell’s median priced call. This sensitivity calculation gives GPT-6.1 Sol (medium) $0.100, Opus 5.5 $0.633 and Sonnet 5.5 $0.303. Opus 5.5 figures are provisional: its cache-read price is under re-check. Highlights mark the observed frontier of these lower-bound costs. Unknown timeout costs can change it; this is not a cost ranking. Claude Haiku 4.5 · Claude Code had no strict pass, so it has no cost per pass.
Sources: Harder tasks head-to-head, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Which is better, Claude Code or Codex CLI?
- Claude Code and Codex CLI share 66 measured metrics and 22 list-price calculations from 12 studies. Claude Code leads on 7 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 4 more. Codex CLI leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 31 ties and 49 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Each run pairs a CLI with a model, so these rows cannot separate the CLI from the model; the contexts name both. Some rows rest on small samples (n = 2 at the smallest).
- How were Claude Code and Codex CLI measured?
- They share 66 measured metrics and 22 list-price calculations from 12 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks; Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks; Prompt caching and run-to-run consistency in Claude Code and Codex CLI; Routing overhead: deterministic policy vs LLM routers vs Jev; Claude Code CLI vs Codex CLI vs the API: latency and tokens; Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks; Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI; How much of an AI bill is thinking? Reasoning tokens by model and effort; Where the seconds go: first text, output speed and prompt size for 6 LLMs; GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Every row names its configuration, its sample size and its interval or range.
- How do Claude Code and Codex CLI compare on time per coding session?
- Claude Code: 23.1 s (Claude Sonnet 5.5 · six small repository tasks with hidden tests; n = 12; run range 18.7 s to 44.5 s). Codex CLI: 113.4 s (GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests; n = 12; run range 78.5 s to 222 s). The run ranges (fastest to slowest) do not overlap (Claude Code 18.7 s to 44.5 s; Codex CLI 78.5 s to 221.9 s). A range is not a confidence interval.
- How do Claude Code and Codex CLI compare on same prompt, 10 times: time per call (Exact number)?
- Claude Code: 6.89 s (Claude Sonnet 5.5 · same prompt repeated 10 times; n = 10; run range 5.8 s to 7.8 s). Codex CLI: 13.4 s (GPT-6.1 Sol · effort medium · same prompt repeated 10 times; n = 10; run range 12.3 s to 18 s). The run ranges (fastest to slowest) do not overlap (Claude Code 5.81 s to 7.81 s; Codex CLI 12.3 s to 18.0 s). A range is not a confidence interval.
- How do Claude Code and Codex CLI compare on same prompt, 10 times: time per call (Code fix)?
- Claude Code: 2.67 s (Claude Sonnet 5.5 · same prompt repeated 10 times; n = 10; run range 2.3 s to 4.3 s). Codex CLI: 11.3 s (GPT-6.1 Sol · effort medium · same prompt repeated 10 times; n = 10; run range 9.1 s to 14.9 s). The run ranges (fastest to slowest) do not overlap (Claude Code 2.32 s to 4.34 s; Codex CLI 9.08 s to 14.8 s). A range is not a confidence interval.
- How do Claude Code and Codex CLI compare on cLI start-up tax on a one-word answer (First model output)?
- Claude Code: 1,461 ms (Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs; n = 5; run range 1.21 s to 2.31 s). Codex CLI: 5,059 ms (default model · CLI start-up, one-word prompt, 5 runs; n = 5; run range 4.39 s to 5.48 s). The run ranges (fastest to slowest) do not overlap (Claude Code 1,206 ms to 2,308 ms; Codex CLI 4,391 ms to 5,478 ms). A range is not a confidence interval. Samples are small (5 runs per side).
- How do Claude Code and Codex CLI compare on cLI start-up tax on a one-word answer (Total wall time)?
- Claude Code: 2,529 ms (Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs; n = 5; run range 2.27 s to 3.38 s). Codex CLI: 5,999 ms (default model · CLI start-up, one-word prompt, 5 runs; n = 5; run range 5.37 s to 6.51 s). The run ranges (fastest to slowest) do not overlap (Claude Code 2,273 ms to 3,382 ms; Codex CLI 5,367 ms to 6,506 ms). A range is not a confidence interval. Samples are small (5 runs per side).
The studies behind this page
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.
Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
36 graded sessions: Claude Code with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol. All passed every hidden test; time, tool calls and diffs differ.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
176 calls on 8 hard tasks at low, medium, high and default effort: Sonnet, Opus and GPT-6.1 Sol. Pass rate with 95% intervals, time, tokens, cost per pass.
Prompt caching and run-to-run consistency in Claude Code and Codex CLI
135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.
Routing overhead: deterministic policy vs LLM routers vs Jev
How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.
Claude Code CLI vs Codex CLI vs the API: latency and tokens
194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.
Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
118 attempts on 8 hard tasks: one call vs an agent loop that runs code in a sandbox. Pass rate, time, tokens and cost.
Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
96 calls. Strict passes, schema vs instructions: Haiku 18/24 vs 0/24, Sonnet 12/12 vs 12/12, GPT-6.1 Sol 12/12 vs 12/12. With intervals.
How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.