66 measured metrics · 22 calculated · 12 studies

Claude CodevsCodex CLI

Claude Code ahead on 7, Codex CLI ahead on 1; 31 ties, 49 unclear. A side is ahead only where the intervals or ranges do not overlap.

The verdict

Claude Code and Codex CLI share 66 measured metrics and 22 list-price calculations from 12 studies. Claude Code leads on 7 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 4 more. Codex CLI leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 31 ties and 49 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Each run pairs a CLI with a model, so these rows cannot separate the CLI from the model; the contexts name both. Some rows rest on small samples (n = 2 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Input tokens per call: what the CLI sends (Cache read), 3.7x (Codex CLI larger).

Watch it build

A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.

Live story · 41 sClaude Code vs Codex CLI: what the measurements say

Claude Code vs Codex CLI: what the measurements say

88 comparison rows from 12 studies: 7 rows favour Claude Code, 1 favour Codex CLI, 80 are ties or unclear. Cost rows are calculations.

Transcript
  1. Comparison · 88 rows · 12 studies. Claude Code vs Codex CLI. A winner only where the 95% intervals or run ranges do not overlap.
  2. 88 comparison rows from 12 studies: Claude Code ahead on 7, Codex CLI ahead on 1. The rest do not separate them. Rows where Claude Code is ahead: 7 (of 88). Rows where Codex CLI is ahead: 1 (of 88). Ties or unclear: 80 (31 ties · 49 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
  3. Coding agents, hidden tests: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 4 rows separates them. Table: Coding agents, hidden tests · Claude Sonnet 5.5 vs GPT-6.1 Sol · n = 12 per side. Source study: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks. Rows shown: Coding sessions that passed every hidden check; Time per coding session; Tool calls per coding session; List-price cost per passing coding session (calculation). Recorded settings: Claude Sonnet 5.5 · six small repository tasks with hidden tests; GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests. Includes a calculation, not a bill or a new run. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
  4. Caching sessions: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 2 of 9 rows separate them. Table: Caching sessions · 5 of 9 rows · Claude Sonnet 5.5 vs GPT-6.1 Sol · n = 10 per side. Source study: Prompt caching and run-to-run consistency in Claude Code and Codex CLI. Rows shown: Same prompt, 10 times: strict pass rate (Exact number); Same prompt, 10 times: strict pass rate (JSON object); Same prompt, 10 times: strict pass rate (Code fix); Same prompt, 10 times: time per call (Exact number); Same prompt, 10 times: time per call (Code fix). Recorded settings: Claude Sonnet 5.5 · same prompt repeated 10 times; GPT-6.1 Sol · effort medium · same prompt repeated 10 times. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  5. Routing overhead: pass rate not measured. 2 of 4 rows separate them. Table: Routing overhead · Claude Haiku 4.5 vs default model · n = 5 per side. Source study: Routing overhead: deterministic policy vs LLM routers vs Jev. Rows shown: CLI start-up tax on a one-word answer (First output event); CLI start-up tax on a one-word answer (First model output); CLI start-up tax on a one-word answer (Total wall time); Input tokens a CLI sends for a one-word answer. Recorded settings: Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs; default model · CLI start-up, one-word prompt, 5 runs. Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
  6. Single call vs agent loop: pass rate 100% vs 63%, Claude Code ahead: 95% intervals separate. 1 of 14 rows separates them. Table: Single call vs agent loop · 5 of 14 rows · Claude Sonnet 5.5 vs GPT-6 Luna · n = 3–24 vs 2–16. Source study: Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks. Rows shown: Strict pass rate: single call vs agent loop on eight hard tasks; Strict passes per task: single call vs agent loop: Interval merge fix; Strict passes per task: single call vs agent loop: DST day-length fix; Strict passes per task: single call vs agent loop: CSV parser; Strict passes per task: single call vs agent loop: Event-loop order. Recorded settings: Claude Sonnet 5.5 · single call; GPT-6 Luna · single call. Caveat: Each row is a CLI + model pair. Claude Code and Codex CLI add their own system prompts and tool schemas, and Codex CLI also loads the account’s user-level instruction file. A gap between Claude and GPT-6 Luna rows is partly the CLI.
  7. No winner where the data shows none. Showing 18 of 88 rows; every row and its reason online.

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Claude Code
  • Codex CLI
  • 95% interval
  • fastest–slowest run (not an interval)
  • where the two overlap
  • hollow: list-price calculation

These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

1 tie · 7 unclear
  • Pass rate on five validated tasks: Claude Code 80% (12/15) (n 15, 95% interval 55%–93%); Codex CLI 100% (15/15) (n 15, 95% interval 80%–100%). Tie.

    Calculation: at these rates, about 33 runs per side would separate them.

  • Total time per call: Claude Code 2.31 s (n 15, run range 2.2 s–7.7 s); Codex CLI 5.65 s (n 15, run range 4.1 s–25.5 s). Unclear.
  • Time to first useful output: Claude Code 1.56 s (n 15, run range 1 s–6.4 s); Codex CLI 5.05 s (n 15, run range 3.4 s–17.8 s). Unclear.
  • Input tokens per call: what the CLI sends (Cache read): Claude Code 1,401 (n 15); Codex CLI 5,180 (n 15). Unclear.
  • Input tokens per call: what the CLI sends (Other input): Claude Code 685 (n 15); Codex CLI 6,943 (n 15). Unclear.
  • Output tokens per call (Output tokens): Claude Code 107 (n 15); Codex CLI 42 (n 15). Unclear.
  • List-price cost per call (calculation), calculation: Claude Code $0.0036 (n 15, run range $0.0034–$0.01); Codex CLI $0.010 (n 15, run range $0.0054–$0.027). Unclear.
  • List-price cost per passing answer (calculation), calculation: Claude Code $0.0062 (n 15); Codex CLI $0.016 (n 15). Unclear.

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

2 ties · 4 unclear
  • Pass rate on eight hard tasks (Strict pass): Claude Code 100% (24/24) (n 24, 95% interval 86%–100%); Codex CLI 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Pass rate on eight hard tasks (Lenient (format misses counted)): Claude Code 100% (24/24) (n 24, 95% interval 86%–100%); Codex CLI 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Total time per call on hard tasks (separate batches): Claude Code 7.75 s (n 24, run range 2.3 s–34.8 s); Codex CLI 13.1 s (n 16, run range 8.5 s–61.6 s). Unclear.
  • Time to first useful output on hard tasks: Claude Code 5.95 s (n 24, run range 0.9 s–30.6 s); Codex CLI 10.2 s (n 16, run range 6.1 s–40.4 s). Unclear.
  • Output tokens per call on hard tasks (Output tokens): Claude Code 1,050 (n 24); Codex CLI 335 (n 16). Unclear.
  • List-price cost per strict pass on hard tasks (calculation), calculation: Claude Code $0.014 (n 24); Codex CLI $0.026 (n 16). Unclear.

Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks

Claude Code 1 · 1 tie · 2 unclear
  • Coding sessions that passed every hidden check: Claude Code 100% (12/12) (n 12, 95% interval 76%–100%); Codex CLI 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
  • Time per coding session: Claude Code 23.1 s (n 12, run range 18.7 s–44.5 s); Codex CLI 113.4 s (n 12, run range 78.5 s–222 s). Claude Code ahead.
  • Tool calls per coding session: Claude Code 7.5 (n 12, run range 3–14); Codex CLI 12.5 (n 12, run range 8–18). Unclear.
  • List-price cost per passing coding session (calculation), calculation: Claude Code $0.085 (n 12); Codex CLI $0.098 (n 12). Unclear.

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks

1 tie · 3 unclear
  • Strict pass rate by effort on eight hard tasks: Claude Code 100% (16/16) (n 16, 95% interval 81%–100%); Codex CLI 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Total time per call by effort on hard tasks: Claude Code 7.63 s (n 16, run range 2.7 s–24 s); Codex CLI 13.1 s (n 16, run range 8.5 s–61.6 s). Unclear.
  • Output tokens per call by effort on hard tasks (Output tokens): Claude Code 770 (n 16); Codex CLI 335 (n 16). Unclear.
  • List-price cost per strict pass by effort (calculation), calculation: Claude Code $0.014 (n 16); Codex CLI $0.026 (n 16). Unclear.

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

Claude Code 2 · 5 ties · 2 unclear
  • Same prompt, 10 times: strict pass rate (Exact number): Claude Code 100% (10/10) (n 10, 95% interval 72%–100%); Codex CLI 100% (10/10) (n 10, 95% interval 72%–100%). Tie.
  • Same prompt, 10 times: strict pass rate (JSON object): Claude Code 100% (10/10) (n 10, 95% interval 72%–100%); Codex CLI 100% (10/10) (n 10, 95% interval 72%–100%). Tie.
  • Same prompt, 10 times: strict pass rate (Code fix): Claude Code 100% (10/10) (n 10, 95% interval 72%–100%); Codex CLI 100% (10/10) (n 10, 95% interval 72%–100%). Tie.
  • Same prompt, 10 times: how many different answers (Exact number): Claude Code 1 (n 10); Codex CLI 1 (n 10). Tie.
  • Same prompt, 10 times: how many different answers (JSON object): Claude Code 1 (n 10); Codex CLI 1 (n 10). Tie.
  • Same prompt, 10 times: how many different answers (Code fix): Claude Code 3 (n 10); Codex CLI 6 (n 10). Unclear.
  • Same prompt, 10 times: time per call (Exact number): Claude Code 6.89 s (n 10, run range 5.8 s–7.8 s); Codex CLI 13.4 s (n 10, run range 12.3 s–18 s). Claude Code ahead.
  • Same prompt, 10 times: time per call (JSON object): Claude Code 2.89 s (n 10, run range 2.7 s–5.3 s); Codex CLI 6.42 s (n 10, run range 5.3 s–8.3 s). Unclear.
  • Same prompt, 10 times: time per call (Code fix): Claude Code 2.67 s (n 10, run range 2.3 s–4.3 s); Codex CLI 11.3 s (n 10, run range 9.1 s–14.9 s). Claude Code ahead.

Routing overhead: deterministic policy vs LLM routers vs Jev

Claude Code 2 · 2 unclear
  • CLI start-up tax on a one-word answer (First output event): Claude Code 563 ms (n 5, run range 519 ms–726 ms); Codex CLI 489 ms (n 5, run range 354 ms–1.3 s). Unclear.
  • CLI start-up tax on a one-word answer (First model output): Claude Code 1,461 ms (n 5, run range 1.21 s–2.31 s); Codex CLI 5,059 ms (n 5, run range 4.39 s–5.48 s). Claude Code ahead.
  • CLI start-up tax on a one-word answer (Total wall time): Claude Code 2,529 ms (n 5, run range 2.27 s–3.38 s); Codex CLI 5,999 ms (n 5, run range 5.37 s–6.51 s). Claude Code ahead.
  • Input tokens a CLI sends for a one-word answer: Claude Code 6,761 (n 5); Codex CLI 17,051 (n 5). Unclear.

Claude Code CLI vs Codex CLI vs the API: latency and tokens

3 unclear
  • Repairing a scheduler: Claude Code vs Codex vs API (Total time): Claude Code 15.0 s (n 3, run range 13.9 s–15.9 s); Codex CLI 61.2 s (n 3, run range 59.9 s–69.5 s). Unclear.
  • Repairing a scheduler: Claude Code vs Codex vs API (First useful output): Claude Code 7.55 s (n 3, run range 6.8 s–7.6 s); Codex CLI 15.6 s (n 3, run range 13.7 s–23 s). Unclear.
  • Output tokens to repair the scheduler (Output tokens): Claude Code 2,227 (n 3); Codex CLI 1,181 (n 3). Unclear.

Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks

Claude Code 1 · 9 ties · 4 unclear
  • Strict pass rate: single call vs agent loop on eight hard tasks: Claude Code 100% (24/24) (n 24, 95% interval 86%–100%); Codex CLI 63% (10/16) (n 16, 95% interval 39%–82%). Claude Code ahead.
  • Strict passes per task: single call vs agent loop: Interval merge fix: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
  • Strict passes per task: single call vs agent loop: DST day-length fix: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
  • Strict passes per task: single call vs agent loop: CSV parser: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
  • Strict passes per task: single call vs agent loop: Event-loop order: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 0% (0/2) (n 2, 95% interval 0%–66%). Tie.

    Calculation: at these rates, about 4 runs per side would separate them.

  • Strict passes per task: single call vs agent loop: Room schedule: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 50% (1/2) (n 2, 95% interval 9.4%–91%). Tie.

    Calculation: at these rates, about 12 runs per side would separate them.

  • Strict passes per task: single call vs agent loop: SemVer regex: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 50% (1/2) (n 2, 95% interval 9.4%–91%). Tie.

    Calculation: at these rates, about 12 runs per side would separate them.

  • Strict passes per task: single call vs agent loop: Money refactor: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 0% (0/2) (n 2, 95% interval 0%–66%). Tie.

    Calculation: at these rates, about 4 runs per side would separate them.

  • Strict passes per task: single call vs agent loop: SQL report: Claude Code 100% (3/3) (n 3, 95% interval 44%–100%); Codex CLI 100% (2/2) (n 2, 95% interval 34%–100%). Tie.
  • Total time per attempt: single call vs agent loop: Claude Code 7.75 s (n 24, run range 2.3 s–34.8 s); Codex CLI 5.16 s (n 16, run range 3.6 s–11.3 s). Unclear.
  • Tokens per attempt: single call vs agent loop (Input tokens (cache reads included)): Claude Code 2,281 (n 24, run range 2,234–2,669); Codex CLI 11,582 (n 16, run range 11,526–11,818). Unclear.
  • Tokens per attempt: single call vs agent loop (Output tokens): Claude Code 1,050 (n 24, run range 176–3,895); Codex CLI 345 (n 16, run range 36–634). Unclear.
  • Tool calls per agent-loop attempt: Claude Code 0 (n 16, run range 0–3); Codex CLI 0 (n 14, run range 0–1). Tie.
  • List-price cost per strict pass: single call vs agent loop (calculation), calculation: Claude Code $0.014 (n 24); Codex CLI $0.0012 (n 16). Unclear.

Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI

Claude Code 1 · 6 ties · 1 unclear
  • Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON), calculation: Claude Code 100% (12/12) (n 12, 95% interval 76%–100%); Codex CLI 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
  • Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss)), calculation: Claude Code 100% (12/12) (n 12, 95% interval 76%–100%); Codex CLI 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
  • What each call produced: strict pass, format miss, wrong values or error (Strict pass): Claude Code 12 (n 12); Codex CLI 12 (n 12). Tie.
  • What each call produced: strict pass, format miss, wrong values or error (Format miss): Claude Code 0 (n 12); Codex CLI 0 (n 12). Tie.
  • What each call produced: strict pass, format miss, wrong values or error (Wrong values): Claude Code 0 (n 12); Codex CLI 0 (n 12). Tie.
  • What each call produced: strict pass, format miss, wrong values or error (Error): Claude Code 0 (n 12); Codex CLI 0 (n 12). Tie.
  • Time per call, instructions vs schema mode: Claude Code 3.52 s (n 12, run range 2.7 s–4.1 s); Codex CLI 6.21 s (n 12, run range 4.2 s–12.3 s). Claude Code ahead.
  • Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call): Claude Code 368 (n 12); Codex CLI 117 (n 12). Unclear.

How much of an AI bill is thinking? Reasoning tokens by model and effort

8 unclear
  • Reasoning share of output tokens per call on hard tasks (calculation), calculation: Claude Code 54.5% (n 24, run range 0%–96%); Codex CLI 46.3% (n 16, run range 11%–87%). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Claude Code $0.0067 (n 24); Codex CLI $0.0023 (n 16). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Claude Code $0.0037 (n 24); Codex CLI $0.0030 (n 16). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Claude Code $0.0040 (n 24); Codex CLI $0.020 (n 16). Unclear.
  • Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass), calculation: Claude Code $0.0059 (n 16); Codex CLI $0.0023 (n 16). Unclear.
  • Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass), calculation: Claude Code $0.014 (n 16); Codex CLI $0.026 (n 16). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Claude Code 54.5% (n 24, run range 0%–96%); Codex CLI 46.3% (n 16, run range 11%–87%). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Claude Code 0% (n 15, run range 0%–73%); Codex CLI 41% (n 15, run range 0%–71%). Unclear.

Where the seconds go: first text, output speed and prompt size for 6 LLMs

1 tie · 9 unclear
  • Time to first text: a 250-line answer, six models: Claude Code 1.96 s (n 4, run range 0.9 s–4.1 s); Codex CLI 3.52 s (n 4, run range 2.8 s–4.4 s). Unclear.
  • Output speed after the first text: visible tokens per second (calculation), calculation: Claude Code 232 (n 4, run range 230–233); Codex CLI 80 (n 4, run range 72–81). Unclear.
  • Output speed in characters per second after the first text (calculation), calculation: Claude Code 517 (n 4, run range 513–519); Codex CLI 323 (n 4, run range 291–327). Unclear.
  • Time to first text as the prompt grows: 1k, calculation: Claude Code 1.45 s (n 3, run range 1.2 s–1.7 s); Codex CLI 3.36 s (n 3, run range 3.4 s–4.8 s). Unclear.
  • Time to first text as the prompt grows: 16k, calculation: Claude Code 1.78 s (n 3, run range 1.6 s–2.1 s); Codex CLI 4.02 s (n 3, run range 3.3 s–4.3 s). Unclear.
  • Time to first text as the prompt grows: 64k, calculation: Claude Code 3.07 s (n 3, run range 1.4 s–3.6 s); Codex CLI 3.93 s (n 3, run range 3.4 s–4.4 s). Unclear.
  • Total time per call by prompt size (1k prompt): Claude Code 1.78 s (n 3, run range 1.6 s–2.1 s); Codex CLI 3.43 s (n 3, run range 3.4 s–4.9 s). Unclear.
  • Total time per call by prompt size (16k prompt): Claude Code 2.10 s (n 3, run range 2 s–2.5 s); Codex CLI 4.14 s (n 3, run range 4 s–4.7 s). Unclear.
  • Total time per call by prompt size (64k prompt): Claude Code 3.44 s (n 3, run range 1.7 s–4.4 s); Codex CLI 3.96 s (n 3, run range 3.5 s–4.4 s). Unclear.
  • Exact lookup answers at the 1k, 16k and 64k prompt-size targets: Claude Code 100% (9/9) (n 9, 95% interval 70%–100%); Codex CLI 100% (9/9) (n 9, 95% interval 70%–100%). Tie.

GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

Codex CLI 1 · 5 ties · 4 unclear
  • Pass rate on 4 harder tasks (Strict pass): Claude Code 38% (6/16) (n 16, 95% interval 18%–61%); Codex CLI 69% (11/16) (n 16, 95% interval 44%–86%). Tie.

    Calculation: at these rates, about 40 runs per side would separate them.

  • Pass rate on 4 harder tasks (Lenient (format misses counted)): Claude Code 38% (6/16) (n 16, 95% interval 18%–61%); Codex CLI 69% (11/16) (n 16, 95% interval 44%–86%). Tie.

    Calculation: at these rates, about 40 runs per side would separate them.

  • Calls that tried a tool although tools were off: Claude Code 31% (5/16) (n 16, 95% interval 14%–56%); Codex CLI 0% (0/16) (n 16, 95% interval 0%–19%). Unclear.
  • Strict pass rate by task: 10x10 nonogram: Claude Code 100% (4/4) (n 4, 95% interval 51%–100%); Codex CLI 75% (3/4) (n 4, 95% interval 30%–95%). Tie.

    Calculation: at these rates, about 27 runs per side would separate them.

  • Strict pass rate by task: Sudoku, 22 givens: Claude Code 0% (0/4) (n 4, 95% interval 0%–49%); Codex CLI 25% (1/4) (n 4, 95% interval 4.6%–70%). Tie.

    Calculation: at these rates, about 26 runs per side would separate them.

  • Strict pass rate by task: 6x6 Skyscrapers: Claude Code 0% (0/4) (n 4, 95% interval 0%–49%); Codex CLI 100% (4/4) (n 4, 95% interval 51%–100%). Codex CLI ahead.
  • Strict pass rate by task: Seeded shuffle output: Claude Code 50% (2/4) (n 4, 95% interval 15%–85%); Codex CLI 75% (3/4) (n 4, 95% interval 30%–95%). Tie.

    Calculation: at these rates, about 54 runs per side would separate them.

  • Total time per call on harder tasks: Claude Code 70.4 s (n 12, run range 4.3 s–210 s); Codex CLI 120.2 s (n 13, run range 46.2 s–273 s). Unclear.
  • Output tokens per call on harder tasks (Output tokens): Claude Code 9,287 (n 12, run range 407–27,921); Codex CLI 4,994 (n 13, run range 2,099–13,413). Unclear.
  • List-price cost per strict pass on harder tasks (calculation), calculation: Claude Code $0.24 (n 16); Codex CLI $0.083 (n 16). Unclear.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

88 rows from 12 studies. Claude Code ahead on 7, Codex CLI ahead on 1; 31 ties, 49 unclear. A side is ahead only where the intervals or ranges do not overlap.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Claude Code

  • List-price cost per call (calculation): $0.0036 vs $0.010. A list-price calculation, not a measured difference. Calculation
  • List-price cost per passing answer (calculation): $0.0062 vs $0.016. A list-price calculation, not a measured difference. Calculation
  • List-price cost per strict pass on hard tasks (calculation): $0.014 vs $0.026. A list-price calculation, not a measured difference. Calculation
  • Time per coding session: 23.1 s vs 113.4 s. The run ranges (fastest to slowest) do not overlap (Claude Code 18.7 s to 44.5 s; Codex CLI 78.5 s to 221.9 s). A range is not a confidence interval.
  • List-price cost per strict pass by effort (calculation): $0.014 vs $0.026. A list-price calculation, not a measured difference. Calculation
  • Same prompt, 10 times: time per call (Exact number): 6.89 s vs 13.4 s. The run ranges (fastest to slowest) do not overlap (Claude Code 5.81 s to 7.81 s; Codex CLI 12.3 s to 18.0 s). A range is not a confidence interval.
  • Same prompt, 10 times: time per call (Code fix): 2.67 s vs 11.3 s. The run ranges (fastest to slowest) do not overlap (Claude Code 2.32 s to 4.34 s; Codex CLI 9.08 s to 14.8 s). A range is not a confidence interval.
  • CLI start-up tax on a one-word answer (First model output): 1,461 ms vs 5,059 ms. The run ranges (fastest to slowest) do not overlap (Claude Code 1,206 ms to 2,308 ms; Codex CLI 4,391 ms to 5,478 ms). A range is not a confidence interval. Samples are small (5 runs per side).
  • CLI start-up tax on a one-word answer (Total wall time): 2,529 ms vs 5,999 ms. The run ranges (fastest to slowest) do not overlap (Claude Code 2,273 ms to 3,382 ms; Codex CLI 5,367 ms to 6,506 ms). A range is not a confidence interval. Samples are small (5 runs per side).
  • Strict pass rate: single call vs agent loop on eight hard tasks: 100% (24/24) vs 63% (10/16). The 95% intervals do not overlap (Claude Code 86% to 100%; Codex CLI 39% to 82%).
  • Time per call, instructions vs schema mode: 3.52 s vs 6.21 s. The run ranges (fastest to slowest) do not overlap (Claude Code 2.67 s to 4.12 s; Codex CLI 4.20 s to 12.3 s). A range is not a confidence interval.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)): $0.0040 vs $0.020. A list-price calculation, not a measured difference. Calculation
  • Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass): $0.014 vs $0.026. A list-price calculation, not a measured difference. Calculation
  • Time to first text as the prompt grows: 1k: 1.45 s vs 3.36 s. A list-price calculation, not a measured difference. Calculation
  • Time to first text as the prompt grows: 16k: 1.78 s vs 4.02 s. A list-price calculation, not a measured difference. Calculation

When to pick Codex CLI

  • List-price cost per strict pass: single call vs agent loop (calculation): $0.0012 vs $0.014. A list-price calculation, not a measured difference. Calculation
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0023 vs $0.0067. A list-price calculation, not a measured difference. Calculation
  • Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass): $0.0023 vs $0.0059. A list-price calculation, not a measured difference. Calculation
  • Strict pass rate by task: 6x6 Skyscrapers: 100% (4/4) vs 0% (0/4). The 95% intervals do not overlap (Claude Code 0% to 49%; Codex CLI 51% to 100%).
  • List-price cost per strict pass on harder tasks (calculation): $0.083 vs $0.24. A list-price calculation, not a measured difference. Calculation

Side by side

The study charts, showing only these two. Open a study for every configuration.

Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI

2 rows. Highest GPT-6.1 Sol (high) · Codex CLI 100% (95% interval 80%–100%, n 15). Lowest Claude Sonnet 5.5 · Claude Code 80% (95% interval 55%–93%, n 15). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 15 per row

Every call counts; failures and timeouts are non-passes

Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 4× real timeMotion reduced: press Replay to animateThe slowest median is 5.7 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 5.7 s (range 4.1 s–25.5 s, n 15). Fastest Claude Sonnet 5.5 · Claude Code 2.3 s (range 2.2 s–7.7 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 3.6× real timeMotion reduced: press Replay to animateThe slowest median is 5.1 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 5.1 s (range 3.4 s–17.8 s, n 15). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1 s–6.4 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

  • Cache read
  • Other input
Bar length is the total; segments are its parts.
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Totals are the sum of the parts shown. Shares are calculated from the same values.

2 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (medium) · Codex CLI 5,180 (n 15). Lowest Claude Sonnet 5.5 · Claude Code 1,401 (n 15). Other input: highest GPT-6.1 Sol (medium) · Codex CLI 6,943 (n 15). Lowest Claude Sonnet 5.5 · Claude Code 685 (n 15).

Notesn = 15 per row

Mean per call, split into prompt-cache reads and other input

The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.

Source: Provider head-to-head: Claude Code models vs Codex efforts

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Claude Code or Codex CLI?
Claude Code and Codex CLI share 66 measured metrics and 22 list-price calculations from 12 studies. Claude Code leads on 7 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 4 more. Codex CLI leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 31 ties and 49 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Each run pairs a CLI with a model, so these rows cannot separate the CLI from the model; the contexts name both. Some rows rest on small samples (n = 2 at the smallest).
How were Claude Code and Codex CLI measured?
They share 66 measured metrics and 22 list-price calculations from 12 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks; Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks; Prompt caching and run-to-run consistency in Claude Code and Codex CLI; Routing overhead: deterministic policy vs LLM routers vs Jev; Claude Code CLI vs Codex CLI vs the API: latency and tokens; Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks; Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI; How much of an AI bill is thinking? Reasoning tokens by model and effort; Where the seconds go: first text, output speed and prompt size for 6 LLMs; GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Every row names its configuration, its sample size and its interval or range.
How do Claude Code and Codex CLI compare on time per coding session?
Claude Code: 23.1 s (Claude Sonnet 5.5 · six small repository tasks with hidden tests; n = 12; run range 18.7 s to 44.5 s). Codex CLI: 113.4 s (GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests; n = 12; run range 78.5 s to 222 s). The run ranges (fastest to slowest) do not overlap (Claude Code 18.7 s to 44.5 s; Codex CLI 78.5 s to 221.9 s). A range is not a confidence interval.
How do Claude Code and Codex CLI compare on same prompt, 10 times: time per call (Exact number)?
Claude Code: 6.89 s (Claude Sonnet 5.5 · same prompt repeated 10 times; n = 10; run range 5.8 s to 7.8 s). Codex CLI: 13.4 s (GPT-6.1 Sol · effort medium · same prompt repeated 10 times; n = 10; run range 12.3 s to 18 s). The run ranges (fastest to slowest) do not overlap (Claude Code 5.81 s to 7.81 s; Codex CLI 12.3 s to 18.0 s). A range is not a confidence interval.
How do Claude Code and Codex CLI compare on same prompt, 10 times: time per call (Code fix)?
Claude Code: 2.67 s (Claude Sonnet 5.5 · same prompt repeated 10 times; n = 10; run range 2.3 s to 4.3 s). Codex CLI: 11.3 s (GPT-6.1 Sol · effort medium · same prompt repeated 10 times; n = 10; run range 9.1 s to 14.9 s). The run ranges (fastest to slowest) do not overlap (Claude Code 2.32 s to 4.34 s; Codex CLI 9.08 s to 14.8 s). A range is not a confidence interval.
How do Claude Code and Codex CLI compare on cLI start-up tax on a one-word answer (First model output)?
Claude Code: 1,461 ms (Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs; n = 5; run range 1.21 s to 2.31 s). Codex CLI: 5,059 ms (default model · CLI start-up, one-word prompt, 5 runs; n = 5; run range 4.39 s to 5.48 s). The run ranges (fastest to slowest) do not overlap (Claude Code 1,206 ms to 2,308 ms; Codex CLI 4,391 ms to 5,478 ms). A range is not a confidence interval. Samples are small (5 runs per side).
How do Claude Code and Codex CLI compare on cLI start-up tax on a one-word answer (Total wall time)?
Claude Code: 2,529 ms (Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs; n = 5; run range 2.27 s to 3.38 s). Codex CLI: 5,999 ms (default model · CLI start-up, one-word prompt, 5 runs; n = 5; run range 5.37 s to 6.51 s). The run ranges (fastest to slowest) do not overlap (Claude Code 2,273 ms to 3,382 ms; Codex CLI 5,367 ms to 6,506 ms). A range is not a confidence interval. Samples are small (5 runs per side).

The studies behind this page

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.