49 measured metrics · 21 calculated · 10 studies

Claude Sonnet 5.5vsGPT-6.1 Sol (Codex CLI)

Claude Sonnet 5.5 ahead on 4, GPT-6.1 Sol (Codex CLI) ahead on 1; 22 ties, 43 unclear. A side is ahead only where the intervals or ranges do not overlap.

The verdict

Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) share 49 measured metrics and 21 list-price calculations from 10 studies. Claude Sonnet 5.5 leads on 4 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 1 more. GPT-6.1 Sol (Codex CLI) leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 22 ties and 43 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Input tokens per call: what the CLI sends (Cache read), 3.7x (GPT-6.1 Sol (Codex CLI) larger).

Watch it build

A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.

Live story · 42 sClaude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI): what the measurements say

Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI): what the measurements say

70 comparison rows from 10 studies: 4 rows favour Sonnet 5.5, 1 favour GPT-6.1 Sol (Codex CLI), 65 are ties or unclear. Cost rows are calculations.

Transcript
  1. Comparison · 70 rows · 10 studies. Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI). A winner only where the 95% intervals or run ranges do not overlap.
  2. 70 comparison rows from 10 studies: Sonnet 5.5 ahead on 4, GPT-6.1 Sol (Codex CLI) ahead on 1. The rest do not separate them. Rows where Sonnet 5.5 is ahead: 4 (of 70). Rows where GPT-6.1 Sol (Codex CLI) is ahead: 1 (of 70). Ties or unclear: 65 (22 ties · 43 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
  3. Coding agents, hidden tests: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 4 rows separates them. Table: Coding agents, hidden tests · Claude Code vs Codex CLI · n = 12 per side. Source study: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks. Rows shown: Coding sessions that passed every hidden check; Time per coding session; Tool calls per coding session; List-price cost per passing coding session (calculation). Recorded settings: Claude Code · six small repository tasks with hidden tests; Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests. Includes a calculation, not a bill or a new run. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
  4. Caching sessions: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 2 of 9 rows separate them. Table: Caching sessions · 5 of 9 rows · Claude Code vs Codex CLI · n = 10 per side. Source study: Prompt caching and run-to-run consistency in Claude Code and Codex CLI. Rows shown: Same prompt, 10 times: strict pass rate (Exact number); Same prompt, 10 times: strict pass rate (JSON object); Same prompt, 10 times: strict pass rate (Code fix); Same prompt, 10 times: time per call (Exact number); Same prompt, 10 times: time per call (Code fix). Recorded settings: Claude Code · same prompt repeated 10 times; Codex CLI · effort medium · same prompt repeated 10 times. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  5. Instructions vs JSON schema: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 8 rows separates them. Table: Instructions vs JSON schema · 5 of 8 rows · Claude Code vs Codex CLI · n = 12 per side. Source study: Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI. Rows shown: Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON); What each call produced: strict pass, format miss, wrong values or error (Strict pass); What each call produced: strict pass, format miss, wrong values or error (Format miss); What each call produced: strict pass, format miss, wrong values or error (Wrong values); Time per call, instructions vs schema mode. Recorded settings: Claude Code · instructions; Codex CLI · effort low · instructions. Includes a calculation, not a bill or a new run. Caveat: The calls repeat only three fixed prompts. Wilson intervals describe call outcomes under a binomial assumption; they do not measure accuracy across unseen tasks. The paired p-values also assume independent pairs and do not remove this limit.
  6. Four harder tasks: pass rate 38% vs 69%, tie: 95% intervals overlap. 1 of 10 rows separates them. Table: Four harder tasks · 5 of 10 rows · Claude Code vs Codex CLI · n = 4–16 per side. Source study: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Rows shown: Pass rate on 4 harder tasks (Strict pass); Pass rate on 4 harder tasks (Lenient (format misses counted)); Calls that tried a tool although tools were off; Strict pass rate by task: 10x10 nonogram; Strict pass rate by task: 6x6 Skyscrapers. Recorded settings: Claude Code; Codex CLI · effort medium. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.
  7. No winner where the data shows none. Showing 19 of 70 rows; every row and its reason online.

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Claude Sonnet 5.5
  • GPT-6.1 Sol (Codex CLI)
  • 95% interval
  • fastest–slowest run (not an interval)
  • where the two overlap
  • hollow: list-price calculation

These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

1 tie · 7 unclear
  • Pass rate on five validated tasks: Claude Sonnet 5.5 80% (12/15) (n 15, 95% interval 55%–93%); GPT-6.1 Sol (Codex CLI) 100% (15/15) (n 15, 95% interval 80%–100%). Tie.

    Calculation: at these rates, about 33 runs per side would separate them.

  • Total time per call: Claude Sonnet 5.5 2.31 s (n 15, run range 2.2 s–7.7 s); GPT-6.1 Sol (Codex CLI) 5.65 s (n 15, run range 4.1 s–25.5 s). Unclear.
  • Time to first useful output: Claude Sonnet 5.5 1.56 s (n 15, run range 1 s–6.4 s); GPT-6.1 Sol (Codex CLI) 5.05 s (n 15, run range 3.4 s–17.8 s). Unclear.
  • Input tokens per call: what the CLI sends (Cache read): Claude Sonnet 5.5 1,401 (n 15); GPT-6.1 Sol (Codex CLI) 5,180 (n 15). Unclear.
  • Input tokens per call: what the CLI sends (Other input): Claude Sonnet 5.5 685 (n 15); GPT-6.1 Sol (Codex CLI) 6,943 (n 15). Unclear.
  • Output tokens per call (Output tokens): Claude Sonnet 5.5 107 (n 15); GPT-6.1 Sol (Codex CLI) 42 (n 15). Unclear.
  • List-price cost per call (calculation), calculation: Claude Sonnet 5.5 $0.0036 (n 15, run range $0.0034–$0.01); GPT-6.1 Sol (Codex CLI) $0.010 (n 15, run range $0.0054–$0.027). Unclear.
  • List-price cost per passing answer (calculation), calculation: Claude Sonnet 5.5 $0.0062 (n 15); GPT-6.1 Sol (Codex CLI) $0.016 (n 15). Unclear.

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

2 ties · 4 unclear
  • Pass rate on eight hard tasks (Strict pass): Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Pass rate on eight hard tasks (Lenient (format misses counted)): Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Total time per call on hard tasks (separate batches): Claude Sonnet 5.5 7.75 s (n 24, run range 2.3 s–34.8 s); GPT-6.1 Sol (Codex CLI) 13.1 s (n 16, run range 8.5 s–61.6 s). Unclear.
  • Time to first useful output on hard tasks: Claude Sonnet 5.5 5.95 s (n 24, run range 0.9 s–30.6 s); GPT-6.1 Sol (Codex CLI) 10.2 s (n 16, run range 6.1 s–40.4 s). Unclear.
  • Output tokens per call on hard tasks (Output tokens): Claude Sonnet 5.5 1,050 (n 24); GPT-6.1 Sol (Codex CLI) 335 (n 16). Unclear.
  • List-price cost per strict pass on hard tasks (calculation), calculation: Claude Sonnet 5.5 $0.014 (n 24); GPT-6.1 Sol (Codex CLI) $0.026 (n 16). Unclear.

Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks

Claude Sonnet 5.5 1 · 1 tie · 2 unclear
  • Coding sessions that passed every hidden check: Claude Sonnet 5.5 100% (12/12) (n 12, 95% interval 76%–100%); GPT-6.1 Sol (Codex CLI) 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
  • Time per coding session: Claude Sonnet 5.5 23.1 s (n 12, run range 18.7 s–44.5 s); GPT-6.1 Sol (Codex CLI) 113.4 s (n 12, run range 78.5 s–222 s). Claude Sonnet 5.5 ahead.
  • Tool calls per coding session: Claude Sonnet 5.5 7.5 (n 12, run range 3–14); GPT-6.1 Sol (Codex CLI) 12.5 (n 12, run range 8–18). Unclear.
  • List-price cost per passing coding session (calculation), calculation: Claude Sonnet 5.5 $0.085 (n 12); GPT-6.1 Sol (Codex CLI) $0.098 (n 12). Unclear.

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks

1 tie · 3 unclear
  • Strict pass rate by effort on eight hard tasks: Claude Sonnet 5.5 100% (16/16) (n 16, 95% interval 81%–100%); GPT-6.1 Sol (Codex CLI) 100% (16/16) (n 16, 95% interval 81%–100%). Tie.
  • Total time per call by effort on hard tasks: Claude Sonnet 5.5 7.63 s (n 16, run range 2.7 s–24 s); GPT-6.1 Sol (Codex CLI) 13.1 s (n 16, run range 8.5 s–61.6 s). Unclear.
  • Output tokens per call by effort on hard tasks (Output tokens): Claude Sonnet 5.5 770 (n 16); GPT-6.1 Sol (Codex CLI) 335 (n 16). Unclear.
  • List-price cost per strict pass by effort (calculation), calculation: Claude Sonnet 5.5 $0.014 (n 16); GPT-6.1 Sol (Codex CLI) $0.026 (n 16). Unclear.

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

Claude Sonnet 5.5 2 · 5 ties · 2 unclear
  • Same prompt, 10 times: strict pass rate (Exact number): Claude Sonnet 5.5 100% (10/10) (n 10, 95% interval 72%–100%); GPT-6.1 Sol (Codex CLI) 100% (10/10) (n 10, 95% interval 72%–100%). Tie.
  • Same prompt, 10 times: strict pass rate (JSON object): Claude Sonnet 5.5 100% (10/10) (n 10, 95% interval 72%–100%); GPT-6.1 Sol (Codex CLI) 100% (10/10) (n 10, 95% interval 72%–100%). Tie.
  • Same prompt, 10 times: strict pass rate (Code fix): Claude Sonnet 5.5 100% (10/10) (n 10, 95% interval 72%–100%); GPT-6.1 Sol (Codex CLI) 100% (10/10) (n 10, 95% interval 72%–100%). Tie.
  • Same prompt, 10 times: how many different answers (Exact number): Claude Sonnet 5.5 1 (n 10); GPT-6.1 Sol (Codex CLI) 1 (n 10). Tie.
  • Same prompt, 10 times: how many different answers (JSON object): Claude Sonnet 5.5 1 (n 10); GPT-6.1 Sol (Codex CLI) 1 (n 10). Tie.
  • Same prompt, 10 times: how many different answers (Code fix): Claude Sonnet 5.5 3 (n 10); GPT-6.1 Sol (Codex CLI) 6 (n 10). Unclear.
  • Same prompt, 10 times: time per call (Exact number): Claude Sonnet 5.5 6.89 s (n 10, run range 5.8 s–7.8 s); GPT-6.1 Sol (Codex CLI) 13.4 s (n 10, run range 12.3 s–18 s). Claude Sonnet 5.5 ahead.
  • Same prompt, 10 times: time per call (JSON object): Claude Sonnet 5.5 2.89 s (n 10, run range 2.7 s–5.3 s); GPT-6.1 Sol (Codex CLI) 6.42 s (n 10, run range 5.3 s–8.3 s). Unclear.
  • Same prompt, 10 times: time per call (Code fix): Claude Sonnet 5.5 2.67 s (n 10, run range 2.3 s–4.3 s); GPT-6.1 Sol (Codex CLI) 11.3 s (n 10, run range 9.1 s–14.9 s). Claude Sonnet 5.5 ahead.

Claude Code CLI vs Codex CLI vs the API: latency and tokens

3 unclear
  • Repairing a scheduler: Claude Code vs Codex vs API (Total time): Claude Sonnet 5.5 15.0 s (n 3, run range 13.9 s–15.9 s); GPT-6.1 Sol (Codex CLI) 61.2 s (n 3, run range 59.9 s–69.5 s). Unclear.
  • Repairing a scheduler: Claude Code vs Codex vs API (First useful output): Claude Sonnet 5.5 7.55 s (n 3, run range 6.8 s–7.6 s); GPT-6.1 Sol (Codex CLI) 15.6 s (n 3, run range 13.7 s–23 s). Unclear.
  • Output tokens to repair the scheduler (Output tokens): Claude Sonnet 5.5 2,227 (n 3); GPT-6.1 Sol (Codex CLI) 1,181 (n 3). Unclear.

Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI

Claude Sonnet 5.5 1 · 6 ties · 1 unclear
  • Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON), calculation: Claude Sonnet 5.5 100% (12/12) (n 12, 95% interval 76%–100%); GPT-6.1 Sol (Codex CLI) 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
  • Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss)), calculation: Claude Sonnet 5.5 100% (12/12) (n 12, 95% interval 76%–100%); GPT-6.1 Sol (Codex CLI) 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
  • What each call produced: strict pass, format miss, wrong values or error (Strict pass): Claude Sonnet 5.5 12 (n 12); GPT-6.1 Sol (Codex CLI) 12 (n 12). Tie.
  • What each call produced: strict pass, format miss, wrong values or error (Format miss): Claude Sonnet 5.5 0 (n 12); GPT-6.1 Sol (Codex CLI) 0 (n 12). Tie.
  • What each call produced: strict pass, format miss, wrong values or error (Wrong values): Claude Sonnet 5.5 0 (n 12); GPT-6.1 Sol (Codex CLI) 0 (n 12). Tie.
  • What each call produced: strict pass, format miss, wrong values or error (Error): Claude Sonnet 5.5 0 (n 12); GPT-6.1 Sol (Codex CLI) 0 (n 12). Tie.
  • Time per call, instructions vs schema mode: Claude Sonnet 5.5 3.52 s (n 12, run range 2.7 s–4.1 s); GPT-6.1 Sol (Codex CLI) 6.21 s (n 12, run range 4.2 s–12.3 s). Claude Sonnet 5.5 ahead.
  • Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call): Claude Sonnet 5.5 368 (n 12); GPT-6.1 Sol (Codex CLI) 117 (n 12). Unclear.

How much of an AI bill is thinking? Reasoning tokens by model and effort

8 unclear
  • Reasoning share of output tokens per call on hard tasks (calculation), calculation: Claude Sonnet 5.5 54.5% (n 24, run range 0%–96%); GPT-6.1 Sol (Codex CLI) 46.3% (n 16, run range 11%–87%). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Claude Sonnet 5.5 $0.0067 (n 24); GPT-6.1 Sol (Codex CLI) $0.0023 (n 16). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Claude Sonnet 5.5 $0.0037 (n 24); GPT-6.1 Sol (Codex CLI) $0.0030 (n 16). Unclear.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Claude Sonnet 5.5 $0.0040 (n 24); GPT-6.1 Sol (Codex CLI) $0.020 (n 16). Unclear.
  • Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass), calculation: Claude Sonnet 5.5 $0.0059 (n 16); GPT-6.1 Sol (Codex CLI) $0.0023 (n 16). Unclear.
  • Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass), calculation: Claude Sonnet 5.5 $0.014 (n 16); GPT-6.1 Sol (Codex CLI) $0.026 (n 16). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Claude Sonnet 5.5 54.5% (n 24, run range 0%–96%); GPT-6.1 Sol (Codex CLI) 46.3% (n 16, run range 11%–87%). Unclear.
  • Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Claude Sonnet 5.5 0% (n 15, run range 0%–73%); GPT-6.1 Sol (Codex CLI) 41% (n 15, run range 0%–71%). Unclear.

Where the seconds go: first text, output speed and prompt size for 6 LLMs

1 tie · 9 unclear
  • Time to first text: a 250-line answer, six models: Claude Sonnet 5.5 1.96 s (n 4, run range 0.9 s–4.1 s); GPT-6.1 Sol (Codex CLI) 3.52 s (n 4, run range 2.8 s–4.4 s). Unclear.
  • Output speed after the first text: visible tokens per second (calculation), calculation: Claude Sonnet 5.5 232 (n 4, run range 230–233); GPT-6.1 Sol (Codex CLI) 80 (n 4, run range 72–81). Unclear.
  • Output speed in characters per second after the first text (calculation), calculation: Claude Sonnet 5.5 517 (n 4, run range 513–519); GPT-6.1 Sol (Codex CLI) 323 (n 4, run range 291–327). Unclear.
  • Time to first text as the prompt grows: 1k, calculation: Claude Sonnet 5.5 1.45 s (n 3, run range 1.2 s–1.7 s); GPT-6.1 Sol (Codex CLI) 3.36 s (n 3, run range 3.4 s–4.8 s). Unclear.
  • Time to first text as the prompt grows: 16k, calculation: Claude Sonnet 5.5 1.78 s (n 3, run range 1.6 s–2.1 s); GPT-6.1 Sol (Codex CLI) 4.02 s (n 3, run range 3.3 s–4.3 s). Unclear.
  • Time to first text as the prompt grows: 64k, calculation: Claude Sonnet 5.5 3.07 s (n 3, run range 1.4 s–3.6 s); GPT-6.1 Sol (Codex CLI) 3.93 s (n 3, run range 3.4 s–4.4 s). Unclear.
  • Total time per call by prompt size (1k prompt): Claude Sonnet 5.5 1.78 s (n 3, run range 1.6 s–2.1 s); GPT-6.1 Sol (Codex CLI) 3.43 s (n 3, run range 3.4 s–4.9 s). Unclear.
  • Total time per call by prompt size (16k prompt): Claude Sonnet 5.5 2.10 s (n 3, run range 2 s–2.5 s); GPT-6.1 Sol (Codex CLI) 4.14 s (n 3, run range 4 s–4.7 s). Unclear.
  • Total time per call by prompt size (64k prompt): Claude Sonnet 5.5 3.44 s (n 3, run range 1.7 s–4.4 s); GPT-6.1 Sol (Codex CLI) 3.96 s (n 3, run range 3.5 s–4.4 s). Unclear.
  • Exact lookup answers at the 1k, 16k and 64k prompt-size targets: Claude Sonnet 5.5 100% (9/9) (n 9, 95% interval 70%–100%); GPT-6.1 Sol (Codex CLI) 100% (9/9) (n 9, 95% interval 70%–100%). Tie.

GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

GPT-6.1 Sol (Codex CLI) 1 · 5 ties · 4 unclear
  • Pass rate on 4 harder tasks (Strict pass): Claude Sonnet 5.5 38% (6/16) (n 16, 95% interval 18%–61%); GPT-6.1 Sol (Codex CLI) 69% (11/16) (n 16, 95% interval 44%–86%). Tie.

    Calculation: at these rates, about 40 runs per side would separate them.

  • Pass rate on 4 harder tasks (Lenient (format misses counted)): Claude Sonnet 5.5 38% (6/16) (n 16, 95% interval 18%–61%); GPT-6.1 Sol (Codex CLI) 69% (11/16) (n 16, 95% interval 44%–86%). Tie.

    Calculation: at these rates, about 40 runs per side would separate them.

  • Calls that tried a tool although tools were off: Claude Sonnet 5.5 31% (5/16) (n 16, 95% interval 14%–56%); GPT-6.1 Sol (Codex CLI) 0% (0/16) (n 16, 95% interval 0%–19%). Unclear.
  • Strict pass rate by task: 10x10 nonogram: Claude Sonnet 5.5 100% (4/4) (n 4, 95% interval 51%–100%); GPT-6.1 Sol (Codex CLI) 75% (3/4) (n 4, 95% interval 30%–95%). Tie.

    Calculation: at these rates, about 27 runs per side would separate them.

  • Strict pass rate by task: Sudoku, 22 givens: Claude Sonnet 5.5 0% (0/4) (n 4, 95% interval 0%–49%); GPT-6.1 Sol (Codex CLI) 25% (1/4) (n 4, 95% interval 4.6%–70%). Tie.

    Calculation: at these rates, about 26 runs per side would separate them.

  • Strict pass rate by task: 6x6 Skyscrapers: Claude Sonnet 5.5 0% (0/4) (n 4, 95% interval 0%–49%); GPT-6.1 Sol (Codex CLI) 100% (4/4) (n 4, 95% interval 51%–100%). GPT-6.1 Sol (Codex CLI) ahead.
  • Strict pass rate by task: Seeded shuffle output: Claude Sonnet 5.5 50% (2/4) (n 4, 95% interval 15%–85%); GPT-6.1 Sol (Codex CLI) 75% (3/4) (n 4, 95% interval 30%–95%). Tie.

    Calculation: at these rates, about 54 runs per side would separate them.

  • Total time per call on harder tasks: Claude Sonnet 5.5 70.4 s (n 12, run range 4.3 s–210 s); GPT-6.1 Sol (Codex CLI) 120.2 s (n 13, run range 46.2 s–273 s). Unclear.
  • Output tokens per call on harder tasks (Output tokens): Claude Sonnet 5.5 9,287 (n 12, run range 407–27,921); GPT-6.1 Sol (Codex CLI) 4,994 (n 13, run range 2,099–13,413). Unclear.
  • List-price cost per strict pass on harder tasks (calculation), calculation: Claude Sonnet 5.5 $0.24 (n 16); GPT-6.1 Sol (Codex CLI) $0.083 (n 16). Unclear.

Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

70 rows from 10 studies. Claude Sonnet 5.5 ahead on 4, GPT-6.1 Sol (Codex CLI) ahead on 1; 22 ties, 43 unclear. A side is ahead only where the intervals or ranges do not overlap.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Claude Sonnet 5.5

  • List-price cost per call (calculation): $0.0036 vs $0.010. A list-price calculation, not a measured difference. Calculation
  • List-price cost per passing answer (calculation): $0.0062 vs $0.016. A list-price calculation, not a measured difference. Calculation
  • List-price cost per strict pass on hard tasks (calculation): $0.014 vs $0.026. A list-price calculation, not a measured difference. Calculation
  • Time per coding session: 23.1 s vs 113.4 s. The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 18.7 s to 44.5 s; GPT-6.1 Sol (Codex CLI) 78.5 s to 221.9 s). A range is not a confidence interval.
  • List-price cost per strict pass by effort (calculation): $0.014 vs $0.026. A list-price calculation, not a measured difference. Calculation
  • Same prompt, 10 times: time per call (Exact number): 6.89 s vs 13.4 s. The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 5.81 s to 7.81 s; GPT-6.1 Sol (Codex CLI) 12.3 s to 18.0 s). A range is not a confidence interval.
  • Same prompt, 10 times: time per call (Code fix): 2.67 s vs 11.3 s. The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 2.32 s to 4.34 s; GPT-6.1 Sol (Codex CLI) 9.08 s to 14.8 s). A range is not a confidence interval.
  • Time per call, instructions vs schema mode: 3.52 s vs 6.21 s. The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 2.67 s to 4.12 s; GPT-6.1 Sol (Codex CLI) 4.20 s to 12.3 s). A range is not a confidence interval.
  • List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)): $0.0040 vs $0.020. A list-price calculation, not a measured difference. Calculation
  • Reasoning cost per strict pass by effort, with the total (calculation) (Total cost per strict pass): $0.014 vs $0.026. A list-price calculation, not a measured difference. Calculation
  • Time to first text as the prompt grows: 1k: 1.45 s vs 3.36 s. A list-price calculation, not a measured difference. Calculation
  • Time to first text as the prompt grows: 16k: 1.78 s vs 4.02 s. A list-price calculation, not a measured difference. Calculation

When to pick GPT-6.1 Sol (Codex CLI)

  • List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0023 vs $0.0067. A list-price calculation, not a measured difference. Calculation
  • Reasoning cost per strict pass by effort, with the total (calculation) (Reasoning cost per strict pass): $0.0023 vs $0.0059. A list-price calculation, not a measured difference. Calculation
  • Strict pass rate by task: 6x6 Skyscrapers: 100% (4/4) vs 0% (0/4). The 95% intervals do not overlap (Claude Sonnet 5.5 0% to 49%; GPT-6.1 Sol (Codex CLI) 51% to 100%).
  • List-price cost per strict pass on harder tasks (calculation): $0.083 vs $0.24. A list-price calculation, not a measured difference. Calculation

Side by side

The study charts, showing only these two. Open a study for every configuration.

Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI

2 rows. Highest GPT-6.1 Sol (high) · Codex CLI 100% (95% interval 80%–100%, n 15). Lowest Claude Sonnet 5.5 · Claude Code 80% (95% interval 55%–93%, n 15). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 15 per row

Every call counts; failures and timeouts are non-passes

Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 4× real timeMotion reduced: press Replay to animateThe slowest median is 5.7 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 5.7 s (range 4.1 s–25.5 s, n 15). Fastest Claude Sonnet 5.5 · Claude Code 2.3 s (range 2.2 s–7.7 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Entrance: medians race at 3.6× real timeMotion reduced: press Replay to animateThe slowest median is 5.1 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

2 rows. Slowest GPT-6.1 Sol (medium) · Codex CLI 5.1 s (range 3.4 s–17.8 s, n 15). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1 s–6.4 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

  • Cache read
  • Other input
Bar length is the total; segments are its parts.
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Totals are the sum of the parts shown. Shares are calculated from the same values.

2 rows, 2 series: Cache read, Other input. Cache read: highest GPT-6.1 Sol (medium) · Codex CLI 5,180 (n 15). Lowest Claude Sonnet 5.5 · Claude Code 1,401 (n 15). Other input: highest GPT-6.1 Sol (medium) · Codex CLI 6,943 (n 15). Lowest Claude Sonnet 5.5 · Claude Code 685 (n 15).

Notesn = 15 per row

Mean per call, split into prompt-cache reads and other input

The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.

Source: Provider head-to-head: Claude Code models vs Codex efforts

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Claude Sonnet 5.5 or GPT-6.1 Sol (Codex CLI)?
Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) share 49 measured metrics and 21 list-price calculations from 10 studies. Claude Sonnet 5.5 leads on 4 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 1 more. GPT-6.1 Sol (Codex CLI) leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 22 ties and 43 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest).
How were Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) measured?
They share 49 measured metrics and 21 list-price calculations from 10 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks; Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks; Prompt caching and run-to-run consistency in Claude Code and Codex CLI; Claude Code CLI vs Codex CLI vs the API: latency and tokens; Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI; How much of an AI bill is thinking? Reasoning tokens by model and effort; Where the seconds go: first text, output speed and prompt size for 6 LLMs; GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Every row names its configuration, its sample size and its interval or range.
How do Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) compare on time per coding session?
Claude Sonnet 5.5: 23.1 s (Claude Code · six small repository tasks with hidden tests; n = 12; run range 18.7 s to 44.5 s). GPT-6.1 Sol (Codex CLI): 113.4 s (Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests; n = 12; run range 78.5 s to 222 s). The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 18.7 s to 44.5 s; GPT-6.1 Sol (Codex CLI) 78.5 s to 221.9 s). A range is not a confidence interval.
How do Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) compare on same prompt, 10 times: time per call (Exact number)?
Claude Sonnet 5.5: 6.89 s (Claude Code · same prompt repeated 10 times; n = 10; run range 5.8 s to 7.8 s). GPT-6.1 Sol (Codex CLI): 13.4 s (Codex CLI · effort medium · same prompt repeated 10 times; n = 10; run range 12.3 s to 18 s). The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 5.81 s to 7.81 s; GPT-6.1 Sol (Codex CLI) 12.3 s to 18.0 s). A range is not a confidence interval.
How do Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) compare on same prompt, 10 times: time per call (Code fix)?
Claude Sonnet 5.5: 2.67 s (Claude Code · same prompt repeated 10 times; n = 10; run range 2.3 s to 4.3 s). GPT-6.1 Sol (Codex CLI): 11.3 s (Codex CLI · effort medium · same prompt repeated 10 times; n = 10; run range 9.1 s to 14.9 s). The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 2.32 s to 4.34 s; GPT-6.1 Sol (Codex CLI) 9.08 s to 14.8 s). A range is not a confidence interval.
How do Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) compare on time per call, instructions vs schema mode?
Claude Sonnet 5.5: 3.52 s (Claude Code · instructions; n = 12; run range 2.7 s to 4.1 s). GPT-6.1 Sol (Codex CLI): 6.21 s (Codex CLI · effort low · instructions; n = 12; run range 4.2 s to 12.3 s). The run ranges (fastest to slowest) do not overlap (Claude Sonnet 5.5 2.67 s to 4.12 s; GPT-6.1 Sol (Codex CLI) 4.20 s to 12.3 s). A range is not a confidence interval.
How do Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) compare on strict pass rate by task: 6x6 Skyscrapers?
Claude Sonnet 5.5: 0% (0/4) (Claude Code; n = 4; 95% interval 0% to 49%). GPT-6.1 Sol (Codex CLI): 100% (4/4) (Codex CLI · effort medium; n = 4; 95% interval 51% to 100%). The 95% intervals do not overlap (Claude Sonnet 5.5 0% to 49%; GPT-6.1 Sol (Codex CLI) 51% to 100%).

The studies behind this page

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.