189 comparisons · 1,571 rows · updated October 7, 2026

Where AI models really differ. Every gap, counted.

The answer

Of 1,571 comparison rows, 64 show a gap the intervals or ranges support. The rest are ties or unclear.

Only 578 rows have an interval or range on both sides, so only those can name a side. 64 of them do: 11% (a calculation).

Every comparison row, by verdict

Each row of every comparison page is one segment of the bar. The data names a side ahead only where the intervals or ranges do not overlap.

64

of 1,571 comparison rows

show a side ahead

One side ahead on 64; 595 ties, 912 unclear. A side is ahead only where the intervals or ranges do not overlap.

Rows by verdict: a side ahead 64, tie 595, unclear 912.

Every gap, by kind

Listed by kind (quality, speed, cost), then by study, then by metric. This is not a ranking.

Quality: 25 gaps

Quality: 25 of 327 rate rows show a gap. They come from 6 studies and 7 pairs.

  • side that is ahead (its own color)
  • side behind
  • 95% interval

Pass rate on eight hard tasks (Lenient (format misses counted))

95% interval · n = 24 per side · Higher is better · Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (24/24) against 67% (16/24) (higher is better; n = 24 per side; 95% interval).
  • Claude Opus 5.5 is ahead of Claude Haiku 4.5: 100% (24/24) against 67% (16/24) (higher is better; n = 24 per side; 95% interval).
  • Claude Fable 5.1 is ahead of Claude Haiku 4.5: 100% (24/24) against 67% (16/24) (higher is better; n = 24 per side; 95% interval).

Pass rate on eight hard tasks (Strict pass)

95% interval · n = 16 to 24 per side · Higher is better · Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (24/24) against 46% (11/24) (higher is better; n = 24 per side; 95% interval).
  • Claude Opus 5.5 is ahead of Claude Haiku 4.5: 100% (24/24) against 46% (11/24) (higher is better; n = 24 per side; 95% interval).
  • GPT-6.1 Sol (Codex CLI) is ahead of Claude Haiku 4.5: 100% (16/16) against 46% (11/24) (higher is better; n = 16 to 24 per side; 95% interval).
  • Claude Fable 5.1 is ahead of Claude Haiku 4.5: 100% (24/24) against 46% (11/24) (higher is better; n = 24 per side; 95% interval).

Same prompt, 10 times: strict pass rate (Exact number)

95% interval · n = 10 per side · Higher is better · Prompt caching and run-to-run consistency in Claude Code and Codex CLI

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (10/10) against 0% (0/10) (higher is better; n = 10 per side; 95% interval).
  • GPT-6.1 Sol (Codex CLI) is ahead of Claude Haiku 4.5: 100% (10/10) against 0% (0/10) (higher is better; n = 10 per side; 95% interval).

Same prompt, 10 times: strict pass rate (JSON object)

95% interval · n = 10 per side · Higher is better · Prompt caching and run-to-run consistency in Claude Code and Codex CLI

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (10/10) against 10% (1/10) (higher is better; n = 10 per side; 95% interval).
  • GPT-6.1 Sol (Codex CLI) is ahead of Claude Haiku 4.5: 100% (10/10) against 10% (1/10) (higher is better; n = 10 per side; 95% interval).

A stale README command: who still ran it?: Raw notes, 60 lines

95% interval · n = 10 to 15 per side · Lower is better · Does memory help Claude Code? 8 kinds of agent memory, tested

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 0% (0/15) against 100% (10/10) (lower is better; n = 10 to 15 per side; 95% interval).

Full pass rate by kind of memory: Handbook, 210 lines

95% interval · n = 10 to 15 per side · Higher is better · Does memory help Claude Code? 8 kinds of agent memory, tested

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (15/15) against 30% (3/10) (higher is better; n = 10 to 15 per side; 95% interval).

Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md

95% interval · n = 10 to 15 per side · Higher is better · Does memory help Claude Code? 8 kinds of agent memory, tested

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 67% (10/15) against 10% (1/10) (higher is better; n = 10 to 15 per side; 95% interval).

Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines

95% interval · n = 10 to 15 per side · Higher is better · Does memory help Claude Code? 8 kinds of agent memory, tested

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (15/15) against 30% (3/10) (higher is better; n = 10 to 15 per side; 95% interval).

Strict pass rate: single call vs agent loop on eight hard tasks

95% interval · n = 16 to 24 per side · Higher is better · Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (24/24) against 46% (11/24) (higher is better; n = 24 per side; 95% interval).
  • Claude Code is ahead of Codex CLI: 100% (24/24) against 63% (10/16) (higher is better; n = 16 to 24 per side; 95% interval).
  • Claude Sonnet 5.5 is ahead of GPT-6 Luna (Codex CLI): 100% (24/24) against 63% (10/16) (higher is better; n = 16 to 24 per side; 95% interval).

Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)

95% interval · n = 12 to 24 per side · Higher is better · Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (12/12) against 0% (0/24) (higher is better; n = 12 to 24 per side; 95% interval). This row is a calculation.
  • GPT-6.1 Sol (Codex CLI) is ahead of Claude Haiku 4.5: 100% (12/12) against 0% (0/24) (higher is better; n = 12 to 24 per side; 95% interval). This row is a calculation.

Pass rate on 4 harder tasks (Lenient (format misses counted))

95% interval · n = 12 to 16 per side · Higher is better · GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

  • Claude Opus 5.5 is ahead of Claude Haiku 4.5: 50% (6/12) against 0% (0/12) (higher is better; n = 12 per side; 95% interval).
  • GPT-6.1 Sol (Codex CLI) is ahead of Claude Haiku 4.5: 69% (11/16) against 0% (0/12) (higher is better; n = 12 to 16 per side; 95% interval).

Pass rate on 4 harder tasks (Strict pass)

95% interval · n = 12 to 16 per side · Higher is better · GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

  • GPT-6.1 Sol (Codex CLI) is ahead of Claude Haiku 4.5: 69% (11/16) against 0% (0/12) (higher is better; n = 12 to 16 per side; 95% interval).

Strict pass rate by task: 6x6 Skyscrapers

95% interval · n = 4 per side · Higher is better · GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks

  • Codex CLI is ahead of Claude Code: 100% (4/4) against 0% (0/4) (higher is better; n = 4 per side; 95% interval).
  • GPT-6.1 Sol (Codex CLI) is ahead of Claude Sonnet 5.5: 100% (4/4) against 0% (0/4) (higher is better; n = 4 per side; 95% interval).

Only a 95% interval is a confidence interval. A run range and a median-to-p95 band are not.

25 rows in 13 metrics. Listed by kind (quality, speed, cost), then by study, then by metric. This is not a ranking.

Speed: 39 gaps

Speed: 39 of 209 time rows show a gap. They come from 9 studies and 14 pairs.

  • side that is ahead (its own color)
  • side behind
  • fastest–slowest run (not an interval)
  • median to p95 (not an interval)

Time per coding session

run range · n = 12 per side · Lower is better · Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks

  • Claude Code is ahead of Codex CLI: 23.1 s against 113.4 s (lower is better; n = 12 per side; run range).
  • Claude Sonnet 5.5 is ahead of GPT-6.1 Sol (Codex CLI): 23.1 s against 113.4 s (lower is better; n = 12 per side; run range).

Same prompt, 10 times: time per call (Code fix)

run range · n = 10 per side · Lower is better · Prompt caching and run-to-run consistency in Claude Code and Codex CLI

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 2.67 s against 5.95 s (lower is better; n = 10 per side; run range).
  • Claude Code is ahead of Codex CLI: 2.67 s against 11.3 s (lower is better; n = 10 per side; run range).
  • Claude Sonnet 5.5 is ahead of GPT-6.1 Sol (Codex CLI): 2.67 s against 11.3 s (lower is better; n = 10 per side; run range).
  • Claude Haiku 4.5 is ahead of GPT-6.1 Sol (Codex CLI): 5.95 s against 11.3 s (lower is better; n = 10 per side; run range).

Same prompt, 10 times: time per call (Exact number)

run range · n = 10 per side · Lower is better · Prompt caching and run-to-run consistency in Claude Code and Codex CLI

  • Claude Code is ahead of Codex CLI: 6.89 s against 13.4 s (lower is better; n = 10 per side; run range).
  • Claude Sonnet 5.5 is ahead of GPT-6.1 Sol (Codex CLI): 6.89 s against 13.4 s (lower is better; n = 10 per side; run range).
  • Claude Haiku 4.5 is ahead of GPT-6.1 Sol (Codex CLI): 5.06 s against 13.4 s (lower is better; n = 10 per side; run range).

Time per routing decision (Model time (API))

median to p95 · n = 82 per side · Lower is better · Jev vs Claude as a router: accuracy and cost

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 1,599 ms against 10,734 ms (lower is better; n = 82 per side; median to p95).

Time per routing decision (Wall time (CLI))

median to p95 · n = 82 per side · Lower is better · Jev vs Claude as a router: accuracy and cost

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 2,598 ms against 12,674 ms (lower is better; n = 82 per side; median to p95).

CLI start-up tax on a one-word answer (First model output)

run range · n = 5 per side · Lower is better · Routing overhead: deterministic policy vs LLM routers vs Jev

  • Claude Code is ahead of Codex CLI: 1,461 ms against 5,059 ms (lower is better; n = 5 per side; run range).

CLI start-up tax on a one-word answer (Total wall time)

run range · n = 5 per side · Lower is better · Routing overhead: deterministic policy vs LLM routers vs Jev

  • Claude Code is ahead of Codex CLI: 2,529 ms against 5,999 ms (lower is better; n = 5 per side; run range).

Time to make one routing decision

median to p95 · n = 82 to 20,000 per side · Lower is better · Routing overhead: deterministic policy vs LLM routers vs Jev

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 2,597 ms against 12,543 ms (lower is better; n = 82 per side; median to p95).
  • Jev 1.13 is ahead of Claude Haiku 4.5: 137 ms against 12,543 ms (lower is better; n = 82 to 246 per side; median to p95).
  • Jev 1.13 is ahead of Claude Sonnet 5.5: 137 ms against 2,597 ms (lower is better; n = 82 to 246 per side; median to p95).
  • Deterministic routing policy is ahead of Jev 1.13: 1.42 µs against 137 ms (lower is better; n = 246 to 20,000 per side; median to p95).
  • Deterministic routing policy is ahead of Claude Sonnet 5.5: 1.42 µs against 2,597 ms (lower is better; n = 82 to 20,000 per side; median to p95).
  • Deterministic routing policy is ahead of Claude Haiku 4.5: 1.42 µs against 12,543 ms (lower is better; n = 82 to 20,000 per side; median to p95).

Where an LLM router’s time goes: model vs CLI (CLI and harness time)

median to p95 · n = 82 per side · Lower is better · Routing overhead: deterministic policy vs LLM routers vs Jev

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 973 ms against 1,698 ms (lower is better; n = 82 per side; median to p95).

Where an LLM router’s time goes: model vs CLI (Model API time)

median to p95 · n = 82 per side · Lower is better · Routing overhead: deterministic policy vs LLM routers vs Jev

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 1,596 ms against 10,508 ms (lower is better; n = 82 per side; median to p95).

CLI vs API: time for a one-line answer (First useful output)

run range · n = 5 per side · Lower is better · Claude Code CLI vs Codex CLI vs the API: latency and tokens

  • GPT-6.1 Sol (OpenAI API) is ahead of GPT-6.1 Sol (Codex CLI): 1.34 s against 3.79 s (lower is better; n = 5 per side; run range).
  • GPT-6 Luna (OpenAI API) is ahead of GPT-6.1 Sol (Codex CLI): 0.82 s against 3.79 s (lower is better; n = 5 per side; run range).
  • GPT-6.1 Sol (OpenAI API) is ahead of GPT-6 Luna (Codex CLI): 1.34 s against 2.79 s (lower is better; n = 5 per side; run range).
  • GPT-6 Luna (OpenAI API) is ahead of GPT-6 Luna (Codex CLI): 0.82 s against 2.79 s (lower is better; n = 5 per side; run range).

CLI vs API: time for a one-line answer (Total time)

run range · n = 5 per side · Lower is better · Claude Code CLI vs Codex CLI vs the API: latency and tokens

  • GPT-6.1 Sol (OpenAI API) is ahead of GPT-6.1 Sol (Codex CLI): 1.52 s against 4.19 s (lower is better; n = 5 per side; run range).
  • GPT-6 Luna (OpenAI API) is ahead of GPT-6.1 Sol (Codex CLI): 0.97 s against 4.19 s (lower is better; n = 5 per side; run range).
  • GPT-6.1 Sol (OpenAI API) is ahead of GPT-6 Luna (Codex CLI): 1.52 s against 3.19 s (lower is better; n = 5 per side; run range).
  • GPT-6 Luna (OpenAI API) is ahead of GPT-6 Luna (Codex CLI): 0.97 s against 3.19 s (lower is better; n = 5 per side; run range).

Total time per attempt: single call vs agent loop

run range · n = 16 to 24 per side · Lower is better · Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks

  • GPT-6 Luna (Codex CLI) is ahead of Claude Haiku 4.5: 5.16 s against 39.0 s (lower is better; n = 16 to 24 per side; run range).

Haiku thinking study: time per routing decision (Model time (API))

median to p95 · n = 82 per side · Lower is better · Does thinking pay for Claude Haiku 4.5? Thinking on vs off

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 1.60 s against 3.79 s (lower is better; n = 82 per side; median to p95).

Haiku thinking study: time per routing decision (Wall time (CLI))

median to p95 · n = 82 per side · Lower is better · Does thinking pay for Claude Haiku 4.5? Thinking on vs off

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 2.60 s against 4.66 s (lower is better; n = 82 per side; median to p95).

Time per call, instructions vs schema mode

run range · n = 12 to 24 per side · Lower is better · Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 3.52 s against 9.52 s (lower is better; n = 12 to 24 per side; run range).
  • Claude Code is ahead of Codex CLI: 3.52 s against 6.21 s (lower is better; n = 12 per side; run range).
  • Claude Sonnet 5.5 is ahead of GPT-6.1 Sol (Codex CLI): 3.52 s against 6.21 s (lower is better; n = 12 per side; run range).

Time per routing decision, by route (Model time (API, CLI-reported))

median to p95 · n = 56 per side · Lower is better · Jev vs Claude routers on unseen decisions: a blind holdout

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 1.49 s against 7.52 s (lower is better; n = 56 per side; median to p95).

Time per routing decision, by route (Wall time)

median to p95 · n = 56 to 168 per side · Lower is better · Jev vs Claude routers on unseen decisions: a blind holdout

  • Claude Sonnet 5.5 is ahead of Claude Haiku 4.5: 2.36 s against 9.44 s (lower is better; n = 56 per side; median to p95).
  • Jev 1.13 is ahead of Claude Haiku 4.5: 0.14 s against 9.44 s (lower is better; n = 56 to 168 per side; median to p95).
  • Jev 1.13 is ahead of Claude Sonnet 5.5: 0.14 s against 2.36 s (lower is better; n = 56 to 168 per side; median to p95).

Only a 95% interval is a confidence interval. A run range and a median-to-p95 band are not.

39 rows in 18 metrics. Listed by kind (quality, speed, cost), then by study, then by metric. This is not a ranking.

Cost: 0 gaps

Cost: none of 825 cost rows shows a gap. 11 have an interval or range on both sides. 148 are list-price calculations, not runs.

Who is ahead of whom

A cell counts the rows where the side in the row is ahead of the side in the column. It counts rows, not wins in a contest. One measurement can appear on several pairs. The diagonal is 0: a side is not compared with itself.

Rowvs Claude Fable 5.11vs Claude Haiku 4.52vs Claude Opus 5.53vs Claude Sonnet 5.54vs GPT-6 Luna (Codex CLI)5vs GPT-6 Luna (OpenAI API)6vs GPT-6.1 Sol (Codex CLI)7vs GPT-6.1 Sol (OpenAI API)8vs Claude Code9vs Codex CLI10vs Deterministic routing policy11vs Jev 1.1312
Claude Fable 5.100000000000
Claude Haiku 4.500000000000
Claude Opus 5.500000000000
Claude Sonnet 5.5000000000
GPT-6 Luna (Codex CLI)00000000000
GPT-6 Luna (OpenAI API)0000000000
GPT-6.1 Sol (Codex CLI)0000000000
GPT-6.1 Sol (OpenAI API)0000000000
Claude Code00000000000
Codex CLI00000000000
Deterministic routing policy000000000
Jev 1.130000000000
  1. vs Claude Fable 5.1
  2. vs Claude Haiku 4.5
  3. vs Claude Opus 5.5
  4. vs Claude Sonnet 5.5
  5. vs GPT-6 Luna (Codex CLI)
  6. vs GPT-6 Luna (OpenAI API)
  7. vs GPT-6.1 Sol (Codex CLI)
  8. vs GPT-6.1 Sol (OpenAI API)
  9. vs Claude Code
  10. vs Codex CLI
  11. vs Deterministic routing policy
  12. vs Jev 1.13

Shade is the value against the largest value in the matrix.

12 sides in 17 pairs. 64 rows with a side ahead. Rows list the side ahead, columns the side behind. Sorted by kind of side, then by name. This is not a ranking.

Where the gaps sit, and where none do

70 rate rows show 100% on both sides, and every one is a tie. A perfect result still has a wide 95% interval: 44% to 100% at n = 3, and 96% to 100% at n = 82. A task set that every model passes cannot show a gap. What benchmark saturation means

Effort: 0 of 172 rows show a gap, across 16 effort comparisons. 31 are ties and 141 are unclear.

Every study, by verdict

Studies in the order of the dataset. A study with no side ahead still has rows: they are ties or unclear.

22 studies with comparison rows; 12 have a row with a side ahead.

Claude vs GPT: is there a real difference?

348 rows from 16 pairs compare an Anthropic side (Claude) with an OpenAI side (GPT). 23 show a gap: 10 on quality, 13 on speed and 0 on cost. Each quality gap names Claude Code and Claude Haiku 4.5 and Claude Sonnet 5.5 and Codex CLI and GPT-6 Luna (Codex CLI) as the side behind.

How to read a gap

  • A side is ahead only when the 95% intervals of a rate do not overlap.
  • For times, a side is ahead only when the fastest-to-slowest ranges (or the median-to-p95 bands) do not overlap, with 5 or more runs per side.
  • A range is not a confidence interval. Only the 95% interval is.
  • A tie means this sample cannot separate the two. It does not mean they are equal.

Limits of this count

  • The 64 rows come from 31 measurements. One measurement can sit against several opponents, so the rows are not independent.
  • Of the 64 gaps, 25 rest on a 95% interval, 23 on a run range (fastest to slowest), 16 on a median-to-p95 band. Only a 95% interval is a confidence interval.
  • The smallest samples are 4 runs per side. 2 rows rest on them.
  • Each side is a model with its route and settings. A gap between two routes is not a gap between two models. Read the context on each pair page.
  • The counts come from the dataset of 2026-10-07. New studies change them.
  • We build Agent. These are our own runs, on our own task sets, and the sets are small.

How we measure

Questions

Are AI models really different?

Of 1,571 comparison rows, 64 show a gap the intervals or ranges support. The rest are ties or unclear. Only 578 rows have an interval or range on both sides, so only those can name a side. 64 of them do: 11% (a calculation).

Is Claude Opus better than Sonnet?

Claude Sonnet 5.5 and Claude Opus 5.5 share 35 measured metrics and 31 list-price calculations from 10 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 16 ties and 50 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).

Is Claude better than GPT?

Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) share 49 measured metrics and 21 list-price calculations from 10 studies. Claude Sonnet 5.5 leads on 4 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 1 more. GPT-6.1 Sol (Codex CLI) leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 22 ties and 43 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest). 348 rows from 16 pairs compare an Anthropic side (Claude) with an OpenAI side (GPT). 23 show a gap: 10 on quality, 13 on speed and 0 on cost. Each quality gap names Claude Code and Claude Haiku 4.5 and Claude Sonnet 5.5 and Codex CLI and GPT-6 Luna (Codex CLI) as the side behind.

Does more reasoning effort change the result?

Effort: 0 of 172 rows show a gap, across 16 effort comparisons. 31 are ties and 141 are unclear.

What do the gaps have in common?

Quality: 25 of 327 rate rows show a gap. They come from 6 studies and 7 pairs. Speed: 39 of 209 time rows show a gap. They come from 9 studies and 14 pairs. Cost: none of 825 cost rows shows a gap. 11 have an interval or range on both sides. 148 are list-price calculations, not runs.

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.