189 comparisons · 1,571 rows · updated October 7, 2026
Where AI models really differ. Every gap, counted.
The answer
Of 1,571 comparison rows, 64 show a gap the intervals or ranges support. The rest are ties or unclear.
Only 578 rows have an interval or range on both sides, so only those can name a side. 64 of them do: 11% (a calculation).
Every comparison row, by verdict
Each row of every comparison page is one segment of the bar. The data names a side ahead only where the intervals or ranges do not overlap.
64
of 1,571 comparison rows
show a side ahead
One side ahead on 64; 595 ties, 912 unclear. A side is ahead only where the intervals or ranges do not overlap.
| Verdict | Rows |
|---|---|
| A side is ahead | 64 |
| Tie | 595 |
| Unclear | 912 |
Rows by verdict: a side ahead 64, tie 595, unclear 912.
Every gap, by kind
Listed by kind (quality, speed, cost), then by study, then by metric. This is not a ranking.
Quality: 25 gaps
Quality: 25 of 327 rate rows show a gap. They come from 6 studies and 7 pairs.
- side that is ahead (its own color)
- side behind
- 95% interval
Pass rate on eight hard tasks (Lenient (format misses counted))
- Claude Haiku 4.5 vs Claude Sonnet 5.50%100%67% (16/24)n 24100% (24/24)n 24Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (24/24) against 67% (16/24) (higher is better; n = 24 per side; 95% interval).
- Claude Haiku 4.5 vs Claude Opus 5.50%100%67% (16/24)n 24100% (24/24)n 24Claude Opus 5.5 aheadClaude Opus 5.5 is ahead of Claude Haiku 4.5: 100% (24/24) against 67% (16/24) (higher is better; n = 24 per side; 95% interval).
- Claude Haiku 4.5 vs Claude Fable 5.10%100%67% (16/24)n 24100% (24/24)n 24Claude Fable 5.1 aheadClaude Fable 5.1 is ahead of Claude Haiku 4.5: 100% (24/24) against 67% (16/24) (higher is better; n = 24 per side; 95% interval).
Pass rate on eight hard tasks (Strict pass)
- Claude Haiku 4.5 vs Claude Sonnet 5.50%100%46% (11/24)n 24100% (24/24)n 24Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (24/24) against 46% (11/24) (higher is better; n = 24 per side; 95% interval).
- Claude Haiku 4.5 vs Claude Opus 5.50%100%46% (11/24)n 24100% (24/24)n 24Claude Opus 5.5 aheadClaude Opus 5.5 is ahead of Claude Haiku 4.5: 100% (24/24) against 46% (11/24) (higher is better; n = 24 per side; 95% interval).
- Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)0%100%46% (11/24)n 24100% (16/16)n 16GPT-6.1 Sol (Codex CLI) aheadGPT-6.1 Sol (Codex CLI) is ahead of Claude Haiku 4.5: 100% (16/16) against 46% (11/24) (higher is better; n = 16 to 24 per side; 95% interval).
- Claude Haiku 4.5 vs Claude Fable 5.10%100%46% (11/24)n 24100% (24/24)n 24Claude Fable 5.1 aheadClaude Fable 5.1 is ahead of Claude Haiku 4.5: 100% (24/24) against 46% (11/24) (higher is better; n = 24 per side; 95% interval).
Same prompt, 10 times: strict pass rate (Exact number)
- Claude Haiku 4.5 vs Claude Sonnet 5.50%100%0% (0/10)n 10100% (10/10)n 10Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (10/10) against 0% (0/10) (higher is better; n = 10 per side; 95% interval).
- Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)0%100%0% (0/10)n 10100% (10/10)n 10GPT-6.1 Sol (Codex CLI) aheadGPT-6.1 Sol (Codex CLI) is ahead of Claude Haiku 4.5: 100% (10/10) against 0% (0/10) (higher is better; n = 10 per side; 95% interval).
Same prompt, 10 times: strict pass rate (JSON object)
- Claude Haiku 4.5 vs Claude Sonnet 5.50%100%10% (1/10)n 10100% (10/10)n 10Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (10/10) against 10% (1/10) (higher is better; n = 10 per side; 95% interval).
- Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)0%100%10% (1/10)n 10100% (10/10)n 10GPT-6.1 Sol (Codex CLI) aheadGPT-6.1 Sol (Codex CLI) is ahead of Claude Haiku 4.5: 100% (10/10) against 10% (1/10) (higher is better; n = 10 per side; 95% interval).
A stale README command: who still ran it?: Raw notes, 60 lines
- Claude Haiku 4.5 vs Claude Sonnet 5.50%100%100% (10/10)n 100% (0/15)n 15Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 0% (0/15) against 100% (10/10) (lower is better; n = 10 to 15 per side; 95% interval).
Full pass rate by kind of memory: Handbook, 210 lines
- Claude Haiku 4.5 vs Claude Sonnet 5.50%100%30% (3/10)n 10100% (15/15)n 15Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (15/15) against 30% (3/10) (higher is better; n = 10 to 15 per side; 95% interval).
Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md
- Claude Haiku 4.5 vs Claude Sonnet 5.50%100%10% (1/10)n 1067% (10/15)n 15Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 67% (10/15) against 10% (1/10) (higher is better; n = 10 to 15 per side; 95% interval).
Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines
- Claude Haiku 4.5 vs Claude Sonnet 5.50%100%30% (3/10)n 10100% (15/15)n 15Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (15/15) against 30% (3/10) (higher is better; n = 10 to 15 per side; 95% interval).
Strict pass rate: single call vs agent loop on eight hard tasks
- Claude Haiku 4.5 vs Claude Sonnet 5.50%100%46% (11/24)n 24100% (24/24)n 24Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (24/24) against 46% (11/24) (higher is better; n = 24 per side; 95% interval).
- Claude Code vs Codex CLI0%100%100% (24/24)n 2463% (10/16)n 16Claude Code aheadClaude Code is ahead of Codex CLI: 100% (24/24) against 63% (10/16) (higher is better; n = 16 to 24 per side; 95% interval).
- Claude Sonnet 5.5 vs GPT-6 Luna (Codex CLI)0%100%100% (24/24)n 2463% (10/16)n 16Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of GPT-6 Luna (Codex CLI): 100% (24/24) against 63% (10/16) (higher is better; n = 16 to 24 per side; 95% interval).
Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)
- Claude Haiku 4.5 vs Claude Sonnet 5.5Calculation0%100%0% (0/24)n 24100% (12/12)n 12Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 100% (12/12) against 0% (0/24) (higher is better; n = 12 to 24 per side; 95% interval). This row is a calculation.
- Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)Calculation0%100%0% (0/24)n 24100% (12/12)n 12GPT-6.1 Sol (Codex CLI) aheadGPT-6.1 Sol (Codex CLI) is ahead of Claude Haiku 4.5: 100% (12/12) against 0% (0/24) (higher is better; n = 12 to 24 per side; 95% interval). This row is a calculation.
Pass rate on 4 harder tasks (Lenient (format misses counted))
- Claude Haiku 4.5 vs Claude Opus 5.50%100%0% (0/12)n 1250% (6/12)n 12Claude Opus 5.5 aheadClaude Opus 5.5 is ahead of Claude Haiku 4.5: 50% (6/12) against 0% (0/12) (higher is better; n = 12 per side; 95% interval).
- Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)0%100%0% (0/12)n 1269% (11/16)n 16GPT-6.1 Sol (Codex CLI) aheadGPT-6.1 Sol (Codex CLI) is ahead of Claude Haiku 4.5: 69% (11/16) against 0% (0/12) (higher is better; n = 12 to 16 per side; 95% interval).
Pass rate on 4 harder tasks (Strict pass)
- Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)0%100%0% (0/12)n 1269% (11/16)n 16GPT-6.1 Sol (Codex CLI) aheadGPT-6.1 Sol (Codex CLI) is ahead of Claude Haiku 4.5: 69% (11/16) against 0% (0/12) (higher is better; n = 12 to 16 per side; 95% interval).
Strict pass rate by task: 6x6 Skyscrapers
- Claude Code vs Codex CLI0%100%0% (0/4)n 4100% (4/4)n 4Codex CLI aheadCodex CLI is ahead of Claude Code: 100% (4/4) against 0% (0/4) (higher is better; n = 4 per side; 95% interval).
- Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)0%100%0% (0/4)n 4100% (4/4)n 4GPT-6.1 Sol (Codex CLI) aheadGPT-6.1 Sol (Codex CLI) is ahead of Claude Sonnet 5.5: 100% (4/4) against 0% (0/4) (higher is better; n = 4 per side; 95% interval).
| Kind | Metric | Study | Side ahead | Side behind | Ahead value | Behind value | Better is | n | Span kind | Calculation |
|---|---|---|---|---|---|---|---|---|---|---|
| quality | Pass rate on eight hard tasks (Lenient (format misses counted)) | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks | Claude Sonnet 5.5 | Claude Haiku 4.5 | 100% (24/24) | 67% (16/24) | higher | 24 | 95% interval | no |
| quality | Pass rate on eight hard tasks (Lenient (format misses counted)) | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks | Claude Opus 5.5 | Claude Haiku 4.5 | 100% (24/24) | 67% (16/24) | higher | 24 | 95% interval | no |
| quality | Pass rate on eight hard tasks (Lenient (format misses counted)) | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks | Claude Fable 5.1 | Claude Haiku 4.5 | 100% (24/24) | 67% (16/24) | higher | 24 | 95% interval | no |
| quality | Pass rate on eight hard tasks (Strict pass) | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks | Claude Sonnet 5.5 | Claude Haiku 4.5 | 100% (24/24) | 46% (11/24) | higher | 24 | 95% interval | no |
| quality | Pass rate on eight hard tasks (Strict pass) | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks | Claude Opus 5.5 | Claude Haiku 4.5 | 100% (24/24) | 46% (11/24) | higher | 24 | 95% interval | no |
| quality | Pass rate on eight hard tasks (Strict pass) | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks | GPT-6.1 Sol (Codex CLI) | Claude Haiku 4.5 | 100% (16/16) | 46% (11/24) | higher | 24 and 16 | 95% interval | no |
| quality | Pass rate on eight hard tasks (Strict pass) | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks | Claude Fable 5.1 | Claude Haiku 4.5 | 100% (24/24) | 46% (11/24) | higher | 24 | 95% interval | no |
| quality | Same prompt, 10 times: strict pass rate (Exact number) | Prompt caching and run-to-run consistency in Claude Code and Codex CLI | Claude Sonnet 5.5 | Claude Haiku 4.5 | 100% (10/10) | 0% (0/10) | higher | 10 | 95% interval | no |
| quality | Same prompt, 10 times: strict pass rate (Exact number) | Prompt caching and run-to-run consistency in Claude Code and Codex CLI | GPT-6.1 Sol (Codex CLI) | Claude Haiku 4.5 | 100% (10/10) | 0% (0/10) | higher | 10 | 95% interval | no |
| quality | Same prompt, 10 times: strict pass rate (JSON object) | Prompt caching and run-to-run consistency in Claude Code and Codex CLI | Claude Sonnet 5.5 | Claude Haiku 4.5 | 100% (10/10) | 10% (1/10) | higher | 10 | 95% interval | no |
| quality | Same prompt, 10 times: strict pass rate (JSON object) | Prompt caching and run-to-run consistency in Claude Code and Codex CLI | GPT-6.1 Sol (Codex CLI) | Claude Haiku 4.5 | 100% (10/10) | 10% (1/10) | higher | 10 | 95% interval | no |
| quality | A stale README command: who still ran it?: Raw notes, 60 lines | Does memory help Claude Code? 8 kinds of agent memory, tested | Claude Sonnet 5.5 | Claude Haiku 4.5 | 0% (0/15) | 100% (10/10) | lower | 10 and 15 | 95% interval | no |
| quality | Full pass rate by kind of memory: Handbook, 210 lines | Does memory help Claude Code? 8 kinds of agent memory, tested | Claude Sonnet 5.5 | Claude Haiku 4.5 | 100% (15/15) | 30% (3/10) | higher | 10 and 15 | 95% interval | no |
| quality | Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md | Does memory help Claude Code? 8 kinds of agent memory, tested | Claude Sonnet 5.5 | Claude Haiku 4.5 | 67% (10/15) | 10% (1/10) | higher | 10 and 15 | 95% interval | no |
| quality | Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines | Does memory help Claude Code? 8 kinds of agent memory, tested | Claude Sonnet 5.5 | Claude Haiku 4.5 | 100% (15/15) | 30% (3/10) | higher | 10 and 15 | 95% interval | no |
| quality | Strict pass rate: single call vs agent loop on eight hard tasks | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks | Claude Sonnet 5.5 | Claude Haiku 4.5 | 100% (24/24) | 46% (11/24) | higher | 24 | 95% interval | no |
| quality | Strict pass rate: single call vs agent loop on eight hard tasks | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks | Claude Code | Codex CLI | 100% (24/24) | 63% (10/16) | higher | 24 and 16 | 95% interval | no |
| quality | Strict pass rate: single call vs agent loop on eight hard tasks | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks | Claude Sonnet 5.5 | GPT-6 Luna (Codex CLI) | 100% (24/24) | 63% (10/16) | higher | 24 and 16 | 95% interval | no |
| quality | Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON) | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI | Claude Sonnet 5.5 | Claude Haiku 4.5 | 100% (12/12) | 0% (0/24) | higher | 24 and 12 | 95% interval | yes |
| quality | Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON) | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI | GPT-6.1 Sol (Codex CLI) | Claude Haiku 4.5 | 100% (12/12) | 0% (0/24) | higher | 24 and 12 | 95% interval | yes |
| quality | Pass rate on 4 harder tasks (Lenient (format misses counted)) | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks | Claude Opus 5.5 | Claude Haiku 4.5 | 50% (6/12) | 0% (0/12) | higher | 12 | 95% interval | no |
| quality | Pass rate on 4 harder tasks (Lenient (format misses counted)) | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks | GPT-6.1 Sol (Codex CLI) | Claude Haiku 4.5 | 69% (11/16) | 0% (0/12) | higher | 12 and 16 | 95% interval | no |
| quality | Pass rate on 4 harder tasks (Strict pass) | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks | GPT-6.1 Sol (Codex CLI) | Claude Haiku 4.5 | 69% (11/16) | 0% (0/12) | higher | 12 and 16 | 95% interval | no |
| quality | Strict pass rate by task: 6x6 Skyscrapers | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks | Codex CLI | Claude Code | 100% (4/4) | 0% (0/4) | higher | 4 | 95% interval | no |
| quality | Strict pass rate by task: 6x6 Skyscrapers | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks | GPT-6.1 Sol (Codex CLI) | Claude Sonnet 5.5 | 100% (4/4) | 0% (0/4) | higher | 4 | 95% interval | no |
Only a 95% interval is a confidence interval. A run range and a median-to-p95 band are not.
25 rows in 13 metrics. Listed by kind (quality, speed, cost), then by study, then by metric. This is not a ranking.
Speed: 39 gaps
Speed: 39 of 209 time rows show a gap. They come from 9 studies and 14 pairs.
- side that is ahead (its own color)
- side behind
- fastest–slowest run (not an interval)
- median to p95 (not an interval)
Time per coding session
- Claude Code vs Codex CLI0s250s23.1 sn 12113.4 sn 12Claude Code aheadClaude Code is ahead of Codex CLI: 23.1 s against 113.4 s (lower is better; n = 12 per side; run range).
- Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)0s250s23.1 sn 12113.4 sn 12Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of GPT-6.1 Sol (Codex CLI): 23.1 s against 113.4 s (lower is better; n = 12 per side; run range).
Same prompt, 10 times: time per call (Code fix)
- Claude Haiku 4.5 vs Claude Sonnet 5.50s10s5.95 sn 102.67 sn 10Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 2.67 s against 5.95 s (lower is better; n = 10 per side; run range).
- Claude Code vs Codex CLI0s20s2.67 sn 1011.3 sn 10Claude Code aheadClaude Code is ahead of Codex CLI: 2.67 s against 11.3 s (lower is better; n = 10 per side; run range).
- Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)0s20s2.67 sn 1011.3 sn 10Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of GPT-6.1 Sol (Codex CLI): 2.67 s against 11.3 s (lower is better; n = 10 per side; run range).
- Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)0s20s5.95 sn 1011.3 sn 10Claude Haiku 4.5 aheadClaude Haiku 4.5 is ahead of GPT-6.1 Sol (Codex CLI): 5.95 s against 11.3 s (lower is better; n = 10 per side; run range).
Same prompt, 10 times: time per call (Exact number)
- Claude Code vs Codex CLI0s20s6.89 sn 1013.4 sn 10Claude Code aheadClaude Code is ahead of Codex CLI: 6.89 s against 13.4 s (lower is better; n = 10 per side; run range).
- Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)0s20s6.89 sn 1013.4 sn 10Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of GPT-6.1 Sol (Codex CLI): 6.89 s against 13.4 s (lower is better; n = 10 per side; run range).
- Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI)0s20s5.06 sn 1013.4 sn 10Claude Haiku 4.5 aheadClaude Haiku 4.5 is ahead of GPT-6.1 Sol (Codex CLI): 5.06 s against 13.4 s (lower is better; n = 10 per side; run range).
Time per routing decision (Model time (API))
- Claude Haiku 4.5 vs Claude Sonnet 5.51 s100 s · log10,734 msn 821,599 msn 82Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 1,599 ms against 10,734 ms (lower is better; n = 82 per side; median to p95).
Time per routing decision (Wall time (CLI))
- Claude Haiku 4.5 vs Claude Sonnet 5.50 ms50 s12,674 msn 822,598 msn 82Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 2,598 ms against 12,674 ms (lower is better; n = 82 per side; median to p95).
CLI start-up tax on a one-word answer (First model output)
- Claude Code vs Codex CLI0 ms10 s1,461 msn 55,059 msn 5Claude Code aheadClaude Code is ahead of Codex CLI: 1,461 ms against 5,059 ms (lower is better; n = 5 per side; run range).
CLI start-up tax on a one-word answer (Total wall time)
- Claude Code vs Codex CLI0 ms10 s2,529 msn 55,999 msn 5Claude Code aheadClaude Code is ahead of Codex CLI: 2,529 ms against 5,999 ms (lower is better; n = 5 per side; run range).
Time to make one routing decision
- Claude Haiku 4.5 vs Claude Sonnet 5.50 ms50 s12,543 msn 822,597 msn 82Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 2,597 ms against 12,543 ms (lower is better; n = 82 per side; median to p95).
- Jev 1.13 vs Claude Haiku 4.5100 ms100 s · log137 msn 24612,543 msn 82Jev 1.13 aheadJev 1.13 is ahead of Claude Haiku 4.5: 137 ms against 12,543 ms (lower is better; n = 82 to 246 per side; median to p95).
- Jev 1.13 vs Claude Sonnet 5.5100 ms10 s · log137 msn 2462,597 msn 82Jev 1.13 aheadJev 1.13 is ahead of Claude Sonnet 5.5: 137 ms against 2,597 ms (lower is better; n = 82 to 246 per side; median to p95).
- Jev 1.13 vs Deterministic routing policy0.001 ms1 s · log137 msn 2461.42 µsn 20000Deterministic routing policy aheadDeterministic routing policy is ahead of Jev 1.13: 1.42 µs against 137 ms (lower is better; n = 246 to 20,000 per side; median to p95).
- Deterministic routing policy vs Claude Sonnet 5.50.001 ms10 s · log1.42 µsn 200002,597 msn 82Deterministic routing policy aheadDeterministic routing policy is ahead of Claude Sonnet 5.5: 1.42 µs against 2,597 ms (lower is better; n = 82 to 20,000 per side; median to p95).
- Deterministic routing policy vs Claude Haiku 4.50.001 ms100 s · log1.42 µsn 2000012,543 msn 82Deterministic routing policy aheadDeterministic routing policy is ahead of Claude Haiku 4.5: 1.42 µs against 12,543 ms (lower is better; n = 82 to 20,000 per side; median to p95).
Where an LLM router’s time goes: model vs CLI (CLI and harness time)
- Claude Haiku 4.5 vs Claude Sonnet 5.50 ms5 s1,698 msn 82973 msn 82Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 973 ms against 1,698 ms (lower is better; n = 82 per side; median to p95).
Where an LLM router’s time goes: model vs CLI (Model API time)
- Claude Haiku 4.5 vs Claude Sonnet 5.51 s100 s · log10,508 msn 821,596 msn 82Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 1,596 ms against 10,508 ms (lower is better; n = 82 per side; median to p95).
CLI vs API: time for a one-line answer (First useful output)
- GPT-6.1 Sol (Codex CLI) vs GPT-6.1 Sol (OpenAI API)0s5s3.79 sn 51.34 sn 5GPT-6.1 Sol (OpenAI API) aheadGPT-6.1 Sol (OpenAI API) is ahead of GPT-6.1 Sol (Codex CLI): 1.34 s against 3.79 s (lower is better; n = 5 per side; run range).
- GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (OpenAI API)0s5s3.79 sn 50.82 sn 5GPT-6 Luna (OpenAI API) aheadGPT-6 Luna (OpenAI API) is ahead of GPT-6.1 Sol (Codex CLI): 0.82 s against 3.79 s (lower is better; n = 5 per side; run range).
- GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (Codex CLI)0s5s1.34 sn 52.79 sn 5GPT-6.1 Sol (OpenAI API) aheadGPT-6.1 Sol (OpenAI API) is ahead of GPT-6 Luna (Codex CLI): 1.34 s against 2.79 s (lower is better; n = 5 per side; run range).
- GPT-6 Luna (Codex CLI) vs GPT-6 Luna (OpenAI API)0s5s2.79 sn 50.82 sn 5GPT-6 Luna (OpenAI API) aheadGPT-6 Luna (OpenAI API) is ahead of GPT-6 Luna (Codex CLI): 0.82 s against 2.79 s (lower is better; n = 5 per side; run range).
CLI vs API: time for a one-line answer (Total time)
- GPT-6.1 Sol (Codex CLI) vs GPT-6.1 Sol (OpenAI API)0s5s4.19 sn 51.52 sn 5GPT-6.1 Sol (OpenAI API) aheadGPT-6.1 Sol (OpenAI API) is ahead of GPT-6.1 Sol (Codex CLI): 1.52 s against 4.19 s (lower is better; n = 5 per side; run range).
- GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (OpenAI API)0s5s4.19 sn 50.97 sn 5GPT-6 Luna (OpenAI API) aheadGPT-6 Luna (OpenAI API) is ahead of GPT-6.1 Sol (Codex CLI): 0.97 s against 4.19 s (lower is better; n = 5 per side; run range).
- GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (Codex CLI)0s5s1.52 sn 53.19 sn 5GPT-6.1 Sol (OpenAI API) aheadGPT-6.1 Sol (OpenAI API) is ahead of GPT-6 Luna (Codex CLI): 1.52 s against 3.19 s (lower is better; n = 5 per side; run range).
- GPT-6 Luna (Codex CLI) vs GPT-6 Luna (OpenAI API)0s5s3.19 sn 50.97 sn 5GPT-6 Luna (OpenAI API) aheadGPT-6 Luna (OpenAI API) is ahead of GPT-6 Luna (Codex CLI): 0.97 s against 3.19 s (lower is better; n = 5 per side; run range).
Total time per attempt: single call vs agent loop
- Claude Haiku 4.5 vs GPT-6 Luna (Codex CLI)1s100s · log39.0 sn 245.16 sn 16GPT-6 Luna (Codex CLI) aheadGPT-6 Luna (Codex CLI) is ahead of Claude Haiku 4.5: 5.16 s against 39.0 s (lower is better; n = 16 to 24 per side; run range).
Haiku thinking study: time per routing decision (Model time (API))
- Claude Haiku 4.5 vs Claude Sonnet 5.50s10s3.79 sn 821.60 sn 82Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 1.60 s against 3.79 s (lower is better; n = 82 per side; median to p95).
Haiku thinking study: time per routing decision (Wall time (CLI))
- Claude Haiku 4.5 vs Claude Sonnet 5.50s10s4.66 sn 822.60 sn 82Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 2.60 s against 4.66 s (lower is better; n = 82 per side; median to p95).
Time per call, instructions vs schema mode
- Claude Haiku 4.5 vs Claude Sonnet 5.50s20s9.52 sn 243.52 sn 12Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 3.52 s against 9.52 s (lower is better; n = 12 to 24 per side; run range).
- Claude Code vs Codex CLI0s20s3.52 sn 126.21 sn 12Claude Code aheadClaude Code is ahead of Codex CLI: 3.52 s against 6.21 s (lower is better; n = 12 per side; run range).
- Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)0s20s3.52 sn 126.21 sn 12Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of GPT-6.1 Sol (Codex CLI): 3.52 s against 6.21 s (lower is better; n = 12 per side; run range).
Time per routing decision, by route (Model time (API, CLI-reported))
- Claude Haiku 4.5 vs Claude Sonnet 5.50s25s7.52 sn 561.49 sn 56Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 1.49 s against 7.52 s (lower is better; n = 56 per side; median to p95).
Time per routing decision, by route (Wall time)
- Claude Haiku 4.5 vs Claude Sonnet 5.50s50s9.44 sn 562.36 sn 56Claude Sonnet 5.5 aheadClaude Sonnet 5.5 is ahead of Claude Haiku 4.5: 2.36 s against 9.44 s (lower is better; n = 56 per side; median to p95).
- Jev 1.13 vs Claude Haiku 4.50.1s100s · log0.14 sn 1689.44 sn 56Jev 1.13 aheadJev 1.13 is ahead of Claude Haiku 4.5: 0.14 s against 9.44 s (lower is better; n = 56 to 168 per side; median to p95).
- Jev 1.13 vs Claude Sonnet 5.50.1s10s · log0.14 sn 1682.36 sn 56Jev 1.13 aheadJev 1.13 is ahead of Claude Sonnet 5.5: 0.14 s against 2.36 s (lower is better; n = 56 to 168 per side; median to p95).
| Kind | Metric | Study | Side ahead | Side behind | Ahead value | Behind value | Better is | n | Span kind | Calculation |
|---|---|---|---|---|---|---|---|---|---|---|
| speed | Time per coding session | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks | Claude Code | Codex CLI | 23.1 s | 113.4 s | lower | 12 | run range | no |
| speed | Time per coding session | Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks | Claude Sonnet 5.5 | GPT-6.1 Sol (Codex CLI) | 23.1 s | 113.4 s | lower | 12 | run range | no |
| speed | Same prompt, 10 times: time per call (Code fix) | Prompt caching and run-to-run consistency in Claude Code and Codex CLI | Claude Sonnet 5.5 | Claude Haiku 4.5 | 2.67 s | 5.95 s | lower | 10 | run range | no |
| speed | Same prompt, 10 times: time per call (Code fix) | Prompt caching and run-to-run consistency in Claude Code and Codex CLI | Claude Code | Codex CLI | 2.67 s | 11.3 s | lower | 10 | run range | no |
| speed | Same prompt, 10 times: time per call (Code fix) | Prompt caching and run-to-run consistency in Claude Code and Codex CLI | Claude Sonnet 5.5 | GPT-6.1 Sol (Codex CLI) | 2.67 s | 11.3 s | lower | 10 | run range | no |
| speed | Same prompt, 10 times: time per call (Code fix) | Prompt caching and run-to-run consistency in Claude Code and Codex CLI | Claude Haiku 4.5 | GPT-6.1 Sol (Codex CLI) | 5.95 s | 11.3 s | lower | 10 | run range | no |
| speed | Same prompt, 10 times: time per call (Exact number) | Prompt caching and run-to-run consistency in Claude Code and Codex CLI | Claude Code | Codex CLI | 6.89 s | 13.4 s | lower | 10 | run range | no |
| speed | Same prompt, 10 times: time per call (Exact number) | Prompt caching and run-to-run consistency in Claude Code and Codex CLI | Claude Sonnet 5.5 | GPT-6.1 Sol (Codex CLI) | 6.89 s | 13.4 s | lower | 10 | run range | no |
| speed | Same prompt, 10 times: time per call (Exact number) | Prompt caching and run-to-run consistency in Claude Code and Codex CLI | Claude Haiku 4.5 | GPT-6.1 Sol (Codex CLI) | 5.06 s | 13.4 s | lower | 10 | run range | no |
| speed | Time per routing decision (Model time (API)) | Jev vs Claude as a router: accuracy and cost | Claude Sonnet 5.5 | Claude Haiku 4.5 | 1,599 ms | 10,734 ms | lower | 82 | median to p95 | no |
| speed | Time per routing decision (Wall time (CLI)) | Jev vs Claude as a router: accuracy and cost | Claude Sonnet 5.5 | Claude Haiku 4.5 | 2,598 ms | 12,674 ms | lower | 82 | median to p95 | no |
| speed | CLI start-up tax on a one-word answer (First model output) | Routing overhead: deterministic policy vs LLM routers vs Jev | Claude Code | Codex CLI | 1,461 ms | 5,059 ms | lower | 5 | run range | no |
| speed | CLI start-up tax on a one-word answer (Total wall time) | Routing overhead: deterministic policy vs LLM routers vs Jev | Claude Code | Codex CLI | 2,529 ms | 5,999 ms | lower | 5 | run range | no |
| speed | Time to make one routing decision | Routing overhead: deterministic policy vs LLM routers vs Jev | Claude Sonnet 5.5 | Claude Haiku 4.5 | 2,597 ms | 12,543 ms | lower | 82 | median to p95 | no |
| speed | Time to make one routing decision | Routing overhead: deterministic policy vs LLM routers vs Jev | Jev 1.13 | Claude Haiku 4.5 | 137 ms | 12,543 ms | lower | 246 and 82 | median to p95 | no |
| speed | Time to make one routing decision | Routing overhead: deterministic policy vs LLM routers vs Jev | Jev 1.13 | Claude Sonnet 5.5 | 137 ms | 2,597 ms | lower | 246 and 82 | median to p95 | no |
| speed | Time to make one routing decision | Routing overhead: deterministic policy vs LLM routers vs Jev | Deterministic routing policy | Jev 1.13 | 1.42 µs | 137 ms | lower | 246 and 20000 | median to p95 | no |
| speed | Time to make one routing decision | Routing overhead: deterministic policy vs LLM routers vs Jev | Deterministic routing policy | Claude Sonnet 5.5 | 1.42 µs | 2,597 ms | lower | 20000 and 82 | median to p95 | no |
| speed | Time to make one routing decision | Routing overhead: deterministic policy vs LLM routers vs Jev | Deterministic routing policy | Claude Haiku 4.5 | 1.42 µs | 12,543 ms | lower | 20000 and 82 | median to p95 | no |
| speed | Where an LLM router’s time goes: model vs CLI (CLI and harness time) | Routing overhead: deterministic policy vs LLM routers vs Jev | Claude Sonnet 5.5 | Claude Haiku 4.5 | 973 ms | 1,698 ms | lower | 82 | median to p95 | no |
| speed | Where an LLM router’s time goes: model vs CLI (Model API time) | Routing overhead: deterministic policy vs LLM routers vs Jev | Claude Sonnet 5.5 | Claude Haiku 4.5 | 1,596 ms | 10,508 ms | lower | 82 | median to p95 | no |
| speed | CLI vs API: time for a one-line answer (First useful output) | Claude Code CLI vs Codex CLI vs the API: latency and tokens | GPT-6.1 Sol (OpenAI API) | GPT-6.1 Sol (Codex CLI) | 1.34 s | 3.79 s | lower | 5 | run range | no |
| speed | CLI vs API: time for a one-line answer (First useful output) | Claude Code CLI vs Codex CLI vs the API: latency and tokens | GPT-6 Luna (OpenAI API) | GPT-6.1 Sol (Codex CLI) | 0.82 s | 3.79 s | lower | 5 | run range | no |
| speed | CLI vs API: time for a one-line answer (First useful output) | Claude Code CLI vs Codex CLI vs the API: latency and tokens | GPT-6.1 Sol (OpenAI API) | GPT-6 Luna (Codex CLI) | 1.34 s | 2.79 s | lower | 5 | run range | no |
| speed | CLI vs API: time for a one-line answer (First useful output) | Claude Code CLI vs Codex CLI vs the API: latency and tokens | GPT-6 Luna (OpenAI API) | GPT-6 Luna (Codex CLI) | 0.82 s | 2.79 s | lower | 5 | run range | no |
| speed | CLI vs API: time for a one-line answer (Total time) | Claude Code CLI vs Codex CLI vs the API: latency and tokens | GPT-6.1 Sol (OpenAI API) | GPT-6.1 Sol (Codex CLI) | 1.52 s | 4.19 s | lower | 5 | run range | no |
| speed | CLI vs API: time for a one-line answer (Total time) | Claude Code CLI vs Codex CLI vs the API: latency and tokens | GPT-6 Luna (OpenAI API) | GPT-6.1 Sol (Codex CLI) | 0.97 s | 4.19 s | lower | 5 | run range | no |
| speed | CLI vs API: time for a one-line answer (Total time) | Claude Code CLI vs Codex CLI vs the API: latency and tokens | GPT-6.1 Sol (OpenAI API) | GPT-6 Luna (Codex CLI) | 1.52 s | 3.19 s | lower | 5 | run range | no |
| speed | CLI vs API: time for a one-line answer (Total time) | Claude Code CLI vs Codex CLI vs the API: latency and tokens | GPT-6 Luna (OpenAI API) | GPT-6 Luna (Codex CLI) | 0.97 s | 3.19 s | lower | 5 | run range | no |
| speed | Total time per attempt: single call vs agent loop | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks | GPT-6 Luna (Codex CLI) | Claude Haiku 4.5 | 5.16 s | 39.0 s | lower | 24 and 16 | run range | no |
| speed | Haiku thinking study: time per routing decision (Model time (API)) | Does thinking pay for Claude Haiku 4.5? Thinking on vs off | Claude Sonnet 5.5 | Claude Haiku 4.5 | 1.60 s | 3.79 s | lower | 82 | median to p95 | no |
| speed | Haiku thinking study: time per routing decision (Wall time (CLI)) | Does thinking pay for Claude Haiku 4.5? Thinking on vs off | Claude Sonnet 5.5 | Claude Haiku 4.5 | 2.60 s | 4.66 s | lower | 82 | median to p95 | no |
| speed | Time per call, instructions vs schema mode | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI | Claude Sonnet 5.5 | Claude Haiku 4.5 | 3.52 s | 9.52 s | lower | 24 and 12 | run range | no |
| speed | Time per call, instructions vs schema mode | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI | Claude Code | Codex CLI | 3.52 s | 6.21 s | lower | 12 | run range | no |
| speed | Time per call, instructions vs schema mode | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI | Claude Sonnet 5.5 | GPT-6.1 Sol (Codex CLI) | 3.52 s | 6.21 s | lower | 12 | run range | no |
| speed | Time per routing decision, by route (Model time (API, CLI-reported)) | Jev vs Claude routers on unseen decisions: a blind holdout | Claude Sonnet 5.5 | Claude Haiku 4.5 | 1.49 s | 7.52 s | lower | 56 | median to p95 | no |
| speed | Time per routing decision, by route (Wall time) | Jev vs Claude routers on unseen decisions: a blind holdout | Claude Sonnet 5.5 | Claude Haiku 4.5 | 2.36 s | 9.44 s | lower | 56 | median to p95 | no |
| speed | Time per routing decision, by route (Wall time) | Jev vs Claude routers on unseen decisions: a blind holdout | Jev 1.13 | Claude Haiku 4.5 | 0.14 s | 9.44 s | lower | 168 and 56 | median to p95 | no |
| speed | Time per routing decision, by route (Wall time) | Jev vs Claude routers on unseen decisions: a blind holdout | Jev 1.13 | Claude Sonnet 5.5 | 0.14 s | 2.36 s | lower | 168 and 56 | median to p95 | no |
Only a 95% interval is a confidence interval. A run range and a median-to-p95 band are not.
39 rows in 18 metrics. Listed by kind (quality, speed, cost), then by study, then by metric. This is not a ranking.
Cost: 0 gaps
Cost: none of 825 cost rows shows a gap. 11 have an interval or range on both sides. 148 are list-price calculations, not runs.
Who is ahead of whom
A cell counts the rows where the side in the row is ahead of the side in the column. It counts rows, not wins in a contest. One measurement can appear on several pairs. The diagonal is 0: a side is not compared with itself.
- vs Claude Fable 5.1
- vs Claude Haiku 4.5
- vs Claude Opus 5.5
- vs Claude Sonnet 5.5
- vs GPT-6 Luna (Codex CLI)
- vs GPT-6 Luna (OpenAI API)
- vs GPT-6.1 Sol (Codex CLI)
- vs GPT-6.1 Sol (OpenAI API)
- vs Claude Code
- vs Codex CLI
- vs Deterministic routing policy
- vs Jev 1.13
Shade is the value against the largest value in the matrix.
| Side ahead | vs Claude Fable 5.1 | vs Claude Haiku 4.5 | vs Claude Opus 5.5 | vs Claude Sonnet 5.5 | vs GPT-6 Luna (Codex CLI) | vs GPT-6 Luna (OpenAI API) | vs GPT-6.1 Sol (Codex CLI) | vs GPT-6.1 Sol (OpenAI API) | vs Claude Code | vs Codex CLI | vs Deterministic routing policy | vs Jev 1.13 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | |
| Claude Haiku 4.5 | 0 | 0 | 0 | 0 | 0 | 2 | 0 | 0 | 0 | 0 | 0 | |
| Claude Opus 5.5 | 0 | 3 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | |
| Claude Sonnet 5.5 | 0 | 21 | 0 | 1 | 0 | 4 | 0 | 0 | 0 | 0 | 0 | |
| GPT-6 Luna (Codex CLI) | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | |
| GPT-6 Luna (OpenAI API) | 0 | 0 | 0 | 0 | 2 | 2 | 0 | 0 | 0 | 0 | 0 | |
| GPT-6.1 Sol (Codex CLI) | 0 | 6 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | |
| GPT-6.1 Sol (OpenAI API) | 0 | 0 | 0 | 0 | 2 | 0 | 2 | 0 | 0 | 0 | 0 | |
| Claude Code | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 7 | 0 | 0 | |
| Codex CLI | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | |
| Deterministic routing policy | 0 | 1 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | |
| Jev 1.13 | 0 | 2 | 0 | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
12 sides in 17 pairs. 64 rows with a side ahead. Rows list the side ahead, columns the side behind. Sorted by kind of side, then by name. This is not a ranking.
- Claude Haiku 4.5 vs Claude Sonnet 5.5 21 of 153 rows
- Claude Haiku 4.5 vs Claude Opus 5.5 3 of 39 rows
- Claude Code vs Codex CLI 8 of 88 rows
- Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI) 5 of 70 rows
- Jev 1.13 vs Claude Haiku 4.5 2 of 23 rows
- Jev 1.13 vs Claude Sonnet 5.5 2 of 23 rows
- GPT-6.1 Sol (Codex CLI) vs GPT-6.1 Sol (OpenAI API) 2 of 8 rows
- Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI) 8 of 56 rows
- Jev 1.13 vs Deterministic routing policy 1 of 7 rows
- Deterministic routing policy vs Claude Sonnet 5.5 1 of 6 rows
- Deterministic routing policy vs Claude Haiku 4.5 1 of 6 rows
- Claude Haiku 4.5 vs GPT-6 Luna (Codex CLI) 1 of 17 rows
- Claude Sonnet 5.5 vs GPT-6 Luna (Codex CLI) 1 of 17 rows
- Claude Haiku 4.5 vs Claude Fable 5.1 2 of 23 rows
- GPT-6.1 Sol (Codex CLI) vs GPT-6 Luna (OpenAI API) 2 of 5 rows
- GPT-6.1 Sol (OpenAI API) vs GPT-6 Luna (Codex CLI) 2 of 5 rows
- GPT-6 Luna (Codex CLI) vs GPT-6 Luna (OpenAI API) 2 of 5 rows
Where the gaps sit, and where none do
70 rate rows show 100% on both sides, and every one is a tie. A perfect result still has a wide 95% interval: 44% to 100% at n = 3, and 96% to 100% at n = 82. A task set that every model passes cannot show a gap. What benchmark saturation means
Effort: 0 of 172 rows show a gap, across 16 effort comparisons. 31 are ties and 141 are unclear.
Every study, by verdict
Studies in the order of the dataset. A study with no side ahead still has rows: they are ties or unclear.
- Agent on SWE-bench Verified vs 11 public models
- Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)
- Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
- Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
- Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
- Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
- Prompt caching and run-to-run consistency in Claude Code and Codex CLI
- Does memory help Claude Code? 8 kinds of agent memory, tested
- Jev vs Claude as a router: accuracy and cost
- Routing overhead: deterministic policy vs LLM routers vs Jev
- Inference provider index: 27 models, 52 providers
- What if every call ran on Opus? Repricing real agent tokens
- Claude Code CLI vs Codex CLI vs the API: latency and tokens
- Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
- Does thinking pay for Claude Haiku 4.5? Thinking on vs off
- Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
- Prompt cache break-even: after how many reuses does a cached prefix cost less?
- Jev vs Claude routers on unseen decisions: a blind holdout
- How much of an AI bill is thinking? Reasoning tokens by model and effort
- Where the seconds go: first text, output speed and prompt size for 6 LLMs
- Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts
- GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
| Study | Rows | Side ahead | Ties | Unclear | Held back | Interval or range on both sides | Calculations |
|---|---|---|---|---|---|---|---|
| Agent on SWE-bench Verified vs 11 public models | 132 | 0 | 66 | 66 | 0 | 66 | 0 |
| Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim) | 3 | 0 | 1 | 2 | 0 | 2 | 1 |
| Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head | 88 | 0 | 13 | 75 | 0 | 44 | 22 |
| Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks | 66 | 7 | 15 | 44 | 0 | 44 | 11 |
| Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks | 16 | 2 | 5 | 9 | 0 | 12 | 4 |
| Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks | 16 | 0 | 4 | 12 | 0 | 8 | 4 |
| Prompt caching and run-to-run consistency in Claude Code and Codex CLI | 40 | 11 | 17 | 12 | 0 | 26 | 2 |
| Does memory help Claude Code? 8 kinds of agent memory, tested | 40 | 4 | 20 | 16 | 0 | 24 | 8 |
| Jev vs Claude as a router: accuracy and cost | 23 | 2 | 18 | 3 | 0 | 20 | 3 |
| Routing overhead: deterministic policy vs LLM routers vs Jev | 43 | 10 | 6 | 27 | 0 | 17 | 24 |
| Inference provider index: 27 models, 52 providers | 621 | 0 | 305 | 316 | 0 | 0 | 0 |
| What if every call ran on Opus? Repricing real agent tokens | 66 | 0 | 0 | 66 | 0 | 0 | 11 |
| Claude Code CLI vs Codex CLI vs the API: latency and tokens | 42 | 8 | 1 | 33 | 0 | 32 | 0 |
| Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks | 56 | 4 | 35 | 17 | 0 | 52 | 4 |
| Does thinking pay for Claude Haiku 4.5? Thinking on vs off | 7 | 2 | 2 | 3 | 0 | 4 | 1 |
| Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI | 32 | 5 | 16 | 11 | 0 | 12 | 8 |
| Prompt cache break-even: after how many reuses does a cached prefix cost less? | 9 | 0 | 0 | 9 | 0 | 0 | 9 |
| Jev vs Claude routers on unseen decisions: a blind holdout | 31 | 4 | 24 | 3 | 0 | 28 | 3 |
| How much of an AI bill is thinking? Reasoning tokens by model and effort | 74 | 0 | 3 | 71 | 0 | 33 | 74 |
| Where the seconds go: first text, output speed and prompt size for 6 LLMs | 91 | 0 | 7 | 84 | 0 | 91 | 49 |
| Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts | 8 | 0 | 0 | 8 | 0 | 0 | 8 |
| GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks | 67 | 5 | 37 | 25 | 0 | 63 | 4 |
22 studies with comparison rows; 12 have a row with a side ahead.
Claude vs GPT: is there a real difference?
348 rows from 16 pairs compare an Anthropic side (Claude) with an OpenAI side (GPT). 23 show a gap: 10 on quality, 13 on speed and 0 on cost. Each quality gap names Claude Code and Claude Haiku 4.5 and Claude Sonnet 5.5 and Codex CLI and GPT-6 Luna (Codex CLI) as the side behind.
- Claude Code vs Codex CLI 8 of 88 rows
- Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI) 5 of 70 rows
- Claude Haiku 4.5 vs GPT-6.1 Sol (Codex CLI) 8 of 56 rows
- Claude Haiku 4.5 vs GPT-6 Luna (Codex CLI) 1 of 17 rows
- Claude Sonnet 5.5 vs GPT-6 Luna (Codex CLI) 1 of 17 rows
- Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI) 0 of 50 rows
- Claude Fable 5.1 vs GPT-6.1 Sol (Codex CLI) 0 of 23 rows
- Claude Haiku 4.5 vs GPT 5.2 0 of 3 rows
How to read a gap
- A side is ahead only when the 95% intervals of a rate do not overlap.
- For times, a side is ahead only when the fastest-to-slowest ranges (or the median-to-p95 bands) do not overlap, with 5 or more runs per side.
- A range is not a confidence interval. Only the 95% interval is.
- A tie means this sample cannot separate the two. It does not mean they are equal.
Limits of this count
- The 64 rows come from 31 measurements. One measurement can sit against several opponents, so the rows are not independent.
- Of the 64 gaps, 25 rest on a 95% interval, 23 on a run range (fastest to slowest), 16 on a median-to-p95 band. Only a 95% interval is a confidence interval.
- The smallest samples are 4 runs per side. 2 rows rest on them.
- Each side is a model with its route and settings. A gap between two routes is not a gap between two models. Read the context on each pair page.
- The counts come from the dataset of 2026-10-07. New studies change them.
- We build Agent. These are our own runs, on our own task sets, and the sets are small.
Questions
Are AI models really different?
Of 1,571 comparison rows, 64 show a gap the intervals or ranges support. The rest are ties or unclear. Only 578 rows have an interval or range on both sides, so only those can name a side. 64 of them do: 11% (a calculation).
Is Claude Opus better than Sonnet?
Claude Sonnet 5.5 and Claude Opus 5.5 share 35 measured metrics and 31 list-price calculations from 10 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 16 ties and 50 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 3 at the smallest).
Is Claude better than GPT?
Claude Sonnet 5.5 and GPT-6.1 Sol (Codex CLI) share 49 measured metrics and 21 list-price calculations from 10 studies. Claude Sonnet 5.5 leads on 4 rows: Time per coding session, 23.1 s vs 113.4 s; Same prompt, 10 times: time per call (Exact number), 6.89 s vs 13.4 s; Same prompt, 10 times: time per call (Code fix), 2.67 s vs 11.3 s; and 1 more. GPT-6.1 Sol (Codex CLI) leads on 1 row: Strict pass rate by task: 6x6 Skyscrapers, 100% (4/4) vs 0% (0/4). On those rows the 95% intervals and run ranges do not overlap; only a 95% interval is a confidence interval. The other rows are 22 ties and 43 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Every row ran the two sides through different routes (for example Claude Code vs Codex CLI), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 3 at the smallest). 348 rows from 16 pairs compare an Anthropic side (Claude) with an OpenAI side (GPT). 23 show a gap: 10 on quality, 13 on speed and 0 on cost. Each quality gap names Claude Code and Claude Haiku 4.5 and Claude Sonnet 5.5 and Codex CLI and GPT-6 Luna (Codex CLI) as the side behind.
Does more reasoning effort change the result?
Effort: 0 of 172 rows show a gap, across 16 effort comparisons. 31 are ties and 141 are unclear.
What do the gaps have in common?
Quality: 25 of 327 rate rows show a gap. They come from 6 studies and 7 pairs. Speed: 39 of 209 time rows show a gap. They come from 9 studies and 14 pairs. Cost: none of 825 cost rows shows a gap. 11 have an interval or range on both sides. 148 are list-price calculations, not runs.