• Methodology
  • Benchmarks
  • Head to head
  • Latency

Most AI model comparisons are ties: 44 of 1,196 rows show a clear gap

44 of 1,196 AI model comparison rows show a gap under our overlap rules. Most gaps are timing rows. No effort row separates quality.

TL;DR

  • Our comparison pages hold 187 pairs and 1,196 rows (dataset of 2026-10-06). Only 44 rows show a gap: the intervals or run ranges do not overlap. In the title, "ties" means "no side is ahead". By strict label, 471 rows are ties and 681 are unclear.
  • 621 rows are provider list prices. A list price has no interval, so no price row can name a side. Only 269 rows meet the evidence rules for naming a side. 44 of those 269 show a gap (16%, a calculation).
  • Speed and route decide most gaps. 29 of the 44 are times. The other 15 are outcome rates, and each involves Claude Haiku 4.5.
  • No row separates Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol on quality. All 28 outcome-rate rows among them are ties. All 130 effort rows are ties or unclear.
  • Are LLM benchmark differences significant? These counts use overlap rules, not significance tests. Check n (the number of runs), intervals and any paired test before you believe "X beats Y".
Live story · 48 sDoes more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.

Transcript
  1. Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
  2. 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  3. All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  4. Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
  5. More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
  6. List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
  7. 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  8. Start at low effort and measure. Every call, interval and cost online.

The count: 44 of 1,196

We read the winner field of every comparison row. The compiler sets it with fixed rules. We judged no row again.

Row typeRowsA gapTieUnclear
Provider list prices (third-party-reported)6210305316
List-price calculations900090
Run results (ours and public)48544166275
All rows1,19644471681

The 1,152 rows without a gap break down like this:

  • 621 price rows. 316 are unclear. 305 are ties, because both sides list the same price.
  • 153 ties, where the 95% intervals overlap.
  • 76 unclear, where run ranges overlap.
  • 161 unclear, with no interval or range (143) or too few runs (18).
  • 141 rows with no better direction, such as a token count: 128 unclear and 13 ties.

Across all rows, 269 rows meet the evidence rules for naming a side. Each has a better direction and a 95% interval (168), a run range (91) or a p50 to p95 band (10). The 44 gaps sit in 15 of the 187 pairs.

What "a gap" means here

The compiler names a side only when a metric has a better direction (a higher pass rate, a lower time) and:

  1. Outcome rates: the 95% Wilson intervals do not overlap. If they overlap, the row is a tie. Higher pass rates are better; lower failure rates are better.
  2. Times: the fastest-to-slowest ranges do not overlap, with 5 or more runs per side.
  3. Routing times: the p50 to p95 bands do not overlap, with 5 or more runs per side.

The run-count rule stops 18 rows. A scheduler repair took a median 15.0 s (13.9 to 15.9 s) with Claude Code. It took 61.2 s (59.9 to 69.5 s) with the Codex CLI. The ranges do not overlap, but each side had only 3 runs. These use Sonnet 5.5 and GPT-6.1 Sol at medium effort, so both model and route differ.

A range is not a confidence interval. A tie means "our overlap rule does not separate them", not "equal". Overlapping individual intervals can still occur with a significant paired test.

What decides: speed and route

Time rows meet our gap rule more often than rate rows. Of 90 eligible time rows with a range or band, 29 show a gap (32%, a calculation). Of 168 rate rows with an interval, 15 do (9%, a calculation). None of the 11 cost rows with eligible ranges shows a gap; these costs are calculations. The 44 gaps are 10 CLI times, 10 routing times, 7 repeated-prompt times, 2 coding-session times and 15 Haiku outcome rates.

  • First output event
  • First model output
  • Total wall time
Entrance: medians race at 4.3× real timeMotion reduced: press Replay to animateThe slowest median is 6 s. The clock runs at the recorded speed.
Claude Code · Claude Haiku 4.5
Codex CLI (default model)

2 rows, 3 series: First output event, First model output, Total wall time. First output event: slowest Claude Code · Claude Haiku 4.5 563 ms (range 519 ms–726 ms, n 5). Fastest Codex CLI (default model) 489 ms (range 354 ms–1.3 s, n 5). All run ranges overlap. First model output: slowest Codex CLI (default model) 5.06 s (range 4.39 s–5.48 s, n 5). Fastest Claude Code · Claude Haiku 4.5 1.46 s (range 1.21 s–2.31 s, n 5). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 5 per row

Median of 5 runs; whiskers = fastest and slowest run

Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.

Source: Routing overhead runs: policy microbenchmark and CLI start-up

CLI start-up. On a one-word answer, Claude Code with Haiku 4.5 reached first model output in a median 1,461 ms (1,206 to 2,308 ms). The Codex CLI, on its default model, took 5,059 ms (4,391 to 5,478 ms). Each side had 5 runs. The models differ, so this row does not separate the CLI from the model.

CLI vs API. GPT-6.1 Sol at high effort answered one line in a median 4.19 s (3.81 to 4.69 s) through the Codex CLI. Through the OpenAI API it took 1.52 s (1.35 to 2.23 s). Each side had 5 runs. All 8 CLI-vs-API gap rows name the API side as faster.

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

Time per decision · log scale: each gridline is 10 times the one before

4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–20000 per row

Median; whiskers = median to 95th percentile

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Routing decisions. A rule-based policy decided in a median 1.42 µs (p95 2.33 µs, 20,000 decisions). Claude Sonnet 5.5 at low effort took 2,597 ms (p95 4,298 ms). Haiku 4.5 with thinking on took 12,543 ms (p95 34,481 ms). Each Claude router made 82 calls through Claude Code.

These are wall times, not model compute times. The policy is an in-process microbenchmark. Jev’s rows use a direct HTTPS route, so routes differ. In all 10 routing rows, the slower median sits above the faster side's p95.

Coding sessions. On 6 hidden-test tasks, Sonnet in Claude Code took a median 23.1 s (18.7 to 44.5 s). Sol in Codex took 113.4 s (78.5 to 221.9 s); each side had 12 sessions. Codex followed the tester’s standing instructions and did extra work, so this compares configured systems.

Repeated prompts. On a code-fix prompt run 10 times, Claude Code with Sonnet 5.5 took a median 2.67 s (2.32 to 4.34 s). The Codex CLI with GPT-6.1 Sol at medium effort took 11.3 s (9.08 to 14.85 s).

What decides on quality: Haiku on these tasks

  • Strict pass
  • Lenient (format misses counted)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 67% (95% interval 47%–82%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 (Strict pass) at 100%: this task set cannot separate them.

Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Each of the 15 outcome-rate rows with a gap involves Claude Haiku 4.5. This does not rank Haiku on all work:

  • Hard set (7 rows). Haiku passed 11/24 strictly (95% interval 28% to 65%). Sonnet 5.5, Opus 5.5 and Fable 5.1 each passed 24/24 (86% to 100%). GPT-6.1 Sol passed 16/16 (81% to 100%). Haiku passed 16/24 under lenient grading (47% to 82%); three rows count this too. There were 30 blocked Codex launch receipts; the pass-rate cells exclude them.
  • Repeated prompts (4 rows). Haiku passed 0/10 on an exact-number prompt (0% to 28%) and 1/10 on a JSON prompt (2% to 40%). Sonnet 5.5 and GPT-6.1 Sol each passed 10/10 on both (72% to 100%). Haiku’s JSON answers were correct in 10/10 calls (95% interval 72% to 100%) but had 9 format misses.
  • Agent memory (4 rows). With a 211-line handbook, Haiku passed 3/10 sessions in full (11% to 60%) and Sonnet 15/15 (80% to 100%). All four rows name Sonnet. The fourth is below.

One row counts a bad outcome. "A stale README command: who still ran it?" counts sessions that ran a test command that fails on Node 25. With 60-line raw notes as memory, Haiku ran it in 10/10 sessions (72% to 100%) and Sonnet in 0/15 (0% to 20%). This chart sets lower as better. The compiler correctly names Sonnet for avoiding the stale command.

What never decides

Every rate is 95% or more
Claude Sonnet 5.5 (low)
Claude Sonnet 5.5 (medium)
Claude Sonnet 5.5 (high)
Claude Sonnet 5.5
Claude Opus 5.5 (low)
Claude Opus 5.5 (medium)
Claude Opus 5.5 (high)
Claude Opus 5.5
GPT-6.1 Sol (low)
GPT-6.1 Sol (medium)
GPT-6.1 Sol (high)

Every interval overlaps every other: this chart does not order these rows.

11 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 16 per row11 of 11 at 100%: this task set cannot separate them.

Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.

Source: Effort ladder: the hard task set at each effort level

  • Quality at the top. All 28 outcome-rate rows among Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol (Codex CLI) are ties. Their only gaps are 3 time rows. This group includes the interim SWE-bench probe, where model and platform build differ together.
  • Effort. All 11 effort-ladder configurations passed 16/16 (81% to 100% each). The 16 effort comparisons hold 130 rows: 31 ties and 99 unclear. Five configurations reuse hard-set receipts; they were not rerun.
  • SWE-bench Verified. Agent resolved 25/33 (59% to 87%). All 66 resolved-rate rows in the study are ties.

Three limits help explain these ties and unclear rows:

  1. The ceiling. 16/16 gives a 95% Wilson interval of 81% to 100%, so two perfect cells overlap. These task sets hit a ceiling. The interval extends to about 81%; it does not measure quality on harder work.
  2. Small n. On five short tasks, Sonnet 5.5 passed 12/15 (55% to 93%) and Opus 5.5 passed 15/15 (80% to 100%). The rates are 20 percentage points apart (a calculation), and still a tie under our overlap rule. The three Sonnet misses were format misses, not wrong final answers.
  3. Overlapping ranges. On the same tasks, Sonnet took a median 2.31 s (2.17 to 7.73 s) and Opus 2.75 s (2.47 to 8.91 s). Each side had 15 calls; these are run ranges.

How to read any "X beats Y" claim

  1. Find n, the interval and any paired test. Without a span, our comparison rule leaves the gap unclear.
  2. Check the overlap and the ceiling. Overlap blocks a winner under our rule; it does not prove equality or replace a paired test. Near 100% on both sides, the task set may have hit a ceiling.
  3. Rank inside one metric. Our leaderboard has no composite score.
  4. Check the route and the date. At high effort, Sol took a median 1.52 s through the API (1.35 to 2.23 s) and 4.19 s through Codex (3.81 to 4.69 s). Each side had 5 runs. Counts change with new studies.

More: how to read AI benchmarks honestly and Wilson confidence intervals.

How we measured

A Python check read the winner field of every row in the public dataset of 2026-10-06. It grouped the 44 gaps by chart and row type. The counts and shares are calculations from the dataset. They describe comparison rows, not independent tests or a share of all AI benchmarks.

Sources: public dataset, short-task receipts, hard-task receipts, CLI/API runs, repeated-prompt receipts, routing summary, coding sessions and memory study.

Caveats

  • The counts are a snapshot. New studies change them.
  • Rows are not independent. The 7 hard-set rows reuse one Haiku result. The 8 CLI-vs-API rows reuse 4 selected sets of 5 runs. Repeated calls on the same tasks do not provide independent task samples.
  • The rule is cautious on purpose. It is not a significance test. A paired test can detect a difference despite overlapping individual intervals.
  • A third party reports the provider prices. We did not measure them.
  • Our task sets are small: 8 hard tasks, 5 short tasks, 6 coding-agent tasks, 5 memory tasks and 33 SWE-bench Verified instances.
  • Disclosure: I build Agent. These comparisons combine our own runs, public SWE-bench runs and reported provider prices.

Run the comparison on your own work

Agent keeps receipts for your own tasks: model, route, tokens, time, cost and validation result. Try Agent and see your numbers.

The data behind this post

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

  • Inference
  • Providers

Inference provider index: 27 models, 52 providers

Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.

265endpoints, 52 providers, 27 models · Provider endpoints in the snapshot · n = 27

34 chartsUpdated October 6, 2026

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.