An AI model leaderboard without a composite score: why, and how to read ours
Our leaderboard lists 46 models, CLIs, routers and providers with their best-supported facts, each with n and an interval. No single score. Here is why.
TL;DR
- Our leaderboard has 46 entries: 21 models, 20 inference providers, 2 CLIs, 2 routers and 1 agent harness. Each one comes from the same dataset as every study page.
- There is no composite score. Our 16 studies differ in task, route, effort, sample size, and measured against calculated values. One weighted number would need weights nobody measured, and it would hide what each value means.
- Each entry shows up to 3 facts per category (quality, speed, cost), one per study, each with its study, n, interval or range and configuration.
- "Best" means best supported, not best looking: a fact with a 95% interval comes first, then one with a run range, then the one with the larger n.
measuredIncounts the studies with at least one measured fact. Calculations do not count.- To rank two options, use one metric at a time on the comparison pages. There, 147 of 162 pass-rate, accuracy and behaviour-rate rows are ties (our count).
New to the terms? Start at /learn, or read how to read AI benchmarks honestly.
Why most leaderboards have one number
A single score is easy to sort and easy to share. Most public AI leaderboards blend several benchmarks into one: an average, an Elo rating or a weighted index. It answers "which model is best?" in one line.
The trouble is what the blend has to assume:
- That the tasks are comparable. A routing decision, a SWE-bench patch and a strict regex are different work.
- That the weights are right. Someone picks them. Usually nobody measured them against what users need.
- That the gaps are real. Two scores 2 points apart often sit inside each other's interval.
- That every number was measured the same way. Some are runs. Some are list-price calculations.
Our data breaks all four.
Why we could not build one honestly
Here is what sits under our 46 entries:
- Different tasks. 33 SWE-bench Verified instances, 8 hard puzzles with strict validators, 5 short tasks, 6 small repositories with hidden tests, 82 typed routing decisions, 1,048 graded arena decisions and 1,620 arena games, prompts repeated 10 times, 5-turn cache sessions, 200 memory sessions.
- Different routes. The same GPT-6.1 Sol model ran through the Codex CLI and through the OpenAI API. On a scheduler repair, the API took a median 17.3 s and the CLI 61.2 s. A blended score would credit or blame the model for the route.
- Different effort. Our effort ladder ran Sonnet, Opus and GPT-6.1 Sol at 3 or 4 effort levels each.
- Ceilings. On the effort ladder, all 11 cells passed 16/16 (95% interval 81% to 100%). A score cannot rank cells that the data cannot separate.
- Small samples. Many cells hold 10 to 24 runs.
- Measured and calculated values. Cost per pass is list price × recorded tokens. It is arithmetic on a run, not a bill.
Add those into one number and the result depends mostly on which studies an entry happened to be in. Claude Sonnet 5.5 appears in 10 studies with measured facts. Claude Opus 5.5 appears in 6. 27 of the 46 entries appear in only 1. A composite would rank coverage, not quality.
How to read an entry
Each entry has four parts.
1. measuredIn. The number of studies with at least one measured fact for this entry. Studies where an entry only appears in a calculation do not count. For example, Agent (our harness) is in 4 studies but measuredIn is 3, because one of them, the cost thought experiments, is a repricing.
2. studies. The list of studies, so you can open each one.
3. factCount. How many facts the dataset holds for this entry. Sonnet 5.5 has 110; GPT-6.1 Sol (Codex CLI) has 74; Opus 5.5 has 63.
4. best. Up to 3 facts per category, one per study:
- Quality: pass rates, accuracy and scores.
- Speed: times.
- Cost: US dollars, measured or calculated.
Tokens and counts are behaviour, not a ranking, so they are left out.
"Best" means best supported
The order inside a category is fixed: a fact with a 95% interval first, then one with a run range, then the larger n. It is not "the highest value first".
Claude Haiku 4.5 shows why that matters. Its quality facts on the leaderboard are:
- Per-question routing accuracy: 94% (183/194), interval 90% to 97%.
- Routing calls that returned a decision: 100% (82/82), interval 96% to 100%.
- SWE-bench Verified, public run on the same 33 instances: 76% (25/33), interval 59% to 87%.
Haiku also passed only 11/24 on our hard tasks and 0/10 on one repeated prompt in the consistency study. Those facts have intervals too, but smaller n, so they are not in the top 3. They are on the Haiku model page and on the comparison pages, where they decide rows.
So read best as "the strongest evidence we hold", not "the best this entry did". It is a summary of support, not a ranking.
Every fact keeps its context
A number without its configuration misleads. So each fact in best carries:
- The study and the chart it came from.
- n, the number of runs or calls.
- The interval or range, and which kind: a 95% interval, a fastest-to-slowest range, or a median-to-95th-percentile span.
- The context: route, effort and task set, for example "Codex CLI · effort medium · eight hard validated tasks".
GPT-6.1 Sol (Codex CLI) at medium effort shows the same 16/16 in two studies, the hard head-to-head and the effort ladder. The ladder reused that cell as a reference, and the context says so.
Where to rank things
Rank within one metric, on the same task, with both sides measured. That is what the comparison pages do. A row names a winner only when the 95% intervals do not overlap, or, for times, when the run ranges do not overlap with at least 5 runs per side.
Under those rules, most quality rows are ties. Across all comparison pages, 147 of 162 pass-rate, accuracy and behaviour-rate rows are ties (our count). The other 15 all separate Claude Haiku 4.5 from a larger model: 7 on the hard set, 4 on repeated prompts and 4 in the memory study. In all 15, Haiku had the worse result. The newest studies added more ties: all three coding agents passed 36 of 36 hidden-test sessions.
The chart above is the clearest case. All 11 effort-ladder cells score 16/16. A composite would still put one cell first. The data says they tie.
The dataset also holds 16 effort comparisons, one for each model and pair of effort levels. Every one of them is tie or unclear.
What a single score would hide
Three results from our data that a blended score would lose:
- Consistent and wrong. Haiku gave the same wrong number 10 times out of 10 on one prompt. See Same prompt, ten answers.
- Price, not quality, separates most options. On the hard set, Sonnet 5.5 at low effort cost $0.0122 per strict pass and Opus 5.5 at high effort $0.0337 (calculations), with the same 16/16. See Does reasoning effort buy quality?
- The cache moves the bill more than the model. Our agent's work would cost about 3.9x as much at Sonnet prices without caching (a calculation). See How much does prompt caching save?
How to use the leaderboard
- Find your option and check
measuredIn. One study is a thin base. - Read the context of each fact. Is it your route and your effort?
- Check the interval. If it is wide, the value is a direction, not a result.
- Open the comparison page for the two options you are choosing between.
- Run your own tasks. Our tasks are not your tasks.
What to read next
- Why we count every failed attempt
- AI coding benchmarks roundup, October 2026
- Learn the terms: intervals, ranges and calculations
- All comparisons: /compare. All models: /models.
Get your own leaderboard
Agent records the model, the route, the effort, the time, the cost and the validation result for every step of your own work. Try Agent and rank options on the tasks you actually run.
Where models differ lists the measured gaps with their original uncertainty and limits.