Explainer · Reading an AI benchmark honestly
How to read AI benchmarks honestly
Definition
Reading an AI benchmark honestly means checking what was measured, on how many cases, under which configuration and with what uncertainty before you accept a ranking. A trustworthy result states its sample size, gives an interval for every rate and a range for every time, counts every failed attempt, separates measured runs from calculations, and names the setup (model, effort, route, harness) behind each number. A table that hides any of these can make a tie look like a win.
Agent team · · 4 min read · Every number is from the public studies
Interactive
Interval playground: when is a gap real?
Move n and k. Each 95% Wilson interval narrows as the runs add up; overlapping intervals are a tie.
Tie: the 95% intervals overlap (87%–95%), so this data cannot separate Jev 1.13 and Claude Sonnet 5.5.
Interval width as n grows, at 90%
±7 points at n = 82
| Configuration | k | n | Rate | 95% Wilson interval |
|---|---|---|---|---|
| Jev 1.13 | 74 | 82 | 90% | 82%–95% |
| Claude Sonnet 5.5 | 77 | 82 | 94% | 87%–97% |
Two configurations: Jev 1.13 74 of 82 (90%, 95% interval 82% to 95%) and Claude Sonnet 5.5 77 of 82 (94%, 87% to 97%). Tie: the 95% intervals overlap (87%–95%), so this data cannot separate Jev 1.13 and Claude Sonnet 5.5.
The interval is the Wilson score interval, the formula behind every rate on this site. Values you set here are a calculation, not a result. The preset loads published counts: see the study.
Source: Methodology: intervals
1. Find n
Every rate is k out of n. "100%" can mean 3 of 3 or 500 of 500. Look for the count of tasks, calls or instances behind each number. If a page does not show it, the number cannot be weighed.
2. Look for the interval, and whether intervals overlap
A rate from a sample has an uncertainty. A 95% Wilson interval shows it. When two intervals overlap, the data does not separate the two systems.
On our eight hard tasks, six configurations passed every call, and their intervals reach down to 86% (24 calls) or 81% (16 calls). They tie. One configuration, Claude Haiku 4.5, passed 11 of 24 strictly, with an interval of 28% to 65% that overlaps none of the others. That gap is real on this task set. See Wilson confidence intervals.
3. Count every failure, and know what kind it was
A benchmark that drops timeouts, crashes or refusals reports a better rate than the system earned. It also matters what kind of failure it was:
Claude Haiku 4.5 had 11 strict passes, 5 format misses (a correct answer in the wrong format, for example inside a code fence) and 8 wrong answers. A format miss is not a pass, but it is a different problem from a wrong answer: a validator or prompt fix may cure it. Read why we count every failed attempt.
4. For times, read the range, not only the median
Latency varies a lot from call to call. A median alone hides that:
On short tasks, the configuration with the lowest median had 1.9 s per call (n = 15), but its slowest call took 9.8 s. The whiskers here are the fastest and slowest call, which is a range, not a confidence interval. When ranges overlap, "faster" describes this run, not a tested ranking.
5. Separate runs from calculations
Many cost figures are not measured bills. They multiply recorded tokens by a list price. That is a fair calculation, but it is not a run, and a repricing of one model's tokens at another model's price is a thought experiment:
This chart prices the tokens Agent recorded on 33 SWE-bench attempts at other models' list prices. Another model would use different tokens and resolve a different set, so these figures bound price sensitivity; they do not predict outcomes. A good page labels each such chart as a calculation.
6. Check the configuration behind each number
The same model gives different results at different efforts, through a CLI or an API, inside different harnesses. "Model A beats model B" is only meaningful when both ran the same tasks, with the same validators, through comparable routes. When routes differ, the gap may be the route. Our comparison rows show both contexts for this reason.
7. Beware the single overall score
A composite score adds numbers that measure different things: pass rates on different task sets, times through different routes, list-price calculations. It needs weights nobody measured, and it hides the configuration behind each part. That is why our leaderboard has no composite score: it shows each system's best-supported values per category, with study, n and interval, and lets you sort on one metric at a time.
8. Ask what the tasks were
- Ceiling: if every strong model scores 100%, the set cannot separate them.
- Contamination: public tasks and fixes may be in training data.
- Fit: a benchmark of short puzzles says little about long agent work, and the reverse.
- Home advantage: a set tuned against one system favors it.
A short checklist
- Is n shown for every number?
- Does every rate have an interval, and every time a range?
- Are failures, timeouts and format misses counted?
- Are calculations labelled as calculations?
- Is the configuration (model, effort, route, harness) named?
- Do the compared systems share the same tasks and validators?
- Is a "winner" claimed only where intervals or ranges do not overlap?
Frequently asked questions
Why do AI benchmark results disagree with each other?
They measure different tasks, through different harnesses and routes, with different sample sizes and attempt policies. Two honest benchmarks can rank the same models differently because they ask different questions.
What sample size makes an AI benchmark trustworthy?
It depends on the gap. A 16-of-16 result still has a 95% interval of about 81% to 100%, so small samples only separate large gaps. Look for the interval rather than a rule of thumb.
Should I trust a leaderboard with one overall score?
Use it with care. A composite mixes metrics from different tasks and setups with chosen weights. Look at the single metric that matches your work, with its n and interval.
Is a list-price cost per task a measured cost?
Usually not. It multiplies recorded tokens by published prices. It is useful for comparison, but it is a calculation, and it ignores discounts, subscriptions and the tokens another model would have used.