Explainer · Wilson confidence interval
Wilson confidence intervals for AI benchmarks
Definition
A Wilson confidence interval (also called the Wilson score interval) is a range around a measured pass rate that shows which true pass rates are consistent with the result, given the number of trials. For a benchmark result of k passes in n attempts, the 95% Wilson interval is the range that would contain the true rate in about 95% of repeated experiments. Unlike the simple "plus or minus" interval, it stays inside 0% to 100% and stays honest at small n and at rates near 0% or 100%.
Agent team · · 4 min read · Every number is from the public studies
Interactive
Interval playground: when is a gap real?
Move n and k. Each 95% Wilson interval narrows as the runs add up; overlapping intervals are a tie.
Tie: the 95% intervals overlap (87%–95%), so this data cannot separate Jev 1.13 and Claude Sonnet 5.5.
Interval width as n grows, at 90%
±7 points at n = 82
| Configuration | k | n | Rate | 95% Wilson interval |
|---|---|---|---|---|
| Jev 1.13 | 74 | 82 | 90% | 82%–95% |
| Claude Sonnet 5.5 | 77 | 82 | 94% | 87%–97% |
Two configurations: Jev 1.13 74 of 82 (90%, 95% interval 82% to 95%) and Claude Sonnet 5.5 77 of 82 (94%, 87% to 97%). Tie: the 95% intervals overlap (87%–95%), so this data cannot separate Jev 1.13 and Claude Sonnet 5.5.
The interval is the Wilson score interval, the formula behind every rate on this site. Values you set here are a calculation, not a result. The preset loads published counts: see the study.
Source: Methodology: intervals
Why a pass rate needs an interval
A pass rate from a benchmark is a sample. Run the same model on another 16 tasks of the same kind and you may get 15 instead of 16. The interval tells you how much of the gap between two models could be this sampling noise.
Two numbers without intervals invite the wrong reading. "Model A: 100%, model B: 94%" sounds like a winner. With intervals, both may be a tie.
The formula
For k passes in n trials, observed rate p = k / n, and z = 1.96 for 95%:
- centre = (p + z² / 2n) / (1 + z² / n)
- half-width = z / (1 + z² / n) × √( p(1 − p) / n + z² / 4n² )
- interval = centre ± half-width
The z² terms pull the centre a little towards 50% and keep the width above zero, even when every trial passed.
A worked example: 16 of 16
In the effort ladder study, every configuration passed 16 of 16 calls. With p = 1 and n = 16:
- z² / n = 3.8416 / 16 = 0.2401
- centre = (1 + 0.12005) / 1.2401 ≈ 0.9032
- half-width = 1.96 / 1.2401 × √(3.8416 / 1024) ≈ 0.0968
- interval ≈ 80.6% to 100%
So a perfect score on 16 calls is consistent with a true pass rate as low as about 81%. That is why the study says it "cannot rule out a difference of up to about 19 points" between effort levels that all scored 16/16.
The simple Wald interval (p ± 1.96 × √(p(1 − p)/n)) gives 100% to 100% for the same result: a zero-width interval from a sample of 16. That is the failure the Wilson interval fixes.
Intervals in practice
On our eight hard tasks:
- The Claude configurations at 24 of 24 have intervals of 86% to 100%.
- GPT-6.1 Sol through the Codex CLI at 16 of 16 has 81% to 100%: the same perfect score, a wider interval, because n is smaller.
- Claude Haiku 4.5 passed 11 of 24 strictly (46%, interval 28% to 65%). Its interval does not overlap the others, so this is a real gap on this task set.
The interval also shows what a small sample cannot do:
In the consistency study, each cell is 10 repetitions of one prompt. A 10 of 10 cell has an interval of 72% to 100%; Haiku's 0 of 10 on the exact-number prompt has 0% to 28%. Ten calls separate "always" from "never", but not 90% from 100%.
How to use intervals
- Overlap means tie. When two 95% intervals overlap, the data does not separate the two rates. Our comparison pages name a winner only when they do not overlap.
- Non-overlap is strong evidence, not proof. It is a conservative check. Two intervals can overlap a little while a formal test still finds a difference; we use the stricter rule.
- More samples narrow the interval. The width shrinks with the square root of n: four times the trials gives about half the width.
- Intervals are for rates. For times we show the median and the fastest-to-slowest range, and say that the range is not a confidence interval.
- Count every attempt. Dropping failed or timed-out runs makes the rate and its interval look better than the system is.
Frequently asked questions
What is the difference between a Wilson interval and a normal (Wald) interval?
The Wald interval is p ± 1.96 × √(p(1 − p)/n). It can go below 0% or above 100% and has zero width at 0% or 100%. The Wilson interval stays inside 0% to 100% and keeps a sensible width at small n and extreme rates, which is where AI benchmarks often are.
What does 16 out of 16 mean with a confidence interval?
It means the 95% Wilson interval is about 81% to 100%. A perfect score on 16 trials is consistent with a true pass rate as low as about 81%, so it does not prove the model never fails.
How many test cases do I need for a benchmark?
It depends on the gap you want to detect. Near 100%, 16 cases leave an interval about 19 points wide and 24 cases about 14 points. To separate two models that differ by a few points you need hundreds of cases, or a paired design on the same cases.
Do overlapping confidence intervals mean there is no difference?
No. They mean this sample cannot show one. A difference may still exist and need more data to see.