Calculation · 95% Wilson intervals
LLM eval sample size calculator: how many runs?
Enter two pass rates. The calculator shows how many runs per side it takes before their 95% intervals stop overlapping. It is arithmetic on your numbers, not a benchmark.
Calculation
Runs per side
795
Runs per side
Calculation: 795 runs on each of 2 sides is 1,590 runs in all.
Calculation: 1,590 runs × $0.004996 = $7.94.
795 runs per side would separate the two rates. This is a calculation.
Calculation: the smallest equal number of runs per side at which the two 95% Wilson intervals do not overlap, if both rates hold.
95% confidence intervals for the two pass rates
Both charts use the 95% Wilson interval. The first chart shows the answer. The second chart shows how the answer comes about.
Two rates at n = 795 per side
n = 795 runs per side. Calculation, not a run.
| Row | Rate | Passes of n | 95% Wilson interval | n |
|---|---|---|---|---|
| Rate A | 90.19% | 717 of 795 | 87.92% to 92.07% | 795 |
| Rate B | 93.96% | 747 of 795 | 92.09% to 95.42% | 795 |
95% Wilson interval (calculation)Rates on a 0–100% axisn = 795 runs per side
Calculation. At n = 795 runs per side, rate A has 717 passes (90.19%) and a 95% interval of 87.92% to 92.07%. Rate B has 747 passes (93.96%) and a 95% interval of 92.09% to 95.42%. The two intervals do not overlap.
Calculation. Each interval is the 95% Wilson interval for k = round(rate × n) passes of n runs. The rates are your inputs. They are not measured results. At n = 795 the gap between the two intervals is 0.02 points. It is small because 795 is the first number of runs at which they stop overlapping. The whole-percent labels can look equal. The Table view shows two decimals.
Both rates stay fixed. The rows show the intervals at more and more runs per side. The overlap band shows where they still touch.
How the intervals narrow as runs grow
Both rates fixed. One row per number of runs per side.
- Rate A
- Rate B
| Row | Rate | Passes of n | 95% Wilson interval | n |
|---|---|---|---|---|
| Rate A: n = 10 | 90% | 9 of 10 | 59.58% to 98.21% | 10 |
| Rate A: n = 25 | 92% | 23 of 25 | 75.03% to 97.78% | 25 |
| Rate A: n = 50 | 90% | 45 of 50 | 78.64% to 95.65% | 50 |
| Rate A: n = 100 | 90% | 90 of 100 | 82.56% to 94.48% | 100 |
| Rate A: n = 250 | 90.4% | 226 of 250 | 86.11% to 93.46% | 250 |
| Rate A: n = 500 | 90.2% | 451 of 500 | 87.28% to 92.51% | 500 |
| Rate A: n = 794 (overlap) | 90.3% | 717 of 794 | 88.05% to 92.17% | 794 |
| Rate A: n = 795 (first apart) | 90.19% | 717 of 795 | 87.92% to 92.07% | 795 |
| Rate A: n = 1,000 | 90.2% | 902 of 1000 | 88.2% to 91.89% | 1,000 |
| Rate A: n = 2,500 | 90.24% | 2256 of 2500 | 89.01% to 91.34% | 2,500 |
| Rate A: n = 5,000 | 90.24% | 4512 of 5000 | 89.39% to 91.03% | 5,000 |
| Rate B: n = 10 | 90% | 9 of 10 | 59.58% to 98.21% | 10 |
| Rate B: n = 25 | 92% | 23 of 25 | 75.03% to 97.78% | 25 |
| Rate B: n = 50 | 94% | 47 of 50 | 83.78% to 97.94% | 50 |
| Rate B: n = 100 | 94% | 94 of 100 | 87.52% to 97.22% | 100 |
| Rate B: n = 250 | 94% | 235 of 250 | 90.34% to 96.33% | 250 |
| Rate B: n = 500 | 94% | 470 of 500 | 91.56% to 95.77% | 500 |
| Rate B: n = 794 (overlap) | 93.95% | 746 of 794 | 92.08% to 95.41% | 794 |
| Rate B: n = 795 (first apart) | 93.96% | 747 of 795 | 92.09% to 95.42% | 795 |
| Rate B: n = 1,000 | 93.9% | 939 of 1000 | 92.24% to 95.22% | 1,000 |
| Rate B: n = 2,500 | 93.92% | 2348 of 2500 | 92.91% to 94.79% | 2,500 |
| Rate B: n = 5,000 | 93.9% | 4695 of 5000 | 93.2% to 94.53% | 5,000 |
95% Wilson interval (calculation)Rates on a 0–100% axisOne n per row, equal on both sides
Calculation. 11 rows, from n = 10 to n = 5,000 runs per side. The intervals first stop overlapping at n = 795. They overlap at n = 794. They do not overlap on 4 of 11 rows.
Calculation. Each row holds both rates fixed and uses k = round(rate × n). Counts are whole numbers, so at a few larger n the intervals can touch again.
Examples from our data
Each card uses values from a published study. The runs-needed numbers are calculations.
Not separated yet
Jev 1.13 (TypeSafe) vs Claude Sonnet 5.5, 82 single-pass routing decisions
Jev 1.13 (TypeSafe): 90% (74/82)
Claude Sonnet 5.5: 94% (77/82)
Jev 1.13 (TypeSafe): 90% (74/82), 95% interval 82% to 95%.
Claude Sonnet 5.5: 94% (77/82), 95% interval 87% to 97%.
Study, n = 82 per side: the intervals overlap, so the study cannot separate them.
Calculation: about 795 runs per side would separate them, 9.7 times the 82 in the study.
Both rates come from the same 82 tasks. This calculation treats them as two separate samples.
Jev used its recorded production run, one pass, on its own cases: a home advantage. These counts do not describe the pooled repeat-call chart.
Calculation: the displayed Wilson intervals use the exact single-pass counts. The planner assumes independent equal-size samples; this paired study needs a paired comparison for inference.
Open the chart in the studyAlready separated
Claude Haiku 4.5 vs Claude Sonnet 5.5, strict pass, 24 calls
Claude Haiku 4.5 · Claude Code: 46% (11/24)
Claude Sonnet 5.5 · Claude Code: 100% (24/24)
Claude Haiku 4.5 · Claude Code: 46% (11/24), 95% interval 28% to 65%.
Claude Sonnet 5.5 · Claude Code: 100% (24/24), 95% interval 86% to 100%.
Study, n = 24 per side: the intervals do not overlap. The study already separates them.
Calculation: about 11 runs per side would be enough at these rates, fewer than the 24 in the study.
Open the chart in the studyNot separated yet
Panel preference, first vs latest attempt, 12 tasks
First scored attempt: 50% (6/12)
Latest attempt: 75% (9/12)
First scored attempt: 50% (6/12), 95% interval 25% to 75%.
Latest attempt: 75% (9/12), 95% interval 47% to 91%.
Study, n = 12 per side: the intervals overlap, so the study cannot separate them.
Calculation: about 54 runs per side would separate them, 4.5 times the 12 in the study.
Both rates come from the same 12 tasks. This calculation treats them as two separate samples.
Open the chart in the studyPerfect score
Perfect scores on a small set: 16 of 16 calls
Claude Sonnet 5.5 (low) · Claude Code: 100% (16/16)
Every configuration passed 16 of 16 calls: 100%.
Calculation: 16 of 16 still allows a true rate as low as 80.64%, a gap of about 19 points.
Calculation: with 24 of 24, the lower end is 86.2%.
Open the chart in the studyInterval width
Agent on 33 SWE-bench Verified instances
Agent resolved: 76% (25/33)
76% (25/33) resolved: 95% interval 59% to 87%, a width of 28.2 points.
Calculation: at the same rate, 132 instances give a width of 14.5 points; 528 instances give a width of 7.3 points.
Open the chart in the study
How the planner makes the number
- For k passes in n runs, the planner builds the 95% Wilson interval. It is the formula the dataset uses for every rate on this site.
- It sets k to round(rate × n) on each side and tries n = 1, 2, 3 and on up to 5,000.
- The answer is the first n at which the two intervals do not overlap. Intervals that touch still overlap.
- The compare pages use the same rule to call a winner. It is stricter than a usual two-sample test, so it can ask for more runs.
What the number does not tell you
- It assumes both rates stay as you typed them. Real runs vary by chance, so a real study can need more runs or fewer. Plan with a margin.
- It is not a power calculation. It does not say how likely a real study is to separate the two rates.
- Counts are whole numbers. At a slightly larger n the intervals can touch again. The ladder chart shows n − 1 and n.
- Both sides get the same number of runs. Each run is independent, and each run has the same chance to pass.
- A paired test on the same tasks can need fewer runs. This calculation does not model it.
- A rate of 0% or 100% still has an interval. 16 passes in 16 runs allow a true rate as low as 81%.
Questions
How many runs do I need to compare two LLMs?
It depends on the gap. For pass rates of 90.24% and 93.9%, about 795 runs per side would separate the two 95% intervals. For 50% and 75%, the number is about 54 runs per side. Both numbers are calculations.
What is a confidence interval for a pass rate?
A 95% confidence interval is a range of true pass rates that fit your result. This page uses the Wilson interval. For 16 passes in 16 runs it is 81% to 100%.
Does a 100% pass rate mean the model never fails?
No. A perfect score on a small set cannot rule out failures. 16 of 16 still allows a true rate as low as 81%. 24 of 24 allows 86%.
Can I use this for any eval?
Use it for a pass or fail rate where each run is independent. It does not cover a score, such as a 1 to 5 rating, or a paired comparison on the same tasks.
Read more
- How to read AI benchmarks honestly — A checklist, with a live interval you can move.
- Wilson confidence intervals for AI benchmarks — The formula and a worked example.
- Methodology — How we make every number on these pages.
- AI cost calculator — Price the runs at list prices.
- Leaderboard — Every system, with its evidence.