Calculation · 95% Wilson intervals

LLM eval sample size calculator: how many runs?

Enter two pass rates. The calculator shows how many runs per side it takes before their 95% intervals stop overlapping. It is arithmetic on your numbers, not a benchmark.

Two pass rates

Example input

Example input: exact single-pass rates on 82 paired routing decisions, from the recorded table in our routing study. Jev ran on its own cases; this is a home advantage. The planner assumes independent samples. The two rates are Jev 1.13 (TypeSafe) and Claude Sonnet 5.5. Replace both numbers with your own.

Example input: the list-price cost of one routing decision with Claude Sonnet 5.5. It is a calculation.

Calculation

Runs per side

795

Runs per side

Calculation: 795 runs on each of 2 sides is 1,590 runs in all.

Calculation: 1,590 runs × $0.004996 = $7.94.

795 runs per side would separate the two rates. This is a calculation.

Calculation: the smallest equal number of runs per side at which the two 95% Wilson intervals do not overlap, if both rates hold.

95% confidence intervals for the two pass rates

Both charts use the 95% Wilson interval. The first chart shows the answer. The second chart shows how the answer comes about.

Two rates at n = 795 per side

n = 795 runs per side. Calculation, not a run.

Calculation
Rate A
Rate B

95% Wilson interval (calculation)Rates on a 0–100% axisn = 795 runs per side

Calculation. At n = 795 runs per side, rate A has 717 passes (90.19%) and a 95% interval of 87.92% to 92.07%. Rate B has 747 passes (93.96%) and a 95% interval of 92.09% to 95.42%. The two intervals do not overlap.

Calculation. Each interval is the 95% Wilson interval for k = round(rate × n) passes of n runs. The rates are your inputs. They are not measured results. At n = 795 the gap between the two intervals is 0.02 points. It is small because 795 is the first number of runs at which they stop overlapping. The whole-percent labels can look equal. The Table view shows two decimals.

Both rates stay fixed. The rows show the intervals at more and more runs per side. The overlap band shows where they still touch.

How the intervals narrow as runs grow

Both rates fixed. One row per number of runs per side.

Calculation
  • Rate A
  • Rate B
n = 10
n = 25
n = 50
n = 100
n = 250
n = 500
n = 794 (overlap)
n = 795 (first apart)
n = 1,000
n = 2,500
n = 5,000

95% Wilson interval (calculation)Rates on a 0–100% axisOne n per row, equal on both sides

Calculation. 11 rows, from n = 10 to n = 5,000 runs per side. The intervals first stop overlapping at n = 795. They overlap at n = 794. They do not overlap on 4 of 11 rows.

Calculation. Each row holds both rates fixed and uses k = round(rate × n). Counts are whole numbers, so at a few larger n the intervals can touch again.

Examples from our data

Each card uses values from a published study. The runs-needed numbers are calculations.

  • Not separated yet

    Jev 1.13 (TypeSafe) vs Claude Sonnet 5.5, 82 single-pass routing decisions

    Jev 1.13 (TypeSafe): 90% (74/82)

    Claude Sonnet 5.5: 94% (77/82)

    Jev 1.13 (TypeSafe): 90% (74/82), 95% interval 82% to 95%.

    Claude Sonnet 5.5: 94% (77/82), 95% interval 87% to 97%.

    Study, n = 82 per side: the intervals overlap, so the study cannot separate them.

    Calculation: about 795 runs per side would separate them, 9.7 times the 82 in the study.

    Both rates come from the same 82 tasks. This calculation treats them as two separate samples.

    Jev used its recorded production run, one pass, on its own cases: a home advantage. These counts do not describe the pooled repeat-call chart.

    Calculation: the displayed Wilson intervals use the exact single-pass counts. The planner assumes independent equal-size samples; this paired study needs a paired comparison for inference.

    Open the chart in the study
  • Already separated

    Claude Haiku 4.5 vs Claude Sonnet 5.5, strict pass, 24 calls

    Claude Haiku 4.5 · Claude Code: 46% (11/24)

    Claude Sonnet 5.5 · Claude Code: 100% (24/24)

    Claude Haiku 4.5 · Claude Code: 46% (11/24), 95% interval 28% to 65%.

    Claude Sonnet 5.5 · Claude Code: 100% (24/24), 95% interval 86% to 100%.

    Study, n = 24 per side: the intervals do not overlap. The study already separates them.

    Calculation: about 11 runs per side would be enough at these rates, fewer than the 24 in the study.

    Open the chart in the study
  • Not separated yet

    Panel preference, first vs latest attempt, 12 tasks

    First scored attempt: 50% (6/12)

    Latest attempt: 75% (9/12)

    First scored attempt: 50% (6/12), 95% interval 25% to 75%.

    Latest attempt: 75% (9/12), 95% interval 47% to 91%.

    Study, n = 12 per side: the intervals overlap, so the study cannot separate them.

    Calculation: about 54 runs per side would separate them, 4.5 times the 12 in the study.

    Both rates come from the same 12 tasks. This calculation treats them as two separate samples.

    Open the chart in the study
  • Perfect score

    Perfect scores on a small set: 16 of 16 calls

    Claude Sonnet 5.5 (low) · Claude Code: 100% (16/16)

    Every configuration passed 16 of 16 calls: 100%.

    Calculation: 16 of 16 still allows a true rate as low as 80.64%, a gap of about 19 points.

    Calculation: with 24 of 24, the lower end is 86.2%.

    Open the chart in the study
  • Interval width

    Agent on 33 SWE-bench Verified instances

    Agent resolved: 76% (25/33)

    76% (25/33) resolved: 95% interval 59% to 87%, a width of 28.2 points.

    Calculation: at the same rate, 132 instances give a width of 14.5 points; 528 instances give a width of 7.3 points.

    Open the chart in the study

How the planner makes the number

  1. For k passes in n runs, the planner builds the 95% Wilson interval. It is the formula the dataset uses for every rate on this site.
  2. It sets k to round(rate × n) on each side and tries n = 1, 2, 3 and on up to 5,000.
  3. The answer is the first n at which the two intervals do not overlap. Intervals that touch still overlap.
  4. The compare pages use the same rule to call a winner. It is stricter than a usual two-sample test, so it can ask for more runs.

What the number does not tell you

  • It assumes both rates stay as you typed them. Real runs vary by chance, so a real study can need more runs or fewer. Plan with a margin.
  • It is not a power calculation. It does not say how likely a real study is to separate the two rates.
  • Counts are whole numbers. At a slightly larger n the intervals can touch again. The ladder chart shows n − 1 and n.
  • Both sides get the same number of runs. Each run is independent, and each run has the same chance to pass.
  • A paired test on the same tasks can need fewer runs. This calculation does not model it.
  • A rate of 0% or 100% still has an interval. 16 passes in 16 runs allow a true rate as low as 81%.

Questions

How many runs do I need to compare two LLMs?

It depends on the gap. For pass rates of 90.24% and 93.9%, about 795 runs per side would separate the two 95% intervals. For 50% and 75%, the number is about 54 runs per side. Both numbers are calculations.

What is a confidence interval for a pass rate?

A 95% confidence interval is a range of true pass rates that fit your result. This page uses the Wilson interval. For 16 passes in 16 runs it is 81% to 100%.

Does a 100% pass rate mean the model never fails?

No. A perfect score on a small set cannot rule out failures. 16 of 16 still allows a true rate as low as 81%. 24 of 24 allows 86%.

Can I use this for any eval?

Use it for a pass or fail rate where each run is independent. It does not cover a score, such as a 1 to 5 rating, or a paired comparison on the same tasks.

Read more

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.