Explainer · Eval sample size
How many runs does an LLM eval need? Sample size, with real intervals
Definition
Eval sample size counts the units behind a rate: tasks, attempts or checks. Choose independent cases for the precision you need. A narrow 95% interval does not guarantee power to detect a gap between systems.
Agent team · · 5 min read · Every number is from the public studies
Interactive
Interval playground: when is a gap real?
Move n and k. Each 95% Wilson interval narrows as the runs add up; overlapping intervals are a tie.
Tie: the 95% intervals overlap (98%–100%), so this data cannot separate Memory checks and Effort-ladder calls.
Interval width as n grows, at 100%
±10 points at n = 15
| Configuration | k | n | Rate | 95% Wilson interval |
|---|---|---|---|---|
| Memory checks | 15 | 15 | 100% | 80%–100% |
| Effort-ladder calls | 176 | 176 | 100% | 98%–100% |
Two configurations: Memory checks 15 of 15 (100%, 95% interval 80% to 100%) and Effort-ladder calls 176 of 176 (100%, 98% to 100%). Tie: the 95% intervals overlap (98%–100%), so this data cannot separate Memory checks and Effort-ladder calls.
The interval is the Wilson score interval, the formula behind every rate on this site. Values you set here are a calculation, not a result. The preset loads published counts: see the study.
Source: Methodology: intervals
What n gave us in our own studies
Each row gives k successes in n counted units and a 95% Wilson interval (method). Width is the difference between its ends, in points (calculation).
| Result | What n counts | k/n | 95% interval | Width |
|---|---|---|---|---|
| Consistency, Sonnet 5.5, one prompt | 1 prompt × 10 repeats | 10/10 | 72.2% to 100% | 27.8 |
| Blind review, AI change preferred, latest attempt | 12 tasks | 9/12 | 46.8% to 91.1% | 44.3 |
| Memory, Sonnet 5.5, curated file, team-knowledge checks | 5 applicable checks × 3 repeats | 15/15 | 79.6% to 100% | 20.4 |
| Memory, Sonnet 5.5, no memory, team-knowledge checks | 5 applicable checks × 3 repeats | 6/15 | 19.8% to 64.3% | 44.4 |
| Effort ladder, each configuration | 8 tasks × 2 repeats | 16/16 | 80.6% to 100% | 19.4 |
| Hard set, four Claude configurations | 8 tasks × 3 repeats | 24/24 | 86.2% to 100% | 13.8 |
| Hard set, Haiku 4.5 | 8 tasks × 3 repeats | 11/24 | 27.9% to 64.9% | 37.0 |
| SWE-bench Verified, Agent | 33 instances | 25/33 | 59.0% to 87.2% | 28.2 |
| Routing, Jev, recorded production run, exact decisions | 82 decisions | 74/82 | 81.9% to 95.0% | 13.1 |
A perfect score leaves uncertainty below 100%. When strong systems all score 100%, the set hits a ceiling.
The curated file is ahead on team-knowledge checks: the 6/15 and 15/15 intervals do not overlap. This is not full-pass performance. The memory study uses one small synthetic repository; its author knew the tasks. Checks and repeats limit independence.
Every interval for the 11 public models overlaps Agent’s on SWE-bench. This sample cannot rank them. Blind review uses latest attempts that learned from earlier ones; preference does not prove correctness.
The smaller campaign slices have wider intervals. The combined result mixes two platform builds.
A planning table: width by n
The table gives the 95% Wilson interval at an observed 50% and an observed 90% (a calculation with z = 1.96).
| n | At 50% | Width | At 90% | Width |
|---|---|---|---|---|
| 10 | 23.7% to 76.3% | 53 | 59.6% to 98.2% | 39 |
| 30 | 33.2% to 66.8% | 34 | 74.4% to 96.5% | 22 |
| 100 | 40.4% to 59.6% | 19 | 82.6% to 94.5% | 12 |
| 300 | 44.4% to 55.6% | 11 | 86.1% to 92.9% | 7 |
| 1,000 | 46.9% to 53.1% | 6 | 88.0% to 91.7% | 4 |
At a fixed rate and large n, 4 times the independent cases give about half the width (calculation). Repeated tasks do not add independent samples of difficulty.
Tasks or repetitions: what n counts
Wilson intervals assume independent results. Repeats share a prompt, checker and difficulty. They measure variation on that task, not wider task coverage.
In our consistency study, Haiku 4.5 ran each prompt 10 times. It scored 0/10 on an exact number (95% interval 0% to 27.8%). JSON scored 1/10 (1.8% to 40.4%); a code fix scored 10/10 (72.2% to 100%). Each cell describes one prompt.
On the hard set, Haiku gave the same strict result on all 3 repeats for 5 of 8 tasks. Three tasks scored 0/3 (95% interval 0% to 56.2%); two scored 3/3 (43.8% to 100%).
If every task shared the overall 11/24 rate and repeats were independent, the expected count would be 2.04 tasks (calculation):
8 × [(11/24)³ + (13/24)³] = 2.04.
The pattern is consistent with clustering by task; it does not measure the amount. The 24-call interval may understate uncertainty across tasks.
Pooling has the same limit. The effort ladder scored 176/176 (95% interval 97.9% to 100%, width 2.1 points; calculation). All calls use the same 8 tasks.
Paired designs: same cases, two systems
A paired design runs two systems on the same cases. The exact McNemar test uses disagreements: cases where only one system is right. It assumes independent case pairs.
Our routing study used 82 cases. Against Jev’s recorded production run, Haiku was right alone on 3 cases and Jev alone on 4 (p = 1). Sonnet was right alone on 4 and Jev alone on 1 (p = 0.375).
With 5 disagreements, the smallest two-sided exact p is 2 × (1/2)⁵ = 0.0625 (calculation). A p below 0.05 needs at least 6 disagreements, all favouring one side at that minimum. More disagreements alone do not guarantee significance.
The 95% intervals overlap: Jev 74/82 (81.9% to 95.0%), Haiku 73/82 (80.4% to 94.1%) and Sonnet 77/82 (86.5% to 97.4%). Neither intervals nor paired tests establish an accuracy difference.
The chart uses a later live Jev run: 221/246 across 3 repeats, with 74, 73 and 74 successes. Its displayed 74/82 rounds to the case scale; its interval uses n = 82. The paired counts use the earlier production run. We revised cases against Jev answers, giving Jev a home advantage.
A checklist before you pick n
- Name the gap. Plan precision and power separately.
- Count independent units. Distinct tasks first, repeats second.
- Watch for a ceiling. Add harder tasks when models score near 100%.
- Pair systems. Report disagreements on shared cases.
- Keep every attempt. Report k/n and intervals. Overlap does not prove equality. See honest benchmark reading and the leaderboard.
Frequently asked questions
How many test cases do I need for an LLM benchmark?
For a Wilson interval no wider than 20 points, 36/40 and 47/94 suffice. For 10 points, 135/150 and 191/382 suffice (calculations assuming independent cases). These examples hold observed rates at exactly 90% and 50%. They do not guarantee comparison power.
Is 100 test cases enough?
At 100 cases, 85/100 (76.7% to 90.7%) and 92/100 (85.0% to 95.9%) have overlapping 95% intervals (calculation). They do not separate these rates.
For exact observed rates of 90% and 95%, the first feasible n with non-overlapping intervals is 440, in multiples of 20 (calculation). The intervals are 396/440: 86.84% to 92.47%, and 418/440: 92.55% to 96.68%. This is not a power calculation.
Should I add tasks or repeat tasks?
Add tasks first. Repeats show variation on selected tasks; there is no universal repeat count.
Why does 16 out of 16 still mean 80.6% to 100%?
For a perfect score, the Wilson lower end is n / (n + 1.96²). At 16/16, 16 / 19.8416 gives 80.6% to 100% (calculation). All 11 effort-ladder configurations scored 16/16: a ceiling. At 73/73, the interval is 95.0% to 100%. This is the first perfect-score n whose lower end reaches 95% (calculation).
Watch the data
Jev vs Claude as a router: accuracy and cost
Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.
Transcript
- Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.
- Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
- Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
- Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
- Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
- Open benchmarks: intervals, sources and every failure kept.