How many runs do you need to compare two AI models? A sample-size table
Calculation: 100 runs can separate an observed 8-point gap at a perfect score. See Wilson intervals, paired tests and limits for LLM eval sample sizes.
TL;DR
- How many test cases does an LLM eval need? It depends on the gap you want to show. Two models need a gap before their 95% intervals stop overlapping. It is about 60 points at 10 runs per model, 29 at 24, 8 at 100 and 2 at 400 (a calculation). This assumes the leader passes every case. These are observed-gap thresholds, not a guarantee that this many runs will detect a true gap.
- A perfect score on a small set proves less than it looks. 16 of 16 passes gives a 95% interval of 80.6% to 100%.
- A paired test can flag a smaller observed gap: 6 points on 100 cases, if the leader passes every case (a calculation).
- When every model scores 100%, add harder tasks.
- Our pages require at least 5 runs per side, and ranges that do not overlap, before they name a faster option. This is a publishing rule, not a sample-size guarantee.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks
All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.
Transcript
- Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
- 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
- All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
- More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
- List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
- 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Start at low effort and measure. Every call, interval and cost online.
What n did to our real intervals
Each rate traces to our public dataset and its source receipts. We show 95% Wilson intervals. Here n counts model calls, agent sessions or instances, as the table shows. Most cells repeat a few tasks.
| Result | What n counts | Passes | 95% interval |
|---|---|---|---|
| Consistency: Sonnet 5.5, one prompt | 1 prompt × 10 | 10 of 10 | 72.2% to 100% |
| Memory: Sonnet 5.5, curated file, full pass | 5 tasks × 3 sessions | 15 of 15 | 79.6% to 100% |
| Effort ladder: 11 configurations, each | 8 tasks × 2 | 16 of 16 | 80.6% to 100% |
| Hard set: 4 Claude configurations, each | 8 tasks × 3 | 24 of 24 | 86.2% to 100% |
| SWE-bench Verified: Agent | 33 instances | 25 of 33 | 59.0% to 87.2% |
| Routing: Jev 1.13, recorded 2026-10-05 run, exact decisions | 82 decisions | 74 of 82 | 81.9% to 95.0% |
A perfect score on 10 runs gives a lower Wilson bound of 72.2%. The 176 pooled effort-ladder calls give 97.9% to 100%. These intervals treat calls as independent; repeats and pooled configurations limit that reading. The observed extremes separate: Haiku 4.5 passed 0 of 10 (0% to 27.8%) on the exact-number prompt. At n = 10, the intervals for 9 of 10 and 10 of 10 overlap.
The LLM benchmark sample-size table (calculation)
This table is a calculation with the Wilson formula (z = 1.96) from our Wilson interval page. The calculations reproduce the real intervals in this post. The 90% column uses p = 0.9 directly; at n = 16 and 24, that is not a possible integer pass count.
| Runs per model (n) | Interval at 100% pass | Hypothetical interval at 90% pass |
|---|---|---|
| 10 | 72.2% to 100% | 59.6% to 98.2% |
| 16 | 80.6% to 100% | 67.0% to 97.6% |
| 24 | 86.2% to 100% | 72.0% to 96.9% |
| 50 | 92.9% to 100% | 78.6% to 95.7% |
| 100 | 96.3% to 100% | 82.6% to 94.5% |
| 200 | 98.1% to 100% | 85.1% to 93.4% |
| 400 | 99.0% to 100% | 86.7% to 92.6% |
The 16-run interval is 19.4 percentage points wide (a calculation). It describes one rate, not a confidence interval for the difference between two configurations.
What gap can each n show?
Our rules call a model "ahead" only when the two 95% intervals do not overlap. Here a clear gap means the lower-scoring side’s upper bound falls below the higher-scoring side’s lower bound.
| Runs per model (n) | Clear gap, leader at 100% (points) | Clear gap, hypothetical leader at 90% (points) | Paired test, leader at 100% (points) |
|---|---|---|---|
| 10 | 60.0 | 60.9 | 60.0 |
| 16 | 43.8 | 46.1 | 37.5 |
| 24 | 29.2 | 36.0 | 25.0 |
| 50 | 16.0 | 22.8 | 12.0 |
| 100 | 8.0 | 14.9 | 6.0 |
| 200 | 4.0 | 9.9 | 3.0 |
| 400 | 2.0 | 6.7 | 1.5 |
Every column is a calculation, rounded to one decimal. The 100% column uses integer pass counts. The 90% column uses hypothetical rates on a 0.05-point grid.
These are thresholds for observed results, not a power analysis. Power is the chance that a test detects a true gap. Planning n also needs expected rates, the task mix and, for paired tests, expected disagreements.
The first two columns give approximate observed-gap thresholds, in percentage points:
- 10 to 16 runs: 44 to 61 points.
- 24 to 50 runs: 16 to 36 points.
- 100 runs: 8 to 15 points.
- 200 runs: 4 to 10 points.
- 400 runs: 2 to 7 points.
Our data illustrates the rule. Haiku's 11 of 24 on the hard set (27.9% to 64.9%) sits 54.2 percentage points below (a calculation) 24 of 24, with no overlap. In the memory study, Sonnet without memory passed 9 of 15 on the full-pass measure (35.7% to 80.2%). That is 40 percentage points below (a calculation) 15 of 15 with a curated file. At 15 runs the calculated threshold is 46.7 percentage points, so the intervals still overlap. On the narrower team-knowledge measure, Sonnet passed 6 of 15 checks without memory (19.8% to 64.3%). With curated memory, it passed 15 of 15 checks (79.6% to 100%). The observed gap is 60 percentage points (a calculation), with no overlap. These 15 checks are not 15 independent tasks.
On 33 SWE-bench Verified instances, the top public run resolved 28 (69.1% to 93.3%). The bottom run resolved 21 (46.6% to 77.8%). The 21.2-percentage-point gap (a calculation) has overlapping intervals.
A paired test can flag a smaller gap
When both models run the same cases, the exact McNemar test counts only the cases where they disagree. If the leader passes every case, 6 cases that only the leader passes give p = 0.03125, two-sided (a calculation). The test assumes independent pairs. This does not meet our separate non-overlap rule for calling a side ahead.
In the recorded 2026-10-05 run in our routing study, Jev 1.13 and Claude Sonnet 5.5 split on 5 of 82 decisions (Sonnet right on 4, Jev on 1). The exact paired test gives p = 0.375 (a calculation): this sample does not establish a difference. The cases were revised against Jev answers, so Jev has a home advantage. This pairing uses the earlier recorded run, not the later three-repeat live run.
When every model scores 100%, add harder tasks
In the effort ladder, all 11 configurations passed 16 of 16 strictly: 176 of 176 together (95% Wilson 97.9% to 100%). Five reference cells reuse hard-study calls; they are not fresh runs. If both configurations keep passing every case, more runs cannot rank them on pass rate. New runs can still expose failures.
A ceiling means the tasks are too easy to rank these models. It does not mean the models are equal.
On five short tasks, Haiku 4.5 passed 15 of 15 (79.6% to 100%), and all nine configurations had overlapping intervals (model head-to-head). On eight hard tasks, Haiku passed 11 of 24 (95% Wilson 27.9% to 64.9%). Four Claude configurations each passed 24 of 24 (86.2% to 100%): the intervals do not overlap. Five Haiku failures were format misses with correct answers, rather than wrong answers.
Timing: medians, ranges and at least 5 runs
For times, we show the median and the fastest-to-slowest range. A range is not a confidence interval. Our comparison pages name a faster side only when the two ranges do not overlap and each side has at least 5 runs. With fewer runs, the row says "unclear".
On the hard set, Sonnet 5.5 took a median 7.75 s per call (2.26 s to 34.79 s, n = 24). Haiku 4.5 took 39.01 s (15.27 s to 75.13 s, n = 24). Haiku’s median is 5.0 times Sonnet’s (a ratio calculation), but the ranges overlap, so the Haiku vs Sonnet page marks the row unclear. More runs can help estimate a median. Adding observations cannot shrink the existing fastest-to-slowest range. One host and one network produced these times.
Best practice for your own evals
- Pick the gap that matters to you. Use the table for observed-gap thresholds, then plan power from expected rates and disagreements.
- Run both models on the same cases. Use a paired test.
- Use many different tasks, not one task many times.
- Count every run, failures included (why).
How we measured
- Real intervals: from our public dataset (generated 2026-10-06). "What n counts" comes from each study's method text.
- Tables: a calculation. The calculations apply the 95% Wilson formula (z = 1.96) and reproduce the cited dataset intervals. The 90% column uses p = 0.9 directly, including fractional counts at n = 16 and 24.
- Gap columns: the most passes the other model can have while its upper bound stays below the leader's lower bound. The 90% column steps the rate by 0.05 points.
- Paired column: an exact two-sided McNemar test: p = 2 × 0.5^b for b disagreements that all favour the leader. The smallest b with p < 0.05 is 6.
Caveats
- Non-overlap of two 95% intervals is stricter than a formal test. We keep the stricter rule on purpose.
- Repeats of one task can be correlated. These Wilson intervals do not account for that dependence or show generalization to unseen tasks. The pooled effort-ladder interval also combines different configurations.
- The paired column assumes that the leader passed every case.
- The tables do not predict your pass rate or guarantee detection. The paired column assumes independent pairs and all disagreements favouring the perfect-scoring side.
What to read next
- Wilson confidence intervals for AI benchmarks
- How to read AI benchmarks honestly
- An AI model leaderboard without a composite score
Measure your own models with receipts
Agent keeps a receipt for each task: model, route, tokens, time, cost and validation result. Try Agent to get your own n and intervals.