Explainer · Eval sample size

How many runs does an LLM eval need? Sample size, with real intervals

Definition

Eval sample size counts the units behind a rate: tasks, attempts or checks. Choose independent cases for the precision you need. A narrow 95% interval does not guarantee power to detect a gap between systems.

Agent team · · 5 min read · Every number is from the public studies

Interactive

Interval playground: when is a gap real?

Move n and k. Each 95% Wilson interval narrows as the runs add up; overlapping intervals are a tie.

Calculated live from your inputs
Load a published result
A: Memory checks

100%15/15

95% Wilson 80%–100%

B: Effort-ladder calls

100%176/176

95% Wilson 98%–100%

Tie: the 95% intervals overlap (98%–100%), so this data cannot separate Memory checks and Effort-ladder calls.

Interval width as n grows, at 100%

±10 points at n = 15

Two configurations: Memory checks 15 of 15 (100%, 95% interval 80% to 100%) and Effort-ladder calls 176 of 176 (100%, 98% to 100%). Tie: the 95% intervals overlap (98%–100%), so this data cannot separate Memory checks and Effort-ladder calls.

The interval is the Wilson score interval, the formula behind every rate on this site. Values you set here are a calculation, not a result. The preset loads published counts: see the study.

Source: Methodology: intervals

What n gave us in our own studies

Each row gives k successes in n counted units and a 95% Wilson interval (method). Width is the difference between its ends, in points (calculation).

ResultWhat n countsk/n95% intervalWidth
Consistency, Sonnet 5.5, one prompt1 prompt × 10 repeats10/1072.2% to 100%27.8
Blind review, AI change preferred, latest attempt12 tasks9/1246.8% to 91.1%44.3
Memory, Sonnet 5.5, curated file, team-knowledge checks5 applicable checks × 3 repeats15/1579.6% to 100%20.4
Memory, Sonnet 5.5, no memory, team-knowledge checks5 applicable checks × 3 repeats6/1519.8% to 64.3%44.4
Effort ladder, each configuration8 tasks × 2 repeats16/1680.6% to 100%19.4
Hard set, four Claude configurations8 tasks × 3 repeats24/2486.2% to 100%13.8
Hard set, Haiku 4.58 tasks × 3 repeats11/2427.9% to 64.9%37.0
SWE-bench Verified, Agent33 instances25/3359.0% to 87.2%28.2
Routing, Jev, recorded production run, exact decisions82 decisions74/8281.9% to 95.0%13.1

A perfect score leaves uncertainty below 100%. When strong systems all score 100%, the set hits a ceiling.

The curated file is ahead on team-knowledge checks: the 6/15 and 15/15 intervals do not overlap. This is not full-pass performance. The memory study uses one small synthetic repository; its author knew the tasks. Checks and repeats limit independence.

Every interval for the 11 public models overlaps Agent’s on SWE-bench. This sample cannot rank them. Blind review uses latest attempts that learned from earlier ones; preference does not prove correctness.

  • Agent
  • Public panel mean, same instances
Campaign 1: 25-instance sample
Campaign 2: 8 compiled-extension instances
Original seed draw of 25
All 33 attempted

4 rows, 2 series: Agent, Public panel mean, same instances. Agent: highest Campaign 2: 8 compiled-extension instances 88% (95% interval 53%–98%, n 8). Lowest Campaign 1: 25-instance sample 72% (95% interval 52%–86%, n 25). All intervals overlap. Public panel mean, same instances: highest Campaign 2: 8 compiled-extension instances 84% (n 8). Lowest Original seed draw of 25 71% (n 25). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 8–33 per row

Agent resolved rate and 95% Wilson interval per declared view

The original draw and "all 33" mix two platform builds.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, SWE-bench campaign rules and sample design

The smaller campaign slices have wider intervals. The combined result mixes two platform builds.

A planning table: width by n

The table gives the 95% Wilson interval at an observed 50% and an observed 90% (a calculation with z = 1.96).

nAt 50%WidthAt 90%Width
1023.7% to 76.3%5359.6% to 98.2%39
3033.2% to 66.8%3474.4% to 96.5%22
10040.4% to 59.6%1982.6% to 94.5%12
30044.4% to 55.6%1186.1% to 92.9%7
1,00046.9% to 53.1%688.0% to 91.7%4

At a fixed rate and large n, 4 times the independent cases give about half the width (calculation). Repeated tasks do not add independent samples of difficulty.

Tasks or repetitions: what n counts

Wilson intervals assume independent results. Repeats share a prompt, checker and difficulty. They measure variation on that task, not wider task coverage.

In our consistency study, Haiku 4.5 ran each prompt 10 times. It scored 0/10 on an exact number (95% interval 0% to 27.8%). JSON scored 1/10 (1.8% to 40.4%); a code fix scored 10/10 (72.2% to 100%). Each cell describes one prompt.

On the hard set, Haiku gave the same strict result on all 3 repeats for 5 of 8 tasks. Three tasks scored 0/3 (95% interval 0% to 56.2%); two scored 3/3 (43.8% to 100%).

If every task shared the overall 11/24 rate and repeats were independent, the expected count would be 2.04 tasks (calculation):

8 × [(11/24)³ + (13/24)³] = 2.04.

The pattern is consistent with clustering by task; it does not measure the amount. The 24-call interval may understate uncertainty across tasks.

  • Strict pass
  • Lenient (format misses counted)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 67% (95% interval 47%–82%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 (Strict pass) at 100%: this task set cannot separate them.

Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Pooling has the same limit. The effort ladder scored 176/176 (95% interval 97.9% to 100%, width 2.1 points; calculation). All calls use the same 8 tasks.

Paired designs: same cases, two systems

A paired design runs two systems on the same cases. The exact McNemar test uses disagreements: cases where only one system is right. It assumes independent case pairs.

Our routing study used 82 cases. Against Jev’s recorded production run, Haiku was right alone on 3 cases and Jev alone on 4 (p = 1). Sonnet was right alone on 4 and Jev alone on 1 (p = 0.375).

With 5 disagreements, the smallest two-sided exact p is 2 × (1/2)⁵ = 0.0625 (calculation). A p below 0.05 needs at least 6 disagreements, all favouring one side at that minimum. More disagreements alone do not guarantee significance.

The 95% intervals overlap: Jev 74/82 (81.9% to 95.0%), Haiku 73/82 (80.4% to 94.1%) and Sonnet 77/82 (86.5% to 97.4%). Neither intervals nor paired tests establish an accuracy difference.

The chart uses a later live Jev run: 221/246 across 3 repeats, with 74, 73 and 74 successes. Its displayed 74/82 rounds to the case scale; its interval uses n = 82. The paired counts use the earlier production run. We revised cases against Jev answers, giving Jev a home advantage.

Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 82 per row

Share of asked cases where every scored question was acceptable

Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

A checklist before you pick n

  • Name the gap. Plan precision and power separately.
  • Count independent units. Distinct tasks first, repeats second.
  • Watch for a ceiling. Add harder tasks when models score near 100%.
  • Pair systems. Report disagreements on shared cases.
  • Keep every attempt. Report k/n and intervals. Overlap does not prove equality. See honest benchmark reading and the leaderboard.

Frequently asked questions

How many test cases do I need for an LLM benchmark?

For a Wilson interval no wider than 20 points, 36/40 and 47/94 suffice. For 10 points, 135/150 and 191/382 suffice (calculations assuming independent cases). These examples hold observed rates at exactly 90% and 50%. They do not guarantee comparison power.

Is 100 test cases enough?

At 100 cases, 85/100 (76.7% to 90.7%) and 92/100 (85.0% to 95.9%) have overlapping 95% intervals (calculation). They do not separate these rates.

For exact observed rates of 90% and 95%, the first feasible n with non-overlapping intervals is 440, in multiples of 20 (calculation). The intervals are 396/440: 86.84% to 92.47%, and 418/440: 92.55% to 96.68%. This is not a power calculation.

Should I add tasks or repeat tasks?

Add tasks first. Repeats show variation on selected tasks; there is no universal repeat count.

Why does 16 out of 16 still mean 80.6% to 100%?

For a perfect score, the Wilson lower end is n / (n + 1.96²). At 16/16, 16 / 19.8416 gives 80.6% to 100% (calculation). All 11 effort-ladder configurations scored 16/16: a ceiling. At 73/73, the interval is 95.0% to 100%. This is the first perfect-score n whose lower end reaches 95% (calculation).

Watch the data

Live story · 34 sJev vs Claude as a router: accuracy and cost

Jev vs Claude as a router: accuracy and cost

Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.

Transcript
  1. Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.
  2. Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
  3. Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
  4. Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
  5. Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
  6. Open benchmarks: intervals, sources and every failure kept.

The data behind this explainer

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.