• Statistics
  • Evaluation
  • Methodology
  • Sample Size

How many runs do you need to compare two AI models? A sample-size table

Calculation: 100 runs can separate an observed 8-point gap at a perfect score. See Wilson intervals, paired tests and limits for LLM eval sample sizes.

TL;DR

  • How many test cases does an LLM eval need? It depends on the gap you want to show. Two models need a gap before their 95% intervals stop overlapping. It is about 60 points at 10 runs per model, 29 at 24, 8 at 100 and 2 at 400 (a calculation). This assumes the leader passes every case. These are observed-gap thresholds, not a guarantee that this many runs will detect a true gap.
  • A perfect score on a small set proves less than it looks. 16 of 16 passes gives a 95% interval of 80.6% to 100%.
  • A paired test can flag a smaller observed gap: 6 points on 100 cases, if the leader passes every case (a calculation).
  • When every model scores 100%, add harder tasks.
  • Our pages require at least 5 runs per side, and ranges that do not overlap, before they name a faster option. This is a publishing rule, not a sample-size guarantee.
Live story · 48 sDoes more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks

All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.

Transcript
  1. Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
  2. 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
  3. All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  4. Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
  5. More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
  6. List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
  7. 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  8. Start at low effort and measure. Every call, interval and cost online.

What n did to our real intervals

Each rate traces to our public dataset and its source receipts. We show 95% Wilson intervals. Here n counts model calls, agent sessions or instances, as the table shows. Most cells repeat a few tasks.

ResultWhat n countsPasses95% interval
Consistency: Sonnet 5.5, one prompt1 prompt × 1010 of 1072.2% to 100%
Memory: Sonnet 5.5, curated file, full pass5 tasks × 3 sessions15 of 1579.6% to 100%
Effort ladder: 11 configurations, each8 tasks × 216 of 1680.6% to 100%
Hard set: 4 Claude configurations, each8 tasks × 324 of 2486.2% to 100%
SWE-bench Verified: Agent33 instances25 of 3359.0% to 87.2%
Routing: Jev 1.13, recorded 2026-10-05 run, exact decisions82 decisions74 of 8281.9% to 95.0%
  • Exact number
  • JSON object
  • Code fix
Claude Haiku 4.5
Claude Sonnet 5.5
GPT-6.1 Sol (medium)

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 10 per row

One series per prompt; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

A perfect score on 10 runs gives a lower Wilson bound of 72.2%. The 176 pooled effort-ladder calls give 97.9% to 100%. These intervals treat calls as independent; repeats and pooled configurations limit that reading. The observed extremes separate: Haiku 4.5 passed 0 of 10 (0% to 27.8%) on the exact-number prompt. At n = 10, the intervals for 9 of 10 and 10 of 10 overlap.

The LLM benchmark sample-size table (calculation)

This table is a calculation with the Wilson formula (z = 1.96) from our Wilson interval page. The calculations reproduce the real intervals in this post. The 90% column uses p = 0.9 directly; at n = 16 and 24, that is not a possible integer pass count.

Runs per model (n)Interval at 100% passHypothetical interval at 90% pass
1072.2% to 100%59.6% to 98.2%
1680.6% to 100%67.0% to 97.6%
2486.2% to 100%72.0% to 96.9%
5092.9% to 100%78.6% to 95.7%
10096.3% to 100%82.6% to 94.5%
20098.1% to 100%85.1% to 93.4%
40099.0% to 100%86.7% to 92.6%

The 16-run interval is 19.4 percentage points wide (a calculation). It describes one rate, not a confidence interval for the difference between two configurations.

What gap can each n show?

Our rules call a model "ahead" only when the two 95% intervals do not overlap. Here a clear gap means the lower-scoring side’s upper bound falls below the higher-scoring side’s lower bound.

Runs per model (n)Clear gap, leader at 100% (points)Clear gap, hypothetical leader at 90% (points)Paired test, leader at 100% (points)
1060.060.960.0
1643.846.137.5
2429.236.025.0
5016.022.812.0
1008.014.96.0
2004.09.93.0
4002.06.71.5

Every column is a calculation, rounded to one decimal. The 100% column uses integer pass counts. The 90% column uses hypothetical rates on a 0.05-point grid.

These are thresholds for observed results, not a power analysis. Power is the chance that a test detects a true gap. Planning n also needs expected rates, the task mix and, for paired tests, expected disagreements.

The first two columns give approximate observed-gap thresholds, in percentage points:

  1. 10 to 16 runs: 44 to 61 points.
  2. 24 to 50 runs: 16 to 36 points.
  3. 100 runs: 8 to 15 points.
  4. 200 runs: 4 to 10 points.
  5. 400 runs: 2 to 7 points.

Our data illustrates the rule. Haiku's 11 of 24 on the hard set (27.9% to 64.9%) sits 54.2 percentage points below (a calculation) 24 of 24, with no overlap. In the memory study, Sonnet without memory passed 9 of 15 on the full-pass measure (35.7% to 80.2%). That is 40 percentage points below (a calculation) 15 of 15 with a curated file. At 15 runs the calculated threshold is 46.7 percentage points, so the intervals still overlap. On the narrower team-knowledge measure, Sonnet passed 6 of 15 checks without memory (19.8% to 64.3%). With curated memory, it passed 15 of 15 checks (79.6% to 100%). The observed gap is 60 percentage points (a calculation), with no overlap. These 15 checks are not 15 independent tasks.

GPT 5.2 (high)
Gemini 3 Flash (high)
GLM 5 (high)
Agent (Sonnet 5.5, full pipeline)
Claude 4.5 Sonnet (high)
Claude 4.5 Haiku (high)
Claude 4.5 Opus (high)
DeepSeek V3.2 (high)
MiniMax M2.5 (high)
Claude 4.6 Opus
Kimi K2.5 (high)
GPT 5 mini

Every interval overlaps every other: this chart does not order these rows.

12 rows. Highest GPT 5.2 (high) 85% (95% interval 69%–93%, n 33). Lowest GPT 5 mini 64% (95% interval 47%–78%, n 33). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 33 per row

Agent vs 11 public mini-SWE-agent v2 runs, one attempt each

Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

On 33 SWE-bench Verified instances, the top public run resolved 28 (69.1% to 93.3%). The bottom run resolved 21 (46.6% to 77.8%). The 21.2-percentage-point gap (a calculation) has overlapping intervals.

A paired test can flag a smaller gap

When both models run the same cases, the exact McNemar test counts only the cases where they disagree. If the leader passes every case, 6 cases that only the leader passes give p = 0.03125, two-sided (a calculation). The test assumes independent pairs. This does not meet our separate non-overlap rule for calling a side ahead.

In the recorded 2026-10-05 run in our routing study, Jev 1.13 and Claude Sonnet 5.5 split on 5 of 82 decisions (Sonnet right on 4, Jev on 1). The exact paired test gives p = 0.375 (a calculation): this sample does not establish a difference. The cases were revised against Jev answers, so Jev has a home advantage. This pairing uses the earlier recorded run, not the later three-repeat live run.

When every model scores 100%, add harder tasks

Every rate is 95% or more
Claude Sonnet 5.5 (low)
Claude Sonnet 5.5 (medium)
Claude Sonnet 5.5 (high)
Claude Sonnet 5.5
Claude Opus 5.5 (low)
Claude Opus 5.5 (medium)
Claude Opus 5.5 (high)
Claude Opus 5.5
GPT-6.1 Sol (low)
GPT-6.1 Sol (medium)
GPT-6.1 Sol (high)

Every interval overlaps every other: this chart does not order these rows.

11 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 16 per row11 of 11 at 100%: this task set cannot separate them.

Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.

Source: Effort ladder: the hard task set at each effort level

In the effort ladder, all 11 configurations passed 16 of 16 strictly: 176 of 176 together (95% Wilson 97.9% to 100%). Five reference cells reuse hard-study calls; they are not fresh runs. If both configurations keep passing every case, more runs cannot rank them on pass rate. New runs can still expose failures.

A ceiling means the tasks are too easy to rank these models. It does not mean the models are equal.

On five short tasks, Haiku 4.5 passed 15 of 15 (79.6% to 100%), and all nine configurations had overlapping intervals (model head-to-head). On eight hard tasks, Haiku passed 11 of 24 (95% Wilson 27.9% to 64.9%). Four Claude configurations each passed 24 of 24 (86.2% to 100%): the intervals do not overlap. Five Haiku failures were format misses with correct answers, rather than wrong answers.

Timing: medians, ranges and at least 5 runs

For times, we show the median and the fastest-to-slowest range. A range is not a confidence interval. Our comparison pages name a faster side only when the two ranges do not overlap and each side has at least 5 runs. With fewer runs, the row says "unclear".

On the hard set, Sonnet 5.5 took a median 7.75 s per call (2.26 s to 34.79 s, n = 24). Haiku 4.5 took 39.01 s (15.27 s to 75.13 s, n = 24). Haiku’s median is 5.0 times Sonnet’s (a ratio calculation), but the ranges overlap, so the Haiku vs Sonnet page marks the row unclear. More runs can help estimate a median. Adding observations cannot shrink the existing fastest-to-slowest range. One host and one network produced these times.

Best practice for your own evals

  1. Pick the gap that matters to you. Use the table for observed-gap thresholds, then plan power from expected rates and disagreements.
  2. Run both models on the same cases. Use a paired test.
  3. Use many different tasks, not one task many times.
  4. Count every run, failures included (why).

How we measured

  • Real intervals: from our public dataset (generated 2026-10-06). "What n counts" comes from each study's method text.
  • Tables: a calculation. The calculations apply the 95% Wilson formula (z = 1.96) and reproduce the cited dataset intervals. The 90% column uses p = 0.9 directly, including fractional counts at n = 16 and 24.
  • Gap columns: the most passes the other model can have while its upper bound stays below the leader's lower bound. The 90% column steps the rate by 0.05 points.
  • Paired column: an exact two-sided McNemar test: p = 2 × 0.5^b for b disagreements that all favour the leader. The smallest b with p < 0.05 is 6.

Caveats

  • Non-overlap of two 95% intervals is stricter than a formal test. We keep the stricter rule on purpose.
  • Repeats of one task can be correlated. These Wilson intervals do not account for that dependence or show generalization to unseen tasks. The pooled effort-ladder interval also combines different configurations.
  • The paired column assumes that the leader passed every case.
  • The tables do not predict your pass rate or guarantee detection. The paired column assumes independent pairs and all disagreements favouring the perfect-scoring side.

Measure your own models with receipts

Agent keeps a receipt for each task: model, route, tokens, time, cost and validation result. Try Agent to get your own n and intervals.

The data behind this post

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.