Explainer · Wilson confidence interval

Wilson confidence intervals for AI benchmarks

Definition

A Wilson confidence interval (also called the Wilson score interval) is a range around a measured pass rate that shows which true pass rates are consistent with the result, given the number of trials. For a benchmark result of k passes in n attempts, the 95% Wilson interval is the range that would contain the true rate in about 95% of repeated experiments. Unlike the simple "plus or minus" interval, it stays inside 0% to 100% and stays honest at small n and at rates near 0% or 100%.

Agent team · · 4 min read · Every number is from the public studies

Interactive

Interval playground: when is a gap real?

Move n and k. Each 95% Wilson interval narrows as the runs add up; overlapping intervals are a tie.

Calculated live from your inputs
Load a published result
A: Jev 1.13

90%74/82

95% Wilson 82%–95%

B: Claude Sonnet 5.5

94%77/82

95% Wilson 87%–97%

Tie: the 95% intervals overlap (87%–95%), so this data cannot separate Jev 1.13 and Claude Sonnet 5.5.

Interval width as n grows, at 90%

±7 points at n = 82

Two configurations: Jev 1.13 74 of 82 (90%, 95% interval 82% to 95%) and Claude Sonnet 5.5 77 of 82 (94%, 87% to 97%). Tie: the 95% intervals overlap (87%–95%), so this data cannot separate Jev 1.13 and Claude Sonnet 5.5.

The interval is the Wilson score interval, the formula behind every rate on this site. Values you set here are a calculation, not a result. The preset loads published counts: see the study.

Source: Methodology: intervals

Why a pass rate needs an interval

A pass rate from a benchmark is a sample. Run the same model on another 16 tasks of the same kind and you may get 15 instead of 16. The interval tells you how much of the gap between two models could be this sampling noise.

Two numbers without intervals invite the wrong reading. "Model A: 100%, model B: 94%" sounds like a winner. With intervals, both may be a tie.

The formula

For k passes in n trials, observed rate p = k / n, and z = 1.96 for 95%:

  • centre = (p + z² / 2n) / (1 + z² / n)
  • half-width = z / (1 + z² / n) × √( p(1 − p) / n + z² / 4n² )
  • interval = centre ± half-width

The z² terms pull the centre a little towards 50% and keep the width above zero, even when every trial passed.

A worked example: 16 of 16

In the effort ladder study, every configuration passed 16 of 16 calls. With p = 1 and n = 16:

  • z² / n = 3.8416 / 16 = 0.2401
  • centre = (1 + 0.12005) / 1.2401 ≈ 0.9032
  • half-width = 1.96 / 1.2401 × √(3.8416 / 1024) ≈ 0.0968
  • interval ≈ 80.6% to 100%

So a perfect score on 16 calls is consistent with a true pass rate as low as about 81%. That is why the study says it "cannot rule out a difference of up to about 19 points" between effort levels that all scored 16/16.

The simple Wald interval (p ± 1.96 × √(p(1 − p)/n)) gives 100% to 100% for the same result: a zero-width interval from a sample of 16. That is the failure the Wilson interval fixes.

Intervals in practice

  • Strict pass
  • Lenient (format misses counted)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 67% (95% interval 47%–82%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 (Strict pass) at 100%: this task set cannot separate them.

Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

On our eight hard tasks:

  • The Claude configurations at 24 of 24 have intervals of 86% to 100%.
  • GPT-6.1 Sol through the Codex CLI at 16 of 16 has 81% to 100%: the same perfect score, a wider interval, because n is smaller.
  • Claude Haiku 4.5 passed 11 of 24 strictly (46%, interval 28% to 65%). Its interval does not overlap the others, so this is a real gap on this task set.

The interval also shows what a small sample cannot do:

  • Exact number
  • JSON object
  • Code fix
Claude Haiku 4.5
Claude Sonnet 5.5
GPT-6.1 Sol (medium)

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 10 per row

One series per prompt; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

In the consistency study, each cell is 10 repetitions of one prompt. A 10 of 10 cell has an interval of 72% to 100%; Haiku's 0 of 10 on the exact-number prompt has 0% to 28%. Ten calls separate "always" from "never", but not 90% from 100%.

How to use intervals

  • Overlap means tie. When two 95% intervals overlap, the data does not separate the two rates. Our comparison pages name a winner only when they do not overlap.
  • Non-overlap is strong evidence, not proof. It is a conservative check. Two intervals can overlap a little while a formal test still finds a difference; we use the stricter rule.
  • More samples narrow the interval. The width shrinks with the square root of n: four times the trials gives about half the width.
  • Intervals are for rates. For times we show the median and the fastest-to-slowest range, and say that the range is not a confidence interval.
  • Count every attempt. Dropping failed or timed-out runs makes the rate and its interval look better than the system is.

Frequently asked questions

What is the difference between a Wilson interval and a normal (Wald) interval?

The Wald interval is p ± 1.96 × √(p(1 − p)/n). It can go below 0% or above 100% and has zero width at 0% or 100%. The Wilson interval stays inside 0% to 100% and keeps a sensible width at small n and extreme rates, which is where AI benchmarks often are.

What does 16 out of 16 mean with a confidence interval?

It means the 95% Wilson interval is about 81% to 100%. A perfect score on 16 trials is consistent with a true pass rate as low as about 81%, so it does not prove the model never fails.

How many test cases do I need for a benchmark?

It depends on the gap you want to detect. Near 100%, 16 cases leave an interval about 19 points wide and 24 cases about 14 points. To separate two models that differ by a few points you need hundreds of cases, or a paired design on the same cases.

Do overlapping confidence intervals mean there is no difference?

No. They mean this sample cannot show one. A difference may still exist and need more data to see.

The data behind this explainer

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.