Explainer · pass@k

pass@k explained: pass@1, pass@k and pass^k with real runs

Definition

pass@k is the chance that at least one of k tries at a task passes its check. pass@1 is the chance that one try passes, and pass^k (with a caret) is the chance that all k tries pass. So pass@k rewards a model that is sometimes right, and pass^k rewards one that is right every time.

Agent team · · 5 min read · Every number is from the public studies

Interactive

Interval playground: when is a gap real?

Move n and k. Each 95% Wilson interval narrows as the runs add up; overlapping intervals are a tie.

Calculated live from your inputs
Load a published result
A: Claude Haiku 4.5

46%11/24

95% Wilson 28%–65%

B: Claude Sonnet 5.5

100%24/24

95% Wilson 86%–100%

Claude Sonnet 5.5 is ahead: the 95% intervals do not overlap.

Interval width as n grows, at 46%

±19 points at n = 24

Two configurations: Claude Haiku 4.5 11 of 24 (46%, 95% interval 28% to 65%) and Claude Sonnet 5.5 24 of 24 (100%, 86% to 100%). Claude Sonnet 5.5 is ahead: the 95% intervals do not overlap.

The interval is the Wilson score interval, the formula behind every rate on this site. Values you set here are a calculation, not a result. The preset loads published counts: see the study.

Source: Methodology: intervals

pass@1 vs pass@k vs pass^k

Each metric answers a different question:

  • pass@1: if I ask once, how often is the answer right?
  • pass@k: if I ask k times, how often is at least one answer right?
  • pass^k: if I ask k times, how often are all k answers right?

A task gets n tries (n at least k), and c of them pass. C(a, b) is the number of ways to choose b items from a:

pass@1 = c / n
pass@k = 1 − C(n − c, k) / C(n, k)
pass^k = C(c, k) / C(n, k)

The pass@k formula is the unbiased estimator from the HumanEval paper (Chen et al. 2021). The pass^k metric comes from the τ-bench paper (Yao et al. 2024).

Compute each value per task, then take the mean over the tasks. At k = 1 all three are the same number. As k grows, pass@k never falls and pass^k never rises.

A worked example from our hard set

In the hard head-to-head, seven configurations ran 8 hard tasks with strict validators. Claude Haiku 4.5 through Claude Code had 3 tries per task. One try is one call. It passed 11 of 24 calls (46%, 95% interval 28% to 65%):

  • Strict pass
  • Lenient (format misses counted)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 67% (95% interval 47%–82%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 (Strict pass) at 100%: this task set cannot separate them.

Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Per task, the passes in 3 tries were 3, 1, 2, 0, 0, 3, 2 and 0. Take the "Fix a time-zone day-length function (DST)" task by hand: Haiku passed 1 of 3 tries (c = 1, n = 3). Calculation:

  • pass@1 = 1/3
  • pass@2 = 1 − C(2, 2) / C(3, 2) = 1 − 1/3 = 2/3
  • pass@3 = 1 − C(2, 3) / C(3, 3) = 1 − 0 = 1
  • pass^3 = C(1, 3) / C(3, 3) = 0

Over all 8 tasks, the same three metrics are below (calculation from the pass matrix). With n = k = 3, a task scores 1 on pass@3 if any try passed, and on pass^3 if all three did:

MetricHaiku 4.5 · Claude Code95% Wilson interval
pass@111/24 calls (46%)28% to 65%
pass@35/8 tasks with at least one pass (63%)31% to 86%
pass^32/8 tasks with three passes (25%)7% to 59%

Eight tasks is a very small n, and the pass@3 and pass^3 intervals overlap. So the example shows how the metrics behave, not a precise rate for Haiku. The 11/24 interval also treats tries on one task as independent, so the real uncertainty can be wider.

The other six configurations passed every call. On Claude Code that is 24/24 (95% interval 86% to 100%). On Codex CLI it is 16/16 (81% to 100%), with 2 tries per task. Every task cell is full, so every metric is 100% there. This is a ceiling: no metric separates those six.

When pass@k fits and when pass^k fits

Use pass@k when you can test every sample and keep one that passes. Examples are code with unit tests, output with a validator, or a person who picks from k drafts. The cost is k calls per task.

Use pass^k when nothing checks the output. An unattended agent, a scheduled job or a pipeline step must pass every time.

Use pass@1 when the user gets one answer and no retry.

On Haiku's hard tasks, the gap between pass@3 (5/8) and pass^3 (2/8) shows inconsistency. It solved 5 tasks at least once, but only 2 of them on every try.

Why more samples do not fix a consistent error

pass@k helps only when the tries differ. When a model makes the same mistake each time, c is 0 and pass@k is 0 for every k. The consistency study asked the same prompt 10 times:

  • Exact number
  • JSON object
  • Code fix
Claude Haiku 4.5
Claude Sonnet 5.5
GPT-6.1 Sol (medium)

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 10 per row

One series per prompt; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

  • Exact number: Haiku 4.5 passed 0/10 (95% interval 0% to 28%). It gave the same wrong number every time: 289 (expected: 282). Consistent is not the same as correct, and pass@10 is 0 (calculation).
  • JSON object: Haiku passed 1/10 strictly (95% interval 2% to 40%), and 9 more replies were correct but in the wrong format. Because one try passed, pass@10 is 100%, but pass^10 is 0% (calculation). A strict parser with no retry rejects 9 of 10.

A format rule fails in the same way. In the five-task head-to-head, Sonnet 5.5 through Claude Code passed 0/3 tries (95% interval 0% to 56%, a calculation) on "Multi-step shift arithmetic". Every reply ended on the expected answer (2292) but added working lines, which the exact-text validator rejects. So pass@3 is 0 (calculation): every try broke the same rule.

On the hard set, 5 of Haiku's 13 non-passes were format misses and 8 were wrong answers:

  • Strict pass
  • Format miss (correct answer, wrong format)
  • Wrong answer
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

One square per call; counts at the right are exact and in legend order.

7 rows, 3 series: Strict pass, Format miss (correct answer, wrong format), Wrong answer. Strict pass: highest Claude Sonnet 5.5 · Claude Code 24 (n 24). Lowest Claude Haiku 4.5 · Claude Code 11 (n 24). Format miss (correct answer, wrong format): highest Claude Haiku 4.5 · Claude Code 5 (n 24). Lowest GPT-6.1 Sol (high) · Codex CLI 0 (n 16).

Notesn 16–24 per row

Counts per configuration: strict passes, format misses and wrong answers

A format miss is a reply that the strict validator rejected (for example, wrapped in a code fence or in prose although the prompt said not to) whose extracted answer passes the same validator. It is not a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

A lenient reading accepts those format misses. Then 7 of 8 tasks had at least one correct answer in 3 tries (calculation; 95% interval 53% to 98%). Strictly, 5 of 8 did.

How to report pass@k

  1. Name the metric and k: pass@1, pass@3 or pass^3.
  2. Give the tries per task (at least k) and the number of tasks.
  3. Use the unbiased estimator.
  4. Give a 95% interval. A count of tasks takes a Wilson interval.
  5. Say if the check is strict or lenient, and report format misses apart from wrong answers.
  6. Say when a set hits a ceiling: when every cell is n of n, no metric separates the models.

Frequently asked questions

What is the difference between pass@1 and pass@k?

The chance that one try passes is pass@1. The chance that at least one of k tries passes is pass@k. On our 8 hard tasks, Claude Haiku 4.5 passed 11/24 calls (pass@1 46%, 95% interval 28% to 65%). It passed 5/8 tasks at least once in 3 tries (pass@3 63%, a calculation, 95% interval 31% to 86%).

Is a higher k always better?

No. A higher k never lowers pass@k, but the model stays the same. Each extra try costs a call and needs a check. A high k can also hide inconsistency: Haiku passed the JSON prompt 1/10 times strictly, yet its pass@10 is 100% (calculation).

What is pass^k?

pass^k is the chance that all k tries pass. It measures reliability for work that nobody checks. For Haiku on the hard set, pass^3 was 2/8 tasks (25%, a calculation, 95% interval 7% to 59%), against 5/8 for pass@3.

How many samples do I need for pass@k?

You need at least k tries per task. We did not measure how the estimate changes with the number of tries. The number of tasks sets the width of the interval. With 8 tasks, a pass@3 of 5/8 still has a 95% interval of 31% to 86%.

Watch the data

Live story · 44 sHaiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.

Transcript
  1. Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
  2. 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
  3. Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
  4. Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  5. Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
  6. Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
  7. Open benchmarks: intervals, sources and every failure kept.

The data behind this explainer

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.