Explainer · pass@k
pass@k explained: pass@1, pass@k and pass^k with real runs
Definition
pass@k is the chance that at least one of k tries at a task passes its check. pass@1 is the chance that one try passes, and pass^k (with a caret) is the chance that all k tries pass. So pass@k rewards a model that is sometimes right, and pass^k rewards one that is right every time.
Agent team · · 5 min read · Every number is from the public studies
Interactive
Interval playground: when is a gap real?
Move n and k. Each 95% Wilson interval narrows as the runs add up; overlapping intervals are a tie.
Claude Sonnet 5.5 is ahead: the 95% intervals do not overlap.
Interval width as n grows, at 46%
±19 points at n = 24
| Configuration | k | n | Rate | 95% Wilson interval |
|---|---|---|---|---|
| Claude Haiku 4.5 | 11 | 24 | 46% | 28%–65% |
| Claude Sonnet 5.5 | 24 | 24 | 100% | 86%–100% |
Two configurations: Claude Haiku 4.5 11 of 24 (46%, 95% interval 28% to 65%) and Claude Sonnet 5.5 24 of 24 (100%, 86% to 100%). Claude Sonnet 5.5 is ahead: the 95% intervals do not overlap.
The interval is the Wilson score interval, the formula behind every rate on this site. Values you set here are a calculation, not a result. The preset loads published counts: see the study.
Source: Methodology: intervals
pass@1 vs pass@k vs pass^k
Each metric answers a different question:
- pass@1: if I ask once, how often is the answer right?
- pass@k: if I ask k times, how often is at least one answer right?
- pass^k: if I ask k times, how often are all k answers right?
A task gets n tries (n at least k), and c of them pass. C(a, b) is the number of ways to choose b items from a:
pass@1 = c / n
pass@k = 1 − C(n − c, k) / C(n, k)
pass^k = C(c, k) / C(n, k)The pass@k formula is the unbiased estimator from the HumanEval paper (Chen et al. 2021). The pass^k metric comes from the τ-bench paper (Yao et al. 2024).
Compute each value per task, then take the mean over the tasks. At k = 1 all three are the same number. As k grows, pass@k never falls and pass^k never rises.
A worked example from our hard set
In the hard head-to-head, seven configurations ran 8 hard tasks with strict validators. Claude Haiku 4.5 through Claude Code had 3 tries per task. One try is one call. It passed 11 of 24 calls (46%, 95% interval 28% to 65%):
Per task, the passes in 3 tries were 3, 1, 2, 0, 0, 3, 2 and 0. Take the "Fix a time-zone day-length function (DST)" task by hand: Haiku passed 1 of 3 tries (c = 1, n = 3). Calculation:
- pass@1 = 1/3
- pass@2 = 1 − C(2, 2) / C(3, 2) = 1 − 1/3 = 2/3
- pass@3 = 1 − C(2, 3) / C(3, 3) = 1 − 0 = 1
- pass^3 = C(1, 3) / C(3, 3) = 0
Over all 8 tasks, the same three metrics are below (calculation from the pass matrix). With n = k = 3, a task scores 1 on pass@3 if any try passed, and on pass^3 if all three did:
| Metric | Haiku 4.5 · Claude Code | 95% Wilson interval |
|---|---|---|
| pass@1 | 11/24 calls (46%) | 28% to 65% |
| pass@3 | 5/8 tasks with at least one pass (63%) | 31% to 86% |
| pass^3 | 2/8 tasks with three passes (25%) | 7% to 59% |
Eight tasks is a very small n, and the pass@3 and pass^3 intervals overlap. So the example shows how the metrics behave, not a precise rate for Haiku. The 11/24 interval also treats tries on one task as independent, so the real uncertainty can be wider.
The other six configurations passed every call. On Claude Code that is 24/24 (95% interval 86% to 100%). On Codex CLI it is 16/16 (81% to 100%), with 2 tries per task. Every task cell is full, so every metric is 100% there. This is a ceiling: no metric separates those six.
When pass@k fits and when pass^k fits
Use pass@k when you can test every sample and keep one that passes. Examples are code with unit tests, output with a validator, or a person who picks from k drafts. The cost is k calls per task.
Use pass^k when nothing checks the output. An unattended agent, a scheduled job or a pipeline step must pass every time.
Use pass@1 when the user gets one answer and no retry.
On Haiku's hard tasks, the gap between pass@3 (5/8) and pass^3 (2/8) shows inconsistency. It solved 5 tasks at least once, but only 2 of them on every try.
Why more samples do not fix a consistent error
pass@k helps only when the tries differ. When a model makes the same mistake each time, c is 0 and pass@k is 0 for every k. The consistency study asked the same prompt 10 times:
- Exact number: Haiku 4.5 passed 0/10 (95% interval 0% to 28%). It gave the same wrong number every time: 289 (expected: 282). Consistent is not the same as correct, and pass@10 is 0 (calculation).
- JSON object: Haiku passed 1/10 strictly (95% interval 2% to 40%), and 9 more replies were correct but in the wrong format. Because one try passed, pass@10 is 100%, but pass^10 is 0% (calculation). A strict parser with no retry rejects 9 of 10.
A format rule fails in the same way. In the five-task head-to-head, Sonnet 5.5 through Claude Code passed 0/3 tries (95% interval 0% to 56%, a calculation) on "Multi-step shift arithmetic". Every reply ended on the expected answer (2292) but added working lines, which the exact-text validator rejects. So pass@3 is 0 (calculation): every try broke the same rule.
On the hard set, 5 of Haiku's 13 non-passes were format misses and 8 were wrong answers:
A lenient reading accepts those format misses. Then 7 of 8 tasks had at least one correct answer in 3 tries (calculation; 95% interval 53% to 98%). Strictly, 5 of 8 did.
How to report pass@k
- Name the metric and k: pass@1, pass@3 or pass^3.
- Give the tries per task (at least k) and the number of tasks.
- Use the unbiased estimator.
- Give a 95% interval. A count of tasks takes a Wilson interval.
- Say if the check is strict or lenient, and report format misses apart from wrong answers.
- Say when a set hits a ceiling: when every cell is n of n, no metric separates the models.
Frequently asked questions
What is the difference between pass@1 and pass@k?
The chance that one try passes is pass@1. The chance that at least one of k tries passes is pass@k. On our 8 hard tasks, Claude Haiku 4.5 passed 11/24 calls (pass@1 46%, 95% interval 28% to 65%). It passed 5/8 tasks at least once in 3 tries (pass@3 63%, a calculation, 95% interval 31% to 86%).
Is a higher k always better?
No. A higher k never lowers pass@k, but the model stays the same. Each extra try costs a call and needs a check. A high k can also hide inconsistency: Haiku passed the JSON prompt 1/10 times strictly, yet its pass@10 is 100% (calculation).
What is pass^k?
pass^k is the chance that all k tries pass. It measures reliability for work that nobody checks. For Haiku on the hard set, pass^3 was 2/8 tasks (25%, a calculation, 95% interval 7% to 59%), against 5/8 for pass@3.
How many samples do I need for pass@k?
You need at least k tries per task. We did not measure how the estimate changes with the number of tries. The number of tasks sets the width of the interval. With 8 tasks, a pass@3 of 5/8 still has a 95% interval of 31% to 86%.
Watch the data
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks
139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.
Transcript
- Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
- 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
- Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
- Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
- Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
- Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
- Open benchmarks: intervals, sources and every failure kept.