Explainer · Benchmark saturation
Benchmark saturation: when every model scores 100%
Definition
Benchmark saturation means most or all models score near the top of a test, so scores stop telling them apart. This is the ceiling effect. A perfect score means every observed call passed. It does not prove equal ability. Our sets hit ceilings while time, tokens and calculated list-price cost still differed.
Agent team · · 5 min read · Every number is from the public studies
Every interval overlaps every other: this chart does not order these rows.
| Item | Strict pass | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 (low) · Claude Code | 100% | 81%–100% | 16 |
| Claude Sonnet 5.5 (medium) · Claude Code | 100% | 81%–100% | 16 |
| Claude Sonnet 5.5 (high) · Claude Code | 100% | 81%–100% | 16 |
| Claude Sonnet 5.5 · Claude Code | 100% | 81%–100% | 16 |
| Claude Opus 5.5 (low) · Claude Code | 100% | 81%–100% | 16 |
| Claude Opus 5.5 (medium) · Claude Code | 100% | 81%–100% | 16 |
| Claude Opus 5.5 (high) · Claude Code | 100% | 81%–100% | 16 |
| Claude Opus 5.5 · Claude Code | 100% | 81%–100% | 16 |
| GPT-6.1 Sol (low) · Codex CLI | 100% | 81%–100% | 16 |
| GPT-6.1 Sol (medium) · Codex CLI | 100% | 81%–100% | 16 |
| GPT-6.1 Sol (high) · Codex CLI | 100% | 81%–100% | 16 |
11 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 16 per row11 of 11 at 100%: this task set cannot separate them.
Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.
Source: Effort ladder: the hard task set at each effort level
The signs of a saturated set
- Perfect cells. All 11 configurations in our effort ladder passed 16/16 each (95% interval 81% to 100%). In our consistency study, 7 of 9 model-and-prompt cells passed 10/10 each (72% to 100%). Haiku 4.5's other 2 cells scored 0/10 (0% to 28%) and 1/10 (2% to 40%). Their intervals do not overlap the perfect cells on the same prompts.
- Overlapping intervals. Even perfect cells have wide 95% Wilson intervals. They do not establish equal ability.
- Only format misses. Our five-task study had 127/130 passes (98%; 95% interval 93.4% to 99.2%). All 3 non-passes came from one configuration on one task. Right answers had extra working lines, which the exact-text validator rejects.
Our CLI latency study also hit a ceiling: 194/194 evaluated runs passed (95% interval 98.1% to 100%). It excludes 18 runs and 18 diagnostic receipts. Pass rate cannot separate these runs.
Our own sets hit the ceiling
On five short tasks, 8 of 9 configurations passed every call: seven scored 15/15 (95% interval 80% to 100%); one scored 10/10 (72% to 100%).
The hard head-to-head followed with eight harder tasks. Still, 6 of 7 configurations passed every call: four scored 24/24 (95% interval 86% to 100%); two scored 16/16 (81% to 100%).
Haiku 4.5 scored 11/24 strictly (46%; 28% to 65%). Its interval overlaps none of the six. The set separates Haiku from them, but cannot separate those six.
The effort ladder reused those 8 tasks across 11 configurations: Sonnet 5.5, Opus 5.5 and GPT-6.1 Sol at different efforts. Six were new; five reused hard-set cells. All scored 16/16 strictly (81% to 100% each). Pooled: 176/176 (95% interval 97.9% to 100%), with 0 format misses and 0 wrong answers. This combines configurations; it is not any one model's interval.
What a ceiling hides
- Uncertainty. A 16/16 interval reaches down to 81%. Calculation: about 19 percentage points below the observed perfect score. This describes one rate, not the difference between models. “No difference found” does not mean “no difference”. More calls help slowly: 24/24 reaches down to 86%.
- Weak spots. Haiku's easy-set 15/15 (95% interval 80% to 100%) overlapped every configuration. The hard set exposed a gap: 11/24 strictly (46%; 28% to 65%). Lenient grading accepts its 5 format misses: 16/24 (47% to 82%).
- Hidden splits. In our memory study, Sonnet 5.5 scored 57/60 or 60/60 on rules the code shows, across conditions. Their 95% intervals are 86% to 98% and 94% to 100%. Team-only checks differed: 6/15 without memory (40%; 20% to 64%), versus 15/15 with an 11-line curated file (80% to 100%). Those intervals do not overlap. Pooling hides the split.
- Selection. In SWE-bench Verified, 14 instances fall in “every panel model solved it”. Each of the 11 panel models scores 14/14 by definition (descriptive 95% Wilson interval 78% to 100%). Bands come from the panel's results. Selection guarantees perfection here; it does not estimate performance on new tasks.
What still differed between configurations
Time, tokens and calculated list-price cost differed at the ceiling:
- Time. Effort-ladder Sonnet 5.5 medians were 5.8 s at low effort and 8.8 s at high (n = 16 each). Single-call ranges: 2.8 to 20 s and 2.9 to 35.8 s. These overlap; medians describe this run, not a ranking. Ranges are not confidence intervals.
- Tokens. The same cells' output medians were 667 and 1,192 tokens (n = 16 each). Single-call ranges overlapped: 176 to 2,263 and 220 to 4,187 tokens.
- Cost. Calculation: all 16 calls' reported tokens × list price, divided by 16 strict passes. Cache reads and writes use separate rates. Low effort cost $0.0122 per pass; high cost $0.0167. Calculated per-call ranges: $0.0051 to $0.0261 and $0.0056 to $0.0453. Subscription calls make this a calculation, not a bill.
- Hard set. The six perfect configurations had different latency and token medians, and calculated costs. Their per-call time ranges overlap. Sonnet's median was 7.7 s (range 2.3 to 34.8 s; n = 24). GPT-6.1 Sol at high effort: 18.1 s (range 11.7 to 92.2 s; n = 16). These are call ranges, not confidence intervals.
Tools were off, on one host. Each CLI adds its prompt and start-up time. Hard-set routes ran on different days; reused ladder cells ran in a different hour. These compare model-and-CLI configurations.
Repeated calls on fixed tasks, and checks within one memory session, are not independent samples of new tasks. Wilson intervals describe these fixed sets.
What to do when a set saturates
- Pilot harder tasks; keep ones models pass only sometimes. This is design advice; our dataset does not test it.
- Show per-task results beside pooled rates.
- Separate format misses from wrong answers. Haiku's 13 hard-set non-passes: 5 format misses, 8 wrong answers.
- Report time, tokens and cost beside pass rate. Give time medians and ranges.
- Show n and the interval beside every 100%. Say “ceiling” when a set hits one.
Frequently asked questions
What does it mean when a benchmark is saturated?
Most models score at or near the maximum, so scores stop separating them. The effort ladder above shows this across 11 configurations. Its pooled interval describes all configurations together.
Does 100% mean two models are equal?
No. Perfect observed scores retain uncertainty. The calculation of about 19 points above concerns one rate, not the difference between cells. Equal scores show both clear this bar.
How do you make a benchmark harder?
Pilot tasks that models sometimes fail. Our eight-task follow-up separated Haiku from six perfect configurations, but could not separate those six. We did not test the pilot step.
What still separates models at the ceiling?
Time, tokens and calculated list-price cost can differ. The Sonnet and GPT-6.1 Sol time ranges above overlap. They are not confidence intervals; the comparison remains unclear, not a ranking.
Watch the data
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks
All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.
Transcript
- Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
- 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
- All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
- More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
- List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
- 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Start at low effort and measure. Every call, interval and cost online.