Explainer · Reading an AI benchmark honestly

How to read AI benchmarks honestly

Definition

Reading an AI benchmark honestly means checking what was measured, on how many cases, under which configuration and with what uncertainty before you accept a ranking. A trustworthy result states its sample size, gives an interval for every rate and a range for every time, counts every failed attempt, separates measured runs from calculations, and names the setup (model, effort, route, harness) behind each number. A table that hides any of these can make a tie look like a win.

Agent team · · 4 min read · Every number is from the public studies

Interactive

Interval playground: when is a gap real?

Move n and k. Each 95% Wilson interval narrows as the runs add up; overlapping intervals are a tie.

Calculated live from your inputs
Load a published result
A: Jev 1.13

90%74/82

95% Wilson 82%–95%

B: Claude Sonnet 5.5

94%77/82

95% Wilson 87%–97%

Tie: the 95% intervals overlap (87%–95%), so this data cannot separate Jev 1.13 and Claude Sonnet 5.5.

Interval width as n grows, at 90%

±7 points at n = 82

Two configurations: Jev 1.13 74 of 82 (90%, 95% interval 82% to 95%) and Claude Sonnet 5.5 77 of 82 (94%, 87% to 97%). Tie: the 95% intervals overlap (87%–95%), so this data cannot separate Jev 1.13 and Claude Sonnet 5.5.

The interval is the Wilson score interval, the formula behind every rate on this site. Values you set here are a calculation, not a result. The preset loads published counts: see the study.

Source: Methodology: intervals

1. Find n

Every rate is k out of n. "100%" can mean 3 of 3 or 500 of 500. Look for the count of tasks, calls or instances behind each number. If a page does not show it, the number cannot be weighed.

2. Look for the interval, and whether intervals overlap

A rate from a sample has an uncertainty. A 95% Wilson interval shows it. When two intervals overlap, the data does not separate the two systems.

On our eight hard tasks, six configurations passed every call, and their intervals reach down to 86% (24 calls) or 81% (16 calls). They tie. One configuration, Claude Haiku 4.5, passed 11 of 24 strictly, with an interval of 28% to 65% that overlaps none of the others. That gap is real on this task set. See Wilson confidence intervals.

3. Count every failure, and know what kind it was

A benchmark that drops timeouts, crashes or refusals reports a better rate than the system earned. It also matters what kind of failure it was:

  • Strict pass
  • Format miss (correct answer, wrong format)
  • Wrong answer
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

One square per call; counts at the right are exact and in legend order.

7 rows, 3 series: Strict pass, Format miss (correct answer, wrong format), Wrong answer. Strict pass: highest Claude Sonnet 5.5 · Claude Code 24 (n 24). Lowest Claude Haiku 4.5 · Claude Code 11 (n 24). Format miss (correct answer, wrong format): highest Claude Haiku 4.5 · Claude Code 5 (n 24). Lowest GPT-6.1 Sol (high) · Codex CLI 0 (n 16).

Notesn 16–24 per row

Counts per configuration: strict passes, format misses and wrong answers

A format miss is a reply that the strict validator rejected (for example, wrapped in a code fence or in prose although the prompt said not to) whose extracted answer passes the same validator. It is not a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Claude Haiku 4.5 had 11 strict passes, 5 format misses (a correct answer in the wrong format, for example inside a code fence) and 8 wrong answers. A format miss is not a pass, but it is a different problem from a wrong answer: a validator or prompt fix may cure it. Read why we count every failed attempt.

4. For times, read the range, not only the median

Latency varies a lot from call to call. A median alone hides that:

Entrance: medians race at 4.5× real timeMotion reduced: press Replay to animateThe slowest median is 6.3 s. The clock runs at the recorded speed.
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

9 rows. Slowest GPT-6.1 Sol (low) · Codex CLI 6.3 s (range 4.7 s–10.5 s, n 10). Fastest Claude Fable 5.1 · Claude Code 1.9 s (range 1.4 s–9.8 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 10–15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

On short tasks, the configuration with the lowest median had 1.9 s per call (n = 15), but its slowest call took 9.8 s. The whiskers here are the fastest and slowest call, which is a range, not a confidence interval. When ranges overlap, "faster" describes this run, not a tested ranking.

5. Separate runs from calculations

Many cost figures are not measured bills. They multiply recorded tokens by a list price. That is a fair calculation, but it is not a run, and a repricing of one model's tokens at another model's price is a thought experiment:

Calculation
Largest value is 310x the smallest; Log shows the small bars.
Claude Fable 5.1
Claude Opus 5
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol
Claude Haiku 4.5
Gemini 3.x Flash
Jev 1.13 (router)

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 8 rows. Highest Claude Fable 5.1 $12.85. Lowest Jev 1.13 (router) $0.042.

Notes

Cost per resolved SWE-bench instance if 162.9M input and 1.8M output tokens had been billed at each model's list price

Calculation, not a run: tokens recorded by Agent on claude-sonnet-5-5 (33 attempts, 25 resolved) times list prices effective 2026-09-21. Another model would use a different number of tokens and resolve a different set. Jev is a routing model and cannot do this work; its bar is a price floor only.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models), Google Gemini list prices, OpenAI list prices, Jev 1.13 list price

This chart prices the tokens Agent recorded on 33 SWE-bench attempts at other models' list prices. Another model would use different tokens and resolve a different set, so these figures bound price sensitivity; they do not predict outcomes. A good page labels each such chart as a calculation.

6. Check the configuration behind each number

The same model gives different results at different efforts, through a CLI or an API, inside different harnesses. "Model A beats model B" is only meaningful when both ran the same tasks, with the same validators, through comparable routes. When routes differ, the gap may be the route. Our comparison rows show both contexts for this reason.

7. Beware the single overall score

A composite score adds numbers that measure different things: pass rates on different task sets, times through different routes, list-price calculations. It needs weights nobody measured, and it hides the configuration behind each part. That is why our leaderboard has no composite score: it shows each system's best-supported values per category, with study, n and interval, and lets you sort on one metric at a time.

8. Ask what the tasks were

  • Ceiling: if every strong model scores 100%, the set cannot separate them.
  • Contamination: public tasks and fixes may be in training data.
  • Fit: a benchmark of short puzzles says little about long agent work, and the reverse.
  • Home advantage: a set tuned against one system favors it.

A short checklist

  • Is n shown for every number?
  • Does every rate have an interval, and every time a range?
  • Are failures, timeouts and format misses counted?
  • Are calculations labelled as calculations?
  • Is the configuration (model, effort, route, harness) named?
  • Do the compared systems share the same tasks and validators?
  • Is a "winner" claimed only where intervals or ranges do not overlap?

Frequently asked questions

Why do AI benchmark results disagree with each other?

They measure different tasks, through different harnesses and routes, with different sample sizes and attempt policies. Two honest benchmarks can rank the same models differently because they ask different questions.

What sample size makes an AI benchmark trustworthy?

It depends on the gap. A 16-of-16 result still has a 95% interval of about 81% to 100%, so small samples only separate large gaps. Look for the interval rather than a rule of thumb.

Should I trust a leaderboard with one overall score?

Use it with care. A composite mixes metrics from different tasks and setups with chosen weights. Look at the single metric that matches your work, with its n and interval.

Is a list-price cost per task a measured cost?

Usually not. It multiplies recorded tokens by published prices. It is useful for comparison, but it is a calculation, and it ignores discounts, subscriptions and the tokens another model would have used.

The data behind this explainer

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.