Why we count every failed attempt: our rules for honest AI benchmarks
How we benchmark AI agents: declare the protocol first, keep every failure, show intervals, publish our own confounds and label interim results.
TL;DR
- Every number on our public benchmark pages comes from one dataset file, with a source for each figure.
- We declare the protocol before the first run, take one attempt per task, and keep every failure, including the ones caused by our own platform.
- We show n and a 95% interval wherever one exists. A gap inside overlapping intervals is not a ranking.
- We label calculations as thought experiments, never as runs.
- Unknown stays unknown. When we did not measure something, the chart says so.
- New rules from our latest runs: a ceiling is not a ranking, we publish a confound we find in our own run, and an interim result says "interim" in its title.
The whole method on one page, with every protocol, counting rule and comparison row: /methodology. The data behind this post: /benchmarks/coding-calibration, plus every other study on /benchmarks.
The problem with most agent benchmarks
You have seen the pattern. A new agent posts a resolve rate a few points above the last one. The fine print, if there is any, says "best of 5 runs" or "on a curated subset" or "excluding infrastructure failures". Nobody shows the interval. Nobody shows the runs that went wrong.
Those numbers are not lies. They are just not useful for a decision. If you are choosing a tool for your team, you need to know what happens on an ordinary day, on the first try, when things break. So we wrote down rules, and we hold our own product to them.
Rule 1: declare the protocol before the first run
Before our SWE-bench campaign started, we wrote down the rules:
- Escalations are graded on what was delivered.
- The gold patch must resolve on our host before an instance counts.
- A blocked instance is replaced with another from the same difficulty band.
- No second attempts.
When six compiled-extension instances could not run in the worker checkout, the replacement rule applied. We did not pick replacements we liked. Later, we ran those compiled instances in a second campaign and published both. You can see every slice side by side.
Campaign 1 resolved 18 of 25 (72%). Campaign 2 resolved 7 of 8 (87.5%). The original seed draw, without replacements, resolved 19 of 25 (76%). All 33 attempted resolved 25 (75.8%). We could have led with the best view. Instead we lead with all 33, and the chart shows the rest.
The protocol also says what a control must show before the first run. In our coding-agent test, every base repository had to fail its hidden checks, and every reference patch had to pass all of them. Both held before the first session. Without that check, a pass could mean a broken test, not a fixed bug. The controls section of /methodology lists what each study holds fixed.
Read the task sets: pass rules, controls and recorded outcomes.
Rule 2: one attempt, every failure kept
"Best of N" hides the cost of the failures, and it hides how often you get a bad result. We take one attempt per task and keep it, whatever happens.
That includes failures that are our fault. Three of the SWE-bench empty patches were platform holds before delivery: missing lint tools, a plan gate that was too literal, and an unanswered question. A friendlier benchmark would drop them as "infrastructure". A user would have seen a failure, so we count a failure.
Our coding calibration is the clearest example. Three real upstream tasks, fastify/session, h3 and uvicorn, run on four platform builds. One attempt per task per build.
The honest summary is not flattering:
- Verified delivery went from 0 of 3 to 1 of 3 across four builds.
- The latest build delivered fastify/session cleanly, in 11.2 minutes.
- On the same build, h3 regressed to an empty patch. A formatting failure was lost across a context fold.
- uvicorn passed its full suite but stayed unverified for two gate-classification reasons.
- The first calibration run produced a passing patch for fastify/session but stopped at its $5 cap before delivery, at $4.89.
We publish this because it is the kind of result that tells you where a tool really is. A single cherry-picked success would tell you nothing.
Rule 3: failed attempts stay in the cost
When we report cost per resolved task, every attempt's cost goes in the numerator. A failed attempt still burns tokens. On SWE-bench, that puts Agent's notional cost at $2.81 per attempt and $3.71 per resolved instance. We put that number next to the public panel, where it does not look cheap. See what one resolved SWE-bench task really costs.
The same rule applies in the model head-to-head. A configuration that fails more often pays more per passing answer.
Rule 4: show the interval, and believe it
A sample of 33 cannot tell 76% from 85%. A sample of 12 cannot tell 75% from 50%. So every rate we publish carries its n and, where one exists, a 95% Wilson interval. When two intervals overlap, we say the data does not rank the systems.
The blind review study shows why this matters.
The headline "critics preferred the AI change on 9 of 12 tasks" is true. Its interval runs from 47% to 91%. The first-attempt figure, 6 of 12, runs from 25% to 75%. We publish both, and we tell you the first attempt is the cleaner estimate, because later attempts had lessons and operator answers. More in do blind AI critics prefer AI pull requests?
For paired comparisons, we go one step further. In the routing study, Sonnet 5.5 beat Jev on the point estimate, 77 against 74 of 82. An exact McNemar test on the same cases gives p = 0.375. We report a tie.
Rule 5: label calculations as calculations
Some of our most interesting charts are not runs. "What if every call ran on Opus?" multiplies recorded tokens by another price list. That is useful for price sensitivity, and useless for predicting outcomes, because a different model would take a different path. Every such chart says "calculation, not a run" in its title or note. See what if every call ran on Opus?
Costs for subscription calls are also calculations. We say "notional" and "list price", never "we paid".
Rule 6: unknown stays unknown
In the first version of the routing study we had not timed Jev, so the latency chart had no Jev point and a note that said so. A live run of 246 calls later filled the gap, with its limits written beside it: one Mac, a direct API call and not a CLI, one 35-second window. A local Clef router has not been run, so the table says "not measured" and gives the reason. We would rather show a gap than fill it with a guess, and fill it when we have the data.
Rule 7: disclose the home advantage
Benchmarks drift toward the system that built them. Our routing case sets were revised against Jev answers. Our blind review critics are mostly from the same model family as the worker. Our calibration tasks were chosen by us. Each of those biases is written in the study's caveats, in plain words, next to the result.
Rule 8: a refusal is data too
The platform has guardrails that turn back tool calls: a plan that breaks the schema, a delivery that breaks the contract, outgoing text that breaks a rule, a loop. We count those too.
Refusals went from 26 on the baseline to 28, 25 and then 19 on the latest build. A refusal is not always a failure. Some catch real problems. Some were platform defects that the next build fixed. Publishing the count lets you see whether a platform is learning or just getting louder.
Rule 9: a ceiling is not a ranking
When every option passes every task, the test has hit its ceiling. It cannot rank the options, however good the result looks.
In the coding-agent test, all 36 sessions passed: 12 of 12 for each agent. A perfect 12 of 12 has a 95% interval of 76% to 100%, so the three agents sit in the same band. The chart says so in its note, and every pass-rate row on the comparison pages is a tie. We then report what does differ (time, tool calls, diff size, cost) and say clearly that it is not quality. The effort ladder hit the same ceiling: 176 of 176 calls passed.
Rule 10: publish the confound you find in your own run
Sometimes the analysis finds a problem that the protocol missed. Then the problem goes on the page, next to the result.
In the coding-agent test, we ran Codex CLI with --ignore-user-config. The session logs showed that it still read the tester's global AGENTS.md in 12 of 12 sessions. 10 of 12 wrote a work log that no task asked for, and 9 reported a commit attempt. No Claude Code session did this.
We did not drop the Codex lane, and we did not rerun it quietly. The confound is the study's first caveat, and every Codex label carries the qualifier "tester's AGENTS.md", so every fact and comparison row built from it shows the confound. The pass results stand, because hidden tests graded the code. The time and diff numbers carry the caveat.
Rule 11: an interim result says "interim"
Long runs stop: usage limits, outages, a broken grader. If a team publishes only when a result looks good, the reader never sees the runs that stopped. So when we publish before a run ends, we follow fixed rules:
- The title, tags and answer say "interim".
- Only graded work counts, and the study text follows the graded count.
- The declared set stays visible. Instances that have not started stay in the table, marked "not started".
Our Opus vs Sonnet probe on SWE-bench Verified is the first study under these rules. A usage gate stopped it after 3 of 8 declared pairs. The page shows all 8, the 2-to-1 result, the exact McNemar p of 1.0 and the build difference between the arms. It also says that even all 8 pairs cannot reach significance on their own.
From protocol to page: how a number gets here
Every number on our public pages travels the same path:
- Run. Each run writes receipts: the model, the route, the tokens, the time, the cost and the validator result.
- Extract. A script copies a sanitized extract into the repository. Prompts and model output are never copied. Local paths, private repository names and token-like strings are refused.
- Compile. A compiler with no network access turns the extracts into one dataset file. Anyone can rebuild it from the repository.
- Validate. Tests fail the build when a rate has no n or interval, when a calculation lacks its label, or when a comparison row names a winner that the intervals do not support.
- Publish. The study pages, the comparison pages, the model pages, the leaderboard and these posts all read that one file. A post can embed a chart by its id, so the post and the study can never disagree.
The methodology page draws this pipeline and adds the details: the counting rules, the three kinds of span (a 95% interval, a run range, a median-to-p95 span), every comparison row by verdict, the honesty rules as a checklist, and the raw data for each source.
How we measured
- One public dataset, compiled from raw run extracts with no network access, holds every number on our benchmark pages and in these posts.
- Each study lists its sources, method and caveats. Each chart lists its sources.
- Rates use 95% Wilson intervals. Paired router comparisons use an exact McNemar test.
- Calibration: fixed model claude-sonnet-5-5, routing off, real cold onboarding, no review replay, no operator answers, offline gates against the merged reference.
- Coding agents: 6 hidden-test tasks × 2 repetitions × 3 agents, controls checked before the first session, nothing retried. Paired SWE-bench probe: 8 instances declared before any Opus run, exact McNemar on the graded pairs.
Caveats
- Small samples everywhere. These are defect-finding runs, not leaderboard entries.
- Builds change between slices. In the calibration, each slice is a different platform build, and the last two remove the caps. No two slices are a matched comparison.
- We choose the tasks. Public OSS task names are published. Private tasks carry neutral labels.
What to read next
- The full methodology: pipeline, protocols, counting rules and every comparison row
- Claude Code vs Codex CLI on hidden tests: when every agent passes, what differs?
- Opus 5.5 vs Sonnet 5.5 on real GitHub issues: an early look
- An AI model leaderboard without a composite score
- Same prompt, ten answers: consistent is not correct
- SWE-bench Verified: Agent vs 11 public models on the same tasks
- Harness vs model: where the gains come from
- What one resolved SWE-bench task really costs
Hold us to it
If you find a number on our pages that does not trace back to the data, tell us. And if you want an AI worker that shows its receipts, failures included, try Agent.