• Methodology
  • Benchmarks
  • Calibration
  • Failures

Why we count every failed attempt: our rules for honest AI benchmarks

How we benchmark AI agents: declare the protocol first, keep every failure, show intervals, publish our own confounds and label interim results.

TL;DR

  • Every number on our public benchmark pages comes from one dataset file, with a source for each figure.
  • We declare the protocol before the first run, take one attempt per task, and keep every failure, including the ones caused by our own platform.
  • We show n and a 95% interval wherever one exists. A gap inside overlapping intervals is not a ranking.
  • We label calculations as thought experiments, never as runs.
  • Unknown stays unknown. When we did not measure something, the chart says so.
  • New rules from our latest runs: a ceiling is not a ranking, we publish a confound we find in our own run, and an interim result says "interim" in its title.

The whole method on one page, with every protocol, counting rule and comparison row: /methodology. The data behind this post: /benchmarks/coding-calibration, plus every other study on /benchmarks.

The problem with most agent benchmarks

You have seen the pattern. A new agent posts a resolve rate a few points above the last one. The fine print, if there is any, says "best of 5 runs" or "on a curated subset" or "excluding infrastructure failures". Nobody shows the interval. Nobody shows the runs that went wrong.

Those numbers are not lies. They are just not useful for a decision. If you are choosing a tool for your team, you need to know what happens on an ordinary day, on the first try, when things break. So we wrote down rules, and we hold our own product to them.

Rule 1: declare the protocol before the first run

Before our SWE-bench campaign started, we wrote down the rules:

  • Escalations are graded on what was delivered.
  • The gold patch must resolve on our host before an instance counts.
  • A blocked instance is replaced with another from the same difficulty band.
  • No second attempts.

When six compiled-extension instances could not run in the worker checkout, the replacement rule applied. We did not pick replacements we liked. Later, we ran those compiled instances in a second campaign and published both. You can see every slice side by side.

  • Agent
  • Public panel mean, same instances
Campaign 1: 25-instance sample
Campaign 2: 8 compiled-extension instances
Original seed draw of 25
All 33 attempted

4 rows, 2 series: Agent, Public panel mean, same instances. Agent: highest Campaign 2: 8 compiled-extension instances 88% (95% interval 53%–98%, n 8). Lowest Campaign 1: 25-instance sample 72% (95% interval 52%–86%, n 25). All intervals overlap. Public panel mean, same instances: highest Campaign 2: 8 compiled-extension instances 84% (n 8). Lowest Original seed draw of 25 71% (n 25). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 8–33 per row

Agent resolved rate and 95% Wilson interval per declared view

The original draw and "all 33" mix two platform builds.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, SWE-bench campaign rules and sample design

Campaign 1 resolved 18 of 25 (72%). Campaign 2 resolved 7 of 8 (87.5%). The original seed draw, without replacements, resolved 19 of 25 (76%). All 33 attempted resolved 25 (75.8%). We could have led with the best view. Instead we lead with all 33, and the chart shows the rest.

The protocol also says what a control must show before the first run. In our coding-agent test, every base repository had to fail its hidden checks, and every reference patch had to pass all of them. Both held before the first session. Without that check, a pass could mean a broken test, not a fixed bug. The controls section of /methodology lists what each study holds fixed.

Read the task sets: pass rules, controls and recorded outcomes.

Rule 2: one attempt, every failure kept

"Best of N" hides the cost of the failures, and it hides how often you get a bad result. We take one attempt per task and keep it, whatever happens.

That includes failures that are our fault. Three of the SWE-bench empty patches were platform holds before delivery: missing lint tools, a plan gate that was too literal, and an unanswered question. A friendlier benchmark would drop them as "infrastructure". A user would have seen a failure, so we count a failure.

Our coding calibration is the clearest example. Three real upstream tasks, fastify/session, h3 and uvicorn, run on four platform builds. One attempt per task per build.

  • Functional pass (offline gates)
  • Verified delivery
  • Pull request opened
Baseline (capped)
Fix wave 1 (capped)
Uncapped, build f0ac3a8a
Uncapped, build 236c0d3f

One square per item of n; 3 strips per row, one per class.

4 rows, 3 series: Functional pass (offline gates), Verified delivery, Pull request opened. Functional pass (offline gates): highest Baseline (capped) 2 (n 3). Lowest Uncapped, build 236c0d3f 1 (n 3). Verified delivery: highest Uncapped, build 236c0d3f 1 (n 3). Lowest Uncapped, build f0ac3a8a 0 (n 3).

Notesn = 3 per row

Tasks per slice: functional pass, verified delivery, pull request opened

One attempt per task per slice. The two capped slices stopped at 20 minutes or $5; the uncapped slices had no ceiling. Each slice is a different platform build, so a change is not a matched improvement.

Source: Coding calibration: fastify/session, h3, uvicorn

The honest summary is not flattering:

  • Verified delivery went from 0 of 3 to 1 of 3 across four builds.
  • The latest build delivered fastify/session cleanly, in 11.2 minutes.
  • On the same build, h3 regressed to an empty patch. A formatting failure was lost across a context fold.
  • uvicorn passed its full suite but stayed unverified for two gate-classification reasons.
  • The first calibration run produced a passing patch for fastify/session but stopped at its $5 cap before delivery, at $4.89.

We publish this because it is the kind of result that tells you where a tool really is. A single cherry-picked success would tell you nothing.

Rule 3: failed attempts stay in the cost

When we report cost per resolved task, every attempt's cost goes in the numerator. A failed attempt still burns tokens. On SWE-bench, that puts Agent's notional cost at $2.81 per attempt and $3.71 per resolved instance. We put that number next to the public panel, where it does not look cheap. See what one resolved SWE-bench task really costs.

The same rule applies in the model head-to-head. A configuration that fails more often pays more per passing answer.

Rule 4: show the interval, and believe it

A sample of 33 cannot tell 76% from 85%. A sample of 12 cannot tell 75% from 50%. So every rate we publish carries its n and, where one exists, a 95% Wilson interval. When two intervals overlap, we say the data does not rank the systems.

The blind review study shows why this matters.

First scored attempt
Latest attempt
Public OSS tasks, latest
Private tasks, latest

Every interval overlaps every other: this chart does not order these rows.

4 rows. Highest Public OSS tasks, latest 80% (95% interval 38%–96%, n 5). Lowest First scored attempt 50% (95% interval 25%–75%, n 12). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 5–12 per row

Share of tasks where the blind panel preferred the AI change

Later attempts had the earlier attempts’ lessons, review replays and, on some tasks, operator answers. They are not independent first tries.

Source: Blind review panel: AI worker change vs merged human change

The headline "critics preferred the AI change on 9 of 12 tasks" is true. Its interval runs from 47% to 91%. The first-attempt figure, 6 of 12, runs from 25% to 75%. We publish both, and we tell you the first attempt is the cleaner estimate, because later attempts had lessons and operator answers. More in do blind AI critics prefer AI pull requests?

For paired comparisons, we go one step further. In the routing study, Sonnet 5.5 beat Jev on the point estimate, 77 against 74 of 82. An exact McNemar test on the same cases gives p = 0.375. We report a tie.

Rule 5: label calculations as calculations

Some of our most interesting charts are not runs. "What if every call ran on Opus?" multiplies recorded tokens by another price list. That is useful for price sensitivity, and useless for predicting outcomes, because a different model would take a different path. Every such chart says "calculation, not a run" in its title or note. See what if every call ran on Opus?

Costs for subscription calls are also calculations. We say "notional" and "list price", never "we paid".

Rule 6: unknown stays unknown

In the first version of the routing study we had not timed Jev, so the latency chart had no Jev point and a note that said so. A live run of 246 calls later filled the gap, with its limits written beside it: one Mac, a direct API call and not a CLI, one 35-second window. A local Clef router has not been run, so the table says "not measured" and gives the reason. We would rather show a gap than fill it with a guess, and fill it when we have the data.

Rule 7: disclose the home advantage

Benchmarks drift toward the system that built them. Our routing case sets were revised against Jev answers. Our blind review critics are mostly from the same model family as the worker. Our calibration tasks were chosen by us. Each of those biases is written in the study's caveats, in plain words, next to the result.

Rule 8: a refusal is data too

The platform has guardrails that turn back tool calls: a plan that breaks the schema, a delivery that breaks the contract, outgoing text that breaks a rule, a loop. We count those too.

Baseline (capped)
Fix wave 1 (capped)
Uncapped, build f0ac3a8a
Uncapped, build 236c0d3f

Hover or focus a bar for its ratio to Uncapped, build 236c0d3f (the highlighted row): a ratio of the two values shown, not a measurement.

4 rows. Highest Fix wave 1 (capped) 28. Lowest Uncapped, build 236c0d3f 19.

Notes

Tool calls the platform turned back (plan schema, delivery contract, outgoing text, loops)

A refusal is a guardrail working, not always a failure: some catch real problems, some were platform defects that the next build fixed.

Source: Coding calibration: fastify/session, h3, uvicorn

Refusals went from 26 on the baseline to 28, 25 and then 19 on the latest build. A refusal is not always a failure. Some catch real problems. Some were platform defects that the next build fixed. Publishing the count lets you see whether a platform is learning or just getting louder.

Rule 9: a ceiling is not a ranking

When every option passes every task, the test has hit its ceiling. It cannot rank the options, however good the result looks.

Every rate is 95% or more
Claude Sonnet 5.5
Claude Opus 5.5
GPT-6.1 Sol (medium, tester’s…

Every interval overlaps every other: this chart does not order these rows.

3 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 12 per row3 of 3 at 100%: this task set cannot separate them.

A pass needs every hidden check · 12 sessions per agent · 95% Wilson intervals

6 tasks × 2 repetitions per agent. All 36 sessions passed, so the chart sits at its ceiling: this task set cannot rank the agents on quality. Before the first session, every base repository failed its hidden checks and every reference patch passed them.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

In the coding-agent test, all 36 sessions passed: 12 of 12 for each agent. A perfect 12 of 12 has a 95% interval of 76% to 100%, so the three agents sit in the same band. The chart says so in its note, and every pass-rate row on the comparison pages is a tie. We then report what does differ (time, tool calls, diff size, cost) and say clearly that it is not quality. The effort ladder hit the same ceiling: 176 of 176 calls passed.

Rule 10: publish the confound you find in your own run

Sometimes the analysis finds a problem that the protocol missed. Then the problem goes on the page, next to the result.

In the coding-agent test, we ran Codex CLI with --ignore-user-config. The session logs showed that it still read the tester's global AGENTS.md in 12 of 12 sessions. 10 of 12 wrote a work log that no task asked for, and 9 reported a commit attempt. No Claude Code session did this.

We did not drop the Codex lane, and we did not rerun it quietly. The confound is the study's first caveat, and every Codex label carries the qualifier "tester's AGENTS.md", so every fact and comparison row built from it shows the confound. The pass results stand, because hidden tests graded the code. The time and diff numbers carry the caveat.

Rule 11: an interim result says "interim"

Long runs stop: usage limits, outages, a broken grader. If a team publishes only when a result looks good, the reader never sees the runs that stopped. So when we publish before a run ends, we follow fixed rules:

  • The title, tags and answer say "interim".
  • Only graded work counts, and the study text follows the graded count.
  • The declared set stays visible. Instances that have not started stay in the table, marked "not started".

Our Opus vs Sonnet probe on SWE-bench Verified is the first study under these rules. A usage gate stopped it after 3 of 8 declared pairs. The page shows all 8, the 2-to-1 result, the exact McNemar p of 1.0 and the build difference between the arms. It also says that even all 8 pairs cannot reach significance on their own.

From protocol to page: how a number gets here

Every number on our public pages travels the same path:

  1. Run. Each run writes receipts: the model, the route, the tokens, the time, the cost and the validator result.
  2. Extract. A script copies a sanitized extract into the repository. Prompts and model output are never copied. Local paths, private repository names and token-like strings are refused.
  3. Compile. A compiler with no network access turns the extracts into one dataset file. Anyone can rebuild it from the repository.
  4. Validate. Tests fail the build when a rate has no n or interval, when a calculation lacks its label, or when a comparison row names a winner that the intervals do not support.
  5. Publish. The study pages, the comparison pages, the model pages, the leaderboard and these posts all read that one file. A post can embed a chart by its id, so the post and the study can never disagree.

The methodology page draws this pipeline and adds the details: the counting rules, the three kinds of span (a 95% interval, a run range, a median-to-p95 span), every comparison row by verdict, the honesty rules as a checklist, and the raw data for each source.

How we measured

  • One public dataset, compiled from raw run extracts with no network access, holds every number on our benchmark pages and in these posts.
  • Each study lists its sources, method and caveats. Each chart lists its sources.
  • Rates use 95% Wilson intervals. Paired router comparisons use an exact McNemar test.
  • Calibration: fixed model claude-sonnet-5-5, routing off, real cold onboarding, no review replay, no operator answers, offline gates against the merged reference.
  • Coding agents: 6 hidden-test tasks × 2 repetitions × 3 agents, controls checked before the first session, nothing retried. Paired SWE-bench probe: 8 instances declared before any Opus run, exact McNemar on the graded pairs.

Caveats

  • Small samples everywhere. These are defect-finding runs, not leaderboard entries.
  • Builds change between slices. In the calibration, each slice is a different platform build, and the last two remove the caps. No two slices are a matched comparison.
  • We choose the tasks. Public OSS task names are published. Private tasks carry neutral labels.

Hold us to it

If you find a number on our pages that does not trace back to the data, tell us. And if you want an AI worker that shows its receipts, failures included, try Agent.

The data behind this post

  • Calibration
  • Coding Agents

Coding calibration: what broke on three real pull requests

One attempt per task, four platform builds, failures kept: how an AI worker did on real fastify/session, h3 and uvicorn issues, and what broke.

1 of 3Verified deliveries, latest build · n = 3

4 chartsUpdated October 5, 2026

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

  • Code Review
  • AI vs human

AI pull requests vs merged human pull requests, judged blind

A blind panel of Claude and GPT critics preferred Agent's change over the merged human change on 9 of 12 real tasks. Votes, scores, caveats.

75% (9/12)Tasks where the panel preferred the AI change (latest attempt) · n = 12

4 chartsUpdated October 5, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.