Public golden task sets

The tasks behind our benchmarks.
See the checks.

Read the task titles, the rules for a pass, the controls and the recorded outcomes. This catalog draws from the public studies. It runs no new tests and shows no prompts or model output.

“Golden” means a curated task and validator catalog. It does not mean an exhaustive test suite, a model ranking or proof that Agent completed a customer goal.

Task sets
8
Task entries
79can overlap across studies
Published check counts
228across 13 tasks
Reference controls passed
14/14only recorded controls

Read a cell before comparing it

  1. Read the pass rule. It differs by set. A strict answer, all hidden tests and verified delivery are separate criteria.
  2. Read the denominator. k/n counts recorded passes and attempts. One attempt is an outcome, not a reliable pass rate.
  3. Read the limits. Perfect cells can have a ceiling. Different builds, routes and public panels can be unmatched.

Unknown means not recorded. Public task IDs stay public. The catalog contains no task prompts, model replies or private run paths.

hard synthetic Updated 2026-10-06

8 hard tasks with strict validators

strict passes of n calls.

Back to sets ↑

What counts as a pass

Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 references wrapped in a fence or prose are flagged as format misses.

Strict pass: the whole reply, trimmed, passes the validator. Every prompt states the output format and says no code fence and no other text. This is stricter than the five-task study, which removed one wrapping fence.

Format miss: the strict check failed, but a lenient extractor (fenced block, outer JSON, first code line, one output line) finds an answer that passes the same validator. Reported apart from wrong answers, never as a pass.

  • Fix an interval-merge function (off-by-one and edge cases)18 validator checksReference passes5/5 plausible wrong answers rejectedWrapped reference flagged as a format miss
  • Fix a time-zone day-length function (DST)36 validator checksReference passes3/3 plausible wrong answers rejectedWrapped reference flagged as a format miss
  • Write a CSV parser (quoted newlines, strict errors)22 validator checksReference passes3/3 plausible wrong answers rejectedWrapped reference flagged as a format miss
  • Predict JavaScript event-loop output order1 validator checksReference passes3/3 plausible wrong answers rejectedWrapped reference flagged as a format miss
  • Solve a multi-constraint room schedule14 validator checksReference passes3/3 plausible wrong answers rejectedWrapped reference flagged as a format miss
  • Write a strict SemVer 2.0.0 regex30 validator checksReference passes3/3 plausible wrong answers rejectedWrapped reference flagged as a format miss
  • Refactor to remove duplication, keep 20 tests green25 validator checksReference passes2/2 plausible wrong answers rejectedWrapped reference flagged as a format miss
  • Write a SQLite reporting query (fan-out, ties, boundaries)8 validator checksReference passes4/4 plausible wrong answers rejectedWrapped reference flagged as a format miss

Counts from the public controls tableBars count checks from zeroMissing controls stay unknown

These controls test the validator before model or agent attempts. A reference pass alone does not prove that every wrong answer will fail.

TaskClaude Sonnet 5.5 · Claude Code1Claude Opus 5.5 · Claude Code2Claude Opus 5.5 (high) · Claude Code3GPT-6.1 Sol (medium) · Codex CLI4Claude Fable 5.1 · Claude Code5GPT-6.1 Sol (high) · Codex CLI6Claude Haiku 4.5 · Claude Code7
Fix an interval-merge function (off-by-one and edge cases)3/33/33/32/23/32/23/3
Fix a time-zone day-length function (DST)3/33/33/32/23/32/21/3
Write a CSV parser (quoted newlines, strict errors)3/33/33/32/23/32/22/3
Predict JavaScript event-loop output order3/33/33/32/23/32/20/3
Solve a multi-constraint room schedule3/33/33/32/23/32/20/3
Write a strict SemVer 2.0.0 regex3/33/33/32/23/32/23/3
Refactor to remove duplication, keep 20 tests green3/33/33/32/23/32/22/3
Write a SQLite reporting query (fan-out, ties, boundaries)3/33/33/32/23/32/20/3
  1. Claude Sonnet 5.5 · Claude Code
  2. Claude Opus 5.5 · Claude Code
  3. Claude Opus 5.5 (high) · Claude Code
  4. GPT-6.1 Sol (medium) · Codex CLI
  5. Claude Fable 5.1 · Claude Code
  6. GPT-6.1 Sol (high) · Codex CLI
  7. Claude Haiku 4.5 · Claude Code

Shade is the share k of n in each cell. * The cell has a note (hover or focus it).

Counts from the source tablen = 2–3 per cell

8 tasks across 7 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.

6 of 7 configurations passed every recorded task. This set cannot separate them on pass rate.

Source: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks · Pass matrix on hard tasks: configuration × task (strict passes) · Validator controls run before the first model call
Study limits and provenance

8 tasks across 7 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.

6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.

Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.

Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.

The Claude and Codex batches ran on different days on the same host, one call at a time per account. Each route used its own subscription.

Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.

CLI timings include CLI start-up and the CLI’s own system prompt. One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled.

Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.

List-price costs are calculations; the calls used a flat subscription.

These tasks also appear in effort ladder.

Source: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks · Pass matrix on hard tasks: configuration × task (strict passes) · Validator controls run before the first model call

coding agents Updated 2026-10-06

6 repository tasks with hidden tests

sessions that passed every hidden check, of n.

Back to sets ↑

What counts as a pass

Controls before the first session (Node v25.2.1): every base repository fails its hidden checks and every reference patch passes all of them.

Pass = every hidden check passes; no partial credit. Recorded per session: wall time, tool calls, tokens as reported, files touched, lines changed against the base commit, commits and edits outside the repository.

  • Pagination fix13 validator checksReference passes (tests 13/13)Starting repository fails (tests 1/13)
  • CLI --top flag11 validator checksReference passes (tests 11/11)Starting repository fails (tests 2/11)
  • Invoice refactor24 validator checksReference passes (tests 24/24)Starting repository fails (tests 21/24)
  • LRU cache16 validator checksReference passes (tests 16/16)Starting repository fails (tests 0/16)
  • Queue race10 validator checksReference passes (tests 10/10)Starting repository fails (tests 4/10)
  • Strict TypeScript typesstatic checks, tsc with a hidden usage file (11 @ts-expect-error probes) and 4 runtime testsReference passes (tests 4/4)Starting repository fails (tests 4/4; static checks fail)

Counts from the public controls tableBars count checks from zeroMissing controls stay unknown

These controls test the validator before model or agent attempts. A reference pass alone does not prove that every wrong answer will fail.

All 18 cells are n of n: this task set cannot separate them.

TaskSonnet 5.5 in Claude Code1Opus 5.5 in Claude Code2GPT-6.1 Sol in Codex CLI3
Pagination fix2/22/22/2
CLI --top flag2/22/22/2
Invoice refactor2/22/22/2
LRU cache2/22/22/2
Queue race2/22/22/2
Strict TypeScript types2/22/22/2
  1. Sonnet 5.5 in Claude Code
  2. Opus 5.5 in Claude Code
  3. GPT-6.1 Sol in Codex CLI

Every cell is at its maximum, so all cells share one calm shade.

Counts from the source tablen = 2 per cell

6 tasks across 3 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.

3 of 3 configurations passed every recorded task. This set cannot separate them on pass rate.

Source: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks · Every task and agent: sessions that passed (of 2) · The six tasks and their controls
Study limits and provenance

6 tasks across 3 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.

Codex CLI ran with the tester’s global AGENTS.md: --ignore-user-config does not switch that file off. 12 of 12 Codex sessions read the tester’s notes, 10 wrote a WORKLOG.md and 9 reported a commit attempt. Its time, tool calls and lines changed include that work. Claude Code ran with setting sources off, and no Claude session did any of it.

Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.

Each run pairs a CLI with a model (Claude Code with Claude models, Codex CLI with GPT-6.1 Sol), so the results cannot separate the CLI from the model.

n = 12 sessions per agent (2 per task). Time ranges are the fastest and slowest sessions, not confidence intervals; the two lanes shared one machine.

Tokens are as each CLI reports them, with different tokenizers and context handling: compare tokens inside Claude Code (Sonnet vs Opus), not across vendors. Costs are list-price calculations on subscription sessions, not invoices.

Source: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks · Every task and agent: sessions that passed (of 2) · The six tasks and their controls

short synthetic Updated 2026-10-05

5 short tasks with deterministic validators

passes of n calls.

Back to sets ↑

What counts as a pass

Protocol declared before the first call: 5 cases with deterministic validators (behavioral checks in a sandbox, canonical JSON or exact text).

No pre-run controls or validator check counts are recorded for this set. Read its pass rule below.

TaskClaude Fable 5.1 · Claude Code1Claude Sonnet 5.5 · Claude Code2Claude Opus 5.5 (high) · Claude Code3Claude Opus 5.5 · Claude Code4Claude Opus 5.5 (low) · Claude Code5Claude Haiku 4.5 · Claude Code6GPT-6.1 Sol (high) · Codex CLI7GPT-6.1 Sol (medium) · Codex CLI8GPT-6.1 Sol (low) · Codex CLI9
Fix a buggy median function3/33/33/33/33/33/33/33/32/2
Extract invoice fields to JSON3/33/33/33/33/33/33/33/32/2
Multi-step shift arithmetic3/30/33/33/33/33/33/33/32/2
Refactor recursion to iteration3/33/33/33/33/33/33/33/32/2
Classify six support tickets3/33/33/33/33/33/33/33/32/2
  1. Claude Fable 5.1 · Claude Code
  2. Claude Sonnet 5.5 · Claude Code
  3. Claude Opus 5.5 (high) · Claude Code
  4. Claude Opus 5.5 · Claude Code
  5. Claude Opus 5.5 (low) · Claude Code
  6. Claude Haiku 4.5 · Claude Code
  7. GPT-6.1 Sol (high) · Codex CLI
  8. GPT-6.1 Sol (medium) · Codex CLI
  9. GPT-6.1 Sol (low) · Codex CLI

Shade is the share k of n in each cell.

Counts from the source tablen = 2–3 per cell

5 tasks across 9 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.

8 of 9 configurations passed every recorded task. This set cannot separate them on pass rate.

Source: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head · Pass matrix: configuration × task
Study limits and provenance

5 tasks across 9 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.

The tasks are short and easy; pass rate saturates. Latency and tokens carry the signal. A harder follow-up with eight tasks and strict validators: /benchmarks/hard-model-head-to-head.

Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.

CLI timings include CLI start-up and the CLI’s own system prompt.

One host, one network, one day.

List-price costs are calculations; the calls used flat subscriptions.

The prompt cache stayed at the provider default, so cache counters differ by route and by call order.

Haiku 4.5 reported reasoning tokens on most calls under the CLI default, which explains much of its extra time and output.

Haiku and Fable are not in the platform runner catalog, so these calls used the runner’s CLI functions directly; each receipt records the model the CLI reported.

Source: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head · Pass matrix: configuration × task

agent memory Updated 2026-10-06

5 tasks for agent memory

full passes of 3 sessions.

Back to sets ↑

What counts as a pass

Grading after the session: hidden tests run in a sandbox, plus deterministic checks on the lines the agent added. Full pass = tests and every applicable check. Every attempt counts.

No pre-run controls or validator check counts are recorded for this set. Read its pass rule below.

TaskNo memory1/init CLAUDE.md2Curated, 11 lines3Raw notes, 60 lines4Dreamed notes5Handbook, 210 lines6Stop hook only7Curated + hook8
Refunds3/33/33/32/33/33/33/33/3
Report decimals0/30/33/33/33/33/33/33/3
Customer phone3/33/33/33/33/33/33/33/3
Late fee0/30/33/33/33/33/30/33/3
Slugify (control)3/33/33/33/33/33/33/33/3
  1. No memory
  2. /init CLAUDE.md
  3. Curated, 11 lines
  4. Raw notes, 60 lines
  5. Dreamed notes
  6. Handbook, 210 lines
  7. Stop hook only
  8. Curated + hook

Shade is the share k of n in each cell.

Counts from the source tablen = 3 per cell

5 tasks across 8 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.

4 of 8 configurations passed every recorded task. This set cannot separate them on pass rate.

Source: Does memory help Claude Code? 8 kinds of agent memory, tested · Full passes per task (Claude Sonnet 5.5, of 3)
Study limits and provenance

5 tasks across 8 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.

One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.

The hook checks the same code rules as the grader. It shows what rules written as code can do; it cannot carry a fact such as the late-fee rate.

n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.

Costs are the CLI's list-price estimates for subscription sessions, not invoices.

In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.

Source: Does memory help Claude Code? 8 kinds of agent memory, tested · Full passes per task (Claude Sonnet 5.5, of 3)

real pull request Updated 2026-10-05

3 real upstream pull requests

one attempt: verified delivery yes or no.

Back to sets ↑

What counts as a pass

Worker and gates run offline with frozen dependencies. Gates run install, lint, typecheck, build and the full suite at base, candidate and reference, plus cross-tests.

"Functional pass" means the frozen patch passed the gates. "Verified delivery" also needs a completed, clean, unassisted run.

No pre-run controls or validator check counts are recorded for this set. Read its pass rule below.

TaskBaseline (capped)1Fix wave 1 (capped)2
fastify/session0/10/1
h3js/h30/10/1
Kludex/uvicorn0/10/1
  1. Baseline (capped)
  2. Fix wave 1 (capped)

Shade is the share k of n in each cell. * The cell has a note (hover or focus it). Totals and other columns are in the Table view.

One attempt per cellA recorded outcome, not a rate

3 tasks across 4 slices. 1 of 12 recorded attempts had verified delivery. One attempt per cell is not a rate.

Source: Coding calibration: what broke on three real pull requests · Every attempt, failures included
Study limits and provenance

3 tasks across 4 slices. 1 of 12 recorded attempts had verified delivery. One attempt per cell is not a rate.

One attempt per cell: these are defect-finding runs, not rates.

Each slice changes the platform build, and the last two also remove the caps. No two slices are a matched comparison.

Costs are list-price estimates for subscription calls, not invoices.

The f0ac3a8a h3 run was verified only after its account-route label was corrected; it counts as unverified here, as declared.

Source: Coding calibration: what broke on three real pull requests · Every attempt, failures included

SWE-bench Verified Updated 2026-10-05

33 SWE-bench Verified instances

how many public models solved the instance (first column) and whether the agent did (second column).

Back to sets ↑

What counts as a pass

Grading uses the official SWE-bench harness and instance images (under amd64 emulation). The gold patch resolved on the host for every instance.

No pre-run controls or validator check counts are recorded for this set. Read its pass rule below.

Each tile is one row; from top: Public panel: models that solved it (of 11), then Agent (Sonnet 5.5, full pipeline): resolved.

11/111/111/110/111/111/111/111/111/111/111/111/111/111/111/111/111/111/111/111/111/111/111/111/111/111/111/110/110/111/110/111/110/111/110/111/110/111/110/111/110/111/19/110/19/110/18/111/17/111/14/110/13/111/13/111/12/111/10/110/10/110/10/111/10/110/1

Shade is the share k of n in each cell. * The cell has a note (hover or focus it).

Public panel: models that resolved the instanceAgent: one attempt per instanceDifferent denominators; no matched comparison

The first column counts public panel models that solved an instance. The second records one Agent attempt. These columns use different denominators and are not a matched comparison.

Source: Agent on SWE-bench Verified vs 11 public models · Every attempt
Study limits and provenance

The first column counts public panel models that solved an instance. The second records one Agent attempt. These columns use different denominators and are not a matched comparison.

n = 33: intervals are wide. This is a defect-finding run, not a ranking.

Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.

Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.

Campaign 1 ended up 19 django, 3 sympy, 2 sphinx and 1 xarray after replacements. Campaign 2 covers the compiled repositories.

Verified issues are public (2015 to 2023) and likely in every model’s training data; contamination is uncontrolled for all systems.

Difficulty bands come from the panel’s own results, so a system outside the panel tends to look better than the panel on hard bands and worse on easy ones (regression to the mean). Read the band chart with that selection effect in mind.

Three campaign-1 empty patches were platform holds before delivery (missing lint tools, a too-literal plan gate, an unanswered question), not wrong fixes. They count as failures here.

Source: Agent on SWE-bench Verified vs 11 public models · Every attempt

routing decisions Updated 2026-10-06

11 routing questions in 4 decision types

correct answers to the question, k of n cases.

Back to sets ↑

What counts as a pass

Cases: the labelled decision suites the platform uses (failure class, message intent, is-it-a-rule, context shape), only the cases production asks a router about.

Scoring: the repository’s own decision-eval runner. Exact means every scored question in a case was acceptable; key accuracy counts each question.

No pre-run controls or validator check counts are recorded for this set. Read its pass rule below.

TaskJev 1.13 (TypeSafe)1Claude Haiku 4.52Claude Sonnet 5.53
Failure class: failure18/1817/1818/18
Message intent: intent20/2020/2020/20
Is it a rule?: kind12/1212/1212/12
Context shape: turn5/65/65/6
Context shape: transcript12/1613/1615/16
Context shape: artifacts4/44/44/4
Context shape: knowledge22/2422/2423/24
Context shape: memories21/2121/2120/21
Context shape: examples10/1010/1010/10
Context shape: scope30/3129/3131/31
Context shape: complexity31/3230/3231/32
  1. Jev 1.13 (TypeSafe)
  2. Claude Haiku 4.5
  3. Claude Sonnet 5.5

Shade is the share k of n in each cell.

Counts from the source tablen = 4–32 per cell

11 tasks across 3 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.

Source: Jev vs Claude as a router: accuracy and cost · Per-question accuracy by decision type (Jev: its recorded production run, one pass)
Study limits and provenance

11 tasks across 3 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.

The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.

Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.

Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.

One sample per decision; production asks a second sample when confidence is low. Confidence-gated coverage is therefore not compared.

82 cases in four small hand-labelled sets: intervals are wide.

Clef / Clef-Flash (local): not measured (No local Clef server was running and installing a 6-20 GB model was out of scope for this run).

Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.

Jev ran as a direct HTTPS call; the Claude routers ran through the Claude Code CLI. These are different routes, so speed and cost compare what a caller pays per decision, not one model against the other.

Jev’s latency is one 35-second window from one Mac over a home network. The API reports no server time. A caller near the API would see less.

These tasks also appear in routing overhead.

Source: Jev vs Claude as a router: accuracy and cost · Per-question accuracy by decision type (Jev: its recorded production run, one pass)

short synthetic Updated 2026-10-05

8 probe tasks for CLI and API latency

passes of all runs, summed over configurations (a calculation).

Back to sets ↑

What counts as a pass

Receipts from the provider explorer: each run records the route (CLI or API), the model, the effort, timings, reported tokens and a deterministic validator result.

Only receipts classified "evaluated" count. Excluded runs (context mismatch, pilot runs, unsupported controls) and diagnostics are kept in the raw file but not charted.

No pre-run controls or validator check counts are recorded for this set. Read its pass rule below.

Calculation

All 8 cells are n of n: this task set cannot separate them.

TaskAll runs of the task1
Fixed 243-token answer43/43
Fixed exact reply76/76
mergeRanges coding repair13/13
PlanJobs scheduler repair9/9
Scratch file edit probe5/5
Synthetic coding task (API prompt)15/15
Synthetic coding task (CLI matched prompt)24/24
Synthetic coding task (CLI parity prompt)9/9
  1. All runs of the task

Every cell is at its maximum, so all cells share one calm shade. * The cell has a note (hover or focus it).

Summed counts: a calculationn = 5–76 per cell

194 of 194 recorded runs passed. These counts sum different configurations; they do not compare them.

Source: Claude Code CLI vs Codex CLI vs the API: latency and tokens · Every evaluated configuration
Study limits and provenance

194 of 194 recorded runs passed. These counts sum different configurations; they do not compare them.

Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.

The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.

The Claude CLI receipts record only the uncached remainder of the input (2 tokens), so the Claude input column is not comparable.

All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.

API costs in the raw file are list-price estimates; CLI runs are subscription calls with no per-call price.

Source: Claude Code CLI vs Codex CLI vs the API: latency and tokens · Every evaluated configuration

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.