- Home
- Benchmarks
- Task sets
Public golden task sets
The tasks behind our benchmarks.
See the checks.
Read the task titles, the rules for a pass, the controls and the recorded outcomes. This catalog draws from the public studies. It runs no new tests and shows no prompts or model output.
“Golden” means a curated task and validator catalog. It does not mean an exhaustive test suite, a model ranking or proof that Agent completed a customer goal.
- Task sets
- 8
- Task entries
- 79can overlap across studies
- Published check counts
- 228across 13 tasks
- Reference controls passed
- 14/14only recorded controls
Read a cell before comparing it
- Read the pass rule. It differs by set. A strict answer, all hidden tests and verified delivery are separate criteria.
- Read the denominator. k/n counts recorded passes and attempts. One attempt is an outcome, not a reliable pass rate.
- Read the limits. Perfect cells can have a ceiling. Different builds, routes and public panels can be unmatched.
Unknown means not recorded. Public task IDs stay public. The catalog contains no task prompts, model replies or private run paths.
Choose a task set
hard synthetic Updated 2026-10-06
8 hard tasks with strict validators
strict passes of n calls.
What counts as a pass
Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 references wrapped in a fence or prose are flagged as format misses.
Strict pass: the whole reply, trimmed, passes the validator. Every prompt states the output format and says no code fence and no other text. This is stricter than the five-task study, which removed one wrapping fence.
Format miss: the strict check failed, but a lenient extractor (fenced block, outer JSON, first code line, one output line) finds an answer that passes the same validator. Reported apart from wrong answers, never as a pass.
- Fix an interval-merge function (off-by-one and edge cases)18 validator checksReference passes5/5 plausible wrong answers rejectedWrapped reference flagged as a format miss
- Fix a time-zone day-length function (DST)36 validator checksReference passes3/3 plausible wrong answers rejectedWrapped reference flagged as a format miss
- Write a CSV parser (quoted newlines, strict errors)22 validator checksReference passes3/3 plausible wrong answers rejectedWrapped reference flagged as a format miss
- Predict JavaScript event-loop output order1 validator checksReference passes3/3 plausible wrong answers rejectedWrapped reference flagged as a format miss
- Solve a multi-constraint room schedule14 validator checksReference passes3/3 plausible wrong answers rejectedWrapped reference flagged as a format miss
- Write a strict SemVer 2.0.0 regex30 validator checksReference passes3/3 plausible wrong answers rejectedWrapped reference flagged as a format miss
- Refactor to remove duplication, keep 20 tests green25 validator checksReference passes2/2 plausible wrong answers rejectedWrapped reference flagged as a format miss
- Write a SQLite reporting query (fan-out, ties, boundaries)8 validator checksReference passes4/4 plausible wrong answers rejectedWrapped reference flagged as a format miss
| Task | Validator checks | Reference answer | Plausible wrong answers rejected | Wrapped reference flagged as format miss |
|---|---|---|---|---|
| Fix an interval-merge function (off-by-one and edge cases) | 18 | passes | 5/5 | yes |
| Fix a time-zone day-length function (DST) | 36 | passes | 3/3 | yes |
| Write a CSV parser (quoted newlines, strict errors) | 22 | passes | 3/3 | yes |
| Predict JavaScript event-loop output order | 1 | passes | 3/3 | yes |
| Solve a multi-constraint room schedule | 14 | passes | 3/3 | yes |
| Write a strict SemVer 2.0.0 regex | 30 | passes | 3/3 | yes |
| Refactor to remove duplication, keep 20 tests green | 25 | passes | 2/2 | yes |
| Write a SQLite reporting query (fan-out, ties, boundaries) | 8 | passes | 4/4 | yes |
Counts from the public controls tableBars count checks from zeroMissing controls stay unknown
These controls test the validator before model or agent attempts. A reference pass alone does not prove that every wrong answer will fail.
- Claude Sonnet 5.5 · Claude Code
- Claude Opus 5.5 · Claude Code
- Claude Opus 5.5 (high) · Claude Code
- GPT-6.1 Sol (medium) · Codex CLI
- Claude Fable 5.1 · Claude Code
- GPT-6.1 Sol (high) · Codex CLI
- Claude Haiku 4.5 · Claude Code
Shade is the share k of n in each cell. * The cell has a note (hover or focus it).
| Task | Claude Sonnet 5.5 · Claude Code | Claude Opus 5.5 · Claude Code | Claude Opus 5.5 (high) · Claude Code | GPT-6.1 Sol (medium) · Codex CLI | Claude Fable 5.1 · Claude Code | GPT-6.1 Sol (high) · Codex CLI | Claude Haiku 4.5 · Claude Code |
|---|---|---|---|---|---|---|---|
| Fix an interval-merge function (off-by-one and edge cases) | 3/3 | 3/3 | 3/3 | 2/2 | 3/3 | 2/2 | 3/3 |
| Fix a time-zone day-length function (DST) | 3/3 | 3/3 | 3/3 | 2/2 | 3/3 | 2/2 | 1/3 |
| Write a CSV parser (quoted newlines, strict errors) | 3/3 | 3/3 | 3/3 | 2/2 | 3/3 | 2/2 | 2/3 (+1 format miss) |
| Predict JavaScript event-loop output order | 3/3 | 3/3 | 3/3 | 2/2 | 3/3 | 2/2 | 0/3 |
| Solve a multi-constraint room schedule | 3/3 | 3/3 | 3/3 | 2/2 | 3/3 | 2/2 | 0/3 (+2 format misses) |
| Write a strict SemVer 2.0.0 regex | 3/3 | 3/3 | 3/3 | 2/2 | 3/3 | 2/2 | 3/3 |
| Refactor to remove duplication, keep 20 tests green | 3/3 | 3/3 | 3/3 | 2/2 | 3/3 | 2/2 | 2/3 |
| Write a SQLite reporting query (fan-out, ties, boundaries) | 3/3 | 3/3 | 3/3 | 2/2 | 3/3 | 2/2 | 0/3 (+2 format misses) |
Counts from the source tablen = 2–3 per cell
8 tasks across 7 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.
6 of 7 configurations passed every recorded task. This set cannot separate them on pass rate.
Study limits and provenance
8 tasks across 7 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.
6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
The Claude and Codex batches ran on different days on the same host, one call at a time per account. Each route used its own subscription.
Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
CLI timings include CLI start-up and the CLI’s own system prompt. One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled.
Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
List-price costs are calculations; the calls used a flat subscription.
These tasks also appear in effort ladder.
Source: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks · Pass matrix on hard tasks: configuration × task (strict passes) · Validator controls run before the first model call
coding agents Updated 2026-10-06
6 repository tasks with hidden tests
sessions that passed every hidden check, of n.
What counts as a pass
Controls before the first session (Node v25.2.1): every base repository fails its hidden checks and every reference patch passes all of them.
Pass = every hidden check passes; no partial credit. Recorded per session: wall time, tool calls, tokens as reported, files touched, lines changed against the base commit, commits and edits outside the repository.
- Pagination fix13 validator checksReference passes (tests 13/13)Starting repository fails (tests 1/13)
- CLI --top flag11 validator checksReference passes (tests 11/11)Starting repository fails (tests 2/11)
- Invoice refactor24 validator checksReference passes (tests 24/24)Starting repository fails (tests 21/24)
- LRU cache16 validator checksReference passes (tests 16/16)Starting repository fails (tests 0/16)
- Queue race10 validator checksReference passes (tests 10/10)Starting repository fails (tests 4/10)
- Strict TypeScript typesstatic checks, tsc with a hidden usage file (11 @ts-expect-error probes) and 4 runtime testsReference passes (tests 4/4)Starting repository fails (tests 4/4; static checks fail)
| Task | Validator checks | Reference answer | Starting repository |
|---|---|---|---|
| Pagination fix | 13 | passes (tests 13/13) | fails (tests 1/13) |
| CLI --top flag | 11 | passes (tests 11/11) | fails (tests 2/11) |
| Invoice refactor | 24 | passes (tests 24/24) | fails (tests 21/24) |
| LRU cache | 16 | passes (tests 16/16) | fails (tests 0/16) |
| Queue race | 10 | passes (tests 10/10) | fails (tests 4/10) |
| Strict TypeScript types | static checks, tsc with a hidden usage file (11 @ts-expect-error probes) and 4 runtime tests | passes (tests 4/4) | fails (tests 4/4; static checks fail) |
Counts from the public controls tableBars count checks from zeroMissing controls stay unknown
These controls test the validator before model or agent attempts. A reference pass alone does not prove that every wrong answer will fail.
All 18 cells are n of n: this task set cannot separate them.
- Sonnet 5.5 in Claude Code
- Opus 5.5 in Claude Code
- GPT-6.1 Sol in Codex CLI
Every cell is at its maximum, so all cells share one calm shade.
| Task | Sonnet 5.5 in Claude Code | Opus 5.5 in Claude Code | GPT-6.1 Sol in Codex CLI |
|---|---|---|---|
| Pagination fix | 2/2 | 2/2 | 2/2 |
| CLI --top flag | 2/2 | 2/2 | 2/2 |
| Invoice refactor | 2/2 | 2/2 | 2/2 |
| LRU cache | 2/2 | 2/2 | 2/2 |
| Queue race | 2/2 | 2/2 | 2/2 |
| Strict TypeScript types | 2/2 | 2/2 | 2/2 |
Counts from the source tablen = 2 per cell
6 tasks across 3 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.
3 of 3 configurations passed every recorded task. This set cannot separate them on pass rate.
Study limits and provenance
6 tasks across 3 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.
Codex CLI ran with the tester’s global AGENTS.md: --ignore-user-config does not switch that file off. 12 of 12 Codex sessions read the tester’s notes, 10 wrote a WORKLOG.md and 9 reported a commit attempt. Its time, tool calls and lines changed include that work. Claude Code ran with setting sources off, and no Claude session did any of it.
Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.
Each run pairs a CLI with a model (Claude Code with Claude models, Codex CLI with GPT-6.1 Sol), so the results cannot separate the CLI from the model.
n = 12 sessions per agent (2 per task). Time ranges are the fastest and slowest sessions, not confidence intervals; the two lanes shared one machine.
Tokens are as each CLI reports them, with different tokenizers and context handling: compare tokens inside Claude Code (Sonnet vs Opus), not across vendors. Costs are list-price calculations on subscription sessions, not invoices.
Source: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks · Every task and agent: sessions that passed (of 2) · The six tasks and their controls
short synthetic Updated 2026-10-05
5 short tasks with deterministic validators
passes of n calls.
What counts as a pass
Protocol declared before the first call: 5 cases with deterministic validators (behavioral checks in a sandbox, canonical JSON or exact text).
No pre-run controls or validator check counts are recorded for this set. Read its pass rule below.
- Claude Fable 5.1 · Claude Code
- Claude Sonnet 5.5 · Claude Code
- Claude Opus 5.5 (high) · Claude Code
- Claude Opus 5.5 · Claude Code
- Claude Opus 5.5 (low) · Claude Code
- Claude Haiku 4.5 · Claude Code
- GPT-6.1 Sol (high) · Codex CLI
- GPT-6.1 Sol (medium) · Codex CLI
- GPT-6.1 Sol (low) · Codex CLI
Shade is the share k of n in each cell.
| Task | Claude Fable 5.1 · Claude Code | Claude Sonnet 5.5 · Claude Code | Claude Opus 5.5 (high) · Claude Code | Claude Opus 5.5 · Claude Code | Claude Opus 5.5 (low) · Claude Code | Claude Haiku 4.5 · Claude Code | GPT-6.1 Sol (high) · Codex CLI | GPT-6.1 Sol (medium) · Codex CLI | GPT-6.1 Sol (low) · Codex CLI |
|---|---|---|---|---|---|---|---|---|---|
| Fix a buggy median function | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 2/2 |
| Extract invoice fields to JSON | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 2/2 |
| Multi-step shift arithmetic | 3/3 | 0/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 2/2 |
| Refactor recursion to iteration | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 2/2 |
| Classify six support tickets | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 2/2 |
Counts from the source tablen = 2–3 per cell
5 tasks across 9 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.
8 of 9 configurations passed every recorded task. This set cannot separate them on pass rate.
Study limits and provenance
5 tasks across 9 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.
The tasks are short and easy; pass rate saturates. Latency and tokens carry the signal. A harder follow-up with eight tasks and strict validators: /benchmarks/hard-model-head-to-head.
Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
CLI timings include CLI start-up and the CLI’s own system prompt.
One host, one network, one day.
List-price costs are calculations; the calls used flat subscriptions.
The prompt cache stayed at the provider default, so cache counters differ by route and by call order.
Haiku 4.5 reported reasoning tokens on most calls under the CLI default, which explains much of its extra time and output.
Haiku and Fable are not in the platform runner catalog, so these calls used the runner’s CLI functions directly; each receipt records the model the CLI reported.
Source: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head · Pass matrix: configuration × task
agent memory Updated 2026-10-06
5 tasks for agent memory
full passes of 3 sessions.
What counts as a pass
Grading after the session: hidden tests run in a sandbox, plus deterministic checks on the lines the agent added. Full pass = tests and every applicable check. Every attempt counts.
No pre-run controls or validator check counts are recorded for this set. Read its pass rule below.
- No memory
- /init CLAUDE.md
- Curated, 11 lines
- Raw notes, 60 lines
- Dreamed notes
- Handbook, 210 lines
- Stop hook only
- Curated + hook
Shade is the share k of n in each cell.
| Task | No memory | /init CLAUDE.md | Curated, 11 lines | Raw notes, 60 lines | Dreamed notes | Handbook, 210 lines | Stop hook only | Curated + hook |
|---|---|---|---|---|---|---|---|---|
| Refunds | 3/3 | 3/3 | 3/3 | 2/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| Report decimals | 0/3 | 0/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| Customer phone | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| Late fee | 0/3 | 0/3 | 3/3 | 3/3 | 3/3 | 3/3 | 0/3 | 3/3 |
| Slugify (control) | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
Counts from the source tablen = 3 per cell
5 tasks across 8 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.
4 of 8 configurations passed every recorded task. This set cannot separate them on pass rate.
Study limits and provenance
5 tasks across 8 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.
One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
The hook checks the same code rules as the grader. It shows what rules written as code can do; it cannot carry a fact such as the late-fee rate.
n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.
Costs are the CLI's list-price estimates for subscription sessions, not invoices.
In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.
Source: Does memory help Claude Code? 8 kinds of agent memory, tested · Full passes per task (Claude Sonnet 5.5, of 3)
real pull request Updated 2026-10-05
3 real upstream pull requests
one attempt: verified delivery yes or no.
What counts as a pass
Worker and gates run offline with frozen dependencies. Gates run install, lint, typecheck, build and the full suite at base, candidate and reference, plus cross-tests.
"Functional pass" means the frozen patch passed the gates. "Verified delivery" also needs a completed, clean, unassisted run.
No pre-run controls or validator check counts are recorded for this set. Read its pass rule below.
- Baseline (capped)
- Fix wave 1 (capped)
Shade is the share k of n in each cell. * The cell has a note (hover or focus it). Totals and other columns are in the Table view.
| Task | Slice | Functional gates | Verified delivery | How it ended |
|---|---|---|---|---|
| fastify/session | Baseline (capped) | verified | no | Escalated to a person |
| fastify/session | Fix wave 1 (capped) | verified | no | Hit the $5 cap |
| fastify/session | Uncapped, build f0ac3a8a | failed | no | Escalated to a person |
| fastify/session | Uncapped, build 236c0d3f | verified | yes | Delivered (in review) |
| h3js/h3 | Baseline (capped) | verified | no | Escalated to a person |
| h3js/h3 | Fix wave 1 (capped) | verified | no | Escalated to a person |
| h3js/h3 | Uncapped, build f0ac3a8a | verified | no | Delivered (in review) |
| h3js/h3 | Uncapped, build 236c0d3f | failed | no | Escalated to a person |
| Kludex/uvicorn | Baseline (capped) | unverified | no | Hit the $5 cap |
| Kludex/uvicorn | Fix wave 1 (capped) | unverified | no | Hit the 20 min cap |
| Kludex/uvicorn | Uncapped, build f0ac3a8a | unverified | no | Delivered (in review) |
| Kludex/uvicorn | Uncapped, build 236c0d3f | unverified | no | Delivered (in review) |
One attempt per cellA recorded outcome, not a rate
3 tasks across 4 slices. 1 of 12 recorded attempts had verified delivery. One attempt per cell is not a rate.
Study limits and provenance
3 tasks across 4 slices. 1 of 12 recorded attempts had verified delivery. One attempt per cell is not a rate.
One attempt per cell: these are defect-finding runs, not rates.
Each slice changes the platform build, and the last two also remove the caps. No two slices are a matched comparison.
Costs are list-price estimates for subscription calls, not invoices.
The f0ac3a8a h3 run was verified only after its account-route label was corrected; it counts as unverified here, as declared.
Source: Coding calibration: what broke on three real pull requests · Every attempt, failures included
SWE-bench Verified Updated 2026-10-05
33 SWE-bench Verified instances
how many public models solved the instance (first column) and whether the agent did (second column).
What counts as a pass
Grading uses the official SWE-bench harness and instance images (under amd64 emulation). The gold patch resolved on the host for every instance.
No pre-run controls or validator check counts are recorded for this set. Read its pass rule below.
Each tile is one row; from top: Public panel: models that solved it (of 11), then Agent (Sonnet 5.5, full pipeline): resolved.
Shade is the share k of n in each cell. * The cell has a note (hover or focus it).
| Instance | Panel solved | Agent outcome |
|---|---|---|
| astropy__astropy-13579 | 11/11 | resolved |
| astropy__astropy-7336 | 11/11 | unresolved |
| django__django-11333 | 11/11 | resolved |
| django__django-12741 | 11/11 | resolved |
| django__django-14373 | 11/11 | resolved |
| django__django-14559 | 11/11 | resolved |
| django__django-14672 | 11/11 | resolved |
| django__django-16333 | 11/11 | resolved |
| django__django-17029 | 11/11 | resolved |
| matplotlib__matplotlib-24970 | 11/11 | resolved |
| pydata__xarray-6721 | 11/11 | resolved |
| scikit-learn__scikit-learn-14894 | 11/11 | resolved |
| sphinx-doc__sphinx-9698 | 11/11 | resolved |
| sympy__sympy-18189 | 11/11 | empty patch |
| astropy__astropy-14096 | 10/11 | resolved |
| django__django-12193 | 10/11 | resolved |
| django__django-13158 | 10/11 | resolved |
| django__django-15572 | 10/11 | resolved |
| django__django-16255 | 10/11 | resolved |
| matplotlib__matplotlib-24149 | 10/11 | resolved |
| scikit-learn__scikit-learn-25973 | 10/11 | resolved |
| django__django-14631 | 9/11 | empty patch |
| sympy__sympy-22456 | 9/11 | empty patch |
| django__django-13925 | 8/11 | resolved |
| django__django-15280 | 7/11 | resolved |
| django__django-15022 | 4/11 | unresolved |
| django__django-11885 | 3/11 | resolved |
| django__django-15732 | 3/11 | resolved |
| sphinx-doc__sphinx-10435 | 2/11 | resolved |
| django__django-10554 | 0/11 | unresolved |
| django__django-14034 | 0/11 | unresolved |
| matplotlib__matplotlib-21568 | 0/11 | resolved |
| sympy__sympy-20428 | 0/11 | unresolved |
Public panel: models that resolved the instanceAgent: one attempt per instanceDifferent denominators; no matched comparison
The first column counts public panel models that solved an instance. The second records one Agent attempt. These columns use different denominators and are not a matched comparison.
Study limits and provenance
The first column counts public panel models that solved an instance. The second records one Agent attempt. These columns use different denominators and are not a matched comparison.
n = 33: intervals are wide. This is a defect-finding run, not a ranking.
Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
Campaign 1 ended up 19 django, 3 sympy, 2 sphinx and 1 xarray after replacements. Campaign 2 covers the compiled repositories.
Verified issues are public (2015 to 2023) and likely in every model’s training data; contamination is uncontrolled for all systems.
Difficulty bands come from the panel’s own results, so a system outside the panel tends to look better than the panel on hard bands and worse on easy ones (regression to the mean). Read the band chart with that selection effect in mind.
Three campaign-1 empty patches were platform holds before delivery (missing lint tools, a too-literal plan gate, an unanswered question), not wrong fixes. They count as failures here.
Source: Agent on SWE-bench Verified vs 11 public models · Every attempt
routing decisions Updated 2026-10-06
11 routing questions in 4 decision types
correct answers to the question, k of n cases.
What counts as a pass
Cases: the labelled decision suites the platform uses (failure class, message intent, is-it-a-rule, context shape), only the cases production asks a router about.
Scoring: the repository’s own decision-eval runner. Exact means every scored question in a case was acceptable; key accuracy counts each question.
No pre-run controls or validator check counts are recorded for this set. Read its pass rule below.
- Jev 1.13 (TypeSafe)
- Claude Haiku 4.5
- Claude Sonnet 5.5
Shade is the share k of n in each cell.
| Task | Jev 1.13 (TypeSafe) | Claude Haiku 4.5 | Claude Sonnet 5.5 |
|---|---|---|---|
| Failure class: failure | 18/18 | 17/18 | 18/18 |
| Message intent: intent | 20/20 | 20/20 | 20/20 |
| Is it a rule?: kind | 12/12 | 12/12 | 12/12 |
| Context shape: turn | 5/6 | 5/6 | 5/6 |
| Context shape: transcript | 12/16 | 13/16 | 15/16 |
| Context shape: artifacts | 4/4 | 4/4 | 4/4 |
| Context shape: knowledge | 22/24 | 22/24 | 23/24 |
| Context shape: memories | 21/21 | 21/21 | 20/21 |
| Context shape: examples | 10/10 | 10/10 | 10/10 |
| Context shape: scope | 30/31 | 29/31 | 31/31 |
| Context shape: complexity | 31/32 | 30/32 | 31/32 |
Counts from the source tablen = 4–32 per cell
11 tasks across 3 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.
Study limits and provenance
11 tasks across 3 configurations. Each cell shows recorded passes of recorded attempts. Missing cells are unknown, not failures.
The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
One sample per decision; production asks a second sample when confidence is low. Confidence-gated coverage is therefore not compared.
82 cases in four small hand-labelled sets: intervals are wide.
Clef / Clef-Flash (local): not measured (No local Clef server was running and installing a 6-20 GB model was out of scope for this run).
Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
Jev ran as a direct HTTPS call; the Claude routers ran through the Claude Code CLI. These are different routes, so speed and cost compare what a caller pays per decision, not one model against the other.
Jev’s latency is one 35-second window from one Mac over a home network. The API reports no server time. A caller near the API would see less.
These tasks also appear in routing overhead.
Source: Jev vs Claude as a router: accuracy and cost · Per-question accuracy by decision type (Jev: its recorded production run, one pass)
short synthetic Updated 2026-10-05
8 probe tasks for CLI and API latency
passes of all runs, summed over configurations (a calculation).
What counts as a pass
Receipts from the provider explorer: each run records the route (CLI or API), the model, the effort, timings, reported tokens and a deterministic validator result.
Only receipts classified "evaluated" count. Excluded runs (context mismatch, pilot runs, unsupported controls) and diagnostics are kept in the raw file but not charted.
No pre-run controls or validator check counts are recorded for this set. Read its pass rule below.
All 8 cells are n of n: this task set cannot separate them.
- All runs of the task
Every cell is at its maximum, so all cells share one calm shade. * The cell has a note (hover or focus it).
| Task | All runs of the task |
|---|---|
| Fixed 243-token answer | 43/43 (3 configurations) |
| Fixed exact reply | 76/76 (6 configurations) |
| mergeRanges coding repair | 13/13 (3 configurations) |
| PlanJobs scheduler repair | 9/9 (3 configurations) |
| Scratch file edit probe | 5/5 (1 configuration) |
| Synthetic coding task (API prompt) | 15/15 (3 configurations) |
| Synthetic coding task (CLI matched prompt) | 24/24 (3 configurations) |
| Synthetic coding task (CLI parity prompt) | 9/9 (1 configuration) |
Summed counts: a calculationn = 5–76 per cell
194 of 194 recorded runs passed. These counts sum different configurations; they do not compare them.
Study limits and provenance
194 of 194 recorded runs passed. These counts sum different configurations; they do not compare them.
Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
The Claude CLI receipts record only the uncached remainder of the input (2 tokens), so the Claude input column is not comparable.
All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
API costs in the raw file are list-price estimates; CLI runs are subscription calls with no per-call price.
Source: Claude Code CLI vs Codex CLI vs the API: latency and tokens · Every evaluated configuration