- Benchmarks
- Methodology
Open data · updated October 7, 2026
How we measure. So you can check.
Every number on these pages comes from one file that anyone can rebuild from the raw extracts. This page shows the pipeline, the protocols, the controls, how we count, which intervals we use, and the rule that decides when one result beats another.
28
Studies, each with its protocol
206
Charts, each with n and a Table view
37
Sources with dates and raw data
1,571
Comparison rows judged by one rule
How it works
From a recorded run to a number on this page
Five stages. Each one refuses to pass on what it cannot check.
Record
28 studies
Runs write receipts to a local store: every attempt, the model, effort and route of each call, its time, tokens and result.
- Every attempt kept, failures too
- Fixed tasks and validators per study
Extract
37 sources
A script copies sanitized extracts from the store into the repository.
- No local paths, emails or org ids
- No tokens or content hashes
- Private tasks get neutral labels
Compile
206 charts
The compiler reads only the committed extracts, with no network, and writes one dataset file.
- 95% Wilson interval for every rate
- Median and range for every timing
- Calculations flagged as calculations
Validate
1,571 rows
Tests check the dataset against its schema and its own rules before anything ships.
- Every fact traces to a chart point or stat
- Comparison rows follow the winner rules
- The dataset must match the extracts
JSON · CSV
Pages, live stories, embeds and downloads all read the same file.
- n and interval kind beside every chart
- A Table view for every chart
- JSON and CSV for every study
| Stage | What it does | Checks |
|---|---|---|
| 1. Record | Runs write receipts to a local store: every attempt, the model, effort and route of each call, its time, tokens and result. (28 studies) | Every attempt kept, failures too; Fixed tasks and validators per study |
| 2. Extract | A script copies sanitized extracts from the store into the repository. (37 sources) | No local paths, emails or org ids; No tokens or content hashes; Private tasks get neutral labels |
| 3. Compile | The compiler reads only the committed extracts, with no network, and writes one dataset file. (206 charts) | 95% Wilson interval for every rate; Median and range for every timing; Calculations flagged as calculations |
| 4. Validate | Tests check the dataset against its schema and its own rules before anything ships. (1,571 rows) | Every fact traces to a chart point or stat; Comparison rows follow the winner rules; The dataset must match the extracts |
| 5. Publish | Pages, live stories, embeds and downloads all read the same file. (JSON · CSV) | n and interval kind beside every chart; A Table view for every chart; JSON and CSV for every study |
Five stages: record, extract, compile, validate and publish. Today they carry 28 studies, 37 sources, 206 charts and 1,571 comparison rows.
Protocols
What each study ran, and how
Each study states its sample, its tasks and validators, and what changed between runs. The full protocol is on the study page. Read the task sets: pass rules, controls and recorded outcomes. Inspect model differences and their uncertainty.
Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
For the same extraction tasks, does enforcing a JSON schema through the CLI change the strict pass rate, the format misses and the wrong values, compared with asking for JSON in the prompt?
Protocol: 9 steps
- The protocol states a pre-call declaration. Its current file birth time is 2026-10-07T01:47:06.955Z. The first counted call began at 2026-10-07T00:43:42.104Z. These files cannot verify a protocol declaration before the Codex calls. The extract keeps every counted call.
- Three prompts each have one exact expected JSON answer and a sandboxed validator. The order task merges lines, drops a cancelled line and applies a discount. The invoice task applies a line discount and then tax, rounded half up. It also includes an untaxed fee and shipping. The meeting task extracts tickets, owners and due dates, resolves relative dates, and skips a decision and a note. We do not publish prompts or replies.
- Two modes use the same task text. Instructions mode adds “Reply with the JSON object only, no code fence, no other text.” Schema mode removes that sentence and uses the CLI’s schema flag. Claude Code uses --json-schema; the harness reads its structured_output field. Codex exec uses --output-schema; the harness reads its final message.
- Repetitions per prompt, in each mode: Haiku: 8; Sonnet: 4; GPT-6.1 Sol: 4. The order alternates by repetition, so time drift falls on both modes. Claude rows use the CLI’s default effort. GPT-6.1 Sol rows use low effort.
- Scoring: a strict pass needs the whole reply to be the right JSON. A format miss is a right answer inside a code fence or prose. Wrong values means any other completed reply, including a fenced reply with a wrong value. An error means the call did not complete; it counts as a fail.
- Rates carry Wilson 95% intervals (a calculation). A side is ahead only when the intervals do not overlap and the paired exact McNemar p is below 0.05. The paired table gives this p-value calculation. Time is a median with the fastest and slowest call: a range, not an interval.
- Controls ran before inference. Each reference answer passes, including through the schema path. A reference inside a code fence counts as a format miss. Four planted wrong answers per prompt fail. One uncounted probe per route (2 here) checked how the harness reads the CLI output. Probes stay outside every cell.
- Each call uses a fresh empty working folder, with tools off, no MCP servers and no session persistence. The harness reads the answer from the CLI output. Claude Code uses its JSON output format in both modes; earlier studies used its stream format. Schema mode adds the schema flag and removes the format sentence.
- Every counted batch completed all planned cells.
Does a new Claude Code session reuse the prompt cache of an earlier one?
In the earlier caching study, a new Claude Code session did not read the cache that an earlier session wrote. Does a fixed working folder change that, and does putting the ledger in the system prompt help?
Protocol: 14 steps
- The protocol claims a declaration at 00:32 UTC. Its surviving file has a birth time of 00:49:27 UTC. The first counted call started at 00:32:33 UTC. We cannot verify a pre-call protocol.
- The run kept every try and made no retries. The run log lists three amendments. They cover the post-hoc Codex rule, a post-hoc pooled calculation and text corrections after a check.
- Claude Code 2.1.286 ran Sonnet 5.5 at its default effort. Tools were off. No MCP servers. No session persistence across processes. Caching was at the provider default.
- A session is one CLI process with 2 turns. Turn 1 asks a quantity lookup over a seeded synthetic stock ledger. The ledger has 100 lines and about 9,600 characters. Turn 2 asks a warehouse lookup. Each question has one reference answer. The checker trims spaces, quotes, backticks, a final period and currency units before comparison.
- Each setup had 3 sessions with 2 turns each (6 calls). A used a new folder each time. B used a fixed folder. Both put the ledger in the first user message. C used a fixed folder and put the ledger in the system prompt.
- Each setup has its own ledger seed. Different seeds prevent full ledger-prefix reuse between setups. Shared CLI-prefix reads remain possible. Sessions of a setup ran back to back. The gap between one session's last call and the next session's first call was 508 to 720 ms.
- A later session is session 2 or 3. The protocol labels it reuse when turn 1 read at least half of its input from the cache. This is a counter-based rule; no token-level trace identifies the ledger. The cache counters are the provider’s: uncached input, cache reads and cache writes (with the 5-minute and 1-hour split).
- List-price cost is a calculation. It prices uncached input at the input price and reads at the cache-read price. It prices 1-hour writes at 2× the input price. The calls ran on a subscription, so nothing here is a bill.
- Codex CLI ran GPT-6.1 Sol at medium effort in setups A and B. Each setup had 3 sessions with 2 turns each. Tools, apps, plugins and web search were off. Each session used an ephemeral read-only thread. The app-server reports input and cached input, but no cache writes. Its own large system prompt could explain a cached count.
- The surviving protocol records a 50% threshold but dates from after the calls. This rule counts 2 of 4 later Codex sessions. The post-hoc 90% rule counts 0 of 4. The respective Wilson 95% intervals are 15% to 85% and 0% to 49%. The highest turn-1 read share was 56% (calculation). No token-level trace identifies the ledger.
- One uncounted probe call per route, with its own ledger seed, checked the driver before the counted calls. The answer checker is the one of the earlier study. Its control test (every reference answer passes, every planted wrong answer fails) ran after the counted calls, because no cache measure depends on it.
- No batch stopped early. Nothing was trimmed. The Claude CLI reported a rate-limit status of "allowed_warning" on 9 of 18 Claude turns. No call was refused and the run did not stop.
- The normalized-answer checker passed 18 of 18 Claude answers (95% interval 82% to 100%).
- The normalized-answer checker passed 12 of 12 Codex answers (95% interval 76% to 100%).
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
On a task set built so that Claude Sonnet 5.5 did not pass it every time, does pass rate separate GPT-6.1 Sol (Codex CLI) from Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 (Claude Code)?
Protocol: 14 steps
- The original protocol file predates the first probe and counted call. This review checked file birth times. Later changes are amendments. This follows the hard head-to-head (/benchmarks/hard-model-head-to-head), which hit a pass-rate ceiling for the strong models.
- The study started with 16 candidate tasks, written in two rounds (12 first, then 4 replacements; code, SQL, reasoning, spec, simulation, numeric). Each is one call with no tools, an exact output format and a deterministic validator that runs in a sandbox without network. Claude Sonnet 5.5 at default effort ran each candidate twice (the pilot, 32 calls). It passed 12 of the 16 candidates 2 of 2. The study dropped them.
- The dropped candidates include 10 of 10 code, SQL, spec, numeric and simulation tasks.
- The protocol declared the selection rule before the pilot. Keep at most 8 tasks: first those Sonnet passed 1 of 2, then 0 of 2. Drop tasks with 2 of 2 passes. Only 1 of the first 12 qualified, below the threshold of 6. The study wrote 4 replacement candidates once and piloted them twice; 3 qualified.
- The counted set is 4 tasks: 10x10 nonogram; Sudoku, 22 givens; 6x6 Skyscrapers; Seeded shuffle output. The study ran no second round of replacements.
- Controls ran in stages. The first 12 candidates had controls before the first probe. After a module-export validator defect, controls ran again and stored pilot replies were re-scored without new calls. Final controls covered all 16 candidates before the replacement pilot and all counted calls. 16/16 references pass. 66/66 wrong answers fail. 16/16 wrapped references are format misses. Independent solvers checked unique reasoning answers. Python integers confirmed the code-reading answer.
- Counted cells: GPT-6.1 Sol (medium) · Codex CLI (16 calls, 4 per task). Claude Opus 5.5 · Claude Code (12 calls, 3 per task). Claude Sonnet 5.5 · Claude Code (16 calls, 4 per task). Claude Haiku 4.5 · Claude Code (12 calls, 3 per task). Rep-major, round-robin order, one call at a time per account, 300 s timeout.
- Strict pass: the whole reply, trimmed, passes the validator. Every prompt states the output format and says no other text. Format miss: the strict check fails, but an extracted answer passes the same validator. The extractor reads fenced blocks, answer lines or grid rows. For exact tasks it reads only the last answer-shaped candidate, not an earlier guess. Reported apart from wrong answers, never as a pass.
- An error or timeout counts as a non-pass. Tool use is a separate flag, not a score. It marks tool-call markup or a CLI tool-call parse error.
- Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn. The Codex runner adds a developer instruction to answer directly and not to call tools. The Claude Code runner adds no such instruction. Claude Code ran with an output-token cap setting of 16,000; the Codex CLI had none.
- Default effort means the effort flag was not passed. GPT-6.1 Sol ran at medium effort.
- Counted calls: Claude Code: 40 counted calls, 40 reached a model; Codex CLI: 16 counted calls, 16 reached a model. The study trimmed no calls. The study repeated no counted key. No run reported a usage or rate limit. The Claude CLI reported its own failed retry on tool-call parse errors; those receipts remain failures. The Claude batch stopped 1 time on a tool-call parse error. A later amendment kept that failure but allowed the lane to continue. The lane repeated no counted key.
- Uncounted probes: 2, kept apart from 32 pilot calls and 56 counted attempts. Call caps, including probes and pilots: Claude Code: 73/105 calls; Codex CLI: 17/33 calls.
- Cost per strict pass is a calculation: total reported-token cost divided by strict passes. It includes failed calls with tokens and prices cache reads and writes. Timeouts report no tokens, so the cost is a lower bound.
Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts
On the 8 hard tasks, what does one correct answer cost, and how long does it take, if you try Claude Haiku 4.5 first and retry or escalate, against Claude Sonnet 5.5 every time?
Protocol: 13 steps
- This study makes no new model call and adds no new raw file. It reuses recorded calls from the hard head-to-head (Haiku 4.5 and Sonnet 5.5 at default effort, 3 calls per task) and the effort ladder (Sonnet 5.5 at low effort, 2 calls per task). All calls ran in Claude Code on the same 8 tasks with the same strict validators.
- The source controls ran before counted inference. Each controls receipt records 8 passing references, 26 rejected wrong answers and 8 detected wrapped format misses. The hard-set protocol outcome says 27 wrong answers; the controls receipt supports 26. Both Claude batches completed their declared call caps: 120 hard-set calls and 80 ladder calls, with no trims or errors. They used a 300 s timeout and a 16,000-output-token setting. The selected calls stayed below that setting. No Claude probe is recorded in these source folders; the hard-set Codex probe is outside this calculation.
- For each task and model, we take the pass rate p. It is strict passes ÷ calls, with its 95% Wilson interval. A format miss is not a pass.
- These are median-input scenarios. A median is not an expected mean. Cost and time can depend on whether a call fails. We do not model that link. We take the cost of one call as the median list-price cost of that task’s calls. List price is reported tokens × price. Cache reads and one-hour cache writes are priced as in the hard head-to-head.
- We take the time of one call as the median total time of that task’s calls. It includes CLI start-up.
- A policy is a list of tries. A try runs only when every earlier try failed. The policies are: Sonnet every time, Haiku once, Haiku with one retry then Sonnet, Haiku with up to 3 tries then Sonnet, and Sonnet at low effort once.
- Per task, let q = 1 − p(Haiku). For Haiku with one retry then Sonnet: cost = C(Haiku) × (1 + q) + C(Sonnet) × q². Time = T(Haiku) × (1 + q) + T(Sonnet) × q². Success = 1 − q² × (1 − p(Sonnet)).
- For Haiku with up to 3 tries then Sonnet: cost = C(Haiku) × (1 + q + q²) + C(Sonnet) × q³. Success = 1 − q³ × (1 − p(Sonnet)). Time uses the same weights.
- Over the 8 tasks, each task is equally likely. Cost per correct answer = total expected cost ÷ total expected correct answers. We compute time the same way, with the tries in sequence.
- We assume three things. A validator detects a failed answer, and running it is free. Tries are independent at each task’s observed rate. A failed last Sonnet try is a miss. Every figure that follows is a calculation.
- For the sensitivity, we set Haiku’s pass rate on all tasks at once to the lower end of its 95% Wilson interval, then to the upper end. A fourth setting counts Haiku’s format misses as passes. The result is a sensitivity range, not a confidence interval.
- A joint extreme moves both models at once: Haiku’s pass rate to the upper end and Sonnet’s to the lower end of each task’s 95% Wilson interval. It is more extreme than a 95% interval on the total, and it is not a forecast.
- For the break-even, we scale Haiku’s call cost (or time) by a factor f and keep all pass rates. We solve for the f at which Haiku with one retry then Sonnet equals Sonnet every time: f = (Σ C(Sonnet) − Σ C(Sonnet) × q²) ÷ Σ C(Haiku) × (1 + q). This is arithmetic, not a forecast.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
For 6 models run through their own coding CLIs (Haiku, Sonnet, Opus, Fable, Sol (low) and Luna (low)), how long until the first text, how fast does text stream after it, and what does a longer prompt add?
Protocol: 8 steps
- The protocol file birth precedes the first probe and counted call. Later edits have no frozen versions. Two probes checked the routes and helped calibrate ledger sizes. They stay outside all cells and charts.
- Part A used 24 calls: 4 per model. The prompt asked for 250 numbers in words, one per line. The check compares the reply with all 250 expected lines. Models: Haiku, Sonnet, Opus, Fable, Sol (low) and Luna (low). Controls ran before the first counted call. The reference passes. All 8/8 wrapped references are format misses. All 11/11 planted wrong answers fail.
- Part B used 36 calls: 3 per size for each model. Models: Haiku, Sonnet, Opus and Sol (low). Each call asks one exact lookup question about a synthetic stock ledger. Ledger sizes: 18 lines for 1k, 327 lines for 16k and 1,315 lines for 64k. The 1k, 16k and 64k labels are approximate targets from probe calibration calculations, using estimated Haiku prefix counts. A new seed changes each ledger. Some fixed text repeats. Cache counts do not identify cached text.
- First text means the first non-empty text after CLI start. The time includes CLI start-up and work before the first word. The table also subtracts the CLI ready time (calculation).
- Output speed is a calculation. Visible tokens equal output tokens minus reported reasoning tokens. When the CLI reports no reasoning count, the calculation uses zero. That does not prove the model did no reasoning. Divide visible tokens by the time from first text to call end. The plain form divides all output tokens by that time. It can overstate visible speed when reasoning comes first.
- The harness uses the hard head-to-head machinery. It disables tools, MCP servers and saved sessions. Claude Code uses default effort; Codex CLI uses low effort. Calls run one at a time on one Mac. The timeout was 300 s. Claude Code had a 16,000-token output cap. Codex had no output-token cap. CLI versions: 2.1.286 (Claude Code) and codex-cli 0.160.0.
- Stop rules require a stop at the first usage-limit text. The gate checks other study markers before each batch. No stored batch stopped or trimmed a cell. No counted call duplicates another. A sequence process ended at the gate. Codex part B ran later. Every counted call completed.
- Cache-read share is a calculation: cache-read tokens divided by reported input tokens for each call. The cache table shows the median, observed range and n for each cell.
Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
On small real repository tasks graded by hidden tests, how do coding-agent CLIs compare when they run with their normal file and shell tools?
Protocol: 6 steps
- Six small Node.js repositories: a pagination bug across two files, a CLI flag with exact error text, a refactor that must keep every quirk, an LRU cache to a written spec, a queue race and strict TypeScript types. Each has 1 to 4 visible tests; the hidden checks live outside the repository and run only after the session ends.
- Controls before the first session (Node v25.2.1): every base repository fails its hidden checks and every reference patch passes all of them.
- Claude Code 2.1.286 headless with the sonnet and opus models at the default effort; tools Bash, Read, Edit, Write, Glob, Grep; no web tools, MCP servers or sub-agents; the OS sandbox on, writes limited to the repository, no network; setting sources off. Codex CLI exec with GPT-6.1 Sol at medium effort, workspace-write sandbox, no network, user config ignored (its version was not recorded).
- 6 tasks × 2 repetitions × 3 agents = 36 sessions with a 10-minute timeout; nothing retried. The Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts.
- Pass = every hidden check passes; no partial credit. Recorded per session: wall time, tool calls, tokens as reported, files touched, lines changed against the base commit, commits and edits outside the repository.
- Gemini CLI 0.16.0 was probed first: it asked for a browser login, so it was not run and is not counted. The protocol was declared after the controls and before the first session.
Does memory help Claude Code? 8 kinds of agent memory, tested
Does project memory make Claude Code do better work, and which kind of memory?
Protocol: 6 steps
- One small Node.js repository with real conventions: money in integer cents, coded errors, a clock helper, a log helper, append-only migrations, a generated report, a changelog rule and a late-fee rate that exists only in memory.
- Five tasks: refunds, a report formatting bug, an optional phone field, a late fee "at our standard rate", and a control task where memory does not matter.
- Eight conditions: no memory, the real /init output, a curated 11-line file, 56 raw dated notes (repeats, one-off events and notes that a later note contradicts), the same notes after one "dreaming" pass, the curated facts inside a 210-line handbook, a Stop hook that blocks finishing while a rule is broken, and curated plus hook.
- Claude Code 2.1.286 headless; Sonnet 3 repetitions per cell, Haiku 2; tools Bash, Read, Edit, Write, Glob, Grep; OS sandbox; the tester's own instruction files excluded and auto memory off (both checked with a probe).
- Grading after the session: hidden tests run in a sandbox, plus deterministic checks on the lines the agent added. Full pass = tests and every applicable check. Every attempt counts.
- The protocol was declared before the first counted session. One amendment, made before any Haiku session, added the Haiku lane. One erratum: the raw notes file holds 56 notes, not 58 as first written. The frozen handbook has 210 lines, not 211 as the protocol first stated. Measurements did not change.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
On 8 hard tasks with strict validators, does a higher effort setting buy a higher pass rate for Sonnet, Opus and GPT-6.1 Sol, and what does it cost in time, tokens and list price per pass?
Protocol: 8 steps
- Protocols declared before the first call, one per route. A follow-up to the hard head-to-head (/benchmarks/hard-model-head-to-head), on the same 8 tasks, validators, controls and CLI flags.
- New cells: Claude Sonnet 5.5 (low) · Claude Code; Claude Sonnet 5.5 (medium) · Claude Code; Claude Sonnet 5.5 (high) · Claude Code; Claude Opus 5.5 (low) · Claude Code; Claude Opus 5.5 (medium) · Claude Code; GPT-6.1 Sol (low) · Codex CLI. 96 calls (8 tasks × 2 repetitions per cell), one call at a time per account; order rep-major, then task, then configuration.
- Reference cells (5): Claude Sonnet 5.5 · Claude Code; Claude Opus 5.5 (high) · Claude Code; Claude Opus 5.5 · Claude Code; GPT-6.1 Sol (medium) · Codex CLI; GPT-6.1 Sol (high) · Codex CLI, from the hard head-to-head. Reused, not rerun. Claude reference cells keep repetitions 1-2 (n = 16) so every cell has the same design; all 24 of their calls passed in the hard study.
- Effort: "--effort <level>" for Claude Code and the thread effort for Codex CLI. "default" means the flag was not passed and the CLI chose the level.
- Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 wrapped references are flagged as format misses.
- Strict pass, format miss and wrong answer as in the hard head-to-head. Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn, 300 s timeout per call.
- Stop rules: stop at the first usage-limit or rate-limit text. No batch stopped early, nothing was trimmed or retried, and no call failed.
- Cost per strict pass: list price × reported tokens for every call in the cell, divided by its strict passes. A calculation.
Does thinking pay for Claude Haiku 4.5? Thinking on vs off
Does extended thinking pay for Claude Haiku 4.5: what does it buy in accuracy, and what does it cost in time and money, on typed routing decisions and on hard tasks?
Protocol: 16 steps
- Routing uses the platform’s labelled decision suites: Failure class (failure-class-2026-10-05.1, 18 cases); Message intent (frontdoor-intent-2026-10-04.2, 20 cases); Is it a rule? (memory-is-rule-2026-10-04.2, 12 cases); Context shape (context-shape-2026-10-05.1, 32 cases).
- Production rules select the cases and leave the scored keys open. This gives 82 decisions and 194 scored questions.
- Each decision uses one Claude Code CLI call. Calls use JSON output, tools off, no MCP and no saved session. They use the production system text, prompt and JSON schema. Calls run one at a time; temperature cannot be set.
- The decision-eval runner scores each reply. Exact means every scored question is acceptable. Per-question accuracy counts each scored question. Errors and unanswered questions count as wrong.
- Thinking off sets MAX_THINKING_TOKENS=0 on the claude process. Two uncounted probes check the setting, one per route. The probes and all new counted calls reported zero thinking tokens.
- Thinking off uses new calls. Thinking on uses the CLI default in the recorded routing run. Sonnet 5.5 uses low effort in that recorded run. Both reference arms are reused; they were not rerun.
- The exact two-sided McNemar test uses case-level pairs that differ (binomial, p = 0.5). Rate intervals are 95% Wilson intervals.
- Timing uses wall time and the CLI’s reported API time. The same code computes nearest-rank p50 and p95 from each arm’s call log. Cost uses reported input, cache and output tokens times list price (a calculation).
- Hard tasks use the 8 unchanged tasks and sandboxed deterministic validators from the hard head-to-head.
- Controls ran before the first new model call, but before protocol creation. All 8 reference answers passed (8 tested). All 26 planted wrong answers failed. All 8 wrapped references were flagged as format misses.
- Strict pass means the whole reply passes. A format miss means a lenient extractor finds a passing answer; it never counts as a strict pass. Errors and incomplete attempts count as failures.
- Thinking off used 24 calls: 8 tasks, three repetitions each. Calls ran one at a time, with no effort flag, tools off, no MCP and one turn. The timeout was 300 s; the output-token cap was 16,000. Thinking on reuses 24 recorded default Haiku receipts.
- Cost per strict pass divides all call costs in a cell by its strict passes. Each call cost uses list price times reported tokens (a calculation).
- Recorded order: controls at 2026-10-06T21:27:56.463Z; protocol creation at 2026-10-06T21:29:01.607477Z; then the two probes. The first new counted call started at 2026-10-06T21:29:57.270Z. Reused reference calls predate both the controls and this protocol.
- Declared caps: 106 new counted calls and 112 calls including probes. Observed: 106 counted calls plus 2 probes; the total stayed within both call caps.
- Attempts: 106 counted calls, 2 uncounted probes. Nothing was trimmed or retried, and no run hit a usage limit. Every attempt is in the published extract.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
On 8 hard tasks with deterministic validators, does pass rate separate the Claude Code models and GPT-6.1 Sol through the Codex CLI, and what do speed, tokens and cost per pass add?
Protocol: 10 steps
- Protocol declared before the first call. A follow-up to the five-task head-to-head (/benchmarks/model-head-to-head), where pass rate hit a ceiling.
- 8 hard tasks, each with a deterministic validator that runs in a sandbox without network: Fix an interval-merge function (off-by-one and edge cases); Fix a time-zone day-length function (DST); Write a CSV parser (quoted newlines, strict errors); Predict JavaScript event-loop output order; Solve a multi-constraint room schedule; Write a strict SemVer 2.0.0 regex; Refactor to remove duplication, keep 20 tests green; Write a SQLite reporting query (fan-out, ties, boundaries).
- Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 references wrapped in a fence or prose are flagged as format misses.
- Configurations that reached a model: Sonnet, Opus, Opus (high), GPT-6.1 Sol (medium), Fable, GPT-6.1 Sol (high), Haiku. Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task).
- Strict pass: the whole reply, trimmed, passes the validator. Every prompt states the output format and says no code fence and no other text. This is stricter than the five-task study, which removed one wrapping fence.
- Format miss: the strict check failed, but a lenient extractor (fenced block, outer JSON, first code line, one output line) finds an answer that passes the same validator. Reported apart from wrong answers, never as a pass.
- Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn, 300 s timeout per call. One call at a time per account.
- Attempts: Claude Code: 120 attempts, 120 reached a model, 0 blocked; Codex CLI: 62 attempts, 32 reached a model, 30 blocked. Nothing was trimmed or retried, and no run hit a usage or rate limit.
- One Codex batch was resumed after its process ended (declared in the protocol before the resume): the resume skipped every task, repetition and effort the batch file already held, so no recorded call was repeated or replaced.
- Cost per strict pass: list price × reported tokens for every call in the configuration (cache reads and writes priced as in the five-task study), divided by its strict passes. A calculation.
How much of an AI bill is thinking? Reasoning tokens by model and effort
Across 378 recorded calls of Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol, how many output tokens come from reasoning (thinking)? What do they cost at list price per call and per strict pass, and do they track time? The calls cover eight hard tasks, five short tasks and several efforts.
Protocol: 10 steps
- This study is a calculation over receipts that already exist. We made no new model call. The inputs are three recorded runs. The hard head-to-head gives 152 completed calls. We leave its 30 pre-inference blocked attempts out of token calculations. They remain failures of the original route attempt, not model answers. The effort ladder gives 96 new calls. The five-task head-to-head gives 130 calls. All current counted calls completed. Recorded errors count as failures when their usage is known. We keep the separate blocked batch on record. The original Codex route was blocked. A later batch ran after a successful uncounted probe. It resumed after two completed calls and skipped them.
- Sum check. Before we compute any share, we test one point. Do the output tokens include the reasoning tokens? We use this accounting assumption and check its consistency with the receipts. We run three tests on each CLI route. Test 1: reasoning never exceeds output. Test 2: within one task and model, one more reasoning token adds about one output token. A slope near 1 supports the assumption. Correlation alone cannot prove the counter semantics. Test 3: we compare characters per remaining token with characters per output token on calls whose reasoning counter is zero. Claude Code: consistent with inclusion, not proof. Reasoning never exceeded output (0 of 290 calls). Within one task and model, each extra reasoning token went with 0.99 extra output tokens (leave-one-task-out range 0.98 to 1.01, 290 calls). Output minus reasoning has 1.94 characters per token. Calls with 0 reasoning have 1.98. For comparison, output including reasoning has only 0.45. Codex CLI: consistent with inclusion, not proof. Reasoning never exceeded output (0 of 88 calls). Within one task and model, each extra reasoning token went with 0.97 extra output tokens (leave-one-task-out range 0.97 to 0.99, 88 calls). Output minus reasoning has 2.99 characters per token. Calls with 0 reasoning have 2.76. For comparison, output including reasoning has only 1.36.
- Not reported is not zero. A route reports reasoning when at least one of its calls reports more than 0. Both routes do (Claude Code 212 of 290 calls, Codex CLI 68 of 88 calls). So a reported 0 is the recorded counter, not proof of no internal reasoning. 98 calls reported 0, and their replies look like visible text only (test 3). If any model attempt lacks usable token counts, this builder returns no study. It never prices unknown usage as zero. 0 calls had no usable count.
- Reasoning share of one call = reasoning tokens ÷ output tokens. A configuration gets two numbers. One is the median of its per-call shares. The other is the pooled share (all reasoning tokens ÷ all output tokens). Ranges are the lowest and highest call. They are not intervals.
- Hard tasks: we use all calls of each configuration in the hard head-to-head. That is 16 to 24 calls: 8 tasks × 3 repetitions for Claude Code and 8 × 2 for Codex CLI. Effort ladder: 11 cells of 8 tasks × 2 repetitions, the design of the effort-ladder study. We reuse its reference cells from the hard head-to-head (Claude repetitions 1-2 only). Short tasks: the five validated tasks (10 to 15 calls per configuration).
- Cost is a calculation at list price, not a bill. The calls ran on flat subscriptions. Reasoning cost = reasoning tokens × the model's output price. Remaining output cost = (output tokens − reasoning tokens) × the same price. We use the rest as a visible-answer estimate. Input cost covers the whole prompt. Cache writes use the one-hour list-price assumption; the receipts do not state the cache lifetime. Prices per million output tokens: Claude Haiku 4.5 $5, Claude Sonnet 5.5 $10, Claude Opus 5.5 $20, Claude Fable 5.1 $50, GPT-6.1 Sol $10. Sources: Anthropic list prices of 2026-09-21, OpenAI of 2026-10-03.
- Cost per strict pass = the list-price cost of all calls in the cell, failures included, ÷ the cell's strict passes. Hard and effort cells use the hard-set strict rule. Short-task cells use their original rule, which can strip a wrapping fence.
- Time link: we use all calls of the hard head-to-head and the effort ladder. For each model and route, we compute the Spearman rank correlation of reasoning tokens with total time (with a leave-one-task-out sensitivity range, not a confidence interval). We also compute the least-squares slope in seconds per 1,000 reasoning tokens. Then we compute the slope within task. We centre each task on its own mean. This controls for differences in task means, not effort or batch effects.
- Isolation, tasks, validators and flags follow the hard head-to-head and the five-task head-to-head. Each call ran in a fresh empty folder with tools off and one turn. The timeout was 300 s per call (180 s for the short tasks). One call ran at a time per account.
- Claude Code calls ran with an output cap of 16,000 tokens in all three runs. The highest output of any analysed Claude call was 9,321, so no call hit the cap. The Codex CLI has no cap setting. Its highest output was 2,569.
Inference provider index: 27 models, 52 providers
For the same model, how much do inference providers differ in price, and what does a gateway such as OpenRouter add over the first-party list price?
Protocol: 6 steps
- One snapshot of OpenRouter’s public, keyless API on 2026-10-06: the model list and the endpoint list of 27 curated models (28 requests). Prices, context, quantization and uptime are copied as the API reported them; a field it did not return stays empty.
- Tier: read from the endpoint tag suffix. Flex, priority, fast, ultrafast and batch are named tiers; a region suffix (us, eu, europe, a cloud region) is regional; anything else is standard. This is a heuristic.
- Per-model charts show the standard tier only, with one bar per provider: its cheapest standard endpoint. The endpoint table lists every endpoint in every tier.
- Blended price = (3 × input + output) ÷ 4, a 3:1 input:output token mix. Spread = most expensive ÷ cheapest blended price among standard-tier providers. Both are calculations on reported prices.
- Markup = OpenRouter list price ÷ first-party list price − 1. First-party prices are the vendors’ published list prices as recorded in the product price table. The effective markup adds the credit-purchase fee for card purchases on the Standard plan.
- Gateway delay (time to first token, total time, billed cost per call): not measured: no key in the environment. A live harness is ready; it sends nothing without a key and has a hard spending cap.
Jev vs Claude as a router: accuracy and cost
Should a small dedicated router or a general LLM make the platform’s typed routing decisions?
Protocol: 6 steps
- Cases: the labelled decision suites the platform uses (failure class, message intent, is-it-a-rule, context shape), only the cases production asks a router about.
- Scoring: the repository’s own decision-eval runner. Exact means every scored question in a case was acceptable; key accuracy counts each question.
- Jev numbers come from a live run on 2026-10-06: 3 repeats of the same 82 decisions (246 counted calls), sent one at a time over HTTPS to the TypeSafe API from one Apple M3 Ultra Mac on a home network. A failed call would count as wrong; there were 0. Latency is client wall time, so the network is inside it, and the API reports no server time. Jev’s earlier recorded production run (2026-10-05, same case versions and runner) scored 74 of 82 and has no per-call latency. Cost is a calculation: reported input tokens × the published price.
- Claude routers ran through the Claude Code CLI with the production system text and schema, one call per decision, one pass over the 82 decisions.
- Stability: 73 of 82 decisions were exact in every repeat, 8 in none and 1 in some (a Context shape case, wrong in repeat 2). Jev can return different probabilities for the same request; the repeats show how much.
- Economics: recorded tokens of the benchmark runs repriced at list prices for each model mix (a calculation).
Jev vs Claude routers on unseen decisions: a blind holdout
Does a router keep its accuracy on typed routing decisions that nobody tuned against its answers?
Protocol: 7 steps
- Cases: 56 new cases, 14 per decision type, in the production case shape. The types are failure class, message intent, is-it-a-rule and context shape. Production asks a router about each case. We score only questions the rule leaves open: 125 labelled questions.
- Blindness: the author says they opened no router answer before the label freeze. File times support this order but cannot prove blindness. A script compared every case with existing cases. The author rewrote 17 close cases before the freeze.
- File times: the protocol file dates from 21:35:45 UTC on 2026-10-06. Case drafts and one uncounted second-labeller probe came first. The first counted labelling call came at 21:35:49 UTC. The freeze came at 21:37:02 UTC. The frozen file and its checksum preceded every router call.
- Second labeller: GPT-6.1 Sol through the Codex CLI, 4 counted calls. It labelled every case without the author labels. Our set accepts its label on 115 of 125 (92%, 95% interval 86% to 96%). It matches our first label on 100 of 125 (80%, 95% interval 72% to 86%). Of these questions, 24 accept two or three labels. The lowest agreement was on “artifacts”; see the caveats. Secondary scoring keeps only questions where our set accepts its label.
- Routers: Jev 1.13 used direct HTTPS with the production request body. It made 168 calls in 3 repetitions, with no retries. Claude Haiku 4.5 used CLI default effort; Claude Sonnet 5.5 used effort low. Both used the Claude Code CLI with the production system text and schema. Each made 56 counted calls, one per decision. Each route had one uncounted probe. Total calls stayed within the caps: Jev 169/200, Claude 114/120, labeller 5/5. No arm stopped early or lost cases.
- Scoring: the repository’s decision-eval runner. Exact means every scored question in a case was acceptable. Key accuracy counts each question. We use Wilson 95% intervals and exact McNemar tests on paired cases. The paired tests have no multiple-test correction. Errors count as wrong.
- Cost: reported tokens × list price per 1,000 decisions, a calculation.
Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)
Does Agent resolve more SWE-bench Verified instances with Claude Opus 5.5 as its brain than with Claude Sonnet 5.5, and at what cost and time?
Protocol: 5 steps
- A paired probe declared before any Opus run: 8 SWE-bench Verified instances from the 33 that Agent attempted with Sonnet, 2 per difficulty band (1 Sonnet miss and 1 Sonnet resolve each), picked by a fixed rule.
- Opus arm: Agent with every model call on Claude Opus 5.5 (routing off), balanced mode, real onboarding, cold start, no cost or time cap, one attempt per instance, on platform build 4f6f4027. Sonnet arm: the earlier campaign attempt on the same instance with the same flags and task spec (builds f0ac3a8a and 236c0d3f).
- Grading: Official SWE-bench harness 5.0.2 with the official instance images, amd64 emulation; the gold patch must resolve on this host first. An escalation is graded as delivered; nobody answers it.
- A usage gate starts an instance only when the subscription has 3% or more left in both its weekly and 5-hour windows. The run ledger resumes after a halt and never repeats an attempt.
- Notional cost: the platform price table at each run commit applied to the recorded tokens. The exact McNemar test reads only the discordant pairs.
Prompt cache break-even: after how many reuses does a cached prefix cost less?
With the cache-write surcharge Anthropic lists, after how many reuses does a cached prompt prefix cost less than no cache? What does that mean for a session of 1 to 20 turns, for each model’s price, and for a workload split across sessions?
Protocol: 13 steps
- This study makes no new model call and has no raw file of its own. It calculates from two inputs. The first is the turn tokens that the caching study recorded in 6 Claude Code sessions (3 on Sonnet 5.5, 3 on Opus 5.5). The second is the list prices in the price sources.
- A prefix of T tokens goes out in k + 1 requests. The first request writes it to the cache. The k later requests, the reuses, read it. Prices are USD per million tokens: input i, cache read r, cache write w.
- No cache: (k + 1) × T × i. A 1-hour write: T × w + k × T × r, where w is the listed 1-hour write price (twice i for every Anthropic model in the price list). A 5-minute write: T × 1.25 × i + k × T × r.
- The 1.25 is the multiplier that the caching study protocol states as Anthropic’s published figure. It is not in the price list. It is an assumption here. No 5-minute write occurred in the recorded sessions.
- Break-even: the cached prefix costs less when k > (w − i) ÷ (i − r). The exact break-even is that fraction. The reuses needed is the smallest non-negative whole number above it. A negative fraction means the first request already costs less, with zero reuses.
- A share f of the prefix can already be in the cache when the first request arrives. That request reads the share and writes the rest: T × (f × r + (1 − f) × w) + k × T × r.
- In the recorded sessions, turn 1 read 1,463 tokens on its first request. The usage counters do not record read/write timing. That is f = 19% (8,778 of 46,978 tokens, 6 sessions). Its origin was not isolated. The dollar figures use f = 0, the whole prefix new, which is the stricter case. The break-even rows show both.
- T is the rounded mean turn-1 input (uncached input + cache reads + cache writes) of a model’s recorded sessions. For Sonnet 5.5 it is 7,831 tokens (range 7,831 to 7,832, n = 3). For Opus 5.5 it is 7,828 tokens (the same in every session, n = 3). A session of N turns is N requests that send the prefix, so k = N − 1. The cost tables give 1,000 sessions of 1, 2, 3, 5, 10 and 20 turns.
- Split case: 10 requests as one 10-turn session (one write, 9 reads) against 10 one-turn sessions (10 writes). Each recorded session ran in a fresh process and a new temporary working folder. Of 4 later sessions, 0 met the reuse proxy. The split case assumes a wholly new prefix with no shared reuse; the receipts do not test that exact case. The reuse proxy follows the caching study’s rule: turn 1 reads more than it writes.
- Opus 5.5 price question: the price list gives a cache read of $0.2 per million (5% of the $4 input price). Another table of the product lists $0.4 (10%). This study did not check the vendor price. Every Opus 5.5 row and the second Opus line show both values.
- Check against the recorded sessions, on the input side only (output excluded). The recorded 5-turn sessions saved 55.4% on Sonnet 5.5 (n = 3; session range 54.3% to 57.1%) and 59.2% on Opus 5.5 (n = 3; range 58.5% to 59.7%). These ranges are not confidence intervals. The formula gives 52.0% and 56.0% with a new prefix, and 59.1% and 63.3% with 19% already cached. The two formula cases bracket the recorded figure for both models. The recorded sessions also wrote the tokens that each turn added.
- With output included, the recorded saving is 50.0% for Sonnet 5.5 (n = 3; session range 47.5% to 54.1%) and 53.1% for Opus 5.5 (n = 3; session range 51.4% to 54.3%). These are list-price calculations. The ranges are not confidence intervals.
- Everything is a list-price calculation, not a bill. The sample behind T is n = 3 sessions per model.
Prompt caching and run-to-run consistency in Claude Code and Codex CLI
When a CLI session reuses a fixed context, how much input comes from the cache, what does that save at list price, and does it change latency? When the same prompt runs 10 times, how much do the pass rate, the answer and the time vary?
Protocol: 10 steps
- Protocols declared before the first call, one per route. Every attempt is kept; nothing was retried.
- Caching: 9 sessions (3 × Claude Sonnet 5.5 · Claude Code, 3 × Claude Opus 5.5 · Claude Code and 3 × GPT-6.1 Sol (medium) · Codex CLI), 5 turns each. A session is one CLI process. Turn 1 sends a seeded synthetic stock ledger plus question 1; turns 2-5 send one short question each (lookups, a count, an arg-max), each with one exact answer. Each turn is one model request.
- A 2-call probe sized the context before the run (not part of any cell). The ledger was larger than declared, so it was cut once, from 170 to 100 lines, and the answers were recomputed, as the protocol allowed.
- The Codex sessions ran before that cut, on the 170-line ledger (16,197 characters vs 9,651). The two routes are reported side by side, never as a like-for-like pair.
- Cache counters as the provider reports them: Claude Code gives uncached input, cache reads and cache writes (with the 5-minute and 1-hour split); the Codex app-server gives input (cached included) and cached input, and no cache writes.
- Cost with and without the cache (Claude only): see the calculation source. Every write in this run was a 1-hour write, so the 5-minute multiplier (an assumption) was not used.
- Consistency: 3 prompts (an exact number, a JSON object with exact keys, a small code fix), each with a deterministic validator; 10 repetitions per prompt for Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI. One-shot calls, one at a time per account. Answer diversity counts distinct normalized answers; the raw extract holds ordinal answer ids, never the text.
- Validator controls ran before inference on each route: every reference answer passes, every plausible wrong answer fails, and a wrapped reference is flagged as a format miss.
- Isolation as in the hard head-to-head: fresh empty working folder, tools off, no MCP servers, no session persistence across processes, provider-default caching. Claude at its default effort; GPT-6.1 Sol at medium.
- No batch stopped early and nothing was trimmed. The Claude CLI reported a rate-limit status of "allowed_warning" on 7 of 30 cache turns; no call was refused and the run did not stop.
Routing overhead: deterministic policy vs LLM routers vs Jev
What delay and what cost does each kind of router add before the real work of a call starts?
Protocol: 6 steps
- Protocol declared before any measurement.
- Deterministic policy: the platform’s production routing decision (plan, model ladder and effort) on its default policy, timed in process on one Apple M3 Ultra Mac with Node 25: 5,000 warm-up calls, then 20,000 timed decisions over 64 synthetic routing contexts, plus a 200,000-call batch for throughput. Database reads and the decision record write are out of scope; the recorded rule-based System One decisions (millisecond resolution, record write included) are shown as a stat.
- LLM routers: the recorded routing runs of the routing study (Claude Sonnet 5.5 (effort low, via Claude Code), 82 calls; Claude Haiku 4.5 (thinking on, via Claude Code), 82 calls), one call at a time. Wall time per call, the API time the CLI reported, and the difference (CLI and harness time). p50 and p95 recomputed from the per-call log.
- Jev: a live run on 2026-10-06. The same 82 typed decisions as the routing study, 3 repeats, 246 counted calls sent one at a time over HTTPS to the TypeSafe API from the same Mac, on a home network. Time per call is client wall time from before the request to after the body was read, so the network is inside it; the API reports no server time. Cost per decision is a calculation: the mean input tokens its API reported × the published price. The recorded production run (82 decisions) has the same cost as a provider-reported figure.
- CLI start-up: Claude Code · Claude Haiku 4.5 5 runs, Codex CLI (default model) 5 runs, a one-word prompt, one call at a time. Times to the first output event, the first model output and exit.
- Per 1,000 tasks (a calculation): decisions per task from 48 recorded bench runs read only (2 empty runs excluded) × cost and median time per decision. Decisions are assumed to wait in line, so the delay is an upper bound.
Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
On 8 hard tasks with strict validators, does an agent loop that may write and run code in a sandbox pass more often than one call with tools off, and what does the loop cost in time, tokens, tool calls and list price per pass?
Protocol: 9 steps
- Protocol declared before the first counted call. The same 8 tasks, prompts, validators and strict grading as the hard head-to-head (/benchmarks/hard-model-head-to-head).
- Single call: the task prompt, tools off, one turn, 300 s timeout. Agent loop: the same prompt plus one paragraph ("You may create and run files in the current folder to test your answer. Your final message must be only the answer, in the format stated above."). The final message is graded exactly like a single call.
- New cells: Claude Haiku 4.5 (agent loop) · Claude Code (24); Claude Sonnet 5.5 (agent loop) · Claude Code (16); GPT-6 Luna (single call) · Codex CLI (16); GPT-6 Luna (agent loop) · Codex CLI (16). Reference cells: Claude Haiku 4.5 (single call) · Claude Code (24); Claude Sonnet 5.5 (single call) · Claude Code (24), from the hard head-to-head. One call or session at a time per account; order rep-major, then task, then configuration.
- Agent-loop sandbox: a fresh empty work folder per attempt, outside the temp folder. Claude Code: tools Bash, Read, Edit, Write, Glob and Grep only, the Claude Code sandbox on (writes only in the work folder, no network, no unsandboxed commands), no MCP servers, no user settings, at most 30 turns, the same 16,000-token output cap per response as the single calls. Codex CLI: exec with the workspace-write sandbox (network off), approvals never. 10-minute limit per session.
- Sandbox probes before the matrix (not counted): Claude Code: network blocked (the command was refused before it ran), write to the parent folder refused, file in the work folder created; Codex CLI: network blocked, write to the parent folder refused, file in the work folder created.
- Controls before inference: 8/8 reference answers pass, 26/26 plausible wrong answers fail, and 8/8 wrapped references are flagged as format misses.
- Audit of every agent-loop transcript: file-edit tool calls, read tool paths and paths in shell commands. An attempt that read a file outside its work folder is contaminated: kept in the raw extract, left out of the rates. Edits outside the work folder that ran must be 0; a call the CLI refused before it ran is counted as an attempt, not an access. The reference answers were locked (no read access) while agents ran.
- Stop rules: stop a route at the first usage-limit or rate-limit message; errors, time-outs and turn-limit stops count as fails. No batch stopped early and nothing was trimmed or retried.
- Cost per strict pass: list price × reported tokens for every attempt in the cell, divided by its strict passes. A calculation.
System One arena: Jev vs Clef and five open decision models, head to head
How good are typed-decision models at decisions with a checkable answer, and which one wins when they play each other?
Protocol: 6 steps
- Seven models with the same API (a JSON state in, a choice or a yes/no with probabilities out): TypeSafe Jev 1.13 over its hosted API, and Cloudflare Clef 27B, Clef-Flash 9B, lev 4B, Kev 4B, Laya and Julia-1 as GGUF files (Q4_K_M or Q8_0, checked against the published SHA-256) in llama.cpp 0.6.0 on one Mac Studio (M3 Ultra, 96 GB).
- Five suites, 1,085 items, written by five agents that never called a model: games solved by search (tic-tac-toe, Connect Four, Nim, Pong intercepts, Sudoku, Wordle, mazes, Hanoi, coin weighing, knight moves), logic and thought experiments (syllogisms, knights and knaves, Monty Hall, base rates, expected value, bias probes, framing pairs), fictional company policies in 20 industries, usability intents on 14 kinds of app, and stress tests (prompt injection, long logs, thresholds, near-identical options, paraphrases).
- Every answer comes from search, arithmetic, logic or a rule written in the item. Each suite has its own verifier that re-derives every answer. A second agent re-solved every item of the first versions of the suites (978) and found no wrong answer; the ambiguities and shortcuts it found were fixed before the freeze, adding 107 items and rewriting others. Consistency-only items (no right answer) are scored only on agreement inside their group.
- Exam pass: every item, every model; choice items also with shuffled options and with options renamed a, b, c. Latency pass: 120 seeded items, one model and one call at a time. Repeat pass: 60 items asked twice.
- Matches: a round robin of tic-tac-toe, Connect Four, Nim and Dots and Boxes (10 games per pairing, colours swapped), real-time Pong (first to 3 points after amendment 2) where each answer takes effect only after its measured latency, and a random and a perfect player as anchors. Elo: K 32, averaged over 200 seeded game orders.
- The protocol, the item hashes and the game counts were declared before the first counted call. Two amendments, both declared before the data they affect, are in the published protocol: the local-server batch size (amendment 1) and the Pong match length (amendment 2). One erratum: the Connect Four reference solver graded some moves in a search-order-dependent way; it was fixed and every move regraded, and no pick, probability, result or rating changed (erratum 2). Every call and every game is kept; errors count as wrong.
Voice agent latency budget: component calculations, not a measured turn
Which decision steps fit inside a voice agent’s turn budget, using only measured times?
Protocol: 9 steps
- Thought experiment on recorded times. This study makes no new model calls. Every time comes from a run that another study published.
- Budgets: 300 ms, 800 ms, 1,500 ms per decision step (assumption: a budget a voice team might set). Nothing in the data says what a product needs; change the budget and the table gives the share and the count for it.
- Deterministic routing policy: the policy timing of the routing overhead run (latencyUs p50 and p95, converted to milliseconds). Rule-based System One decision: the recorded decision times of that run (millisecond resolution, record write excluded).
- Jev 1.13: this study checks the per-call rows against the live-run summary. The rows record client wall time over HTTPS. They give n, median and p95 (latencyMs calls). The first call uses a fresh connection (coldFirstCall). This calculation subtracts the later-call median (callsWithoutColdFirst) from the first-call time. The run starts at run.startedAt and ends at run.endedAt. The client uses one Apple Silicon Mac on a home network.
- Claude Sonnet 5.5 and Claude Haiku 4.5 as routers: median and interpolated p95 from the routing extract (latencyMs wallP50 and wallP95). The overhead extract supplies n and the observed limits, through the Claude Code CLI, one call at a time. CLI start-up: first model output on a one-word prompt, from the routing overhead run (firstModelEventMs; 5 runs per CLI).
- Model first output: this study uses the matched provider explorer cohort for the fixed exact reply. Each configuration has 5 runs with firstUsefulMs timings. Routes are the OpenAI API and the Codex CLI. This study also selects the Claude Code configuration with the lowest observed first-output median in the 5-task head-to-head. Its timing sample includes completed validator failures. Failed or untimed calls have no successful first-output time.
- Slow end: the 95th percentile where a step has 30 or more runs, and the slowest run where it has fewer. The slowest run is stricter than a 95th percentile.
- Calculations: share of a budget = time ÷ budget, at the median and at the slow end. A step fits when its time is at or below the budget. Runs that fit back to back = the budget ÷ the time, rounded down, counted up to 1,000. A router in front of a model adds the two medians and the two slow ends, as a scenario. Neither sum is a measured percentile or a guaranteed upper bound for the pair.
- The charts separate steps with a median under 1,000 ms from the rest so each axis stays readable. Separate charts show median-to-p95 bands and observed min-to-max ranges. A range maximum is not a p95.
What does routing a million AI requests a day cost? A calculation from measured runs
At 10,000 to 10 million routing decisions a day, what does each router cost per day, what in-flight and waiting scenarios do median and p95 times give?
Protocol: 6 steps
- This is a calculation, not a run. We made no new model call and ran nothing at scale. Every input is a number from a recorded run in this dataset. The volumes (10,000, 100,000, 1 million, 10 million decisions a day) are parameters, not measurements.
- Jev cost per 1,000 is a calculation from the routing study’s provider-reported total cost and counted calls. Claude cost comes from its token totals and list prices, with one-hour cache writes. Jev costs $0.0337 per 1,000, a calculation from the reported production run cost (n = 82). Sonnet 5.5 at effort low and Haiku 4.5 ran through the Claude Code CLI. Their cost is list price × the tokens the CLI reported (Sonnet n = 82, Haiku n = 82). The policy makes no model call, so it costs $0. The live Jev run gives $0.0337 per 1,000, a calculation from reported input tokens and the published price. The two figures agree at this precision.
- Claude median and p95 times come from the routing receipts; we checked them against the per-call rows. The overhead summary uses different quantiles for those calls, so we do not mix its median or p95 into this calculation. Run ranges and policy times come from the overhead results. The policy is the production routing decision, timed in process over 20,000 decisions (median 1.42 µs, 95th percentile 2.33 µs, maximum 2538.21 µs; minimum unavailable). Sonnet 5.5 and Haiku 4.5 are wall time per call through the CLI, one call at a time, from the recorded routing runs (Sonnet n = 82, Haiku n = 82). Jev is the live run: 246 calls (82 typed decisions × 3 repetitions) over direct HTTPS from one Mac, one call at a time, as client wall time.
- Daily cost = decisions a day × cost per decision. Decisions per second = decisions a day ÷ 86,400. Decisions in flight = decisions per second × time per decision in seconds (median/p95 scenarios; Little’s law requires the mean). The median time gives the bar. The 95th percentile time gives the whisker. Waiting hours a day = decisions a day × time per decision ÷ 3,600, if each request waits for its decision. Cost per year = cost per day × 365.
- The scenarios use 1,000,000 model calls a day. In the first, a router decides every call. In the second, it decides only the System One decisions: 7 ÷ 49.5 = 14.1% of calls. That ratio uses the medians of 48 recorded bench runs. Routing was off in those runs, so each model call counts as one decision a router could make. In the third, the policy decides every call in process.
- Policy capacity is the measured 469,409 decisions per second in one process, set against the rate each volume needs. The measurement leaves out the database reads and the decision-record write of the production decision.
Agent on SWE-bench Verified vs 11 public models
How does Agent, a full worker pipeline on one model, do on SWE-bench Verified next to public single-model runs on the very same instances?
Protocol: 6 steps
- Sample: 25 of the 500 Verified instances, stratified by public difficulty (seed 20261004). Difficulty is how many of the 11 public mini-SWE-agent v2 runs solved the instance.
- Campaign 1 ran 25 instances on platform build f0ac3a8a. Six compiled-extension instances could not import in the worker checkout, so the declared rule replaced them in the same band.
- Campaign 2 ran those compiled instances (plus two replacement candidates) on build 236c0d3f after a sandbox fix.
- One attempt per instance. No retries, no operator answers: an escalation is graded on what was delivered.
- Grading uses the official SWE-bench harness and instance images (under amd64 emulation). The gold patch resolved on the host for every instance.
- Agent runs its full pipeline (onboarding, research, plan, act, verify, review) on claude-sonnet-5-5 through a subscription CLI. Costs are list-price estimates of the recorded tokens.
AI pull requests vs merged human pull requests, judged blind
When critics cannot see which change came from a person, do they prefer the AI worker’s pull request or the one the maintainers merged?
Protocol: 5 steps
- Each task is a real merged pull request: the issue as the ticket, the repository at the base commit, the merged change as the human reference.
- The Agent worker solves the ticket in a sandbox without seeing the reference. Its pull request and the merged one form a pair.
- Critic models (Claude Opus 5, Claude Fable 5, Claude Sonnet, and on some pairs Codex or GPT 5.5) review both changes without labels, once in each order.
- Decision rule 2: the AI change wins on a strict majority of verdicts and no panel veto (serious issues flagged by at least two critics).
- Every scored attempt is kept. A re-scoring of the same pair under the current rule replaces the earlier scoring.
Claude Code CLI vs Codex CLI vs the API: latency and tokens
How much time and how many tokens does a coding CLI add on top of the model, and how do Claude Code and Codex compare on the same repair task?
Protocol: 4 steps
- Receipts from the provider explorer: each run records the route (CLI or API), the model, the effort, timings, reported tokens and a deterministic validator result.
- First useful output is the first streamed text that belongs to the answer. Total time runs from launch to exit, including CLI start-up.
- Only receipts classified "evaluated" count. Excluded runs (context mismatch, pilot runs, unsupported controls) and diagnostics are kept in the raw file but not charted.
- The matched cohort ran every configuration back to back on the same host with the same prompt.
Coding calibration: what broke on three real pull requests
On three real upstream issues, does the AI worker deliver a verified change, and what stops it when it does not?
Protocol: 5 steps
- Tasks: fastify/session #348, h3js/h3 #1533 and Kludex/uvicorn #3036, each frozen at the base commit with the merged change as reference.
- Fixed model claude-sonnet-5-5, routing off, balanced mode, real cold onboarding, no review replay, no operator answers.
- Worker and gates run offline with frozen dependencies. Gates run install, lint, typecheck, build and the full suite at base, candidate and reference, plus cross-tests.
- "Functional pass" means the frozen patch passed the gates. "Verified delivery" also needs a completed, clean, unassisted run.
- One attempt per task per slice. Every stop and failure is kept; nothing is rerun or replaced.
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
On short tasks with strict validators, how do the Claude Code models and efforts compare with Codex on pass rate, speed and tokens?
Protocol: 6 steps
- Protocol declared before the first call: 5 cases with deterministic validators (behavioral checks in a sandbox, canonical JSON or exact text).
- Claude Code: Haiku, Sonnet, Opus and Fable at default effort, plus Opus at low and high, 3 repetitions. Codex: GPT-6.1 Sol at low, medium and high, 2 repetitions.
- Isolation: fresh empty working folder, tools off, no MCP servers, no session persistence, one turn. One call at a time per account.
- Every attempt is kept. Nothing is retried. A cell that did not run is shown as trimmed.
- All 130 planned calls ran. Nothing was trimmed or retried, and no run hit a usage or rate limit.
- Cost per passing answer: list price × reported tokens for every call in the configuration (cache reads and writes priced separately), divided by its passes.
What if every call ran on Opus? Repricing real agent tokens
Agent recorded every token it used on 33 SWE-bench instances. What would the same tokens cost at other models’ list prices, and what did caching save?
Protocol: 4 steps
- Tokens: the sum over every model call in the run telemetry of all 33 SWE-bench attempts (input, cache reads, one-hour cache writes, output).
- Prices: list prices recorded in the product price table, effective 2026-09-21 (Jev 2026-09-23, OpenAI rows 2026-10-03).
- Formula: uncached input × input price + cache reads × cache-read price + cache writes × write price + output × output price. Anthropic one-hour writes cost twice the input price; other vendors’ writes are priced as plain input.
- Cost per resolved keeps every attempt’s cost in the numerator and divides by the 25 instances Agent resolved.
Controls
What we hold fixed so a gap means something
01 One file, no typed numbers
A page, a story or an export never types a number by hand. Each one reads dataset.json, the file you can download below.
02 Same series, same study
Both sides of a comparison row come from the same chart series or stat in the same study: same tasks, same validators, same route rules.
03 The fairest pair
When a model ran in several configurations, the row uses its preferred one, then the same effort on both sides, then the default effort. The row prints both contexts.
04 Like with like
A model is never compared with the CLI that ran it. Providers are compared only with providers. Effort comparisons change the effort and nothing else.
05 Calculations stay calculations
Repricing recorded tokens at another list price is a thought experiment, never a run. It is drawn hollow, carries a Calculation chip and never becomes a comparison row.
06 Private stays private
Prompts and model output are never copied. Private tasks publish only their language and kind; public task ids stay.
Counting rules
What counts as a pass, a miss and an attempt
01 Every attempt counts
Failures, timeouts and wrong answers stay in the denominator. Nothing is retried quietly until it passes.
02 Format misses apart from wrong answers
A strict validator fails a correct answer in the wrong format; we report both readings. On the hard set: 5 of 13 non-passes (8 wrong answers).
03 Blocked is not scored, and is shown
An attempt stopped before any model call is not a model result, so it is left out of rates and reported on its own. Hard set: 30 (Codex CLI; 0 model calls).
04 Unknown stays unknown
A study without raw input is left out. A latency nobody measured says "not measured", never a placeholder.
05 No direction, no winner
Tokens, calls and counts describe behaviour; more is not better or worse. Their rows are a tie (same value) or unclear.
06 Zero is not a sample
A point with n = 0 makes no fact and no comparison row.
Intervals
Three span kinds, three glyphs
A span next to a number is one of three kinds, and each is drawn differently, so a range of runs never passes for a confidence interval.
95% Wilson interval
For rates. A tinted capsule with end caps. Two configurations whose capsules overlap are a tie.
Example: 91% (139/152) strict passes on the hard set, 95% Wilson 86%–95%.
fastest–slowest run (not an interval)
For run times. A thin track from the fastest to the slowest run, without caps. It is not a confidence interval.
Example: Claude Sonnet 5.5 (low) · Claude Code, median 5.8 s, fastest 2.8 s to slowest 20 s, n = 16.
median to p95 (not an interval)
For timings with many runs. A one-sided track from the median to a tick at the 95th percentile. Not a confidence interval.
Example: Jev 1.13 (TypeSafe), median 137 ms, p95 196 ms, n = 246.
Interactive
Interval playground: when is a gap real?
Move n and k. Each 95% Wilson interval narrows as the runs add up; overlapping intervals are a tie.
Tie: the 95% intervals overlap (87%–95%), so this data cannot separate Jev 1.13 and Claude Sonnet 5.5.
Interval width as n grows, at 90%
±7 points at n = 82
| Configuration | k | n | Rate | 95% Wilson interval |
|---|---|---|---|---|
| Jev 1.13 | 74 | 82 | 90% | 82%–95% |
| Claude Sonnet 5.5 | 77 | 82 | 94% | 87%–97% |
Two configurations: Jev 1.13 74 of 82 (90%, 95% interval 82% to 95%) and Claude Sonnet 5.5 77 of 82 (94%, 87% to 97%). Tie: the 95% intervals overlap (87%–95%), so this data cannot separate Jev 1.13 and Claude Sonnet 5.5.
The interval is the Wilson score interval, the formula behind every rate on this site. Values you set here are a calculation, not a result. The preset loads published counts: see the study.
Plan a sample size for two pass rates. This calculation holds each rate fixed; it is not statistical power or a paired test.
Comparison rules
When one result beats another
A side wins a row only when the two 95% intervals, or the two run ranges, do not overlap. Across all 189 comparisons, that leaves most rows without a winner, and the pages say so.
- 64Separated (4%)the intervals or run ranges do not overlap
- 595Tie (38%)the 95% intervals overlap, or the values are the same
- 912Unclear (58%)run ranges overlap, too few runs, or no interval
1,571 rows in 189 comparisons; 250 of them are list-price calculations and are counted apart from measured rows in every verdict.
01 Separated intervals
A winner needs two 95% intervals, or two run ranges, that do not overlap, and a unit with a better direction: higher rate or score, lower time or cost.
02 Enough runs
A range-based winner needs at least 5 runs per side. Below 10 runs per side the basis says the samples are small.
03 Overlap is a tie
Overlapping 95% intervals give a tie. Overlapping run ranges, or no interval at all, give unclear: the gap is stated, not tested.
04 Median above p95
A p50-p95 basis says the median of one side is above the 95th percentile of the other, and that this is not a confidence interval.
05 No composite score
The studies differ in task, route, effort and sample size. Readers rank within one metric, where these rules already say when a gap is real.
Enforced by the leaderboard, which sorts by evidence, not by a score
Honesty rules
Nine rules every chart follows
The rules are code: tests fail a build that breaks them.
01 n is visible
Every value with an n shows it in the label or the interval line, not only in a tooltip.
Enforced by the honesty tests (rule 1)
02 The interval kind is written
"95% Wilson interval", "fastest–slowest run (not an interval)" or "median to p95", next to the chart.
Enforced by the honesty tests (rule 2)
03 Calculations are hollow
A calculation is drawn hollow, carries a Calculation chip and states its formula. Thought experiments get a framed banner.
Enforced by the honesty tests (rule 3)
04 No fake precision
Labels use the published text. Count-ups never show more decimals than the data; ratios use two significant figures.
Enforced by the honesty tests (rule 4)
05 Honest axes
Bars start at zero. Rates sit on a fixed 0–100% axis. A log scale says "log scale".
Enforced by the honesty tests (rule 5)
06 Winners come from the data
The page never computes a winner, a rank score or a composite; it shows the row verdict.
Enforced by the honesty tests (rule 6)
07 Ceilings are called out
When three or more rows sit at the top with overlapping intervals, the chart says the task set cannot separate them.
Enforced by the honesty tests (rule 7)
08 Kinds never share an unmarked axis
Measured, calculated and reported values are never mixed on one axis without a mark.
Enforced by the honesty tests (rule 8)
09 Same numbers everywhere
Live stories and exports show the same numbers, notes and interval labels as the page.
Enforced by the story tests
Raw data
Check it yourself
Every source has a date. Our own runs link the sanitized extract the compiler read; vendor prices link the vendor page.
Our recorded runs 23
JSON schema vs instructions
Prompt cache across sessions
LLM speed anatomy
Harder tasks head-to-head
Provider head-to-head, hard set: eight hard tasks with strict validators
Effort ladder: the hard task set at each effort level
System One arena: typed-decision models on checkable decisions and in head-to-head games
Agent memory study: 8 kinds of project memory on Claude Code
Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks
Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim)
Caching sessions and repeated prompts (Claude Code and Codex CLI)
Jev live run: 246 timed calls on the 82 routing decisions
Routing overhead runs: policy microbenchmark and CLI start-up
Single call vs agent loop
Haiku thinking on vs off
Routing on unseen holdout decisions
Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)
Coding calibration: fastify/session, h3, uvicorn
Provider head-to-head: Claude Code models vs Codex efforts
Routing runs: Jev router vs LLM routing
Agent on SWE-bench Verified, campaign 1 (25 instances)
Provider explorer receipts: CLI vs API
Blind review panel: AI worker change vs merged human change
Calculations from stated formulas 6
Reasoning token bill (calculation)
Voice-agent latency budget calculation
Routing at scale calculation
Routing overhead per 1,000 tasks (calculation)
Cost with and without the prompt cache (calculation)
Repricing calculation
Vendor and gateway price lists 6
Third-party published results 1
Declared rules 1
SWE-bench campaign rules and sample design
Per-study downloads
- Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLIJSONCSV
- Does a new Claude Code session reuse the prompt cache of an earlier one?JSONCSV
- GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasksJSONCSV
- Is Claude Haiku cheaper? Retry and escalate, calculated on real receiptsJSONCSV
- Where the seconds go: first text, output speed and prompt size for 6 LLMsJSONCSV
- Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasksJSONCSV
- Does memory help Claude Code? 8 kinds of agent memory, testedJSONCSV
- Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasksJSONCSV
- Does thinking pay for Claude Haiku 4.5? Thinking on vs offJSONCSV
- Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasksJSONCSV
- How much of an AI bill is thinking? Reasoning tokens by model and effortJSONCSV
- Inference provider index: 27 models, 52 providersJSONCSV
- Jev vs Claude as a router: accuracy and costJSONCSV
- Jev vs Claude routers on unseen decisions: a blind holdoutJSONCSV
- Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)JSONCSV
- Prompt cache break-even: after how many reuses does a cached prefix cost less?JSONCSV
- Prompt caching and run-to-run consistency in Claude Code and Codex CLIJSONCSV
- Routing overhead: deterministic policy vs LLM routers vs JevJSONCSV
- Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasksJSONCSV
- System One arena: Jev vs Clef and five open decision models, head to headJSONCSV
- Voice agent latency budget: component calculations, not a measured turnJSONCSV
- What does routing a million AI requests a day cost? A calculation from measured runsJSONCSV
- Agent on SWE-bench Verified vs 11 public modelsJSONCSV
- AI pull requests vs merged human pull requests, judged blindJSONCSV
- Claude Code CLI vs Codex CLI vs the API: latency and tokensJSONCSV
- Coding calibration: what broke on three real pull requestsJSONCSV
- Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-headJSONCSV
- What if every call ran on Opus? Repricing real agent tokensJSONCSV