Open data · October 7, 2026

AI benchmarks you can check.

Claude, Codex and routers, measured on real tasks: pass rates with intervals, latency, tokens and cost per passing answer. Every number links to its sample, its method and its raw data.

28 studies · 206 charts · 37 sources · JSON and CSV for every study

  1. Structured Output96(72 Claude Code, 24 Codex CLI) 
  2. Prompt Caching0of 2 (95% interval 0% to 66%)n = 2
  3. Head to head39%(22/56)n = 56
  4. Thought experiment46%(11/24)n = 24
  5. Latency1.8x n = 23
  6. Claude Code4.9×n = 12
  7. Claude Code40%(6/15)n = 15 Live story
  8. Effort100%(96/96)n = 96 Live story
  9. Claude Haiku87%(71/82)n = 82
  10. Head to head91%(139/152)n = 152 Live story
  11. Thought experiment92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%)n = 24
  12. Inference265endpoints, 52 providers, 27 modelsn = 27 Live story
  13. Routing90%(74/82)n = 82 Live story
  14. Routing82%(46/56)n = 56
  15. SWE-bench67%(2/3)n = 3
  16. Prompt Caching2reuses (the 3rd request)Calculation
  17. Prompt Caching50%($0.1350 vs $0.2698)Calculation Live story
  18. Routing1.42µsp95 2.33 µs, p99 3.04 µsn = 20,000 Live story
  19. Agent Loop54%(13/24)n = 24
  20. Jev1,085 n = 1,085 Live story
  21. Voice Agent14stepsn = 14
  22. Thought experiment$0.0337 n = 82
  23. SWE-bench76%(25/33)n = 33 Live story
  24. Code Review75%(9/12)n = 12 Live story
  25. Claude Code100%(194/194)n = 194 Live story
  26. Calibration1 of 3n = 3
  27. Head to head98%(127/130)n = 130 Live story
  28. Thought experiment162.9M n = 33 Live story

Studies

Each study answers one question. Newest first. The glyph is the study’s lead chart.

  • Prompt Caching
  • Cache Reuse

Does a new Claude Code session reuse the prompt cache of an earlier one?

30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.

0of 2 (95% interval 0% to 66%) · Later sessions with at least 50% of turn-1 input cached, A: new folder each time · n = 2

4 chartsUpdated October 7, 2026

Includes calculations
  • Thought experiment
  • Claude Haiku

Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts

Haiku 4.5 first, then retry or escalate to Sonnet 5.5? A calculation on 64 recorded calls across 8 hard tasks: cost and time per correct answer.

46% (11/24)Strict passes, Claude Haiku 4.5 · Claude Code, eight hard tasks · n = 24

6 chartsUpdated October 7, 2026

Live story
  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

  • Claude Haiku
  • Extended Thinking

Does thinking pay for Claude Haiku 4.5? Thinking on vs off

Claude Haiku 4.5 with extended thinking on and off: 82 routing decisions and 8 hard tasks. Accuracy with 95% intervals, time and cost.

87% (71/82)Claude Haiku 4.5 (thinking off): exact routing decisions · n = 82

6 chartsUpdated October 6, 2026

Live story
  • Head to head
  • Hard Tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.

91% (139/152)Calls that passed strictly (hard set) · n = 152

7 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Live story
  • Inference
  • Providers

Inference provider index: 27 models, 52 providers

Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.

265endpoints, 52 providers, 27 models · Provider endpoints in the snapshot · n = 27

34 chartsUpdated October 6, 2026

Live storyIncludes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Jev

Jev vs Claude routers on unseen decisions: a blind holdout

Jev 1.13, Claude Haiku 4.5 and Sonnet 5.5 on 56 routing decisions written blind and frozen first: exact rate with 95% intervals, tuned vs unseen.

82% (46/56)Jev 1.13 (TypeSafe): exact on unseen decisions · n = 56

6 chartsUpdated October 6, 2026

Includes calculations
  • Prompt Caching
  • Break Even

Prompt cache break-even: after how many reuses does a cached prefix cost less?

A calculation on list prices: reuses before a cached prompt prefix costs less, per Claude and GPT model, and cost per 1,000 sessions of 1 to 20 turns.

2reuses (the 3rd request) · Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation)

3 chartsUpdated October 6, 2026

Live story
  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Live story
  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Live story
  • Jev
  • Clef

System One arena: Jev vs Clef and five open decision models, head to head

Jev 1.13, Clef 27B, Clef-Flash and four small open decision models on 1,085 checkable decisions and in head-to-head games, Pong included.

1,085Decision items, written by agents that never called a model (1,048 graded, 37 consistency-only) · n = 1,085

20 chartsUpdated October 6, 2026

Live story
  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Live story
  • Code Review
  • AI vs human

AI pull requests vs merged human pull requests, judged blind

A blind panel of Claude and GPT critics preferred Agent's change over the merged human change on 9 of 12 real tasks. Votes, scores, caveats.

75% (9/12)Tasks where the panel preferred the AI change (latest attempt) · n = 12

4 chartsUpdated October 5, 2026

Live story
  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

  • Calibration
  • Coding Agents

Coding calibration: what broke on three real pull requests

One attempt per task, four platform builds, failures kept: how an AI worker did on real fastify/session, h3 and uvicorn issues, and what broke.

1 of 3Verified deliveries, latest build · n = 3

4 chartsUpdated October 5, 2026

Live story
  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Live storyIncludes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

Explore the data

The same numbers, arranged by pair, by model, by metric, by provider and by budget.

  • Compare two side by side (189)

    Every pair measured on at least two shared metrics, with the rows the data separates and the rows it does not.

  • Models, CLIs and routers

    One page per system with every value the studies recorded for it, its configuration and its interval.

  • Leaderboard, metric by metric

    No composite score. Pick a metric and see which systems the intervals can and cannot separate.

  • Inference providers

    Every listed provider’s price for each model, and how far the prices spread for the same model.

  • Which model should I use?

    Pick a job, then cost or speed. See measured configurations, sample sizes and what the data cannot separate.

  • AI cost calculator

    Price a month of work from recorded token mixes at each vendor’s list price, with and without prompt caching.

Popular pairs

How we keep the numbers honest

  • Sample size on every number

    Each rate shows n and a 95% Wilson interval. Timings show the median and the observed range.

  • No ranking the data cannot carry

    When intervals overlap, we say the data does not separate the systems, even when our own product is in the row.

  • Calculations are labelled

    "What if every call ran on Opus" is a calculation from recorded tokens and list prices, never a run. The page says so.

  • Open data, failures kept

    Every study downloads as JSON and CSV. Failed and capped attempts stay in the data.

Read the task sets: pass rules, controls and recorded outcomes. Want everything at once? Download the full dataset (28 studies, schema agent-public-bench@1).

From the blog

All posts

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.