- Benchmarks
- Changelog
Latest change October 10, 2026
What changed, day by day.
New and updated studies, posts, explainers, sources and dataset builds. The log is built from the dates the content carries, so it cannot drift from what the pages show.
Changelog RSS Blog RSSHow we measure
192
Changes logged
12
Days with a change
28
Studies published
113
Posts and explainers
Activity
Changes per day
One square per change. Days without a change are left out.
- 1Feb 17February 17, 2026: 1 change
- 2Sep 21September 21, 2026: 2 changes
- 1Sep 23September 23, 2026: 1 change
- 2Sep 28September 28, 2026: 2 changes
- 1Oct 2October 2, 2026: 1 change
- 3Oct 3October 3, 2026: 3 changes
- 2Oct 4October 4, 2026: 2 changes
- 21Oct 5October 5, 2026: 21 changes
- 65Oct 6October 6, 2026: 65 changes
- 81Oct 7October 7, 2026: 81 changes
- 3Oct 8October 8, 2026: 3 changes
- 10Oct 10October 10, 2026: 10 changes
- Studies
- Posts and explainers
- Sources
- Dataset builds
The log
10 changes
- New post
AI week, October 2–8, 2026: research, workers and interfaces
Seven dated AI updates from October 2–8, 2026, with practical limits for research, subagents, API credits, interfaces and decision models.
- New post
Claude Frontier Academy and the hard part of production AI
Anthropic's Frontier Academy plan points to a production skill gap: ownership, recovery, security and evidence matter after the demo.
- New post
Claude Haiku 5.5 subagents: a lead-worker pattern worth testing
Haiku 5.5 is positioned for focused coding subagents. Here is a practical lead-worker workflow with evidence, escalation and review.
- New post
Claude Max included API credits: what the allowance covers
Eligible Claude Max plans may include monthly Claude Platform API credits. Learn the amounts, claim path, expiry and interactive-limit boundary.
- New post
Claude vs Codex for a personal AI workflow: choose the route that holds up
A personal Claude-versus-Codex choice, with practical route comparisons that avoid universal benchmark claims.
- New post
How to measure AI subscriptions when your work spans accounts and machines
A practical way to track AI allowances, local activity and reset dates when several agents run across a laptop and a remote Mac.
Show 4 more changes on Oct 10
- New post
Intelligent UI: when the question becomes the interface
OpenAI's Intelligent UI points to answers with charts and controls. A Tiersel build shows why context, computation and evidence must stay separate.
- New post
Liquid AI d1 decision models: small choices inside a larger workflow
Liquid AI's open-weight d1-3B targets structured decisions in one forward pass. Learn where a small decision model fits and what the latency means.
- New post
My 96 GB Mac Studio still hit a memory warning
A real Mac Studio memory warning stopped parallel AI work. The dialog reported 625.49 GB for two apps; here is what the screenshots prove.
- New post
OpenAI's math repository makes verification part of the release
The OpenAI math repository lists 719 manuscripts and 372 result families. Its formalizations, withdrawals and history show how evidence changes.
- New post
3 changes
- Post updated
Claude Code vs Codex CLI vs the API: latency, time to first token and hidden prompts
194 timed runs. Codex CLI took 3.5x as long as the OpenAI API for a one-line answer and sent 19,551 input tokens instead of 17. What a coding CLI adds.
- Post updated
Claude Sonnet 5.5 vs Opus 5.5: when is Opus worth the price?
Sonnet 5.5 and Opus 5.5 tied on every quality test we ran, easy, hard and agentic. Opus cost 1.6x to 2.6x per unit of work. Where the gap comes from.
- Post updated
How much does prompt caching actually save? Measured in Claude Code and Codex
Claude Code read 97% of later-turn input from the cache. At list price that halved a 5-turn session and cut an agent bill about 3.9x. Turn 1 costs more.
- Post updated
81 changes
- Dataset rebuilt
28 studies, 206 charts, 37 sources
The one file every page, story and download reads.
- New study
Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
96 (72 Claude Code, 24 Codex CLI)Counted calls in this study (every one counted)
For the same extraction tasks, does enforcing a JSON schema through the CLI change the strict pass rate, the format misses and the wrong values, compared with asking for JSON in the prompt?
- New study
Does a new Claude Code session reuse the prompt cache of an earlier one?
0 of 2 (95% interval 0% to 66%)Later sessions with at least 50% of turn-1 input cached, A: new folder each time
In the earlier caching study, a new Claude Code session did not read the cache that an earlier session wrote. Does a fixed working folder change that, and does putting the ledger in the system prompt help?
- New study
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
39% (22/56)Counted calls that passed strictly (harder set)
On a task set built so that Claude Sonnet 5.5 did not pass it every time, does pass rate separate GPT-6.1 Sol (Codex CLI) from Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 (Claude Code)?
- New study
Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts
46% (11/24)Strict passes, Claude Haiku 4.5 · Claude Code, eight hard tasks
On the 8 hard tasks, what does one correct answer cost, and how long does it take, if you try Claude Haiku 4.5 first and retry or escalate, against Claude Sonnet 5.5 every time?
- New study
Where the seconds go: first text, output speed and prompt size for 6 LLMs
1.8x · n = 23The same 4,327-character reply: most tokens ÷ fewest tokens across models (calculation)
For 6 models run through their own coding CLIs (Haiku, Sonnet, Opus, Fable, Sol (low) and Luna (low)), how long until the first text, how fast does text stream after it, and what does a longer prompt add?
Show 75 more changes on Oct 7
- New post
A latency budget for voice agents: which LLM steps fit in one turn?
136.5 ms for Jev, 0.82 s for a small-model API, 2.79 s to 3.79 s for Codex CLI: which steps fit a voice agent latency budget? A thought experiment.
- New post
A voice agent latency budget, with measured times: what fits in one turn?
Rules and Jev 1.13 fit every budget we assumed; a Claude router through a CLI fits none. 14 measured steps vs 300, 800 and 1,500 ms. A thought experiment.
- New post
AI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
- New post
AI coding cost per developer: a formula built on recorded work
AI coding cost per developer: a monthly budget formula using recorded tokens and list-price calculations. Replace our task volume and pass rate with yours.
- New post
Best LLM for JSON output? Haiku, Sonnet and GPT-6.1 Sol, 10 runs each
Best LLM for JSON output? On one prompt, Sonnet and GPT-6.1 Sol passed 10/10 (95%: 72–100%); Haiku passed 1/10 (2–40%). CLI results, not JSON mode.
- New post
Claude Code cost per task: a price ladder from one decision to one agent run
$0.087 per Claude Code coding session at API list prices, $0.0036 per short call, $2.81 per full agent attempt. Calculations on recorded tokens, not bills.
- New post
Claude Code Stop hook, tested: it enforced the rules, used 1.6x the median input and lacked a fact
Claude Code Stop hook: 60/60 Sonnet code-rule checks (95% interval 94–100%). Median input was 1.6x no memory (calculation). Limits and raw data.
- New post
Claude Fable 5.1 vs Opus 5.5 vs Sonnet 5.5: speed, tokens and price tested
24 of 24: Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 each passed every hard task. Fable cost 3.3x Opus and 6.5x Sonnet per pass (list-price calculation).
- New post
Claude Haiku 4.5 vs Sonnet 5.5: all 80 comparison rows, and where the small model loses
Haiku 4.5 vs Sonnet 5.5 on 80 rows: Sonnet ahead on 14, Haiku on none, 31 ties. Hard tasks 11/24 vs 24/24, plus speed, memory and price.
- New post
Claude tokens per second and time to first token: six models timed, and what a longer prompt adds
60 timed calls: time to first text for Claude Haiku, Sonnet, Opus, Fable, GPT-6.1 Sol and Luna, their output speed, and what a 64k prompt adds.
- New post
Decision models play each other: Jev vs Clef at Connect Four, Nim and Pong
Seven decision models played 1,620 games. Only Jev and Clef-Flash clearly beat random; speed alone won nothing.
- New post
Devin's $0.60 per task and our $3.71 per resolved task are different numbers
Devin reports $0.60 per task. Our list-price calculation gives $3.71 per resolved SWE-bench task. Compare the units with a table and buyer checklist.
- New post
Does a JSON schema stop format misses? Haiku went from 0/24 to 18/24
96 calls on 3 JSON tasks. Haiku 4.5 passed 0/24 with instructions and 18/24 with a JSON schema. Sonnet 5.5 and GPT-6.1 Sol passed 12/12 either way.
- New post
Does Claude Code reuse the prompt cache across sessions? A fixed-folder test
Claude Code showed near-full cache reads in later fixed-folder sessions (4/4; 95% interval 51–100%). New folders: 0/2 (0–66%). Exploratory, 30 calls.
- New post
Does LLM routing save money? The saving, the router and the net
Routing would save 2.8% ($3.01) on 2,362 recorded calls (a calculation). A Sonnet router on every call costs about $11.80, so the net is a loss.
- New post
Format misses vs wrong answers: your LLM eval may be failing right answers
5 of 13 failed calls on our hard set were right answers in the wrong format. How to grade LLM output strictly, test validators first and report both numbers.
- New post
GPT-6.1 Sol vs Claude Opus 5.5 and Sonnet 5.5 on harder tasks: only one strict pair separates
GPT-6.1 Sol vs Claude Opus and Sonnet on four selected harder tasks. Pass intervals overlap for these models; only Sol vs Haiku separates on strict passes.
- New post
GPT-6.1 Sol vs Claude Sonnet 5.5 vs Opus 5.5: every row we measured
Of 70 comparison rows for GPT-6.1 Sol, Claude Sonnet 5.5 and Opus 5.5, only 2 have a winner (speed). All 15 pass-rate rows tie. Tokens, price and route differ.
- New post
How fast is Jev? 136.5 ms per routing decision, measured live
Jev router latency over direct HTTPS: median 136.5 ms, p95 195.7 ms, range 100.9–297.3 ms, n = 246. Claude routers used a different CLI route.
- New post
How long does an AI coding agent take per task? Minutes, calls and where the time goes
Agent took a median 9.6 minutes per SWE-bench attempt (n = 33, range 1.6 to 54.1). Compare single calls, repairs and full tasks with limits.
- New post
How many runs do you need to compare two AI models? A sample-size table
Calculation: 100 runs can separate an observed 8-point gap at a perfect score. See Wilson intervals, paired tests and limits for LLM eval sample sizes.
- New post
How much of your AI bill is thinking tokens? Claude and GPT-6.1 Sol, measured
Thinking tokens: 46% to 92% of output per median call, 9% to 80% of pooled list-price cost. Calculation over recorded Claude and GPT-6.1 Sol calls.
- New post
How to turn off extended thinking in Claude Code, and what we measured for Haiku 4.5
Set MAX_THINKING_TOKENS=0, then check both counters. Our Haiku 4.5 test covers routing and hard tasks, with timings, costs and uncertainty.
- New post
Is Claude Haiku cheaper if you retry or escalate to Sonnet? A calculation on real receipts
Haiku 4.5 first, one retry, then Sonnet 5.5? On 8 hard tasks, the calculation gives 4.3x the cost and 8.7x the time per correct answer versus Sonnet alone.
- New post
Is Claude Haiku cheaper than Sonnet? Cost per correct answer, with retries
Haiku 4.5 lists at half the price of Sonnet 5.5, yet calculated cost per correct hard answer was 4.7x higher. Retry maths, two exceptions and every source.
- New post
Jev in fair mode: 2nd of 9 at Tron and Snake, last of 9 at Othello
The same 150 ms for every answer: Jev ranks 2 of 9 in Tron and Snake, 3 of 8 in Pong, 9 of 9 in Othello.
- New post
LLM API pricing comparison, October 2026: Claude vs GPT vs Gemini per million tokens
21 LLM API prices per million tokens, October 2026: Claude, GPT, Gemini and more. Blended prices span 100x. Price per token is not price per task.
- New post
Most AI model comparisons are ties: 44 of 1,196 rows show a clear gap
44 of 1,196 AI model comparison rows show a gap under our overlap rules. Most gaps are timing rows. No effort row separates quality.
- New post
Opus at low effort or Sonnet at high effort? A bigger model that thinks less, tested
Opus 5.5 at low effort and Sonnet 5.5 at high effort both passed 16/16. Opus low cost 1.3x as much per pass (calculation). Sonnet at low effort cost least.
- New post
Plan for p95, not the median: LLM tail latency in our runs
4.30 s p95 against a 2.60 s median for Claude Sonnet 5.5; 34.5 s against 12.5 s for Haiku 4.5. Measured LLM tail latency and what to do about it.
- New post
The cheapest LLM for classification: 100,000 decisions a day, and why Haiku cost more than Sonnet
Cost calculations for 100,000 routing decisions a day: Jev $3.37, Sonnet $732.40, Haiku $892.40. Tested on 82 cases; not general classification.
- New post
The cheapest way to run an AI coding agent: 7 levers from measured runs
7 levers that may cut an AI coding agent's bill, sized from our data: prompt cache 3.9x, Fable/Sonnet cost per pass 6.5x, and 5 more. List-price calculations.
- New post
What does an agent loop cost? 1.3 to 8.7 times the median tokens (calculation)
Agent loop cost on 8 hard tasks: 1.3–8.7 times the median tokens and up to 2.1 times the price per strict pass (calculations), with no clear pass-rate gain.
- New post
What does routing a million AI requests a day cost? Rules vs Jev vs Claude
What 10,000 to 10 million routing decisions a day cost with rules, Jev and Claude routers, how many run at once and how long they make requests wait.
- New post
What is the best AI model for coding? Our data says four models tie
4 models tied at the top of our hard coding set: Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol. Only Haiku 4.5 separated. A tier list built from intervals.
- New post
What thinking costs: reasoning tokens per call for Claude and GPT-6.1 Sol
Haiku 4.5 spent a median 4,556 reasoning tokens per hard call, Sonnet 5.5 585, GPT-6.1 Sol 150. List price per 1,000 calls: $22.78, $5.85, $1.50 (calculation).
- New post
When does prompt caching pay off? The Anthropic cache break-even, calculated
A 1-hour Anthropic cache write pays back after 2 reuses; a 5-minute write (assumed 1.25x) after 1. Break-even by model, cost per 1,000 sessions.
- New post
Which Claude model is fastest? It depends on the task, and on thinking
1.94 s was the lowest median on short calls (Fable 5.1). 7.75 s on hard calls (Sonnet 5.5). Haiku 4.5 took 4.43 s and 39.01 s with default thinking.
- New post
Which Claude model should you use? A task-by-task guide from our measurements
Which Claude model fits your job? Sonnet passed 24/24 hard calls (95% interval 86%–100%) at $0.01435 per pass (calculation). The top models hit a ceiling.
- New post
Why a paired test: comparing two LLMs on the same cases with McNemar
Only 5 to 7 of 82 routing cases split Claude and Jev. The exact McNemar test uses just those cases: p = 0.375 and p = 1. The method, with the arithmetic.
- New post
Why AI coding agents fail on real pull requests: every failure from 12 attempts
32 AI coding agent misses from our own runs, sorted into 8 classes: wrong answers, gates, caps, lost context and more. Counts, not rates.
- New post
Why is Claude Code slow? Where the seconds go in a coding CLI call, and what to change
A one-word Claude Code call took a median 2.5 s, with 1.7 s outside the model (n = 5). Where the rest goes: thinking, effort, route. Measured splits and fixes.
- New post
Would majority voting fix it? A self-consistency thought experiment on 90 real calls
Calculation on 90 recorded calls: a vote over 10 Haiku replies returns 289, not 282. Repeated errors and format misses limit self-consistency voting.
- New post
Your AI agent says it is done. Is it? Claimed vs verified in our runs
6 of 6 Haiku 4.5 sessions invented a late-fee rate and reported done. Sonnet 5.5 flagged the gap in 9 of 9. Claimed vs verified, with small n.
- Post updated
Does letting an AI run code help? A single call vs an agent loop on 8 hard tasks
Does tool use improve LLM accuracy? 8 hard tasks, 120 attempts (118 scored): single call vs agent loop for Haiku, Sonnet and GPT-6 Luna. No clear gain.
- Post updated
How to estimate your AI coding bill from real token mixes
Estimate AI coding costs from recorded token mixes: an agent task at $2.64, a hard call at $0.014, a routing decision at $0.005. Calculations, limits stated.
- Post updated
Jev vs Claude routers on unseen decisions: a blind holdout test
56 new routing decisions, written blind and frozen before any router call. Jev 1.13 vs Claude Haiku 4.5 vs Sonnet 5.5: accuracy, 95% intervals, cost, speed.
- Post updated
Jev vs Clef and five open decision models: 1,085 checkable decisions, tested
Seven decision models, 1,085 checkable decisions. The hosted model led; size bought less than you think.
- Post updated
What one resolved SWE-bench task really costs an AI coding agent
A full agent pipeline spent a notional $3.71 per resolved SWE-bench instance. Where the money went, what caching saved, and the public panel range.
- New explainer
Benchmark saturation: when every model scores 100%
Benchmark saturation: models reach the top score and the test stops telling them apart. We measured it on our own sets and show which other measures differ.
- New explainer
Claude Code hooks, explained with a measured Stop hook
A Claude Code hook runs a command at an event such as Stop. What a measured Stop hook enforced, what it missed and what it cost, in 200 graded sessions.
- New explainer
Cost per correct answer: the LLM price that counts failures
LLM cost per correct answer is all spend, failed calls included, divided by passes. How to compute it, and why a lower token price can cost more per task.
- New explainer
How many runs does an LLM eval need? Sample size, with real intervals
How many test cases does an LLM eval need? Real 95% intervals from our studies, a planning table for 10 to 1,000 cases, and why 16/16 spans 80.6% to 100%.
- New explainer
LLM pricing per million tokens, explained with real token counts
LLM price per million tokens: how input, output, cache read and cache write rates build a bill, shown on 33 recorded attempts. The bill is a calculation.
- New explainer
Memory consolidation (dreaming) for AI agents, measured
Claude Code dreaming, measured: what a memory consolidation pass kept and dropped, and how Haiku 4.5 and Sonnet 5.5 sessions differed, with n.
- New explainer
Open-weight models: why the same model has many prices
An open-weight model has public weights, so many providers host it. Learn how to compare reported provider prices and endpoint limits.
- New explainer
p95 latency, explained: median, tail and range for LLM calls
p95 latency is the time at or below which about 95% of calls fall. Compare median, tail and range with measured LLM timings and sample limits.
- New explainer
pass@k explained: pass@1, pass@k and pass^k with real runs
pass@k is the chance that at least one of k tries passes. How it differs from pass@1 and pass^k, worked through on real runs of 8 hard tasks.
- New explainer
Reasoning tokens, explained: the hidden output you pay for
Reasoning tokens are the hidden output a model writes before it answers. How many each model and effort used on hard tasks, and the cost in time and money.
- New explainer
Strict grading and format misses: when a right answer fails
Strict vs lenient LLM grading on 152 calls: 139 strict passes and 5 format misses. See both rates, their 95% intervals and how to report them.
- New explainer
The McNemar test: compare two models on the same cases
The McNemar test compares two LLMs on the same cases and reads only their disagreements. Worked example: 82 routing decisions, exact p = 1 and p = 0.375.
- New explainer
The Pareto frontier: reading LLM cost against quality
A Pareto frontier lists LLM options that no rival matches on all axes and beats on one. Read cost-quality charts from our head-to-heads.
- New explainer
Time to first token (TTFT), explained with CLI and API timings
Time to first token is the wait before an LLM starts to answer. What adds to it, and measured CLI and API timings with n and run ranges.
- New explainer
What is an agent harness? The code around the model, measured
An agent harness is the code around a model: the loop, tools, prompts, memory and checks. What it changes in time and tokens, measured on real runs.
- New explainer
What is an inference provider? Bedrock, Vertex, Azure and first-party prices
What is an inference provider? Our snapshot shows the same standard price for Claude Sonnet 5.5 on Bedrock, Vertex, Azure and Anthropic.
- New explainer
What is CLAUDE.md? Agent memory files, measured
CLAUDE.md is the file Claude Code reads at the start of a session. What it holds, what /init writes, how long it should be and what it costs, measured.
- New explainer
What is context engineering? What to put in the context, measured
Context engineering is choosing what a model reads on each call. What to add, what to remove and what it costs, with measured Claude Code and Codex CLI runs.
- New explainer
Why the same prompt gives different answers: LLM nondeterminism, measured
Why the same prompt gives different answers, measured on 90 repeated calls: the wording moves, the answer moves less, and one wrong answer came back 10 times.
- Source recorded
Frozen tasks selected with a Sonnet pilot. Fresh counted calls retain failures, format misses and timeouts. Replies and expected answers are not published.
- Source recorded
Receipts copied from a run of the same three JSON extraction prompts, each asked with instructions only (mode I) and with the CLI’s JSON schema mode (mode S). Prompts, model output and failure reasons are not published; a failed check is named, never quoted. Every attempt is kept. Both routes ran one call at a time. The current protocol file does not verify pre-call registration; see protocolAudit. Calls repeat three fixed hand-made prompts, including a known format-miss case. Both routes used a shared Mac.
- Source recorded
Receipts of the speed anatomy study: one prompt that asks for 250 numbers in words (output speed, 24 calls) and a seeded synthetic ledger at three sizes with one lookup question (prompt size, 36 calls). Every attempt is kept, failures included. A new seed for every ledger call; no ledger text, prompt or model output is copied.
- Source recorded
Sanitized cache-session receipts: one CLI process per session, 2 turns each, a seeded synthetic ledger (a different seed per condition) and 2 short questions with exact answers. Conditions: A a new temporary working folder per session, B one fixed folder, C a fixed folder with the ledger in the system prompt (Claude Code only). The Codex app-server reports cached input only, so its rows have no cache-write count. Probes are single uncounted calls with their own ledger seed. The answers and the ledger itself are not published. Session 1 was the first use of each setup’s ledger; sessions 2 and 3 reused that ledger. Different seeds prevent full ledger-prefix reuse between setups, but shared CLI-prefix reads remain possible. Correctness uses the reference checker after it trims spaces, surrounding quotes and backticks, a final period and currency units. The surviving protocol file dates from after the counted calls; pre-call declaration is not verified.
- Source recorded
Reasoning token bill (calculation)
Reported reasoning tokens priced at the recorded list prices. A calculation, not a new run.
- Source recorded
Recorded routing costs and times scaled to assumed daily volumes. No load test ran; median and p95 scenarios are not measured mean concurrency.
- Source recorded
Voice-agent latency budget calculation
Recorded decision and first-output times compared with assumed 300, 800 and 1,500 ms budgets. No complete voice turn ran.
- Dataset rebuilt
65 changes
- New study
Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
4.9× · n = 12Median time, GPT-6.1 Sol in Codex CLI vs Sonnet 5.5 in Claude Code (ratio of medians)
On small real repository tasks graded by hidden tests, how do coding-agent CLIs compare when they run with their normal file and shell tools?
- New study
Does memory help Claude Code? 8 kinds of agent memory, tested
40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5)
Does project memory make Claude Code do better work, and which kind of memory?
- New study
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
100% (96/96)New effort-ladder calls that passed strictly
On 8 hard tasks with strict validators, does a higher effort setting buy a higher pass rate for Sonnet, Opus and GPT-6.1 Sol, and what does it cost in time, tokens and list price per pass?
- New study
Does thinking pay for Claude Haiku 4.5? Thinking on vs off
87% (71/82)Claude Haiku 4.5 (thinking off): exact routing decisions
Does extended thinking pay for Claude Haiku 4.5: what does it buy in accuracy, and what does it cost in time and money, on typed routing decisions and on hard tasks?
- New study
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
91% (139/152)Calls that passed strictly (hard set)
On 8 hard tasks with deterministic validators, does pass rate separate the Claude Code models and GPT-6.1 Sol through the Codex CLI, and what do speed, tokens and cost per pass add?
- New study
How much of an AI bill is thinking? Reasoning tokens by model and effort
92% (Claude Haiku 4.5 · Claude Code; range 76% to 99%) · n = 24Highest median reasoning share of output tokens, hard tasks (calculation)
Across 378 recorded calls of Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol, how many output tokens come from reasoning (thinking)? What do they cost at list price per call and per strict pass, and do they track time? The calls cover eight hard tasks, five short tasks and several efforts.
Show 59 more changes on Oct 6
- New study
Inference provider index: 27 models, 52 providers
265 endpoints, 52 providers, 27 models · n = 27Provider endpoints in the snapshot
For the same model, how much do inference providers differ in price, and what does a gateway such as OpenRouter add over the first-party list price?
- New study
Jev vs Claude routers on unseen decisions: a blind holdout
82% (46/56)Jev 1.13 (TypeSafe): exact on unseen decisions
Does a router keep its accuracy on typed routing decisions that nobody tuned against its answers?
- New study
Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)
67% (2/3)Claude Opus 5.5 in Agent: resolved, interim (3 of 8 pairs graded)
Does Agent resolve more SWE-bench Verified instances with Claude Opus 5.5 as its brain than with Claude Sonnet 5.5, and at what cost and time?
- New study
Prompt cache break-even: after how many reuses does a cached prefix cost less?
2 reuses (the 3rd request)Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation)
With the cache-write surcharge Anthropic lists, after how many reuses does a cached prompt prefix cost less than no cache? What does that mean for a session of 1 to 20 turns, for each model’s price, and for a workload split across sessions?
- New study
Prompt caching and run-to-run consistency in Claude Code and Codex CLI
50% ($0.1350 vs $0.2698)List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)
When a CLI session reuses a fixed context, how much input comes from the cache, what does that save at list price, and does it change latency? When the same prompt runs 10 times, how much do the pass rate, the answer and the time vary?
- New study
Routing overhead: deterministic policy vs LLM routers vs Jev
1.42 µs (p95 2.33 µs, p99 3.04 µs) · n = 20,000Deterministic routing policy: median decision time
What delay and what cost does each kind of router add before the real work of a call starts?
- New study
54% (13/24)Claude Haiku 4.5 strict pass rate, agent loop
On 8 hard tasks with strict validators, does an agent loop that may write and run code in a sandbox pass more often than one call with tools off, and what does the loop cost in time, tokens, tool calls and list price per pass?
- New study
System One arena: Jev vs Clef and five open decision models, head to head
1,085Decision items, written by agents that never called a model (1,048 graded, 37 consistency-only)
How good are typed-decision models at decisions with a checkable answer, and which one wins when they play each other?
- New study
Voice agent latency budget: component calculations, not a measured turn
14 steps · n = 14Decision and first-output steps compared
Which decision steps fit inside a voice agent’s turn budget, using only measured times?
- New study
What does routing a million AI requests a day cost? A calculation from measured runs
$0.0337 · n = 82Cost per 1,000 routing decisions, Jev (calculation)
At 10,000 to 10 million routing decisions a day, what does each router cost per day, what in-flight and waiting scenarios do median and p95 times give?
- Study updated
Jev vs Claude as a router: accuracy and cost
7 charts, 7 key numbers.
- New post
AI coding benchmarks roundup, October 2026: sixteen studies, every number in one place
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
- New post
An AI model leaderboard without a composite score: why, and how to read ours
Our leaderboard lists 46 models, CLIs, routers and providers with their best-supported facts, each with n and an interval. No single score. Here is why.
- New post
Claude Code vs Codex CLI: the hidden context tax and what it does to speed
Codex CLI sent a median 12,124 input tokens per call; Claude Code sent 2,130. What a coding CLI's hidden context costs, and how much time the route adds.
- New post
Claude Sonnet 5.5 vs Opus 5.5: when is Opus worth the price?
Sonnet 5.5 and Opus 5.5 tied on every quality test we ran, easy, hard and agentic. Opus cost 1.6x to 2.6x per unit of work. Where the gap comes from.
- New post
Claude vs Codex on hard tasks: GPT-6.1 Sol joins the hard set
GPT-6.1 Sol in the Codex CLI passed 16/16 hard tasks at medium and high effort. Sonnet, Opus and Fable passed 24/24. What separates them: time and cost.
- New post
Does CLAUDE.md help? We tested 8 kinds of agent memory on Claude Code
200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a handbook and a Stop hook. Memory mattered where the repo was silent.
- New post
Does extended thinking pay for Claude Haiku 4.5? We turned it off and measured
Claude Haiku 4.5, thinking off: 2.7x faster (a calculation from two medians), no routing accuracy gap shown. Hard tasks: 4/24 passed, thinking on 11/24.
- New post
Does letting an AI run code help? A single call vs an agent loop on 8 hard tasks
Does tool use improve LLM accuracy? 8 hard tasks, 120 attempts (118 scored): single call vs agent loop for Haiku, Sonnet and GPT-6 Luna. No clear gain.
- New post
Does reasoning effort buy quality? Claude and Codex on hard tasks, low to high
176 calls on 8 hard tasks: Sonnet, Opus and GPT-6.1 Sol at low, medium and high effort. Every cell passed 16/16. Higher effort cost more tokens and money.
- New post
How much does prompt caching actually save? Measured in Claude Code and Codex
Claude Code read 97% of later-turn input from the cache. At list price that halved a 5-turn session and cut an agent bill about 3.9x. Turn 1 costs more.
- New post
How to estimate your AI coding bill from real token mixes
Estimate AI coding costs from recorded token mixes: an agent task at $2.64, a hard call at $0.014, a routing decision at $0.005. Calculations, limits stated.
- New post
Jev vs Claude Haiku vs Claude Sonnet as a router: an honest comparison
Every row of our Jev, Haiku and Sonnet router comparisons: accuracy ties, Jev is far cheaper and quicker per call over its API, on a different route.
- New post
Jev vs Claude routers on unseen decisions: a blind holdout test
56 new routing decisions, written blind and frozen before any router call. Jev 1.13 vs Claude Haiku 4.5 vs Sonnet 5.5: accuracy, 95% intervals, cost, speed.
- New post
Jev vs Clef and five open decision models: 1,085 checkable decisions, tested
Seven decision models, 1,085 checkable decisions. The hosted model led; size bought less than you think.
- New post
OpenRouter vs going direct: what the gateway really costs
OpenRouter charged the vendor's per-token price on 10 of 10 models we checked. The cost is a 5.5% credit fee. What we know, and what is still unmeasured.
- New post
Opus 5.5 vs Sonnet 5.5 on real GitHub issues: an early look, limits first
Interim: 3 of 8 paired SWE-bench Verified issues graded. Opus 5.5 resolved 2, Sonnet 5.5 1; McNemar p = 1.0. Opus cost 2.6x, a calculation. Limits first.
- New post
Same prompt, ten answers: how consistent are Claude and Codex?
3 prompts, 10 runs each, on Haiku, Sonnet and GPT-6.1 Sol. 7 of 9 cells passed 10/10. Haiku gave the same wrong number 10 times: consistent is not correct.
- New post
The cheapest place to run open models right now (October 2026 snapshot)
DeepSeek V4, gpt-oss-120b, Llama, Kimi K3 and GLM 5.3 across 52 providers. Prices differ up to 12.6x. Snapshot 2026-10-06, with the caveats.
- New post
What does a router cost you? Rules vs Jev vs an LLM router
A rule-based router decides in 1.42 µs for $0, Jev in 136.5 ms, Sonnet in 2.60 s. What each adds per 1,000 tasks, in delay and in dollars.
- New post
When tasks get hard: Claude Haiku vs Sonnet vs Opus vs Fable (and where Codex is)
120 calls on 8 hard tasks with strict validators. Sonnet, Opus and Fable passed 24/24; Haiku 11/24. Speed, tokens and cost per pass. Codex results: see update.
- Post updated
Why we count every failed attempt: our rules for honest AI benchmarks
How we benchmark AI agents: declare the protocol first, keep every failure, show intervals, publish our own confounds and label interim results.
- New explainer
How to read AI benchmarks honestly
A checklist for reading AI model benchmarks: sample size, intervals, ranges, failures, calculations vs runs, and why one overall score hides what was measured.
- New explainer
Inference gateways and OpenRouter, explained
An inference gateway puts many model providers behind one API. How OpenRouter prices compare with first-party prices, what its fee adds, and what is unmeasured.
- New explainer
LLM as judge: how blind review works, and where it fails
LLM-as-judge uses models to grade outputs. How a blind, order-swapped review panel works, which biases to expect, and what our AI vs human PR panel found.
- New explainer
Prompt caching, explained with measured sessions
Prompt caching reuses an unchanged prompt prefix at a lower price: how cache reads and writes are billed, how much a CLI session reads from cache, the saving.
- New explainer
Quantization and cheap inference, explained
Quantization stores model weights in fewer bits (fp8, fp4) so providers can serve them for less. What it changes, how to spot it in prices, and what to test.
- New explainer
SWE-bench Verified scores coding agents on real GitHub issues with hidden tests. What a resolved rate means, and why 33 instances give wide intervals.
- New explainer
Tokens per call and the CLI context tax
A coding CLI wraps every request in its own system prompt and tools. How many input tokens that adds per call, what it costs in time, and when it matters.
- New explainer
An LLM router picks the model, effort and context for each request. What routers exist, what a decision costs in time and money, and how to judge one.
- New explainer
Reasoning effort sets how much a model thinks before it answers. What low, medium and high change in quality, time, tokens and cost, measured on hard tasks.
- New explainer
Wilson confidence intervals for AI benchmarks
A Wilson interval shows the range of true pass rates that fit a benchmark result. The formula, a worked example, and why 16 of 16 still means 81% to 100%.
- Source recorded
Agent memory study: 8 kinds of project memory on Claude Code
One small Node.js repository, 5 tasks, 8 memory conditions (none, /init, curated, raw notes, dreamed notes, long handbook, Stop hook, curated + hook). Claude Code 2.1.286 headless: Sonnet lane 3 repetitions, Haiku lane 2. Hidden tests and deterministic convention checks; protocol declared before the first session; every attempt kept. The exact memory files and three dreaming passes are published.
- Source recorded
Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim)
8 Verified instances declared before the first run, 2 per difficulty band. One Opus attempt each on platform build 4f6f4027, paired with the earlier Sonnet attempt on the same instance; official grader. Interim: a usage gate stopped the campaign after 3 instances; the other 5 resume after the reset on 2026-10-09.
- Source recorded
Caching sessions and repeated prompts (Claude Code and Codex CLI)
Part 1: 5-turn CLI sessions over a fixed synthetic ledger, with the cache counters each provider reports per turn. Part 2: three prompts with deterministic validators, 10 repetitions per model. Declared protocols, validator controls before inference, every attempt kept; answers are published as ordinal ids, never as text.
- Source recorded
Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks
Six small Node.js repositories with hidden tests; controls before the first session (every base fails, every reference passes). Claude Code 2.1.286 with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol at medium effort, 2 repetitions per task, OS sandboxes without network, protocol declared before the first session, every session kept. Gemini CLI was probed and not run (browser login).
- Source recorded
Cost with and without the prompt cache (calculation)
Recorded tokens per turn × Anthropic list prices. With the cache: uncached input at the input price, cache reads at the cache-read price, 1-hour cache writes at twice the input price, 5-minute writes at 1.25 times (an assumption; none occurred). Without a cache: every input token at the input price. Output is priced the same in both. Not a bill.
- Source recorded
Effort ladder: the hard task set at each effort level
The eight hard tasks of the hard head-to-head at low, medium and high effort (Claude Sonnet 5.5 and Claude Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI), 8 tasks × 2 repetitions per cell, declared protocols, every attempt kept. Efforts the hard head-to-head already ran are reused from its receipts as reference cells.
- Source recorded
Receipts of the Haiku thinking study: Claude Haiku 4.5 through Claude Code with thinking off (MAX_THINKING_TOKENS=0) against the recorded thinking-on arms. Routing arms carry per-arm totals computed from each arm’s per-call log and per-case results; hard-task receipts are every attempt of the thinking-off run. Every attempt is kept, failures included. The thinking-on hard-task receipts live in the hard head-to-head extract.
- Source recorded
Jev live run: 246 timed calls on the 82 routing decisions
Jev 1.13 called over HTTPS, 3 repeats of the same 82 typed decisions, one call at a time, from one Mac over a home network: client wall time, with the network inside it. The API reports no server time. Cost per 1,000 decisions is a calculation from the reported input tokens and the published price. The case sets were revised against Jev answers, so Jev has a home advantage.
- Source recorded
OpenRouter states that inference is billed at the provider list price and that its fee is charged when credits are bought (5.5% on Standard by card, $0.80 minimum; 8% on Business; 5% by crypto). Page fetched 2026-10-06.
- Source recorded
OpenRouter public API: models and provider endpoints (snapshot)
Prices, context, quantization and uptime per provider endpoint as reported by OpenRouter’s public, keyless API on 2026-10-06. Third-party-reported, not measured by Agent. Latency and throughput were not returned.
- Source recorded
Provider head-to-head, hard set: eight hard tasks with strict validators
Eight hard tasks with sandboxed deterministic validators and pre-inference controls, declared protocol, every attempt kept. Format misses are recorded apart from wrong answers.
- Source recorded
Routing on unseen holdout decisions
Frozen unseen routing decisions; every router call retained. Costs are calculations from recorded tokens and list prices.
- Source recorded
Routing overhead per 1,000 tasks (calculation)
Decisions per task from recorded bench runs multiplied by the cost and the median time per decision. Decisions are assumed to wait in line, so the delay is an upper bound. A calculation, not a run.
- Source recorded
Routing overhead runs: policy microbenchmark and CLI start-up
In-process timing of the deterministic routing policy (20,000 timed decisions), CLI start-up with a one-word prompt (5 runs per CLI), and decision counts read from recorded bench runs. LLM router timings are reused from the routing runs.
- Source recorded
Receipts of the single call vs agent loop study (public-runs/single-call-vs-agent-loop). Every attempt is kept, failures and contaminated attempts included. Reference single-call cells (Claude Haiku 4.5 and Claude Sonnet 5.5, default effort) are read from raw/provider-h2h-hard.
- Source recorded
System One arena: typed-decision models on checkable decisions and in head-to-head games
Jev 1.13 (TypeSafe API) and six open System One models in llama.cpp 0.6.0 on one Mac Studio (Clef 27B and Clef-Flash 9B Q4_K_M, lev 4B and Kev 4B Q4_K_M, Laya and Julia-1 Q8_0). 1,085 model-blind items in five suites with independent verifiers; exam, latency and repeat passes. Round-robin games (tic-tac-toe, Connect Four, Nim, Dots and Boxes, real-time Pong) against each other and a random and a perfect player, with the featured series and a speed ladder; declared amendments 1 and 2 in the published protocol. Protocol, item hashes and game counts declared before the first counted call; every call and game kept.
- New study
21 changes
- New study
Agent on SWE-bench Verified vs 11 public models
76% (25/33)Agent resolved, all 33 attempted instances
How does Agent, a full worker pipeline on one model, do on SWE-bench Verified next to public single-model runs on the very same instances?
- New study
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
98% (127/130)Calls that passed their validator
On short tasks with strict validators, how do the Claude Code models and efforts compare with Codex on pass rate, speed and tokens?
- New study
Jev vs Claude as a router: accuracy and cost
90% (74/82)Jev 1.13 (TypeSafe): exact decisions
Should a small dedicated router or a general LLM make the platform’s typed routing decisions?
- New study
What if every call ran on Opus? Repricing real agent tokens
162.9M · n = 33Input tokens recorded
Agent recorded every token it used on 33 SWE-bench instances. What would the same tokens cost at other models’ list prices, and what did caching save?
- Study updated
AI pull requests vs merged human pull requests, judged blind
4 charts, 7 key numbers.
- Study updated
Claude Code CLI vs Codex CLI vs the API: latency and tokens
5 charts, 5 key numbers.
Show 15 more changes on Oct 5
- Study updated
Coding calibration: what broke on three real pull requests
4 charts, 4 key numbers.
- New post
Claude Code vs Codex CLI vs the API: latency, time to first token and hidden prompts
194 timed runs. Codex CLI took 3.5x as long as the OpenAI API for a one-line answer and sent 19,551 input tokens instead of 17. What a coding CLI adds.
- New post
Claude Haiku vs Sonnet vs Opus vs Fable vs Codex: 130 timed calls, head to head
130 timed calls on five validated tasks. Pass rate, latency, hidden input tokens and cost per passing answer for Claude Code models and Codex CLI.
- New post
Do blind AI critics prefer AI pull requests over human ones?
Blind Claude and GPT critics preferred an AI worker's change over the merged human change on 9 of 12 real tasks, and 6 of 12 on the first try. Caveats inside.
- New post
Harness vs model: where do AI coding agent gains really come from?
Is it the model or the harness? Our data shows the harness clearly moves speed, tokens and cost. Whether it moves accuracy, our samples cannot yet say.
- New post
Jev vs Claude Haiku and Sonnet as a router: tied on accuracy, 148x to 265x cheaper
82 typed routing decisions. Jev, Claude Haiku 4.5 and Sonnet 5.5 tie on accuracy, but Jev costs $0.0337 per 1,000 decisions against $5 to $8.92.
- New post
SWE-bench Verified: an agent pipeline vs 11 public models on the same 33 tasks
Agent resolved 25 of 33 SWE-bench Verified instances (76%). Eleven public models solved 21 to 28 of the same ones. Why that is a tie, not a win.
- New post
What if every call ran on Opus? Repricing real agent tokens across models
We repriced 162.9M recorded agent tokens at Haiku, Sonnet, Opus, Fable, Gemini Flash and GPT prices. A calculation, not a run, with clear limits.
- New post
What one resolved SWE-bench task really costs an AI coding agent
A full agent pipeline spent a notional $3.71 per resolved SWE-bench instance. Where the money went, what caching saved, and the public panel range.
- New post
Why we count every failed attempt: our rules for honest AI benchmarks
How we benchmark AI agents: declare the protocol first, keep every failure, show intervals, publish our own confounds and label interim results.
- Source recorded
Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)
The 6 compiled-extension instances that campaign 1 could not run, plus 2 replacement candidates. One attempt each, platform build 236c0d3f.
- Source recorded
Coding calibration: fastify/session, h3, uvicorn
Three real upstream tasks, one attempt per task per platform slice, offline gates against the merged reference.
- Source recorded
Provider head-to-head: Claude Code models vs Codex efforts
Five short tasks with deterministic validators, declared protocol, every attempt kept.
- Source recorded
Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.
- Source recorded
Routing runs: Jev router vs LLM routing
Routing decisions recorded per case and arm.
- New study
2 changes
- Source recorded
Agent on SWE-bench Verified, campaign 1 (25 instances)
Stratified sample of 25 Verified instances (seed 20261004), one attempt each, official grading harness. Fixed model claude-sonnet-5-5, platform build f0ac3a8a.
- Source recorded
SWE-bench campaign rules and sample design
Rules declared before the first run: escalations are graded as delivered, gold must resolve on the host, blocked instances are replaced in the same difficulty band, no second attempts. The excluded and replaced instances are listed in the exclusions extract.
- Source recorded
3 changes
- New study
Claude Code CLI vs Codex CLI vs the API: latency and tokens
100% (194/194)Evaluated runs that passed their validator
How much time and how many tokens does a coding CLI add on top of the model, and how do Claude Code and Codex compare on the same repair task?
- Source recorded
Token prices as listed by the vendor on 2026-10-03.
- Source recorded
Provider explorer receipts: CLI vs API
230 imported receipts for short fixed tasks over Claude Code CLI, Codex CLI and the OpenAI API, with time to first useful output, total time, tokens and validation.
- New study
1 change
- New study
Coding calibration: what broke on three real pull requests
1 of 3Verified deliveries, latest build
On three real upstream issues, does the AI worker deliver a verified change, and what stops it when it does not?
- New study
2 changes
- New study
AI pull requests vs merged human pull requests, judged blind
75% (9/12)Tasks where the panel preferred the AI change (latest attempt)
When critics cannot see which change came from a person, do they prefer the AI worker’s pull request or the one the maintainers merged?
- Source recorded
Blind review panel: AI worker change vs merged human change
Each pair is judged by 3 or 4 critic models without labels, in both orders. 5 public OSS tasks and 7 private tasks. Private task names are replaced by neutral labels.
- New study
1 change
- Source recorded
Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free.
- Source recorded
2 changes
- Source recorded
Anthropic list prices (Claude models)
Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.
- Source recorded
Gemini 3.x Flash prices as listed by the vendor on 2026-09-21. The vendor announced a doubling from 2027-01-01.
- Source recorded
1 change
- Source recorded
SWE-bench Verified leaderboard, mini-SWE-agent v2 runs
Public per-instance results of 11 models under mini-SWE-agent 2.0.0 (bash only, one attempt). Costs are API list prices as published.
- Source recorded