Prompt caching and run-to-run consistency in Claude Code and Codex CLI
When a CLI session reuses a fixed context, how much input comes from the cache, what does that save at list price, and does it change latency? When the same prompt runs 10 times, how much do the pass rate, the answer and the time vary?
Published · 6 charts · Download the data or a carousel
50%
The answer
Caching: inside one Claude Code session, turns 2-5 read 97% of their input from the cache on average (turn 1: 19%, the CLI's own prefix). At list price, a calculation, all recorded turns cost Sonnet $0.1350 vs $0.2698 (50% less) and Opus $0.2551 vs $0.5442 (53% less) without the cache. Turn 1 costs more with the cache, because a 1-hour cache write costs twice the input price. A new session did not reuse the cache of an earlier one: on turn 1, all 4 later sessions wrote the ledger to the cache again. The cause was not tested. The cache showed no clear speed effect: median turn time was Sonnet 1.6 s on turn 1 vs 1.6 s on turns 2-5 and Opus 1.9 s on turn 1 vs 2.4 s on turns 2-5, and the fastest-to-slowest ranges overlap. Codex CLI (GPT-6.1 Sol) read 99% of later-turn input from its cache on a larger context; its app-server reports no cache writes, so no cost is calculated for it. Consistency: 7 of 9 model-and-prompt cells passed all 10 repetitions (95% interval 72% to 100%). Haiku passed 0/10 on the exact-number prompt; Haiku passed 1/10 on the JSON prompt (9 more were correct but in the wrong format). Haiku gave the same wrong answer every time (289; expected 282): consistent is not the same as correct. The code-fix prompt gave 6 different code bodies for Haiku, 3 different code bodies for Sonnet and 6 different code bodies for GPT-6.1 Sol (medium).
Live story
Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.
Prompt caching and consistency: what the cache saves, and how much answers vary
Calculation at list price: the cache cut a 5-question session 50% on Sonnet 5.5 and 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every repetition.
Transcript
- Caching and consistency · 135 calls. What the cache saves, and how much answers vary. Five-question sessions on a fixed context. Then the same prompt, 10 times.
- 135 calls: 45 cache turns, 90 repeated prompts. Every call counted. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- Claude Code turn 1 reads 19% from the cache, the CLI’s own prefix. Turns 2–5 read 90%–99%. Codex CLI: 98%–99%. Chart: Share of input read from the cache · mean of 3 sessions per turn (n = 3 each). Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- A calculation: Sonnet 5.5 $0.1350 with the cache vs $0.2698 without, 50% less. Opus 5.5 $0.2551 vs $0.5442, 53% less. Chart: List-price cost of the recorded sessions · calculation (n = 15 each). Calculation, not a run. Caveat: Costs are list-price calculations; the calls used flat subscriptions.
- No clear speed effect: Sonnet 5.5 1.6 s on turn 1 vs 1.6 s later; Opus 5.5 1.9 s vs 2.4 s. The ranges overlap. Sonnet 5.5: median turn 1 vs turns 2–5: 1.6 s vs 1.6 s (ranges 1.6–1.8 s and 1.4–5.6 s · n = 3 and 12). Opus 5.5: median turn 1 vs turns 2–5: 1.9 s vs 2.4 s (ranges 1.8–4.4 s and 1.6–12.7 s · n = 3 and 12). Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- 7 of 9 cells passed 10/10. Haiku 4.5: 0/10 on the exact number (10 wrong, 1 distinct answer), 1/10 on JSON (9 format misses). Chart: Same prompt, 10 times · strict passes · a 10/10 is 72%–100% at 95% (n = 10 each). Caveat: Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
- Consistent is not correct: Haiku 4.5 gave the same wrong number all 10 times. Code fix, distinct correct bodies: Haiku 4.5 6, Sonnet 5.5 3, GPT-6.1 Sol (medium) 6. Chart: Same prompt, 10 times · distinct answers (n = 10 each). Caveat: The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
- Cache the fixed context; check answers, not agreement. Every call online.
Key numbers
53%
List-price saving from the cache over 15 turns, Claude Opus 5.5 · Claude Code (calculation)
($0.2551 vs $0.5442)
0 of 4
Later sessions whose first turn read the ledger from an earlier session’s cache
n = 4
7 of 9
Model-and-prompt cells that passed all 10 repetitions
n = 9
135
Calls in this study (every one counted)
(45 cache turns, 90 repeated prompts)
Calculate when repeated cache reads repay the first write
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
- With the cache, as recorded
- Without a cache: every input token at the input price (square)
Gap labels, Without a cache: every input token at the input price vs With the cache, as recorded: Without a cache: every input token at the input price is x% higher (+) or lower (−) than With the cache, as recorded, calculated from the two values shown (the change counted from With the cache, as recorded’s value).
| Item | With the cache, as recorded | Without a cache: every input token at the input price | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.14 | $0.27 | 15 |
| Claude Opus 5.5 · Claude Code | $0.26 | $0.54 | 15 |
List-price calculation, not a run. 2 rows, 2 series: With the cache, as recorded, Without a cache: every input token at the input price. With the cache, as recorded: highest Claude Opus 5.5 · Claude Code $0.26 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.14 (n 15). Without a cache: every input token at the input price: highest Claude Opus 5.5 · Claude Code $0.54 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.27 (n 15).
Notesn = 15 per row
All recorded turns per model; the same reported tokens priced two ways
Calculation, not a bill: the calls ran on a subscription. Cache reads at the cache-read price, 1-hour cache writes at 2× the input price (every write in this run was a 1-hour write). Codex CLI is not priced here: it reports no cache-write count.
Sources: Caching sessions and repeated prompts (Claude Code and Codex CLI), Cost with and without the prompt cache (calculation), Anthropic list prices (Claude models)
- Turn 1 (writes the ledger to the cache)
- Turns 2-5 (read the ledger from the cache) (square)
Gap labels, Turns 2-5 (read the ledger from the cache) vs Turn 1 (writes the ledger to the cache): Turns 2-5 (read the ledger from the cache) is x% higher (+) or lower (−) than Turn 1 (writes the ledger to the cache), calculated from the two values shown (the change counted from Turn 1 (writes the ledger to the cache)’s value); lines are the fastest–slowest run (not an interval).
| Item | Turn 1 (writes the ledger to the cache) | Turns 2-5 (read the ledger from the cache) | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1.6 s | 1.6 s | Turn 1 (writes the ledger to the cache): 1.6 s–1.8 s; Turns 2-5 (read the ledger from the cache): 1.4 s–5.6 s | 3 |
| Claude Opus 5.5 · Claude Code | 1.9 s | 2.4 s | Turn 1 (writes the ledger to the cache): 1.8 s–4.4 s; Turns 2-5 (read the ledger from the cache): 1.6 s–12.7 s | 3 |
2 rows, 2 series: Turn 1 (writes the ledger to the cache), Turns 2-5 (read the ledger from the cache). Turn 1 (writes the ledger to the cache): slowest Claude Opus 5.5 · Claude Code 1.9 s (range 1.8 s–4.4 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1.6 s–1.8 s, n 3). All run ranges overlap. Turns 2-5 (read the ledger from the cache): slowest Claude Opus 5.5 · Claude Code 2.4 s (range 1.6 s–12.7 s, n 12). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1.4 s–5.6 s, n 12). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 3–12 per row
Median; whiskers = fastest and slowest turn
Whiskers are a range (fastest and slowest turn), not a confidence interval. Turns ask different questions: the slow later turns are the counting question, which produced the most output.
Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)
- Exact number
- JSON object
- Code fix
| Item | Exact number | JSON object | Code fix | 95% interval | n |
|---|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 0% | 10% | 100% | Exact number: 0%–28%; JSON object: 1.8%–40%; Code fix: 72%–100% | 10 |
| Claude Sonnet 5.5 · Claude Code | 100% | 100% | 100% | Exact number: 72%–100%; JSON object: 72%–100%; Code fix: 72%–100% | 10 |
| GPT-6.1 Sol (medium) · Codex CLI | 100% | 100% | 100% | Exact number: 72%–100%; JSON object: 72%–100%; Code fix: 72%–100% | 10 |
3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 10 per row
One series per prompt; whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.
Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)
Exact number
JSON object
Code fix
One panel per series, all on the same axis.
| Item | Exact number | JSON object | Code fix | n |
|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 1 | 1 | 6 | 10 |
| Claude Sonnet 5.5 · Claude Code | 1 | 1 | 3 | 10 |
| GPT-6.1 Sol (medium) · Codex CLI | 1 | 1 | 6 | 10 |
3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: all at 1. JSON object: all at 1.
Notesn = 10 per row
Distinct normalized answers over 10 repetitions (1 = the same answer every time)
Normalized: the last number for the exact-number prompt; key-sorted JSON for the JSON prompt; for the code fix, the code without fences, comments, whitespace, quote style and line-end semicolons. One distinct answer is not the same as a correct answer: an answer can be the same and wrong every time. Different code can be equally correct.
Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)
- Exact number
- JSON object
- Code fix
| Item | Exact number | JSON object | Code fix | Range (lowest–highest run) | n |
|---|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 5.1 s | 7 s | 6 s | Exact number: 4.4 s–6.2 s; JSON object: 5.3 s–12.3 s; Code fix: 4.9 s–7.3 s | 10 |
| Claude Sonnet 5.5 · Claude Code | 6.9 s | 2.9 s | 2.7 s | Exact number: 5.8 s–7.8 s; JSON object: 2.7 s–5.3 s; Code fix: 2.3 s–4.3 s | 10 |
| GPT-6.1 Sol (medium) · Codex CLI | 13.4 s | 6.4 s | 11.3 s | Exact number: 12.3 s–18 s; JSON object: 5.3 s–8.3 s; Code fix: 9.1 s–14.9 s | 10 |
3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: slowest GPT-6.1 Sol (medium) · Codex CLI 13.4 s (range 12.3 s–18 s, n 10). Fastest Claude Haiku 4.5 · Claude Code 5.1 s (range 4.4 s–6.2 s, n 10). Not all run ranges overlap. JSON object: slowest Claude Haiku 4.5 · Claude Code 7 s (range 5.3 s–12.3 s, n 10). Fastest Claude Sonnet 5.5 · Claude Code 2.9 s (range 2.7 s–5.3 s, n 10). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 10 per row
Median; whiskers = fastest and slowest of 10 calls
Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.
Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)
Tables
Cache counters per turn (mean over sessions)
| Configuration | Turn | Sessions | Exact answers | Input tokens (all) | Read from cache | Written to cache | Read share | Median time (s) | USD with cache (calculation) | USD without cache (calculation) |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1 | 3 | 3/3 | 7,831 | 1,463 | 6,366 | 19% | 1.6 s | $0.026 | $0.016 |
| Claude Sonnet 5.5 · Claude Code | 2 | 3 | 3/3 | 7,889 | 7,829 | 58 | 99% | 1.5 s | $0.0019 | $0.016 |
| Claude Sonnet 5.5 · Claude Code | 3 | 3 | 3/3 | 7,970 | 7,887 | 81 | 99% | 1.5 s | $0.002 | $0.016 |
| Claude Sonnet 5.5 · Claude Code | 4 | 3 | 3/3 | 8,032 | 7,968 | 62 | 99% | 5.2 s | $0.0096 | $0.024 |
| Claude Sonnet 5.5 · Claude Code | 5 | 3 | 3/3 | 8,861 | 8,030 | 829 | 91% | 1.8 s | $0.0058 | $0.019 |
| Claude Opus 5.5 · Claude Code | 1 | 3 | 3/3 | 7,828 | 1,463 | 6,363 | 19% | 1.9 s | $0.051 | $0.031 |
| Claude Opus 5.5 · Claude Code | 2 | 3 | 3/3 | 7,886 | 7,826 | 58 | 99% | 1.6 s | $0.0026 | $0.032 |
| Claude Opus 5.5 · Claude Code | 3 | 3 | 3/3 | 7,991 | 7,884 | 105 | 99% | 1.8 s | $0.003 | $0.033 |
| Claude Opus 5.5 · Claude Code | 4 | 3 | 3/3 | 8,075 | 7,989 | 84 | 99% | 6.1 s | $0.018 | $0.048 |
| Claude Opus 5.5 · Claude Code | 5 | 3 | 3/3 | 8,941 | 8,073 | 866 | 90% | 2.1 s | $0.0097 | $0.037 |
| GPT-6.1 Sol (medium) · Codex CLI | 1 | 3 | 3/3 | 18,774 | 10,411 | not reported | 55% | 3.6 s | — | — |
| GPT-6.1 Sol (medium) · Codex CLI | 2 | 3 | 3/3 | 18,803 | 18,560 | not reported | 99% | 2.4 s | — | — |
| GPT-6.1 Sol (medium) · Codex CLI | 3 | 3 | 3/3 | 18,848 | 18,645 | not reported | 99% | 3.1 s | — | — |
| GPT-6.1 Sol (medium) · Codex CLI | 4 | 3 | 3/3 | 18,880 | 18,688 | not reported | 99% | 6 s | — | — |
| GPT-6.1 Sol (medium) · Codex CLI | 5 | 3 | 3/3 | 19,097 | 18,688 | not reported | 98% | 2.4 s | — | — |
Every consistency cell (10 repetitions each)
Ten repeats of the same prompt per cell: every run as a mark
- Strict pass
- Format miss (right answer, wrong format)
- Wrong answer
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
| Configuration | Prompt | Strict passes | 95% interval | Format misses | Wrong answers | Distinct answers | Distinct raw replies | Median time (s) | Interquartile range (s) | Coefficient of variation |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | Exact number | 0/10 | 0% to 28% | 0 | 10 | 1 | 10 | 5.1 s | 1.1 s | 0.13 |
| Claude Haiku 4.5 · Claude Code | JSON object | 1/10 | 2% to 40% | 9 | 0 | 1 | 3 | 7 s | 2.2 s | 0.28 |
| Claude Haiku 4.5 · Claude Code | Code fix | 10/10 | 72% to 100% | 0 | 0 | 6 | 6 | 6 s | 1.1 s | 0.13 |
| Claude Sonnet 5.5 · Claude Code | Exact number | 10/10 | 72% to 100% | 0 | 0 | 1 | 1 | 6.9 s | 1 s | 0.09 |
| Claude Sonnet 5.5 · Claude Code | JSON object | 10/10 | 72% to 100% | 0 | 0 | 1 | 1 | 2.9 s | 0.5 s | 0.28 |
| Claude Sonnet 5.5 · Claude Code | Code fix | 10/10 | 72% to 100% | 0 | 0 | 3 | 3 | 2.7 s | 1.2 s | 0.26 |
| GPT-6.1 Sol (medium) · Codex CLI | Exact number | 10/10 | 72% to 100% | 0 | 0 | 1 | 1 | 13.4 s | 1.5 s | 0.14 |
| GPT-6.1 Sol (medium) · Codex CLI | JSON object | 10/10 | 72% to 100% | 0 | 0 | 1 | 1 | 6.4 s | 1.8 s | 0.17 |
| GPT-6.1 Sol (medium) · Codex CLI | Code fix | 10/10 | 72% to 100% | 0 | 0 | 6 | 6 | 11.3 s | 2.3 s | 0.15 |
Marks are counts from the table, sorted by outcome (run order is not recorded)95% Wilson interval on strict passesn = 10 runs per cell
9 cells: 3 configurations by 3 prompts, 10 runs each. Each row shows strict passes, format misses and wrong answers as marks, the distinct answers and the median time.
Method
- Protocols declared before the first call, one per route. Every attempt is kept; nothing was retried.
- Caching: 9 sessions (3 × Claude Sonnet 5.5 · Claude Code, 3 × Claude Opus 5.5 · Claude Code and 3 × GPT-6.1 Sol (medium) · Codex CLI), 5 turns each. A session is one CLI process. Turn 1 sends a seeded synthetic stock ledger plus question 1; turns 2-5 send one short question each (lookups, a count, an arg-max), each with one exact answer. Each turn is one model request.
- A 2-call probe sized the context before the run (not part of any cell). The ledger was larger than declared, so it was cut once, from 170 to 100 lines, and the answers were recomputed, as the protocol allowed.
- The Codex sessions ran before that cut, on the 170-line ledger (16,197 characters vs 9,651). The two routes are reported side by side, never as a like-for-like pair.
- Cache counters as the provider reports them: Claude Code gives uncached input, cache reads and cache writes (with the 5-minute and 1-hour split); the Codex app-server gives input (cached included) and cached input, and no cache writes.
- Cost with and without the cache (Claude only): see the calculation source. Every write in this run was a 1-hour write, so the 5-minute multiplier (an assumption) was not used.
- Consistency: 3 prompts (an exact number, a JSON object with exact keys, a small code fix), each with a deterministic validator; 10 repetitions per prompt for Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI. One-shot calls, one at a time per account. Answer diversity counts distinct normalized answers; the raw extract holds ordinal answer ids, never the text.
- Validator controls ran before inference on each route: every reference answer passes, every plausible wrong answer fails, and a wrapped reference is flagged as a format miss.
- Isolation as in the hard head-to-head: fresh empty working folder, tools off, no MCP servers, no session persistence across processes, provider-default caching. Claude at its default effort; GPT-6.1 Sol at medium.
- No batch stopped early and nothing was trimmed. The Claude CLI reported a rate-limit status of "allowed_warning" on 7 of 30 cache turns; no call was refused and the run did not stop.
Caveats
- Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- Costs are list-price calculations; the calls used flat subscriptions.
- Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
- The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
- Claude rows ran at default effort and GPT-6.1 Sol at medium; Codex CLI adds its own system prompt and tool schemas. Rows across routes compare route + model pairs.
Sources
Caching sessions and repeated prompts (Claude Code and Codex CLI)
Part 1: 5-turn CLI sessions over a fixed synthetic ledger, with the cache counters each provider reports per turn. Part 2: three prompts with deterministic validators, 10 repetitions per model. Declared protocols, validator controls before inference, every attempt kept; answers are published as ordinal ids, never as text.
Cost with and without the prompt cache (calculation)
Recorded tokens per turn × Anthropic list prices. With the cache: uncached input at the input price, cache reads at the cache-read price, 1-hour cache writes at twice the input price, 5-minute writes at 1.25 times (an assumption; none occurred). Without a cache: every input token at the input price. Output is priced the same in both. Not a bill.
Anthropic list prices (Claude models)
Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “Prompt caching and run-to-run consistency in Claude Code and Codex CLI”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/caching-consistency.
Explainers that cite this study
Read the methods and terms in the context of these recorded results.
- Benchmark saturation: when every model scores 100%
- How many runs does an LLM eval need? Sample size, with real intervals
- p95 latency, explained: median, tail and range for LLM calls
- Why the same prompt gives different answers: LLM nondeterminism, measured
- LLM pricing per million tokens, explained with real token counts
- pass@k explained: pass@1, pass@k and pass^k with real runs
- Prompt caching, explained with measured sessions
- Strict grading and format misses: when a right answer fails
- Time to first token (TTFT), explained with CLI and API timings
- What is context engineering? What to put in the context, measured
- Wilson confidence intervals for AI benchmarks
Models and comparisons in this study
Write-ups on this study
AI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
AI coding cost per developer: a formula built on recorded work
AI coding cost per developer: a monthly budget formula using recorded tokens and list-price calculations. Replace our task volume and pass rate with yours.
Best LLM for JSON output? Haiku, Sonnet and GPT-6.1 Sol, 10 runs each
Best LLM for JSON output? On one prompt, Sonnet and GPT-6.1 Sol passed 10/10 (95%: 72–100%); Haiku passed 1/10 (2–40%). CLI results, not JSON mode.
Claude Code cost per task: a price ladder from one decision to one agent run
$0.087 per Claude Code coding session at API list prices, $0.0036 per short call, $2.81 per full agent attempt. Calculations on recorded tokens, not bills.
Claude Haiku 4.5 vs Sonnet 5.5: all 80 comparison rows, and where the small model loses
Haiku 4.5 vs Sonnet 5.5 on 80 rows: Sonnet ahead on 14, Haiku on none, 31 ties. Hard tasks 11/24 vs 24/24, plus speed, memory and price.
Does a JSON schema stop format misses? Haiku went from 0/24 to 18/24
96 calls on 3 JSON tasks. Haiku 4.5 passed 0/24 with instructions and 18/24 with a JSON schema. Sonnet 5.5 and GPT-6.1 Sol passed 12/12 either way.
Does Claude Code reuse the prompt cache across sessions? A fixed-folder test
Claude Code showed near-full cache reads in later fixed-folder sessions (4/4; 95% interval 51–100%). New folders: 0/2 (0–66%). Exploratory, 30 calls.
Format misses vs wrong answers: your LLM eval may be failing right answers
5 of 13 failed calls on our hard set were right answers in the wrong format. How to grade LLM output strictly, test validators first and report both numbers.
GPT-6.1 Sol vs Claude Sonnet 5.5 vs Opus 5.5: every row we measured
Of 70 comparison rows for GPT-6.1 Sol, Claude Sonnet 5.5 and Opus 5.5, only 2 have a winner (speed). All 15 pass-rate rows tie. Tokens, price and route differ.
How many runs do you need to compare two AI models? A sample-size table
Calculation: 100 runs can separate an observed 8-point gap at a perfect score. See Wilson intervals, paired tests and limits for LLM eval sample sizes.
Is Claude Haiku cheaper than Sonnet? Cost per correct answer, with retries
Haiku 4.5 lists at half the price of Sonnet 5.5, yet calculated cost per correct hard answer was 4.7x higher. Retry maths, two exceptions and every source.
Most AI model comparisons are ties: 44 of 1,196 rows show a clear gap
44 of 1,196 AI model comparison rows show a gap under our overlap rules. Most gaps are timing rows. No effort row separates quality.
Plan for p95, not the median: LLM tail latency in our runs
4.30 s p95 against a 2.60 s median for Claude Sonnet 5.5; 34.5 s against 12.5 s for Haiku 4.5. Measured LLM tail latency and what to do about it.
The cheapest way to run an AI coding agent: 7 levers from measured runs
7 levers that may cut an AI coding agent's bill, sized from our data: prompt cache 3.9x, Fable/Sonnet cost per pass 6.5x, and 5 more. List-price calculations.
What is the best AI model for coding? Our data says four models tie
4 models tied at the top of our hard coding set: Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol. Only Haiku 4.5 separated. A tier list built from intervals.
When does prompt caching pay off? The Anthropic cache break-even, calculated
A 1-hour Anthropic cache write pays back after 2 reuses; a 5-minute write (assumed 1.25x) after 1. Break-even by model, cost per 1,000 sessions.
Which Claude model is fastest? It depends on the task, and on thinking
1.94 s was the lowest median on short calls (Fable 5.1). 7.75 s on hard calls (Sonnet 5.5). Haiku 4.5 took 4.43 s and 39.01 s with default thinking.
Which Claude model should you use? A task-by-task guide from our measurements
Which Claude model fits your job? Sonnet passed 24/24 hard calls (95% interval 86%–100%) at $0.01435 per pass (calculation). The top models hit a ceiling.
Would majority voting fix it? A self-consistency thought experiment on 90 real calls
Calculation on 90 recorded calls: a vote over 10 Haiku replies returns 289, not 282. Repeated errors and format misses limit self-consistency voting.
AI coding benchmarks roundup, October 2026: sixteen studies, every number in one place
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
An AI model leaderboard without a composite score: why, and how to read ours
Our leaderboard lists 46 models, CLIs, routers and providers with their best-supported facts, each with n and an interval. No single score. Here is why.
Does CLAUDE.md help? We tested 8 kinds of agent memory on Claude Code
200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a handbook and a Stop hook. Memory mattered where the repo was silent.
How much does prompt caching actually save? Measured in Claude Code and Codex
Claude Code read 97% of later-turn input from the cache. At list price that halved a 5-turn session and cut an agent bill about 3.9x. Turn 1 costs more.
Same prompt, ten answers: how consistent are Claude and Codex?
3 prompts, 10 runs each, on Haiku, Sonnet and GPT-6.1 Sol. 7 of 9 cells passed 10/10. Haiku gave the same wrong number 10 times: consistent is not correct.
More studies
All benchmarksHow much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.
Does a new Claude Code session reuse the prompt cache of an earlier one?
30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.