• Prompt Caching
  • Consistency
  • Variance
  • Claude Code
  • Codex CLI
  • Claude Haiku
  • Claude Sonnet
  • Claude Opus
  • GPT-6.1 Sol
  • Calculation

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

When a CLI session reuses a fixed context, how much input comes from the cache, what does that save at list price, and does it change latency? When the same prompt runs 10 times, how much do the pass rate, the answer and the time vary?

Published · 6 charts · Download the data or a carousel

50%

Calculation

List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation) · ($0.1350 vs $0.2698)

Calculation from recorded tokens and list prices; not a bill.

The answer

Caching: inside one Claude Code session, turns 2-5 read 97% of their input from the cache on average (turn 1: 19%, the CLI's own prefix). At list price, a calculation, all recorded turns cost Sonnet $0.1350 vs $0.2698 (50% less) and Opus $0.2551 vs $0.5442 (53% less) without the cache. Turn 1 costs more with the cache, because a 1-hour cache write costs twice the input price. A new session did not reuse the cache of an earlier one: on turn 1, all 4 later sessions wrote the ledger to the cache again. The cause was not tested. The cache showed no clear speed effect: median turn time was Sonnet 1.6 s on turn 1 vs 1.6 s on turns 2-5 and Opus 1.9 s on turn 1 vs 2.4 s on turns 2-5, and the fastest-to-slowest ranges overlap. Codex CLI (GPT-6.1 Sol) read 99% of later-turn input from its cache on a larger context; its app-server reports no cache writes, so no cost is calculated for it. Consistency: 7 of 9 model-and-prompt cells passed all 10 repetitions (95% interval 72% to 100%). Haiku passed 0/10 on the exact-number prompt; Haiku passed 1/10 on the JSON prompt (9 more were correct but in the wrong format). Haiku gave the same wrong answer every time (289; expected 282): consistent is not the same as correct. The code-fix prompt gave 6 different code bodies for Haiku, 3 different code bodies for Sonnet and 6 different code bodies for GPT-6.1 Sol (medium).

Live story

Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.

Live story · 49 sPrompt caching and consistency: what the cache saves, and how much answers vary

Prompt caching and consistency: what the cache saves, and how much answers vary

Calculation at list price: the cache cut a 5-question session 50% on Sonnet 5.5 and 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every repetition.

Transcript
  1. Caching and consistency · 135 calls. What the cache saves, and how much answers vary. Five-question sessions on a fixed context. Then the same prompt, 10 times.
  2. 135 calls: 45 cache turns, 90 repeated prompts. Every call counted. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
  3. Claude Code turn 1 reads 19% from the cache, the CLI’s own prefix. Turns 2–5 read 90%–99%. Codex CLI: 98%–99%. Chart: Share of input read from the cache · mean of 3 sessions per turn (n = 3 each). Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  4. A calculation: Sonnet 5.5 $0.1350 with the cache vs $0.2698 without, 50% less. Opus 5.5 $0.2551 vs $0.5442, 53% less. Chart: List-price cost of the recorded sessions · calculation (n = 15 each). Calculation, not a run. Caveat: Costs are list-price calculations; the calls used flat subscriptions.
  5. No clear speed effect: Sonnet 5.5 1.6 s on turn 1 vs 1.6 s later; Opus 5.5 1.9 s vs 2.4 s. The ranges overlap. Sonnet 5.5: median turn 1 vs turns 2–5: 1.6 s vs 1.6 s (ranges 1.6–1.8 s and 1.4–5.6 s · n = 3 and 12). Opus 5.5: median turn 1 vs turns 2–5: 1.9 s vs 2.4 s (ranges 1.8–4.4 s and 1.6–12.7 s · n = 3 and 12). Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
  6. 7 of 9 cells passed 10/10. Haiku 4.5: 0/10 on the exact number (10 wrong, 1 distinct answer), 1/10 on JSON (9 format misses). Chart: Same prompt, 10 times · strict passes · a 10/10 is 72%–100% at 95% (n = 10 each). Caveat: Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
  7. Consistent is not correct: Haiku 4.5 gave the same wrong number all 10 times. Code fix, distinct correct bodies: Haiku 4.5 6, Sonnet 5.5 3, GPT-6.1 Sol (medium) 6. Chart: Same prompt, 10 times · distinct answers (n = 10 each). Caveat: The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
  8. Cache the fixed context; check answers, not agreement. Every call online.

Key numbers

53%

List-price saving from the cache over 15 turns, Claude Opus 5.5 · Claude Code (calculation)

($0.2551 vs $0.5442)

0 of 4

Later sessions whose first turn read the ledger from an earlier session’s cache

n = 4

7 of 9

Model-and-prompt cells that passed all 10 repetitions

n = 9

135

Calls in this study (every one counted)

(45 cache turns, 90 repeated prompts)

Calculate when repeated cache reads repay the first write

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

  • Claude Sonnet 5.5 · Claude Code
  • Claude Opus 5.5 · Claude Code
  • GPT-6.1 Sol (medium) · Codex CLI*

* The note below the chart says what this route does not report.

5 turn in the sessions, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (medium) · Codex CLI. Claude Sonnet 5.5 · Claude Code: highest Turn 2 99% (n 3). Lowest Turn 1 19% (n 3). Claude Opus 5.5 · Claude Code: highest Turn 2 99% (n 3). Lowest Turn 1 19% (n 3).

Notesn = 3 per row

Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each

Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Share card (PNG)
Calculation
  • With the cache, as recorded
  • Without a cache: every input token at the input price (square)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code

Gap labels, Without a cache: every input token at the input price vs With the cache, as recorded: Without a cache: every input token at the input price is x% higher (+) or lower (−) than With the cache, as recorded, calculated from the two values shown (the change counted from With the cache, as recorded’s value).

List-price calculation, not a run. 2 rows, 2 series: With the cache, as recorded, Without a cache: every input token at the input price. With the cache, as recorded: highest Claude Opus 5.5 · Claude Code $0.26 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.14 (n 15). Without a cache: every input token at the input price: highest Claude Opus 5.5 · Claude Code $0.54 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.27 (n 15).

Notesn = 15 per row

All recorded turns per model; the same reported tokens priced two ways

Calculation, not a bill: the calls ran on a subscription. Cache reads at the cache-read price, 1-hour cache writes at 2× the input price (every write in this run was a 1-hour write). Codex CLI is not priced here: it reports no cache-write count.

Sources: Caching sessions and repeated prompts (Claude Code and Codex CLI), Cost with and without the prompt cache (calculation), Anthropic list prices (Claude models)

Share card (PNG)
  • Turn 1 (writes the ledger to the cache)
  • Turns 2-5 (read the ledger from the cache) (square)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code

Gap labels, Turns 2-5 (read the ledger from the cache) vs Turn 1 (writes the ledger to the cache): Turns 2-5 (read the ledger from the cache) is x% higher (+) or lower (−) than Turn 1 (writes the ledger to the cache), calculated from the two values shown (the change counted from Turn 1 (writes the ledger to the cache)’s value); lines are the fastest–slowest run (not an interval).

2 rows, 2 series: Turn 1 (writes the ledger to the cache), Turns 2-5 (read the ledger from the cache). Turn 1 (writes the ledger to the cache): slowest Claude Opus 5.5 · Claude Code 1.9 s (range 1.8 s–4.4 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1.6 s–1.8 s, n 3). All run ranges overlap. Turns 2-5 (read the ledger from the cache): slowest Claude Opus 5.5 · Claude Code 2.4 s (range 1.6 s–12.7 s, n 12). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1.4 s–5.6 s, n 12). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 3–12 per row

Median; whiskers = fastest and slowest turn

Whiskers are a range (fastest and slowest turn), not a confidence interval. Turns ask different questions: the slow later turns are the counting question, which produced the most output.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Share card (PNG)
  • Exact number
  • JSON object
  • Code fix
Claude Haiku 4.5
Claude Sonnet 5.5
GPT-6.1 Sol (medium)

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 10 per row

One series per prompt; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Share card (PNG)

Exact number

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

JSON object

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Code fix

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

One panel per series, all on the same axis.

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: all at 1. JSON object: all at 1.

Notesn = 10 per row

Distinct normalized answers over 10 repetitions (1 = the same answer every time)

Normalized: the last number for the exact-number prompt; key-sorted JSON for the JSON prompt; for the code fix, the code without fences, comments, whitespace, quote style and line-end semicolons. One distinct answer is not the same as a correct answer: an answer can be the same and wrong every time. Different code can be equally correct.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Share card (PNG)
  • Exact number
  • JSON object
  • Code fix
Entrance: medians race at 9.6× real timeMotion reduced: press Replay to animateThe slowest median is 13.4 s. The clock runs at the recorded speed.
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: slowest GPT-6.1 Sol (medium) · Codex CLI 13.4 s (range 12.3 s–18 s, n 10). Fastest Claude Haiku 4.5 · Claude Code 5.1 s (range 4.4 s–6.2 s, n 10). Not all run ranges overlap. JSON object: slowest Claude Haiku 4.5 · Claude Code 7 s (range 5.3 s–12.3 s, n 10). Fastest Claude Sonnet 5.5 · Claude Code 2.9 s (range 2.7 s–5.3 s, n 10). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 10 per row

Median; whiskers = fastest and slowest of 10 calls

Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Share card (PNG)

Tables

Cache counters per turn (mean over sessions)

ConfigurationTurnSessionsExact answersInput tokens (all)Read from cacheWritten to cacheRead shareMedian time (s)USD with cache (calculation)USD without cache (calculation)
Claude Sonnet 5.5 · Claude Code133/37,8311,4636,36619%1.6 s$0.026$0.016
Claude Sonnet 5.5 · Claude Code233/37,8897,8295899%1.5 s$0.0019$0.016
Claude Sonnet 5.5 · Claude Code333/37,9707,8878199%1.5 s$0.002$0.016
Claude Sonnet 5.5 · Claude Code433/38,0327,9686299%5.2 s$0.0096$0.024
Claude Sonnet 5.5 · Claude Code533/38,8618,03082991%1.8 s$0.0058$0.019
Claude Opus 5.5 · Claude Code133/37,8281,4636,36319%1.9 s$0.051$0.031
Claude Opus 5.5 · Claude Code233/37,8867,8265899%1.6 s$0.0026$0.032
Claude Opus 5.5 · Claude Code333/37,9917,88410599%1.8 s$0.003$0.033
Claude Opus 5.5 · Claude Code433/38,0757,9898499%6.1 s$0.018$0.048
Claude Opus 5.5 · Claude Code533/38,9418,07386690%2.1 s$0.0097$0.037
GPT-6.1 Sol (medium) · Codex CLI133/318,77410,411not reported55%3.6 s——
GPT-6.1 Sol (medium) · Codex CLI233/318,80318,560not reported99%2.4 s——
GPT-6.1 Sol (medium) · Codex CLI333/318,84818,645not reported99%3.1 s——
GPT-6.1 Sol (medium) · Codex CLI433/318,88018,688not reported99%6 s——
GPT-6.1 Sol (medium) · Codex CLI533/319,09718,688not reported98%2.4 s——

Every consistency cell (10 repetitions each)

Ten repeats of the same prompt per cell: every run as a mark

  • Strict pass
  • Format miss (right answer, wrong format)
  • Wrong answer

Claude Haiku 4.5 · Claude Code

Claude Sonnet 5.5 · Claude Code

GPT-6.1 Sol (medium) · Codex CLI

Marks are counts from the table, sorted by outcome (run order is not recorded)95% Wilson interval on strict passesn = 10 runs per cell

9 cells: 3 configurations by 3 prompts, 10 runs each. Each row shows strict passes, format misses and wrong answers as marks, the distinct answers and the median time.

Method

  1. Protocols declared before the first call, one per route. Every attempt is kept; nothing was retried.
  2. Caching: 9 sessions (3 × Claude Sonnet 5.5 · Claude Code, 3 × Claude Opus 5.5 · Claude Code and 3 × GPT-6.1 Sol (medium) · Codex CLI), 5 turns each. A session is one CLI process. Turn 1 sends a seeded synthetic stock ledger plus question 1; turns 2-5 send one short question each (lookups, a count, an arg-max), each with one exact answer. Each turn is one model request.
  3. A 2-call probe sized the context before the run (not part of any cell). The ledger was larger than declared, so it was cut once, from 170 to 100 lines, and the answers were recomputed, as the protocol allowed.
  4. The Codex sessions ran before that cut, on the 170-line ledger (16,197 characters vs 9,651). The two routes are reported side by side, never as a like-for-like pair.
  5. Cache counters as the provider reports them: Claude Code gives uncached input, cache reads and cache writes (with the 5-minute and 1-hour split); the Codex app-server gives input (cached included) and cached input, and no cache writes.
  6. Cost with and without the cache (Claude only): see the calculation source. Every write in this run was a 1-hour write, so the 5-minute multiplier (an assumption) was not used.
  7. Consistency: 3 prompts (an exact number, a JSON object with exact keys, a small code fix), each with a deterministic validator; 10 repetitions per prompt for Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI. One-shot calls, one at a time per account. Answer diversity counts distinct normalized answers; the raw extract holds ordinal answer ids, never the text.
  8. Validator controls ran before inference on each route: every reference answer passes, every plausible wrong answer fails, and a wrapped reference is flagged as a format miss.
  9. Isolation as in the hard head-to-head: fresh empty working folder, tools off, no MCP servers, no session persistence across processes, provider-default caching. Claude at its default effort; GPT-6.1 Sol at medium.
  10. No batch stopped early and nothing was trimmed. The Claude CLI reported a rate-limit status of "allowed_warning" on 7 of 30 cache turns; no call was refused and the run did not stop.

Caveats

  • Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
  • Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  • Costs are list-price calculations; the calls used flat subscriptions.
  • Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
  • The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
  • Claude rows ran at default effort and GPT-6.1 Sol at medium; Codex CLI adds its own system prompt and tool schemas. Rows across routes compare route + model pairs.

Sources

  • Caching sessions and repeated prompts (Claude Code and Codex CLI)

    Our recorded runs ·

    Part 1: 5-turn CLI sessions over a fixed synthetic ledger, with the cache counters each provider reports per turn. Part 2: three prompts with deterministic validators, 10 repetitions per model. Declared protocols, validator controls before inference, every attempt kept; answers are published as ordinal ids, never as text.

    Raw data: caching-consistency/caching.json, caching-consistency/consistency.json

  • Cost with and without the prompt cache (calculation)

    Calculation ·

    Recorded tokens per turn × Anthropic list prices. With the cache: uncached input at the input price, cache reads at the cache-read price, 1-hour cache writes at twice the input price, 5-minute writes at 1.25 times (an assumption; none occurred). Without a cache: every input token at the input price. Output is priced the same in both. Not a bill.

    Raw data: caching-consistency/caching.json

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Prompt caching and run-to-run consistency in Claude Code and Codex CLI”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/caching-consistency.

Explainers that cite this study

Read the methods and terms in the context of these recorded results.

Models and comparisons in this study

More studies

All benchmarks
Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

  • Prompt Caching
  • Cache Reuse

Does a new Claude Code session reuse the prompt cache of an earlier one?

30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.

0of 2 (95% interval 0% to 66%) · Later sessions with at least 50% of turn-1 input cached, A: new folder each time · n = 2

4 chartsUpdated October 7, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.