Explainer · Prompt caching

Prompt caching, explained with measured sessions

Definition

Prompt caching is a feature of model APIs that stores the processed form of a prompt prefix, so that a later request that starts with the same prefix reads it from the cache instead of processing it again. A cache read is billed at a fraction of the normal input price, and the first request that writes the prefix may cost more than normal input. Caching saves money when the same long context, such as a system prompt, tool definitions or a document, is sent many times.

Agent team · · Updated · 4 min read · Every number is from the public studies

Interactive

What the cache does, turn by turn

Each block is one turn's input. The filled part was read from the cache instead of being processed again.

MeasuredCost: calculation

List-price cost of this session (calculation)

Without a cache$0.27
With the cache$0.14

Claude Sonnet 5.5 · Claude Code: turn 1 read 19% of its input from the cache; turns 2 to 5 read 91% or more (n = 3 sessions). List-price calculation of the session: $0.14 with the cache, $0.27 without.

Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.

Source: Prompt caching and consistency study

How it works

A model turns input tokens into internal state before it writes the first output token. For a long prompt this is most of the work. When the start of the next prompt is identical, the provider can keep that state and reuse it.

Three rules follow from this:

  1. Only an identical prefix is reused. One changed character early in the prompt ends the reuse at that point. Put the stable parts (instructions, tools, documents) first and the changing parts (the question) last.
  2. The cache expires. Providers keep an entry for a limited time. A write with a longer lifetime costs more.
  3. Reads and writes have their own prices. In our calculations for Anthropic models, a cache read is priced at the cache-read rate and a 1-hour cache write at 2× the input price.

How much of the input comes from the cache?

In our caching study, each session asked five questions about the same synthetic ledger. Turn 1 sends the ledger; turns 2 to 5 send one short question each.

  • Claude Sonnet 5.5 · Claude Code
  • Claude Opus 5.5 · Claude Code
  • GPT-6.1 Sol (medium) · Codex CLI*

* The note below the chart says what this route does not report.

5 turn in the sessions, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (medium) · Codex CLI. Claude Sonnet 5.5 · Claude Code: highest Turn 2 99% (n 3). Lowest Turn 1 19% (n 3). Claude Opus 5.5 · Claude Code: highest Turn 2 99% (n 3). Lowest Turn 1 19% (n 3).

Notesn = 3 per row

Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each

Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

  • Inside one Claude Code session, turns 2 to 5 read 97% of their input from the cache on average.
  • Turn 1 read 19%: that is the CLI's own prefix (its system prompt and tools), which was already cached.
  • The Codex CLI (GPT-6.1 Sol) read 99% of later-turn input from its cache, on a larger context. Its app-server reports no cache writes, so we do not price it.

Each point is a mean over 3 sessions, so treat the exact shares as this run's values.

What it saves

Calculation
  • With the cache, as recorded
  • Without a cache: every input token at the input price (square)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code

Gap labels, Without a cache: every input token at the input price vs With the cache, as recorded: Without a cache: every input token at the input price is x% higher (+) or lower (−) than With the cache, as recorded, calculated from the two values shown (the change counted from With the cache, as recorded’s value).

List-price calculation, not a run. 2 rows, 2 series: With the cache, as recorded, Without a cache: every input token at the input price. With the cache, as recorded: highest Claude Opus 5.5 · Claude Code $0.26 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.14 (n 15). Without a cache: every input token at the input price: highest Claude Opus 5.5 · Claude Code $0.54 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.27 (n 15).

Notesn = 15 per row

All recorded turns per model; the same reported tokens priced two ways

Calculation, not a bill: the calls ran on a subscription. Cache reads at the cache-read price, 1-hour cache writes at 2× the input price (every write in this run was a 1-hour write). Codex CLI is not priced here: it reports no cache-write count.

Sources: Caching sessions and repeated prompts (Claude Code and Codex CLI), Cost with and without the prompt cache (calculation), Anthropic list prices (Claude models)

The same recorded tokens, priced twice (a list-price calculation; the calls ran on a subscription):

  • Claude Sonnet 5.5: $0.1350 with the cache vs $0.2698 without it, 50% less.
  • Claude Opus 5.5: $0.2551 vs $0.5442, 53% less.

Turn 1 costs more with the cache than without, because the 1-hour write costs twice the input price. The saving comes from the later turns. A one-question session would pay the write and never collect the reads.

At agent scale the effect is larger. Over 33 SWE-bench attempts, 94.0% of the input tokens Agent recorded were cache reads. At Sonnet 5.5 list prices those tokens cost $87.23; without caching the same tokens would cost $343.33 (a calculation, not a run):

Calculation
  • With caching (as recorded)
  • Without caching (square)
In chart order.
Claude Haiku 4.5
Claude Sonnet 5.5
Claude Opus 5.5
Claude Fable 5.1

Gap labels, Without caching vs With caching (as recorded): Without caching is x% higher (+) or lower (−) than With caching (as recorded), calculated from the two values shown (the change counted from With caching (as recorded)’s value).

List-price calculation, not a run. 4 rows, 2 series: With caching (as recorded), Without caching. With caching (as recorded): highest Claude Fable 5.1 $321. Lowest Claude Haiku 4.5 $43.61. Without caching: highest Claude Fable 5.1 $1,717. Lowest Claude Haiku 4.5 $172.

Notes

The same recorded tokens with and without cache pricing

94.0% of recorded input tokens were cache reads. Calculation, not a run.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)

What caching does not do

  • It did not make calls faster here. Median turn time was 1.6 s on turn 1 and 1.6 s on turns 2 to 5 for Sonnet, and 1.9 s vs 2.4 s for Opus. The fastest-to-slowest ranges overlap, and the later turns asked different questions. Short prompts leave little time to save.
  • It did not carry over between sessions. On turn 1, all 4 later sessions wrote the ledger to the cache again, so a new session did not reuse the cache of an earlier one. We did not test the cause.
  • It does not change the answer. The model sees the same tokens; caching changes the bill, not the reasoning.

How to get the most from it

  • Keep the prefix stable: fixed instructions, tool definitions and documents first, in the same order, every time.
  • Send the changing part last.
  • Reuse one session for related questions instead of opening a new one for each.
  • Check the usage fields your provider returns (cache read and cache write tokens) to confirm the hits.
  • Price a workload with and without caching in the AI cost calculator.
  • Find the break-even turn for your prefix, reads and writes in the prompt caching calculator.

Frequently asked questions

How much does prompt caching save?

It depends on how often the same prefix repeats. In our five-question Claude Code sessions, the list-price saving was 50% for Sonnet 5.5 and 53% for Opus 5.5. Over 33 recorded SWE-bench attempts, where 94.0% of input was cache reads, the Sonnet 5.5 bill was $87.23 with caching and would be $343.33 without it (a calculation).

Does prompt caching make responses faster?

Not in our sessions. Median turn times for the first turn and the later turns were close, and their ranges overlap. Very long prompts can show a speed gain; these prompts were short.

Why does the first request cost more with caching?

The first request writes the prefix to the cache, and a cache write is priced above normal input (a 1-hour write at 2× the input price in our calculation). The saving comes only when later requests read that prefix.

Does a new session reuse an earlier session's cache?

Not in our runs: all 4 later Claude Code sessions wrote the ledger to the cache again on their first turn. We did not test why.

Watch the data

Live story · 49 sPrompt caching and consistency: what the cache saves, and how much answers vary

Prompt caching and consistency: what the cache saves, and how much answers vary

Calculation at list price: the cache cut a 5-question session 50% on Sonnet 5.5 and 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every repetition.

Transcript
  1. Caching and consistency · 135 calls. What the cache saves, and how much answers vary. Five-question sessions on a fixed context. Then the same prompt, 10 times.
  2. 135 calls: 45 cache turns, 90 repeated prompts. Every call counted. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
  3. Claude Code turn 1 reads 19% from the cache, the CLI’s own prefix. Turns 2–5 read 90%–99%. Codex CLI: 98%–99%. Chart: Share of input read from the cache · mean of 3 sessions per turn (n = 3 each). Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  4. A calculation: Sonnet 5.5 $0.1350 with the cache vs $0.2698 without, 50% less. Opus 5.5 $0.2551 vs $0.5442, 53% less. Chart: List-price cost of the recorded sessions · calculation (n = 15 each). Calculation, not a run. Caveat: Costs are list-price calculations; the calls used flat subscriptions.
  5. No clear speed effect: Sonnet 5.5 1.6 s on turn 1 vs 1.6 s later; Opus 5.5 1.9 s vs 2.4 s. The ranges overlap. Sonnet 5.5: median turn 1 vs turns 2–5: 1.6 s vs 1.6 s (ranges 1.6–1.8 s and 1.4–5.6 s · n = 3 and 12). Opus 5.5: median turn 1 vs turns 2–5: 1.9 s vs 2.4 s (ranges 1.8–4.4 s and 1.6–12.7 s · n = 3 and 12). Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
  6. 7 of 9 cells passed 10/10. Haiku 4.5: 0/10 on the exact number (10 wrong, 1 distinct answer), 1/10 on JSON (9 format misses). Chart: Same prompt, 10 times · strict passes · a 10/10 is 72%–100% at 95% (n = 10 each). Caveat: Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
  7. Consistent is not correct: Haiku 4.5 gave the same wrong number all 10 times. Code fix, distinct correct bodies: Haiku 4.5 6, Sonnet 5.5 3, GPT-6.1 Sol (medium) 6. Chart: Same prompt, 10 times · distinct answers (n = 10 each). Caveat: The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
  8. Cache the fixed context; check answers, not agreement. Every call online.

The data behind this explainer

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.