Explainer · Prompt caching
Prompt caching, explained with measured sessions
Definition
Prompt caching is a feature of model APIs that stores the processed form of a prompt prefix, so that a later request that starts with the same prefix reads it from the cache instead of processing it again. A cache read is billed at a fraction of the normal input price, and the first request that writes the prefix may cost more than normal input. Caching saves money when the same long context, such as a system prompt, tool definitions or a document, is sent many times.
Agent team · · Updated · 4 min read · Every number is from the public studies
Interactive
What the cache does, turn by turn
Each block is one turn's input. The filled part was read from the cache instead of being processed again.
List-price cost of this session (calculation)
| Turn | Claude Sonnet 5.5 · Claude Code | Claude Opus 5.5 · Claude Code | GPT-6.1 Sol (medium) · Codex CLI | n (sessions) |
|---|---|---|---|---|
| Turn 1 | 19% | 19% | 55% | 3 |
| Turn 2 | 99% | 99% | 99% | 3 |
| Turn 3 | 99% | 99% | 99% | 3 |
| Turn 4 | 99% | 99% | 99% | 3 |
| Turn 5 | 91% | 90% | 98% | 3 |
Claude Sonnet 5.5 · Claude Code: turn 1 read 19% of its input from the cache; turns 2 to 5 read 91% or more (n = 3 sessions). List-price calculation of the session: $0.14 with the cache, $0.27 without.
Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.
How it works
A model turns input tokens into internal state before it writes the first output token. For a long prompt this is most of the work. When the start of the next prompt is identical, the provider can keep that state and reuse it.
Three rules follow from this:
- Only an identical prefix is reused. One changed character early in the prompt ends the reuse at that point. Put the stable parts (instructions, tools, documents) first and the changing parts (the question) last.
- The cache expires. Providers keep an entry for a limited time. A write with a longer lifetime costs more.
- Reads and writes have their own prices. In our calculations for Anthropic models, a cache read is priced at the cache-read rate and a 1-hour cache write at 2× the input price.
How much of the input comes from the cache?
In our caching study, each session asked five questions about the same synthetic ledger. Turn 1 sends the ledger; turns 2 to 5 send one short question each.
- Inside one Claude Code session, turns 2 to 5 read 97% of their input from the cache on average.
- Turn 1 read 19%: that is the CLI's own prefix (its system prompt and tools), which was already cached.
- The Codex CLI (GPT-6.1 Sol) read 99% of later-turn input from its cache, on a larger context. Its app-server reports no cache writes, so we do not price it.
Each point is a mean over 3 sessions, so treat the exact shares as this run's values.
What it saves
The same recorded tokens, priced twice (a list-price calculation; the calls ran on a subscription):
- Claude Sonnet 5.5: $0.1350 with the cache vs $0.2698 without it, 50% less.
- Claude Opus 5.5: $0.2551 vs $0.5442, 53% less.
Turn 1 costs more with the cache than without, because the 1-hour write costs twice the input price. The saving comes from the later turns. A one-question session would pay the write and never collect the reads.
At agent scale the effect is larger. Over 33 SWE-bench attempts, 94.0% of the input tokens Agent recorded were cache reads. At Sonnet 5.5 list prices those tokens cost $87.23; without caching the same tokens would cost $343.33 (a calculation, not a run):
What caching does not do
- It did not make calls faster here. Median turn time was 1.6 s on turn 1 and 1.6 s on turns 2 to 5 for Sonnet, and 1.9 s vs 2.4 s for Opus. The fastest-to-slowest ranges overlap, and the later turns asked different questions. Short prompts leave little time to save.
- It did not carry over between sessions. On turn 1, all 4 later sessions wrote the ledger to the cache again, so a new session did not reuse the cache of an earlier one. We did not test the cause.
- It does not change the answer. The model sees the same tokens; caching changes the bill, not the reasoning.
How to get the most from it
- Keep the prefix stable: fixed instructions, tool definitions and documents first, in the same order, every time.
- Send the changing part last.
- Reuse one session for related questions instead of opening a new one for each.
- Check the usage fields your provider returns (cache read and cache write tokens) to confirm the hits.
- Price a workload with and without caching in the AI cost calculator.
- Find the break-even turn for your prefix, reads and writes in the prompt caching calculator.
Frequently asked questions
How much does prompt caching save?
It depends on how often the same prefix repeats. In our five-question Claude Code sessions, the list-price saving was 50% for Sonnet 5.5 and 53% for Opus 5.5. Over 33 recorded SWE-bench attempts, where 94.0% of input was cache reads, the Sonnet 5.5 bill was $87.23 with caching and would be $343.33 without it (a calculation).
Does prompt caching make responses faster?
Not in our sessions. Median turn times for the first turn and the later turns were close, and their ranges overlap. Very long prompts can show a speed gain; these prompts were short.
Why does the first request cost more with caching?
The first request writes the prefix to the cache, and a cache write is priced above normal input (a 1-hour write at 2× the input price in our calculation). The saving comes only when later requests read that prefix.
Does a new session reuse an earlier session's cache?
Not in our runs: all 4 later Claude Code sessions wrote the ledger to the cache again on their first turn. We did not test why.
Watch the data
Prompt caching and consistency: what the cache saves, and how much answers vary
Calculation at list price: the cache cut a 5-question session 50% on Sonnet 5.5 and 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every repetition.
Transcript
- Caching and consistency · 135 calls. What the cache saves, and how much answers vary. Five-question sessions on a fixed context. Then the same prompt, 10 times.
- 135 calls: 45 cache turns, 90 repeated prompts. Every call counted. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- Claude Code turn 1 reads 19% from the cache, the CLI’s own prefix. Turns 2–5 read 90%–99%. Codex CLI: 98%–99%. Chart: Share of input read from the cache · mean of 3 sessions per turn (n = 3 each). Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- A calculation: Sonnet 5.5 $0.1350 with the cache vs $0.2698 without, 50% less. Opus 5.5 $0.2551 vs $0.5442, 53% less. Chart: List-price cost of the recorded sessions · calculation (n = 15 each). Calculation, not a run. Caveat: Costs are list-price calculations; the calls used flat subscriptions.
- No clear speed effect: Sonnet 5.5 1.6 s on turn 1 vs 1.6 s later; Opus 5.5 1.9 s vs 2.4 s. The ranges overlap. Sonnet 5.5: median turn 1 vs turns 2–5: 1.6 s vs 1.6 s (ranges 1.6–1.8 s and 1.4–5.6 s · n = 3 and 12). Opus 5.5: median turn 1 vs turns 2–5: 1.9 s vs 2.4 s (ranges 1.8–4.4 s and 1.6–12.7 s · n = 3 and 12). Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- 7 of 9 cells passed 10/10. Haiku 4.5: 0/10 on the exact number (10 wrong, 1 distinct answer), 1/10 on JSON (9 format misses). Chart: Same prompt, 10 times · strict passes · a 10/10 is 72%–100% at 95% (n = 10 each). Caveat: Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
- Consistent is not correct: Haiku 4.5 gave the same wrong number all 10 times. Code fix, distinct correct bodies: Haiku 4.5 6, Sonnet 5.5 3, GPT-6.1 Sol (medium) 6. Chart: Same prompt, 10 times · distinct answers (n = 10 each). Caveat: The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
- Cache the fixed context; check answers, not agreement. Every call online.