How much does prompt caching actually save? Measured in Claude Code and Codex
Claude Code read 97% of later-turn input from the cache. At list price that halved a 5-turn session and cut an agent bill about 3.9x. Turn 1 costs more.
TL;DR
- In a 5-turn Claude Code session with a fixed context, turns 2 to 5 read on average 97% of their input from the cache. Turn 1 read 19%: that is the CLI's own prefix.
- At list price (a calculation), the cache cut the session cost about in half: Sonnet 5.5 $0.1350 vs $0.2698 (50% less) and Opus 5.5 $0.2551 vs $0.5442 (53% less), over 15 turns each.
- Turn 1 costs more with the cache. A 1-hour cache write is priced at twice the input price. In our sessions the write paid for itself on turn 2 (our arithmetic).
- On a long agent workload, the effect is larger. Our agent's SWE-bench run read 94.0% of 162.9M input tokens from the cache. At Sonnet list price that is $87.23 with the cache and $343.33 without, about 3.9x (a calculation).
- No cross-session reuse: on turn 1, 0 of 4 later sessions read the ledger from an earlier session's cache. We did not test why.
- No clear speed effect. Median turn times were close, and the ranges overlap.
- The Codex CLI read 99% of later-turn input from its cache, but it reports no cache writes, so we do not price it.
Study: /benchmarks/caching-consistency. Earlier repricing: /benchmarks/cost-thought-experiments.
The question
Every vendor says prompt caching saves money (how prompt caching works). The real question is how much, on your kind of work. That depends on three things:
- How much of each request repeats the one before it.
- What the provider charges to write the cache, not only to read it.
- Whether the cache survives between sessions.
We measured the first and the third, and priced the second with list prices. Then we put the result next to the cache share our agent recorded on SWE-bench Verified.
What we ran
- 9 sessions, 5 turns each, 45 turns in all: 3 sessions each for Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI.
- Turn 1 sends a seeded synthetic stock ledger plus a question. Turns 2 to 5 send one short question each: a lookup, a count, an arg-max. Each question has one exact answer. All 45 turns gave the exact answer.
- Cache counters come from the provider, as reported: Claude Code gives uncached input, cache reads and cache writes; the Codex app-server gives input and cached input, and no writes.
- Costs are list price × recorded tokens, a calculation. Every write in this run was a 1-hour write.
The Codex sessions ran on a larger ledger (16,197 characters against 9,651), so we show the two routes side by side, not as a like-for-like pair.
How much comes from the cache
Share of input read from the cache, mean over 3 sessions:
| Turn | Sonnet 5.5 | Opus 5.5 | GPT-6.1 Sol (Codex CLI) |
|---|---|---|---|
| 1 | 19% | 19% | 55% |
| 2 | 99% | 99% | 99% |
| 3 | 99% | 99% | 99% |
| 4 | 99% | 99% | 99% |
| 5 | 91% | 90% | 98% |
Turn 1 already hits the cache for a part of its input: about 1,463 tokens on both Claude models. That is the part every Claude Code call starts with. The ledger itself, about 6,366 tokens, is written to the cache on turn 1 and read back on every later turn. Turn 5 dips because it wrote about 830 to 870 new tokens to the cache.
What that saves at list price
The chart prices the same recorded tokens two ways: as recorded, and with every input token at the full input price.
- Sonnet 5.5: $0.1350 with the cache, $0.2698 without: 50% less.
- Opus 5.5: $0.2551 with the cache, $0.5442 without: 53% less.
Both are totals over 15 turns (3 sessions × 5 turns). Opus saves a little more because its input costs twice Sonnet's, while its cache-read price is the same.
Turn 1 costs more
The cache does not save on the first turn. It costs more:
| Turn 1, mean per session | With the cache | Without | Our ratio |
|---|---|---|---|
| Sonnet 5.5 | $0.0258 | $0.0157 | 1.6x |
| Opus 5.5 | $0.0513 | $0.0314 | 1.6x |
The reason is the write price. A 1-hour cache write costs twice the input price ($4 per million tokens on Sonnet, $8 on Opus). So you pay extra to store the ledger, and you get it back only when a later turn reads it.
When the write pays back
Our arithmetic on the per-turn means in the study table:
- Sonnet: turn 1 cost $0.0101 more with the cache. Turn 2 alone saved $0.0140. So the cache paid back on turn 2.
- Opus: turn 1 cost $0.0199 more. Turn 2 saved $0.0295. Same result.
- Over turns 2 to 5, the cache cut the cost by 74% on Sonnet and 78% on Opus.
A one-shot call with a large, unique context is the bad case: you pay the write and never read it. A loop that re-sends the same context is the good case.
The agent-scale picture
The 5-turn sessions are small. Our agent's recorded SWE-bench Verified run is a real long workload: 162.9M input tokens, 94.0% of them cache reads, and 1.8M output tokens across 33 attempts.
Priced at list price, a calculation, not a run:
- Sonnet 5.5: $87.23 with caching, $343.33 without: about 3.9x (our ratio), or 75% less.
- Opus 5.5: $143.83 with caching, $686.66 without.
The savings are larger here than in our sessions for a simple reason: an agent loop re-reads its growing context on every model call, many more than 5 times. The more turns re-send the same prefix, the closer the cost falls to the cache-read price.
No reuse across sessions
You might expect a second session with the same ledger to read it from the first session's cache. It did not. On turn 1, all 4 later sessions wrote the ledger to the cache again (stat: 0 of 4 reused).
We did not test the cause. The CLI may add per-process context before the user message, which would change the prefix. Do not read this as a provider property. Read it as: in this CLI version, plan for one cache write per session.
Does the cache make calls faster?
Not clearly in our data:
- Sonnet 5.5: turn 1 median 1.64 s (1.58 to 1.79 s, n = 3), turns 2 to 5 median 1.61 s (1.35 to 5.63 s, n = 12).
- Opus 5.5: turn 1 median 1.90 s (1.78 to 4.36 s), turns 2 to 5 median 2.40 s (1.63 to 12.67 s).
The ranges overlap, and the context is small. Caching may help latency on much larger contexts. We did not see it here, so we do not claim it.
Codex: high cache share, no price
The Codex CLI read 99% of later-turn input from its cache on average, and 55% on turn 1. Its app-server reports cached input but no cache writes, so we cannot price the write side, and we do not calculate a saving for it.
How to use this
- Keep the prefix stable. Put the large, fixed part of the prompt first, and the part that changes last.
- Count the write. A 1-hour write costs 2x input. If a context is used once, caching it costs more.
- Expect one write per session. In our runs, a new CLI session did not reuse an earlier session's cache.
- Check the cache-read price, not only the input price. It decides the cost of long agent loops. For example, Opus 5.5 lists at twice Sonnet's input price but the same cache-read price.
- Run your own mix in the AI cost calculator, with and without caching.
Caveats
- 3 sessions per model. The figures describe this CLI version and this context size.
- Calculations. All costs are list price × recorded tokens. The calls ran on flat subscriptions.
- Different routes. The Codex sessions used a larger ledger; compare the routes side by side only.
- Repricing is not a run. At another model's price, the SWE-bench agent would have made different calls.
What to read next
- Prompt caching calculator: find the break-even turn
- How to estimate your AI coding bill
- What if every call ran on Opus?
- Claude Sonnet vs Opus: when is Opus worth the price?
- Claude Code vs Codex CLI: the hidden context tax
See your own cache share
Agent records cache reads and writes for every model call, so you can see how much of your bill the cache already saves. Try Agent and check your own numbers.