• Prompt Caching
  • LLM pricing
  • Claude Code
  • Codex CLI

How much does prompt caching actually save? Measured in Claude Code and Codex

Claude Code read 97% of later-turn input from the cache. At list price that halved a 5-turn session and cut an agent bill about 3.9x. Turn 1 costs more.

TL;DR

  • In a 5-turn Claude Code session with a fixed context, turns 2 to 5 read on average 97% of their input from the cache. Turn 1 read 19%: that is the CLI's own prefix.
  • At list price (a calculation), the cache cut the session cost about in half: Sonnet 5.5 $0.1350 vs $0.2698 (50% less) and Opus 5.5 $0.2551 vs $0.5442 (53% less), over 15 turns each.
  • Turn 1 costs more with the cache. A 1-hour cache write is priced at twice the input price. In our sessions the write paid for itself on turn 2 (our arithmetic).
  • On a long agent workload, the effect is larger. Our agent's SWE-bench run read 94.0% of 162.9M input tokens from the cache. At Sonnet list price that is $87.23 with the cache and $343.33 without, about 3.9x (a calculation).
  • No cross-session reuse: on turn 1, 0 of 4 later sessions read the ledger from an earlier session's cache. We did not test why.
  • No clear speed effect. Median turn times were close, and the ranges overlap.
  • The Codex CLI read 99% of later-turn input from its cache, but it reports no cache writes, so we do not price it.

Study: /benchmarks/caching-consistency. Earlier repricing: /benchmarks/cost-thought-experiments.

The question

Every vendor says prompt caching saves money (how prompt caching works). The real question is how much, on your kind of work. That depends on three things:

  1. How much of each request repeats the one before it.
  2. What the provider charges to write the cache, not only to read it.
  3. Whether the cache survives between sessions.

We measured the first and the third, and priced the second with list prices. Then we put the result next to the cache share our agent recorded on SWE-bench Verified.

What we ran

  • 9 sessions, 5 turns each, 45 turns in all: 3 sessions each for Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI.
  • Turn 1 sends a seeded synthetic stock ledger plus a question. Turns 2 to 5 send one short question each: a lookup, a count, an arg-max. Each question has one exact answer. All 45 turns gave the exact answer.
  • Cache counters come from the provider, as reported: Claude Code gives uncached input, cache reads and cache writes; the Codex app-server gives input and cached input, and no writes.
  • Costs are list price × recorded tokens, a calculation. Every write in this run was a 1-hour write.

The Codex sessions ran on a larger ledger (16,197 characters against 9,651), so we show the two routes side by side, not as a like-for-like pair.

How much comes from the cache

  • Claude Sonnet 5.5 · Claude Code
  • Claude Opus 5.5 · Claude Code
  • GPT-6.1 Sol (medium) · Codex CLI*

* The note below the chart says what this route does not report.

5 turn in the sessions, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (medium) · Codex CLI. Claude Sonnet 5.5 · Claude Code: highest Turn 2 99% (n 3). Lowest Turn 1 19% (n 3). Claude Opus 5.5 · Claude Code: highest Turn 2 99% (n 3). Lowest Turn 1 19% (n 3).

Notesn = 3 per row

Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each

Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Share of input read from the cache, mean over 3 sessions:

TurnSonnet 5.5Opus 5.5GPT-6.1 Sol (Codex CLI)
119%19%55%
299%99%99%
399%99%99%
499%99%99%
591%90%98%

Turn 1 already hits the cache for a part of its input: about 1,463 tokens on both Claude models. That is the part every Claude Code call starts with. The ledger itself, about 6,366 tokens, is written to the cache on turn 1 and read back on every later turn. Turn 5 dips because it wrote about 830 to 870 new tokens to the cache.

What that saves at list price

Calculation
  • With the cache, as recorded
  • Without a cache: every input token at the input price (square)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code

Gap labels, Without a cache: every input token at the input price vs With the cache, as recorded: Without a cache: every input token at the input price is x% higher (+) or lower (−) than With the cache, as recorded, calculated from the two values shown (the change counted from With the cache, as recorded’s value).

List-price calculation, not a run. 2 rows, 2 series: With the cache, as recorded, Without a cache: every input token at the input price. With the cache, as recorded: highest Claude Opus 5.5 · Claude Code $0.26 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.14 (n 15). Without a cache: every input token at the input price: highest Claude Opus 5.5 · Claude Code $0.54 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.27 (n 15).

Notesn = 15 per row

All recorded turns per model; the same reported tokens priced two ways

Calculation, not a bill: the calls ran on a subscription. Cache reads at the cache-read price, 1-hour cache writes at 2× the input price (every write in this run was a 1-hour write). Codex CLI is not priced here: it reports no cache-write count.

Sources: Caching sessions and repeated prompts (Claude Code and Codex CLI), Cost with and without the prompt cache (calculation), Anthropic list prices (Claude models)

The chart prices the same recorded tokens two ways: as recorded, and with every input token at the full input price.

  • Sonnet 5.5: $0.1350 with the cache, $0.2698 without: 50% less.
  • Opus 5.5: $0.2551 with the cache, $0.5442 without: 53% less.

Both are totals over 15 turns (3 sessions × 5 turns). Opus saves a little more because its input costs twice Sonnet's, while its cache-read price is the same.

Turn 1 costs more

The cache does not save on the first turn. It costs more:

Turn 1, mean per sessionWith the cacheWithoutOur ratio
Sonnet 5.5$0.0258$0.01571.6x
Opus 5.5$0.0513$0.03141.6x

The reason is the write price. A 1-hour cache write costs twice the input price ($4 per million tokens on Sonnet, $8 on Opus). So you pay extra to store the ledger, and you get it back only when a later turn reads it.

When the write pays back

Our arithmetic on the per-turn means in the study table:

  • Sonnet: turn 1 cost $0.0101 more with the cache. Turn 2 alone saved $0.0140. So the cache paid back on turn 2.
  • Opus: turn 1 cost $0.0199 more. Turn 2 saved $0.0295. Same result.
  • Over turns 2 to 5, the cache cut the cost by 74% on Sonnet and 78% on Opus.

A one-shot call with a large, unique context is the bad case: you pay the write and never read it. A loop that re-sends the same context is the good case.

The agent-scale picture

Calculation
  • With caching (as recorded)
  • Without caching (square)
In chart order.
Claude Haiku 4.5
Claude Sonnet 5.5
Claude Opus 5.5
Claude Fable 5.1

Gap labels, Without caching vs With caching (as recorded): Without caching is x% higher (+) or lower (−) than With caching (as recorded), calculated from the two values shown (the change counted from With caching (as recorded)’s value).

List-price calculation, not a run. 4 rows, 2 series: With caching (as recorded), Without caching. With caching (as recorded): highest Claude Fable 5.1 $321. Lowest Claude Haiku 4.5 $43.61. Without caching: highest Claude Fable 5.1 $1,717. Lowest Claude Haiku 4.5 $172.

Notes

The same recorded tokens with and without cache pricing

94.0% of recorded input tokens were cache reads. Calculation, not a run.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)

The 5-turn sessions are small. Our agent's recorded SWE-bench Verified run is a real long workload: 162.9M input tokens, 94.0% of them cache reads, and 1.8M output tokens across 33 attempts.

Priced at list price, a calculation, not a run:

  • Sonnet 5.5: $87.23 with caching, $343.33 without: about 3.9x (our ratio), or 75% less.
  • Opus 5.5: $143.83 with caching, $686.66 without.

The savings are larger here than in our sessions for a simple reason: an agent loop re-reads its growing context on every model call, many more than 5 times. The more turns re-send the same prefix, the closer the cost falls to the cache-read price.

No reuse across sessions

You might expect a second session with the same ledger to read it from the first session's cache. It did not. On turn 1, all 4 later sessions wrote the ledger to the cache again (stat: 0 of 4 reused).

We did not test the cause. The CLI may add per-process context before the user message, which would change the prefix. Do not read this as a provider property. Read it as: in this CLI version, plan for one cache write per session.

Does the cache make calls faster?

  • Turn 1 (writes the ledger to the cache)
  • Turns 2-5 (read the ledger from the cache) (square)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code

Gap labels, Turns 2-5 (read the ledger from the cache) vs Turn 1 (writes the ledger to the cache): Turns 2-5 (read the ledger from the cache) is x% higher (+) or lower (−) than Turn 1 (writes the ledger to the cache), calculated from the two values shown (the change counted from Turn 1 (writes the ledger to the cache)’s value); lines are the fastest–slowest run (not an interval).

2 rows, 2 series: Turn 1 (writes the ledger to the cache), Turns 2-5 (read the ledger from the cache). Turn 1 (writes the ledger to the cache): slowest Claude Opus 5.5 · Claude Code 1.9 s (range 1.8 s–4.4 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1.6 s–1.8 s, n 3). All run ranges overlap. Turns 2-5 (read the ledger from the cache): slowest Claude Opus 5.5 · Claude Code 2.4 s (range 1.6 s–12.7 s, n 12). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1.4 s–5.6 s, n 12). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 3–12 per row

Median; whiskers = fastest and slowest turn

Whiskers are a range (fastest and slowest turn), not a confidence interval. Turns ask different questions: the slow later turns are the counting question, which produced the most output.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Not clearly in our data:

  • Sonnet 5.5: turn 1 median 1.64 s (1.58 to 1.79 s, n = 3), turns 2 to 5 median 1.61 s (1.35 to 5.63 s, n = 12).
  • Opus 5.5: turn 1 median 1.90 s (1.78 to 4.36 s), turns 2 to 5 median 2.40 s (1.63 to 12.67 s).

The ranges overlap, and the context is small. Caching may help latency on much larger contexts. We did not see it here, so we do not claim it.

Codex: high cache share, no price

The Codex CLI read 99% of later-turn input from its cache on average, and 55% on turn 1. Its app-server reports cached input but no cache writes, so we cannot price the write side, and we do not calculate a saving for it.

How to use this

  1. Keep the prefix stable. Put the large, fixed part of the prompt first, and the part that changes last.
  2. Count the write. A 1-hour write costs 2x input. If a context is used once, caching it costs more.
  3. Expect one write per session. In our runs, a new CLI session did not reuse an earlier session's cache.
  4. Check the cache-read price, not only the input price. It decides the cost of long agent loops. For example, Opus 5.5 lists at twice Sonnet's input price but the same cache-read price.
  5. Run your own mix in the AI cost calculator, with and without caching.

Caveats

  • 3 sessions per model. The figures describe this CLI version and this context size.
  • Calculations. All costs are list price × recorded tokens. The calls ran on flat subscriptions.
  • Different routes. The Codex sessions used a larger ledger; compare the routes side by side only.
  • Repricing is not a run. At another model's price, the SWE-bench agent would have made different calls.

See your own cache share

Agent records cache reads and writes for every model call, so you can see how much of your bill the cache already saves. Try Agent and check your own numbers.

The data behind this post

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.