When does prompt caching pay off? The Anthropic cache break-even, calculated
A 1-hour Anthropic cache write pays back after 2 reuses; a 5-minute write (assumed 1.25x) after 1. Break-even by model, cost per 1,000 sessions.
Prompt caching and consistency: what the cache saves, and how much answers vary
Calculation at list price: the cache cut a 5-question session 50% on Sonnet 5.5 and 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every repetition.
Transcript
- Caching and consistency · 135 calls. What the cache saves, and how much answers vary. Five-question sessions on a fixed context. Then the same prompt, 10 times.
- 135 calls: 45 cache turns, 90 repeated prompts. Every call counted. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- Claude Code turn 1 reads 19% from the cache, the CLI’s own prefix. Turns 2–5 read 90%–99%. Codex CLI: 98%–99%. Chart: Share of input read from the cache · mean of 3 sessions per turn (n = 3 each). Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- A calculation: Sonnet 5.5 $0.1350 with the cache vs $0.2698 without, 50% less. Opus 5.5 $0.2551 vs $0.5442, 53% less. Chart: List-price cost of the recorded sessions · calculation (n = 15 each). Calculation, not a run. Caveat: Costs are list-price calculations; the calls used flat subscriptions.
- No clear speed effect: Sonnet 5.5 1.6 s on turn 1 vs 1.6 s later; Opus 5.5 1.9 s vs 2.4 s. The ranges overlap. Sonnet 5.5: median turn 1 vs turns 2–5: 1.6 s vs 1.6 s (ranges 1.6–1.8 s and 1.4–5.6 s · n = 3 and 12). Opus 5.5: median turn 1 vs turns 2–5: 1.9 s vs 2.4 s (ranges 1.8–4.4 s and 1.6–12.7 s · n = 3 and 12). Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- 7 of 9 cells passed 10/10. Haiku 4.5: 0/10 on the exact number (10 wrong, 1 distinct answer), 1/10 on JSON (9 format misses). Chart: Same prompt, 10 times · strict passes · a 10/10 is 72%–100% at 95% (n = 10 each). Caveat: Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
- Consistent is not correct: Haiku 4.5 gave the same wrong number all 10 times. Code fix, distinct correct bodies: Haiku 4.5 6, Sonnet 5.5 3, GPT-6.1 Sol (medium) 6. Chart: Same prompt, 10 times · distinct answers (n = 10 each). Caveat: The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
- Cache the fixed context; check answers, not agreement. Every call online.
TL;DR
- This is a calculation, not a run. We made no new model call. We took the tokens of our recorded caching sessions and the list prices. Then we worked out when a cached prompt prefix costs less than sending it again.
- A 1-hour cache write pays back after 2 reuses, which is the 3rd request that sends the prefix. The answer is the same on Claude Haiku 4.5, Sonnet 5.5, Opus 5.5 and Fable 5.1. Each lists a 1-hour write at twice its input price. The exact break-even is 1.03 to 1.11 reuses.
- A 5-minute write needs 1 reuse. That row uses an assumed write price of 1.25 times input. No 5-minute write occurred in our recorded sessions, so we did not test the assumption.
- Used once, the cache costs 2.0 times as much as no cache (1-hour write). At 2 turns it still costs more (5% more on Sonnet 5.5). At 10 turns it costs 71.0% less: $45.42 against $156.62 per 1,000 sessions on our Sonnet 5.5 prefix.
- Splitting a job can hurt (calculation). Ten full prefix writes cost 6.9 times one write plus nine reads on Sonnet 5.5. This assumes no shared reuse; it does not replay the partly cached recorded sessions.
- GPT-6.1 Sol and GPT-6 Luna list no write surcharge in our price list, so they save from the first reuse.
- One price is open. The Opus 5.5 cache read is $0.2 per million in one product table and $0.4 in another. The 10-turn saving is 75.5% or 71.0%. We did not check the vendor price.
The study with every input: /benchmarks/prompt-cache-break-even.
The question
Prompt caching is not free to switch on. The first request that sends a prefix writes it to the cache, and Anthropic lists that write at more than the normal input price. Later requests read the prefix at a fraction of the input price. So a prefix that you send once costs more with the cache. A prefix that you send often costs less.
Where is the crossing point? For the basics, see Prompt caching, explained. For the measured saving in our sessions, see How much does prompt caching actually save?. This post asks the next question: how many reuses do you need, and what does that mean for your session length and your model?
A prefix is the start of a prompt that stays the same between requests, such as a system prompt, tool definitions or a document. A reuse is a later request that sends the same prefix. A session of N turns is N requests, so it has N − 1 reuses.
The rule in one line
Prices are USD per million tokens: input i, cache read r and cache write w. Divide the token-price products below by 1,000,000 to get USD. For a prefix of T tokens sent in k + 1 requests:
- No cache: (k + 1) × T × i
- 1-hour write: T × w + k × T × r, where w is the listed 1-hour write price (2 × i for every Anthropic model in our price list)
- 5-minute write: T × 1.25 × i + k × T × r
The cached prefix costs less when k > (w − i) ÷ (i − r).
Take Sonnet 5.5: input $2, cache read $0.2, 1-hour write $4. The break-even is (4 − 2) ÷ (2 − 0.2) = 1.11 reuses. One reuse is not enough, so you need two. Per million prefix tokens (our arithmetic):
| Reuses | No cache | 1-hour cache | Cheaper |
|---|---|---|---|
| 0 (used once) | $2.00 | $4.00 | no cache |
| 1 | $4.00 | $4.20 | no cache |
| 2 | $6.00 | $4.40 | cache |
The 1.25 is the multiplier that our caching study protocol states as Anthropic's published figure. It is not in our price list, so the 5-minute rows are an assumption.
Break-even by model
Every Anthropic model in our price list has a 1-hour write at twice its input price. The break-even then depends only on how small the read price is next to the input price. That is why the models agree:
| Model | Input | Cache read | 1-hour write | Exact break-even | Reuses needed | 5-minute write (assumed): exact | Needed |
|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | $1 | $0.1 | $2 | 1.11 | 2 | 0.28 | 1 |
| Claude Sonnet 5.5 | $2 | $0.2 | $4 | 1.11 | 2 | 0.28 | 1 |
| Claude Opus 5.5 (read $0.2) | $4 | $0.2 | $8 | 1.05 | 2 | 0.26 | 1 |
| Claude Opus 5.5 (read $0.4) | $4 | $0.4 | $8 | 1.11 | 2 | 0.28 | 1 |
| Claude Fable 5.1 | $10 | $0.25 | $20 | 1.03 | 2 | 0.26 | 1 |
| GPT-6.1 Sol | $2 | $0.1 | $2 (no surcharge) | 0 | 1 | 0 | 1 |
| GPT-6 Luna | $0.1 | $0.01 | $0.1 (no surcharge) | 0 | 1 | 0 | 1 |
Calculation from list prices dated 2026-09-21 (Anthropic) and 2026-10-03 (OpenAI). "Reuses needed" is the smallest whole number of reuses at which the cached prefix costs less. The model changes the dollars, not the number of reuses.
Part of the prefix may already be cached. As a calculation on recorded Claude Code counters, turn 1 read 19% of its input from the cache (8,778 of 46,978 tokens, 6 sessions). This share ranged from 18.680% to 18.689% per session (a range, not a confidence interval). The counters do not identify its origin. It costs the read price, not the write price. With that share already cached, the 1-hour break-even drops to 0.65 to 0.72 reuses, so 1 reuse is enough. That fits the recorded sessions: in 6 of 6 (95% Wilson interval 61% to 100%), the calculated total input cost with the cache fell below the cost without it on turn 2.
What it means per session length
We priced the recorded turn-1 prefix: 7,831 tokens, the rounded mean in the Sonnet 5.5 sessions (n = 3; range 7,831 to 7,832). Opus 5.5 had 7,828 tokens in every session (n = 3). The chart shows 1,000 sessions at each length with a 1-hour write and the whole prefix new:
| Turns | No cache | 1-hour cache | Saving |
|---|---|---|---|
| 1 | $15.66 | $31.32 | the cache costs 2.0 times as much |
| 2 | $31.32 | $32.89 | the cache costs 5% more |
| 3 | $46.99 | $34.46 | 26.7% |
| 5 | $78.31 | $37.59 | 52.0% |
| 10 | $156.62 | $45.42 | 71.0% |
| 20 | $313.24 | $61.08 | 80.5% |
USD per 1,000 sessions on the Sonnet 5.5 prefix. Output and the tokens each turn adds are left out.
Opus 5.5 costs about twice as much per token, so its no-cache dollars roughly double. Its cache-read price changes the cached cost and saving. For 1,000 ten-turn sessions on the 7,828-token Opus 5.5 prefix:
| Opus 5.5 cache read | No cache | 1-hour cache | Saving |
|---|---|---|---|
| $0.2 per million | $313.12 | $76.71 | 75.5% |
| $0.4 per million | $313.12 | $90.80 | 71.0% |
At 20 turns, the same 1,000 Opus sessions cost $626.24 with no cache and $92.37 or $122.12 with the cache, by the read price.
The two ends are plain. A one-turn session loses money with the cache. A long session saves most of the prefix cost in this calculation: at 20 turns the saving passes 80% on every Claude row we priced.
The trap: one job split across sessions in new folders
The saving needs the reads. Each earlier session ran in a fresh process and a new temporary folder. Of 4 later sessions, 0 read more tokens than they wrote on turn 1 (95% Wilson interval 0% to 49%). This is a reuse proxy. They still read some cached tokens. The counters cannot identify cache provenance or the cause of a miss.
In a follow-up, 4 of 4 later fixed-folder sessions met a different proxy: at least half of turn-1 input was cached. This pools two setups after the run (95% Wilson interval 51% to 100%). The new-folder setup had 0 of 2 (0% to 66%). These intervals overlap. Folder, ledger seed and call order changed together, so this does not isolate a folder effect. All setups used Sonnet 5.5 in Claude Code, with 3 sessions per setup. Does Claude Code reuse the prompt cache across sessions? has the detail.
We priced the same 10 requests two ways, assuming a wholly new prefix and no reuse between sessions (calculation):
| 10 requests as | Sonnet 5.5 | Opus 5.5 (read $0.2) |
|---|---|---|
| No cache | $156.62 | $313.12 |
| One 10-turn session, 1-hour cache | $45.42 | $76.71 |
| Ten 1-turn sessions in new folders, 1-hour cache | $313.24 | $626.24 |
USD per 1,000 workloads of 10 requests. Ten one-turn sessions in new folders cost 6.9 times one ten-turn session on Sonnet 5.5 and 8.2 times on Opus 5.5 (read $0.2). In this calculation they cost twice the no-cache price. Each session pays a full write and collects no read. The recorded first turns were partly cached, so this is not their measured cost. With the assumed 5-minute write the same split costs $195.78 on Sonnet 5.5 against $156.62 with no cache, which is 25% more (calculation). If every later request read the whole prefix, the ten sessions would cost the same as one ten-turn session (calculation).
Does the formula match the recorded sessions?
We checked the formula against the five-turn Claude Code sessions we recorded. On the input side only, the list-price calculation saved 55.4% on Sonnet 5.5 (n = 3; session range 54.3% to 57.1%). Opus 5.5 saved 59.2% at a $0.2 read price (n = 3; range 58.5% to 59.7%). These are ranges, not confidence intervals. The formula gives 52.0% and 56.0% with a new prefix. It gives 59.1% and 63.3% with the recorded cached share (8,778 ÷ 46,978, rounded to 19%). The recorded figure sits between the two formula cases for both models.
With output included, the recorded saving is 50.0% on Sonnet 5.5 (n = 3; range 47.5% to 54.1%) and 53.1% on Opus 5.5 (n = 3; range 51.4% to 54.3%). These list-price calculations use recorded tokens; the ranges are not confidence intervals:
The recorded sessions also wrote 58 to 1,117 new tokens per turn after turn 1, which the formula leaves out. Only the input-side savings, 55.4% and 59.2%, fall between the two formula cases. The savings with output fall below both cases.
What about GPT-6.1 Sol and GPT-6 Luna?
Our price list has no write surcharge for them, so a write costs plain input. The cache never costs more, and the first reuse already saves money. At 2 turns that is 47.5% on GPT-6.1 Sol and 45.0% on GPT-6 Luna; at 10 turns it is 85.5% and 81.0% (calculation).
One caution. Another table of our product lists a write price of 1.25 times input for both models. At that price the break-even is 0.26 reuses (Sol) and 0.28 (Luna). They still need 1 reuse, and a prefix used once costs 1.25 times as much. We did not check the vendor price. We did not price reuse on GPT; the caching study measured Codex CLI cache reads on a larger context and did not price them.
What to do with this
These are suggestions from the numbers. We have not tested all of them.
- Cache a prefix that you expect to send at least 3 times inside its cache lifetime (2 reuses). With a 5-minute write, at least 2 times. Do not cache a prompt that you send once: it costs 2.0 times as much with a 1-hour write.
- Test a fixed working folder when a script starts many short sessions. Compare the counters with those from new folders. The follow-up intervals overlap, and the test did not isolate a folder effect. Check your own receipts.
- Check the gap between requests. The formula assumes that every reuse arrives before the entry expires. If it does not, you pay the write again, and the saving shrinks.
- Read your own receipts. Look for cache-write and cache-read tokens in each response. Calculate the write surcharge and read saving to check whether the cache costs you money.
- Resolve the price you quote. For Opus 5.5, check the cache-read price with the vendor before you copy a number from any table, ours included.
- Run your own mix in the AI cost calculator.
How we calculated
- Inputs: the turn tokens that our caching study recorded in 6 Claude Code sessions (3 on Sonnet 5.5, 3 on Opus 5.5), and the list prices in our price sources. No new call, no new raw file.
- T: the rounded mean turn-1 input (uncached input + cache reads + cache writes) of each model's sessions.
- Dollar figures: 1,000 sessions of N turns = 1,000 × the formula with k = N − 1, the whole prefix new (the stricter case).
- Break-even: the exact fraction (w − i) ÷ (i − r), and the smallest whole number of reuses above it.
- Share already cached: use the unrounded recorded share f = 8,778 ÷ 46,978 (about 19%). A share f of the prefix is read, not written, on the first request: T × (f × r + (1 − f) × w) + k × T × r.
- Check: the formula against the recorded 5-turn sessions, input side only.
Caveats
- A calculation, not a run and not a bill. The recorded calls used flat subscriptions. List prices can change.
- The 1.25 is an assumption. None of the 30 recorded Claude turns wrote a 5-minute entry (95% Wilson interval 0% to 11%). The 5-minute rows are untested.
- The formula prices the shared prefix only. Real turns add tokens and produce output, which costs the same with or without the cache.
- Every reuse must arrive before expiry. Late reuses pay the write again.
- The split case assumes ten full writes and no shared reuse. It is not the cost of the partly cached recorded sessions. Neither reuse proxy identifies which tokens came from an earlier session.
- The source protocols are retrospective. The earlier protocol file was created at 14:43:55 UTC, after the first counted session at 14:35:39 UTC on 2026-10-06. The follow-up file was created at 00:49:27 UTC, after its first call at 00:32:33 UTC on 2026-10-07. File times do not support their pre-call declarations.
- Warning records disagree. Seven earlier Claude turns have a provider warning status, but the outcome log says there was no warning. The receipts do not resolve this. The follow-up has 9 warning statuses in 18 Claude turns and no refused call.
- The task set hits a ceiling. All 30 earlier Claude answers passed (95% Wilson interval 89% to 100%). These synthetic lookups do not rank model quality or prove a cache speed effect. The ledger was shortened after a two-call probe before counted sessions.
- One ledger, one CLI version, 3 sessions per model. A different prefix changes the dollars. It does not change the break-even reuses, which depend on the price ratios only.
- Haiku 4.5 and Fable 5.1 are not in the recorded cells. Their rows use the listed prices. The 19% already-cached share is a what-if for them.
- Two prices are open: the Opus 5.5 cache read ($0.2 or $0.4 per million) and the GPT write price (none, or 1.25 times input). We did not check the vendor.
What to read next
- Prompt caching, explained with measured sessions
- How much does prompt caching actually save?
- Does Claude Code reuse the prompt cache across sessions?
- What if every call ran on Opus?
- The caching and consistency study
See your own cache reads and writes
Agent records tokens, cache reads, cache writes and cost for every step of every task, so you can check a break-even against your own receipts. Try Agent.
Disclosure: I build Agent, the product behind these benchmarks. The recorded sessions in this post ran Claude Code on our own synthetic ledger, not on customer work.