{"i":20,"study":{"slug":"prompt-cache-break-even","title":"Prompt cache break-even: after how many reuses does a cached prefix cost less?","seoTitle":"Prompt cache break-even by model and session length","description":"A calculation on list prices: reuses before a cached prompt prefix costs less, per Claude and GPT model, and cost per 1,000 sessions of 1 to 20 turns.","question":"With the cache-write surcharge Anthropic lists, after how many reuses does a cached prompt prefix cost less than no cache? What does that mean for a session of 1 to 20 turns, for each model’s price, and for a workload split across sessions?","answer":"Calculation: a wholly new prefix with a 1-hour write costs less from the 3rd request, after 2 reuses. All four Claude price rows list a 1-hour write at twice the input price. Exact break-even is 1.03 to 1.11 reuses. With 19% already cached, it needs 1 reuse. This uses a pooled share from n = 6 sessions; session range 18.680% to 18.689%. A 5-minute write (an assumed 1.25 times the input price) needs 1 reuse. GPT-6.1 Sol and GPT-6 Luna list no write surcharge in the price list, so they save from the first reuse. They still do at the 1.25× write that another product table lists. For 1,000 10-turn sessions, the Sonnet 1-hour prefix cost is $45.42 against $156.62 without caching: 71.0% less. The prefix is a rounded mean of 7,831 tokens (n = 3; range 7,831 to 7,832). A one-turn session costs 2.0 times as much with the cache ($31.32 against $15.66 per 1,000). Each recorded session ran in a new temporary folder. Of 4 later sessions, 0 met the reuse proxy (95% Wilson 0% to 49%). The split calculation assumes ten full writes; it is not a replay of those partly cached sessions. The Opus 5.5 cache-read price is open: $0.2 or $0.4 per million. At $0.4, the break-even moves from 1.05 to 1.11 reuses, and the 10-turn saving from 75.5% to 71.0%.","date":"2026-10-06","updated":"2026-10-06","tags":["prompt-caching","break-even","thought-experiment","calculation","llm-pricing","claude-haiku","claude-sonnet","claude-opus","claude-fable","gpt-6-1-sol","gpt-6-luna"],"caveats":["The retained Claude protocol file was created at 14:43:55 UTC, after the first counted session at 14:35:39 UTC on 2026-10-06. Its declaration says 14:25 UTC, but file times do not verify that claim. Treat this as a retrospective protocol record.","The source has 30 attempted Claude turns, 0 failed turns and 0 turns without token usage. Failed turns with usage remain in cost totals. Missing usage cannot be priced. No quality rate or cache-caused speed effect is claimed.","7 counted turns record a provider rate-limit warning status. The retained outcome log says there was no warning. The protocol calls for a stop on limit text in an error or warning. These receipts do not establish whether that text appeared; this discrepancy remains unresolved.","One shared Mac and one synthetic ledger supplied these tokens. The ledger was shortened once after a two-call probe and before the counted sessions. All 30 counted answers passed, so this task set hits a quality ceiling. This calculation does not rank model quality or speed.","A calculation, not a run and not a bill. The recorded calls used flat subscriptions. The prices are list prices dated 2026-09-21 (Anthropic) and 2026-10-03 (OpenAI); they can change.","The formula prices the shared prefix only. A real turn also writes new tokens to the cache at the write price. In the recorded sessions that was 58 to 1,117 tokens per turn after turn 1. A real turn also produces output, which costs the same with or without the cache. The check against the recorded sessions shows the size of this effect.","The formula assumes that every reuse arrives before the cache entry expires. An entry lasts 5 minutes or 1 hour, by write type. When the gap between requests is longer, the write repeats and the saving shrinks.","The recorded sessions wrote 1-hour entries only (0 of 30 turns wrote a 5-minute entry). The 5-minute rows use an assumed 1.25 multiplier that this study did not test.","Only 0 of 4 later sessions met the read-more-than-write proxy (95% Wilson 0% to 49%). They still read some cached tokens. This proxy does not establish cache provenance. Each recorded session ran in a fresh process and a new temporary working folder. The caching study protocol names that folder as one untested hypothesis for the miss, and this study did not test it. The split case assumes no shared reuse and a wholly new prefix. It does not price the partly cached recorded first turns. A script that keeps one working folder may read an earlier session’s cache: check your own receipts before you use the split case.","The 19% already-cached share is not a constant. We recorded it on Sonnet 5.5 and Opus 5.5 sessions. For Haiku 4.5 and Fable 5.1 it is a what-if. The one Haiku 4.5 probe session (outside every cell) read 0 tokens from the cache on turn 1.","The Opus 5.5 cache-read price is open: $0.2 or $0.4 per million. This study shows both. Resolve the price before you quote an Opus figure.","The dollar figures depend on T: one synthetic ledger in one CLI version, 3 sessions per model. A different prefix changes the dollars. It does not change the break-even reuses, which depend on the price ratios only.","The GPT rows follow the price list, which has no write surcharge. Another table of the product lists a write price of 1.25× input for both models. At that price they need 1 reuse (exact break-even 0.26 and 0.28), and a prefix used once costs 1.25 times as much. We did not check the vendor price. The Codex app-server reports no cache writes, and this study did not test reuse for GPT. The caching study measured Codex reads on a larger context."],"sourceIds":["calc-cache-pricing","agent-caching-consistency","price-anthropic","price-openai"],"stats":{"$k":["id","label","value","unit","display","note","n"],"$r":[["cache-break-even-1h-reuses","Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation)",2,"count","2 reuses (the 3rd request)","Same for Haiku 4.5, Sonnet 5.5, Opus 5.5 and Fable 5.1. Exact break-even 1.03 to 1.11 reuses.","\u0001"],["cache-break-even-1h-reuses-recorded","Reuses before a 1-hour cached prefix costs less, 19% of it already cached as a pooled recorded share (calculation)",1,"count","1 reuse (the 2nd request)","Exact break-even 0.65 to 0.72 reuses. The recorded Claude Code sessions paid back on turn 2 (see the recorded-payback stat).","\u0001"],["cache-break-even-5m-reuses","Reuses before a 5-minute cached prefix costs less, whole prefix new (calculation, assumed 1.25× write)",1,"count","1 reuse (the 2nd request)","Exact break-even 0.26 to 0.28 reuses. The 1.25 multiplier is an assumption.","\u0001"],["cache-break-even-no-surcharge-reuses","Reuses before a cached prefix costs less when the list price has no write surcharge (GPT-6.1 Sol and GPT-6 Luna; calculation)",1,"count","1 reuse (saves from the first read)","Exact break-even 0 reuses: a write is plain input, so the first read is already cheaper than sending the prefix again. Another table of the product lists a write price of 1.25× input for both models. At that price the exact break-even is 0.26 reuses (GPT-6.1 Sol) and 0.28 reuses (GPT-6 Luna), so 1 reuse still pays. We did not check the vendor price.","\u0001"],["cache-break-even-prefix-sonnet","Mean turn-1 input in the Sonnet 5.5 sessions (calculation)",7831,"tokens","7,831 tokens","Rounded mean calculation from 3 sessions (range 7,831 to 7,832): uncached input + cache reads + cache writes on turn 1.",3],["cache-break-even-prefix-opus","Mean turn-1 input in the Opus 5.5 sessions (calculation)",7828,"tokens","7,828 tokens","Rounded mean calculation from 3 sessions (the same in every session).",3],["cache-break-even-precached-share","Share of turn-1 input read from the cache (calculation; origin not isolated)",0.1869,"rate","19% (8,778 of 46,978 turn-1 tokens)","Calculation across 6 sessions; session share range 18.680% to 18.689%. This is not a binomial pass rate. This part cost the read price on turn 1, not the write price.",6],["cache-break-even-recorded-payback","Turn at which a recorded Claude Code session’s total input cost with the cache first fell below its cost with no cache (calculation on recorded tokens)",2,"count","2 turns (6 of 6 sessions)","Input side only, 1-hour writes at the list price, output left out. After turn 1 the cache had cost more, as the formula says for a prefix used once.",6],["cache-break-even-once-penalty","Cost of caching a prefix that is used once, as a multiple of no cache (1-hour write; calculation)",2,"ratio","2.0x","A 5-minute write: 1.3x (assumption). Same for every Anthropic model in the price list.","\u0001"],["cache-break-even-sonnet-10-turns","Saving from a 1-hour cache over 10 turns, Sonnet 5.5 prefix (calculation)",0.71,"rate","71.0% ($45.42 vs $156.62 per 1,000 sessions)","Prefix of 7,831 tokens; whole prefix new; output left out.","\u0001"],["cache-break-even-opus-10-turns","Saving from a 1-hour cache over 10 turns, Opus 5.5 prefix, cache read $0.2 per M (calculation)",0.755,"rate","75.5% ($76.71 vs $313.12 per 1,000 sessions)","Prefix of 7,828 tokens; whole prefix new; output left out.","\u0001"],["cache-break-even-opus-read-price","Saving from a 1-hour cache over 10 turns, Opus 5.5 prefix, cache read $0.4 per M (calculation)",0.71,"rate","71.0% ($90.80 vs $313.12 per 1,000 sessions)","Break-even 1.11 reuses at $0.4 per M against 1.05 at $0.2 per M. The vendor price was not checked.","\u0001"],["cache-break-even-split-sonnet","10 one-turn sessions with no shared reuse with a 1-hour cache, as a multiple of one 10-turn session, Sonnet 5.5 prefix (calculation)",6.9,"ratio","6.9x ($313.24 vs $45.42 per 1,000 workloads)","With no cache the same 10 requests cost $156.62. Each recorded session ran in a new temporary folder, and 0 of 4 met the read-more-than-write proxy. The split case assumes a wholly new prefix with no shared reuse.","\u0001"],["cache-break-even-cross-session","Later sessions whose first turn read more tokens than it wrote (reuse proxy)",0,"count","0 of 4","95% Wilson 0% to 49%, n = 4 later sessions. Same proxy as the caching study: reads exceed writes. It cannot identify cache provenance. The cause was not tested.",4],["cache-break-even-5m-writes-recorded","Recorded Claude turns that wrote a 5-minute cache entry",0,"count","0 of 30","Every recorded write was a 1-hour write, so the 1.25 multiplier of the 5-minute rows was not tested.",30],["cache-break-even-check-sonnet","Saving over 5 turns on the input side, recorded vs the formula (calculation), Sonnet 5.5",0.5535,"rate","55.4% recorded (formula 52.0% with a new prefix, 59.1% with 19% already cached)","Calculation on recorded tokens and list prices; session saving range 54.3% to 57.1%, not a confidence interval. The recorded figure includes new turn tokens; the formula does not.",3],["cache-break-even-check-opus","Saving over 5 turns on the input side, recorded vs the formula (calculation), Opus 5.5",0.5918,"rate","59.2% recorded (formula 56.0% with a new prefix, 63.3% with 19% already cached)","Calculation on recorded tokens and list prices; session saving range 58.5% to 59.7%, not a confidence interval. Opus 5.5 cache read at $0.2 per M.",3]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds","xLabel"],"$r":[["cache-break-even-reads","Reuses before a cached prefix costs less, by model and write type (calculation)","The break-even point: reuses at which the cached and the uncached cost are equal, at list prices","grouped-bar","score","Reuses at break-even",{"$k":["name","points"],"$r":[["1-hour write (2× input), whole prefix new",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",1.11],["Claude Sonnet 5.5",1.11],["Claude Opus 5.5 (cache read $0.2 per M)",1.05],["Claude Opus 5.5 (cache read $0.4 per M)",1.11],["Claude Fable 5.1",1.03]]}],["1-hour write, pooled n = 6 session share, 19% already cached (as recorded)",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",0.72],["Claude Sonnet 5.5",0.72],["Claude Opus 5.5 (cache read $0.2 per M)",0.67],["Claude Opus 5.5 (cache read $0.4 per M)",0.72],["Claude Fable 5.1",0.65]]}],["5-minute write (1.25× input, an assumption)",{"$k":["label","value"],"$r":[["Claude Haiku 4.5",0.28],["Claude Sonnet 5.5",0.28],["Claude Opus 5.5 (cache read $0.2 per M)",0.26],["Claude Opus 5.5 (cache read $0.4 per M)",0.28],["Claude Fable 5.1",0.26]]}]]},"Calculation, not a run: break-even reuses = (write price − input price) ÷ (input price − cache-read price). A cached prefix costs less once the reuses pass that point. On a new prefix, a 1-hour write needs 2 reuses (the 3rd request). With 19% already cached it needs 1 reuse, and a 5-minute write needs 1 reuse. The 1.25 of the 5-minute write is an assumption. The caching study protocol states it as Anthropic’s published figure. No 5-minute write occurred in the recorded sessions. We recorded the 19% on Sonnet 5.5 and Opus 5.5 sessions. For Haiku 4.5 and Fable 5.1 it is a what-if. GPT-6.1 Sol and GPT-6 Luna list no write surcharge, so their break-even is 0 reuses and the chart leaves them out. The price list gives Opus 5.5 a cache read of $0.2 per million. Another table of the product lists $0.4. We did not check the vendor price, so the chart shows both.",["calc-cache-pricing","price-anthropic","price-openai"],"\u0001"],["cache-break-even-cost-curve","Cost of a reused prefix with and without the cache, by session length (calculation)","Prefix of 7,831 tokens (Sonnet 5.5) and 7,828 tokens (Opus 5.5), the rounded mean turn-1 input; n = 3 Sonnet sessions (range 7,831 to 7,832) and 3 Opus sessions (the same in every session); USD per 1,000 sessions","line","usd","USD per 1,000 sessions (list price)",{"$k":["name","points"],"$r":[["Claude Sonnet 5.5 (no cache)",{"$k":["label","value"],"$r":[["1 turn",15.66],["2 turns",31.32],["3 turns",46.99],["5 turns",78.31],["10 turns",156.62],["20 turns",313.24]]}],["Claude Sonnet 5.5 (1-hour cache write)",{"$k":["label","value"],"$r":[["1 turn",31.32],["2 turns",32.89],["3 turns",34.46],["5 turns",37.59],["10 turns",45.42],["20 turns",61.08]]}],["Claude Opus 5.5 (no cache)",{"$k":["label","value"],"$r":[["1 turn",31.31],["2 turns",62.62],["3 turns",93.94],["5 turns",156.56],["10 turns",313.12],["20 turns",626.24]]}],["Claude Opus 5.5 (1-hour cache write, read $0.2 per M)",{"$k":["label","value"],"$r":[["1 turn",62.62],["2 turns",64.19],["3 turns",65.76],["5 turns",68.89],["10 turns",76.71],["20 turns",92.37]]}],["Claude Opus 5.5 (1-hour cache write, read $0.4 per M)",{"$k":["label","value"],"$r":[["1 turn",62.62],["2 turns",65.76],["3 turns",68.89],["5 turns",75.15],["10 turns",90.8],["20 turns",122.12]]}]]},"Calculation, not a bill. The chart multiplies list prices by a prefix of 7,831 tokens (Sonnet) or 7,828 tokens (Opus). It uses a 1-hour write at 2× input and treats the whole prefix as new. It leaves out output and the tokens that each turn adds. The cache costs more at 1 to 2 turns and less from 3. The second Opus line uses a $0.4 cache read, which another table of the product lists. We did not check the vendor price. The table adds the 5-minute write.",["calc-cache-pricing","agent-caching-consistency","price-anthropic","price-openai"],"Turns in the session (requests that send the prefix)"],["cache-break-even-session-split","One 10-turn session or ten 1-turn sessions: cost with and without the cache (calculation)","USD per 1,000 workloads of 10 requests that send the same prefix; each recorded session started in a new temporary folder, n = 4 later sessions; 0 met the read-more-than-write proxy, 95% Wilson 0% to 49%","grouped-bar","usd","USD per 1,000 workloads of 10 requests (list price)",{"$k":["name","points"],"$r":[["No cache (the same either way)",{"$k":["label","value"],"$r":[["Claude Sonnet 5.5",156.62],["Claude Opus 5.5 (cache read $0.2 per M)",313.12],["Claude Opus 5.5 (cache read $0.4 per M)",313.12]]}],["One 10-turn session, 1-hour cache",{"$k":["label","value"],"$r":[["Claude Sonnet 5.5",45.42],["Claude Opus 5.5 (cache read $0.2 per M)",76.71],["Claude Opus 5.5 (cache read $0.4 per M)",90.8]]}],["Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix)",{"$k":["label","value"],"$r":[["Claude Sonnet 5.5",313.24],["Claude Opus 5.5 (cache read $0.2 per M)",626.24],["Claude Opus 5.5 (cache read $0.4 per M)",626.24]]}]]},"Calculation, not a bill: the same prefix, prices and 1-hour write as the cost-curve chart. Each recorded session ran in a new temporary working folder. On turn 1, 0 of 4 later sessions read more tokens than they wrote. This proxy cannot identify which earlier call supplied a cache entry. The split calculation assumes no shared reuse and a wholly new prefix, so it charges ten full writes. With reuse across sessions they would cost the same as one 10-turn session. The cause of the missing reuse was not tested here. The caching study protocol names the random temporary folder as one untested hypothesis. The table adds the 5-minute variants.",["calc-cache-pricing","agent-caching-consistency","price-anthropic","price-openai"],"\u0001"]]},"related":["caching-consistency","cost-thought-experiments"]}}