• Prompt Caching
  • Break Even
  • Thought experiment
  • Calculation
  • LLM pricing
  • Claude Haiku
  • Claude Sonnet
  • Claude Opus
  • Claude Fable
  • GPT-6.1 Sol
  • Gpt 6 Luna

Prompt cache break-even: after how many reuses does a cached prefix cost less?

With the cache-write surcharge Anthropic lists, after how many reuses does a cached prompt prefix cost less than no cache? What does that mean for a session of 1 to 20 turns, for each model’s price, and for a workload split across sessions?

Published · 3 charts · Download the data or a carousel

2

Calculation

Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation) · reuses (the 3rd request)

Same for Haiku 4.5, Sonnet 5.5, Opus 5.5 and Fable 5.1. Exact break-even 1.03 to 1.11 reuses.

The answer

Calculation: a wholly new prefix with a 1-hour write costs less from the 3rd request, after 2 reuses. All four Claude price rows list a 1-hour write at twice the input price. Exact break-even is 1.03 to 1.11 reuses. With 19% already cached, it needs 1 reuse. This uses a pooled share from n = 6 sessions; session range 18.680% to 18.689%. A 5-minute write (an assumed 1.25 times the input price) needs 1 reuse. GPT-6.1 Sol and GPT-6 Luna list no write surcharge in the price list, so they save from the first reuse. They still do at the 1.25× write that another product table lists. For 1,000 10-turn sessions, the Sonnet 1-hour prefix cost is $45.42 against $156.62 without caching: 71.0% less. The prefix is a rounded mean of 7,831 tokens (n = 3; range 7,831 to 7,832). A one-turn session costs 2.0 times as much with the cache ($31.32 against $15.66 per 1,000). Each recorded session ran in a new temporary folder. Of 4 later sessions, 0 met the reuse proxy (95% Wilson 0% to 49%). The split calculation assumes ten full writes; it is not a replay of those partly cached sessions. The Opus 5.5 cache-read price is open: $0.2 or $0.4 per million. At $0.4, the break-even moves from 1.05 to 1.11 reuses, and the 10-turn saving from 75.5% to 71.0%.

Key numbers

1

Reuses before a 1-hour cached prefix costs less, 19% of it already cached as a pooled recorded share (calculation)

reuse (the 2nd request)

1

Reuses before a 5-minute cached prefix costs less, whole prefix new (calculation, assumed 1.25× write)

reuse (the 2nd request)

1

Reuses before a cached prefix costs less when the list price has no write surcharge (GPT-6.1 Sol and GPT-6 Luna; calculation)

reuse (saves from the first read)

7,831 tokens

Mean turn-1 input in the Sonnet 5.5 sessions (calculation)

n = 3

7,828 tokens

Mean turn-1 input in the Opus 5.5 sessions (calculation)

n = 3

19%

Share of turn-1 input read from the cache (calculation; origin not isolated)

(8,778 of 46,978 turn-1 tokens) · n = 6

2

Turn at which a recorded Claude Code session’s total input cost with the cache first fell below its cost with no cache (calculation on recorded tokens)

turns (6 of 6 sessions) · n = 6

2.0x

Cost of caching a prefix that is used once, as a multiple of no cache (1-hour write; calculation)

71.0%

Saving from a 1-hour cache over 10 turns, Sonnet 5.5 prefix (calculation)

($45.42 vs $156.62 per 1,000 sessions)

75.5%

Saving from a 1-hour cache over 10 turns, Opus 5.5 prefix, cache read $0.2 per M (calculation)

($76.71 vs $313.12 per 1,000 sessions)

71.0%

Saving from a 1-hour cache over 10 turns, Opus 5.5 prefix, cache read $0.4 per M (calculation)

($90.80 vs $313.12 per 1,000 sessions)

6.9x

10 one-turn sessions with no shared reuse with a 1-hour cache, as a multiple of one 10-turn session, Sonnet 5.5 prefix (calculation)

($313.24 vs $45.42 per 1,000 workloads)

0 of 4

Later sessions whose first turn read more tokens than it wrote (reuse proxy)

n = 4

0 of 30

Recorded Claude turns that wrote a 5-minute cache entry

n = 30

55.4%

Saving over 5 turns on the input side, recorded vs the formula (calculation), Sonnet 5.5

recorded (formula 52.0% with a new prefix, 59.1% with 19% already cached) · n = 3

59.2%

Saving over 5 turns on the input side, recorded vs the formula (calculation), Opus 5.5

recorded (formula 56.0% with a new prefix, 63.3% with 19% already cached) · n = 3

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Thought experiment: not a run. These values reprice recorded tokens at list prices. No model was called again.

Calculation

1-hour write (2× input), whole prefix new

Claude Haiku 4.5
Claude Sonnet 5.5
Claude Opus 5.5 (cache read $0.2 per M)
Claude Opus 5.5 (cache read $0.4 per M)
Claude Fable 5.1

1-hour write, pooled n = 6 session share, 19% already cached (as recorded)

Claude Haiku 4.5
Claude Sonnet 5.5
Claude Opus 5.5 (cache read $0.2 per M)
Claude Opus 5.5 (cache read $0.4 per M)
Claude Fable 5.1

5-minute write (1.25× input, an assumption)

Claude Haiku 4.5
Claude Sonnet 5.5
Claude Opus 5.5 (cache read $0.2 per M)
Claude Opus 5.5 (cache read $0.4 per M)
Claude Fable 5.1

One panel per series, all on the same axis.

List-price calculation, not a run. 5 rows, 3 series: 1-hour write (2× input), whole prefix new, 1-hour write, pooled n = 6 session share, 19% already cached (as recorded), 5-minute write (1.25× input, an assumption). 1-hour write (2× input), whole prefix new: highest Claude Haiku 4.5 1.11. Lowest Claude Fable 5.1 1.03. 1-hour write, pooled n = 6 session share, 19% already cached (as recorded): highest Claude Haiku 4.5 0.72. Lowest Claude Fable 5.1 0.65.

Notes

The break-even point: reuses at which the cached and the uncached cost are equal, at list prices

Calculation, not a run: break-even reuses = (write price − input price) ÷ (input price − cache-read price). A cached prefix costs less once the reuses pass that point. On a new prefix, a 1-hour write needs 2 reuses (the 3rd request). With 19% already cached it needs 1 reuse, and a 5-minute write needs 1 reuse. The 1.25 of the 5-minute write is an assumption. The caching study protocol states it as Anthropic’s published figure. No 5-minute write occurred in the recorded sessions. We recorded the 19% on Sonnet 5.5 and Opus 5.5 sessions. For Haiku 4.5 and Fable 5.1 it is a what-if. GPT-6.1 Sol and GPT-6 Luna list no write surcharge, so their break-even is 0 reuses and the chart leaves them out. The price list gives Opus 5.5 a cache read of $0.2 per million. Another table of the product lists $0.4. We did not check the vendor price, so the chart shows both.

Sources: Cost with and without the prompt cache (calculation), Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)
Calculation
  • Claude Sonnet 5.5 (no cache)
  • Claude Sonnet 5.5 (1-hour cache write)
  • Claude Opus 5.5 (no cache)
  • Claude Opus 5.5 (1-hour cache write, read $0.2 per M)
  • Claude Opus 5.5 (1-hour cache write, read $0.4 per M)

List-price calculation, not a run. 6 rows, 5 series: Claude Sonnet 5.5 (no cache), Claude Sonnet 5.5 (1-hour cache write), Claude Opus 5.5 (no cache), Claude Opus 5.5 (1-hour cache write, read $0.2 per M), Claude Opus 5.5 (1-hour cache write, read $0.4 per M). Claude Sonnet 5.5 (no cache): highest 20 turns $313. Lowest 1 turn $15.66. Claude Sonnet 5.5 (1-hour cache write): highest 20 turns $61.08. Lowest 1 turn $31.32.

Notes

Prefix of 7,831 tokens (Sonnet 5.5) and 7,828 tokens (Opus 5.5), the rounded mean turn-1 input; n = 3 Sonnet sessions (range 7,831 to 7,832) and 3 Opus sessions (the same in every session); USD per 1,000 sessions

Calculation, not a bill. The chart multiplies list prices by a prefix of 7,831 tokens (Sonnet) or 7,828 tokens (Opus). It uses a 1-hour write at 2× input and treats the whole prefix as new. It leaves out output and the tokens that each turn adds. The cache costs more at 1 to 2 turns and less from 3. The second Opus line uses a $0.4 cache read, which another table of the product lists. We did not check the vendor price. The table adds the 5-minute write.

Sources: Cost with and without the prompt cache (calculation), Caching sessions and repeated prompts (Claude Code and Codex CLI), Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)
Calculation

No cache (the same either way)

Claude Sonnet 5.5
Claude Opus 5.5 (cache read $0.2 per M)
Claude Opus 5.5 (cache read $0.4 per M)

One 10-turn session, 1-hour cache

Claude Sonnet 5.5
Claude Opus 5.5 (cache read $0.2 per M)
Claude Opus 5.5 (cache read $0.4 per M)

Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix)

Claude Sonnet 5.5
Claude Opus 5.5 (cache read $0.2 per M)
Claude Opus 5.5 (cache read $0.4 per M)

One panel per series, all on the same axis.

List-price calculation, not a run. 3 rows, 3 series: No cache (the same either way), One 10-turn session, 1-hour cache, Ten 1-turn sessions, 1-hour cache (assumed no reuse of the new prefix). No cache (the same either way): highest Claude Opus 5.5 (cache read $0.2 per M) $313. Lowest Claude Sonnet 5.5 $157. One 10-turn session, 1-hour cache: highest Claude Opus 5.5 (cache read $0.4 per M) $90.80. Lowest Claude Sonnet 5.5 $45.42.

Notes

USD per 1,000 workloads of 10 requests that send the same prefix; each recorded session started in a new temporary folder, n = 4 later sessions; 0 met the read-more-than-write proxy, 95% Wilson 0% to 49%

Calculation, not a bill: the same prefix, prices and 1-hour write as the cost-curve chart. Each recorded session ran in a new temporary working folder. On turn 1, 0 of 4 later sessions read more tokens than they wrote. This proxy cannot identify which earlier call supplied a cache entry. The split calculation assumes no shared reuse and a wholly new prefix, so it charges ten full writes. With reuse across sessions they would cost the same as one 10-turn session. The cause of the missing reuse was not tested here. The caching study protocol names the random temporary folder as one untested hypothesis. The table adds the 5-minute variants.

Sources: Cost with and without the prompt cache (calculation), Caching sessions and repeated prompts (Claude Code and Codex CLI), Anthropic list prices (Claude models), OpenAI list prices

Share card (PNG)

Tables

Break-even reuses by model (calculation)

Priced asInput $/MCache read $/M1-hour write $/M5-minute write $/M (assumed 1.25×)Break-even, 1-hour, new prefixReuses needed, 1-hour, new prefixBreak-even, 1-hour, 19% already cachedReuses needed, 1-hour, 19% already cachedBreak-even, 5-minuteReuses needed, 5-minuteNote
Claude Haiku 4.5$1.00$0.1$2.00$1.251.1120.7210.281Anthropic list price, 1-hour write at 2× input
Claude Sonnet 5.5$2.00$0.2$4.00$2.501.1120.7210.281Anthropic list price, 1-hour write at 2× input
Claude Opus 5.5 (cache read $0.2 per M)$4.00$0.2$8.00$5.001.0520.6710.261Anthropic list price, 1-hour write at 2× input
Claude Opus 5.5 (cache read $0.4 per M)$4.00$0.4$8.00$5.001.1120.7210.281Cache-read price from another product table; the vendor price was not checked
Claude Fable 5.1$10.00$0.25$20.00$12.501.0320.6510.261Anthropic list price, 1-hour write at 2× input
GPT-6.1 Sol$2.00$0.1$2.00$2.0001-0.19001No write surcharge in the price list: a write is plain input, so the cache saves from the first read. Another product table lists 1.25× input as the write price: break-even 0.26 reuses
GPT-6 Luna$0.1$0.01$0.1$0.101-0.19001No write surcharge in the price list: a write is plain input, so the cache saves from the first read. Another product table lists 1.25× input as the write price: break-even 0.28 reuses

Saving from the cache by session length (calculation; negative = the cache costs more)

Priced asWrite1 turn2 turns3 turns5 turns10 turns20 turns
Claude Haiku 4.51-hour write, new prefix-100%-5%27%52%71%81%
Claude Haiku 4.55-minute write (assumed 1.25×)-25%33%52%67%79%84%
Claude Sonnet 5.51-hour write, new prefix-100%-5%27%52%71%81%
Claude Sonnet 5.55-minute write (assumed 1.25×)-25%33%52%67%79%84%
Claude Opus 5.5 (cache read $0.2 per M)1-hour write, new prefix-100%-2.5%30%56%76%85%
Claude Opus 5.5 (cache read $0.2 per M)5-minute write (assumed 1.25×)-25%35%55%71%83%89%
Claude Opus 5.5 (cache read $0.4 per M)1-hour write, new prefix-100%-5%27%52%71%81%
Claude Opus 5.5 (cache read $0.4 per M)5-minute write (assumed 1.25×)-25%33%52%67%79%84%
Claude Fable 5.11-hour write, new prefix-100%-1.3%32%58%78%88%
Claude Fable 5.15-minute write (assumed 1.25×)-25%36%57%73%85%91%
GPT-6.1 Solno write surcharge0%48%63%76%86%90%
GPT-6 Lunano write surcharge0%45%60%72%81%86%

Cost per 1,000 sessions by session length (calculation)

Priced asPrefix (tokens)TurnsNo cache (USD)1-hour write (USD)5-minute write, assumed 1.25× (USD)Saving, 1-hourSaving, 5-minute
Claude Sonnet 5.57,8311$15.66$31.32$19.58-100%-25%
Claude Sonnet 5.57,8312$31.32$32.89$21.14-5%33%
Claude Sonnet 5.57,8313$46.99$34.46$22.7127%52%
Claude Sonnet 5.57,8315$78.31$37.59$25.8452%67%
Claude Sonnet 5.57,83110$157$45.42$33.6771%79%
Claude Sonnet 5.57,83120$313$61.08$49.3481%84%
Claude Opus 5.5 (cache read $0.2 per M)7,8281$31.31$62.62$39.14-100%-25%
Claude Opus 5.5 (cache read $0.2 per M)7,8282$62.62$64.19$40.71-2.5%35%
Claude Opus 5.5 (cache read $0.2 per M)7,8283$93.94$65.76$42.2730%55%
Claude Opus 5.5 (cache read $0.2 per M)7,8285$157$68.89$45.4056%71%
Claude Opus 5.5 (cache read $0.2 per M)7,82810$313$76.71$53.2376%83%
Claude Opus 5.5 (cache read $0.2 per M)7,82820$626$92.37$68.8985%89%
Claude Opus 5.5 (cache read $0.4 per M)7,8281$31.31$62.62$39.14-100%-25%
Claude Opus 5.5 (cache read $0.4 per M)7,8282$62.62$65.76$42.27-5%33%
Claude Opus 5.5 (cache read $0.4 per M)7,8283$93.94$68.89$45.4027%52%
Claude Opus 5.5 (cache read $0.4 per M)7,8285$157$75.15$51.6652%67%
Claude Opus 5.5 (cache read $0.4 per M)7,82810$313$90.80$67.3271%79%
Claude Opus 5.5 (cache read $0.4 per M)7,82820$626$122$98.6381%84%

10 requests as one session or as 10 one-turn sessions with no shared reuse, per 1,000 workloads (calculation)

Priced asNo cache (USD)One 10-turn session, 1-hour (USD)10 one-turn sessions with no shared reuse, 1-hour (USD)One 10-turn session, 5-minute (USD)10 one-turn sessions with no shared reuse, 5-minute (USD)10 one-turn sessions with no shared reuse vs one session, 1-hour
Claude Sonnet 5.5$157$45.42$313$33.67$1966.9x
Claude Opus 5.5 (cache read $0.2 per M)$313$76.71$626$53.23$3918.2x
Claude Opus 5.5 (cache read $0.4 per M)$313$90.80$626$67.32$3916.9x

Recorded inputs and a check of the formula (calculation)

Recorded sessionsSessionsTurn-1 input, rounded mean (tokens)Turn-1 prefix, range (tokens)Already cached at turn 1, mean (tokens)Written on turn 1, mean (tokens)Input saving, session range (not a confidence interval)Recorded saving, input side, 5 turnsFormula, whole prefix new, 5 turnsFormula, 19% already cached, 5 turnsSaving with output, session range (not a confidence interval)Recorded saving, with output, 5 turns
Claude Sonnet 5.5 · Claude Code37,8317,831 to 7,8321,4636,36654.3% to 57.1%55%52%59%47.5% to 54.1%50%
Claude Opus 5.5 · Claude Code37,8287,8281,4636,36358.5% to 59.7%59%56%63%51.4% to 54.3%53%

Method

  1. This study makes no new model call and has no raw file of its own. It calculates from two inputs. The first is the turn tokens that the caching study recorded in 6 Claude Code sessions (3 on Sonnet 5.5, 3 on Opus 5.5). The second is the list prices in the price sources.
  2. A prefix of T tokens goes out in k + 1 requests. The first request writes it to the cache. The k later requests, the reuses, read it. Prices are USD per million tokens: input i, cache read r, cache write w.
  3. No cache: (k + 1) × T × i. A 1-hour write: T × w + k × T × r, where w is the listed 1-hour write price (twice i for every Anthropic model in the price list). A 5-minute write: T × 1.25 × i + k × T × r.
  4. The 1.25 is the multiplier that the caching study protocol states as Anthropic’s published figure. It is not in the price list. It is an assumption here. No 5-minute write occurred in the recorded sessions.
  5. Break-even: the cached prefix costs less when k > (w − i) ÷ (i − r). The exact break-even is that fraction. The reuses needed is the smallest non-negative whole number above it. A negative fraction means the first request already costs less, with zero reuses.
  6. A share f of the prefix can already be in the cache when the first request arrives. That request reads the share and writes the rest: T × (f × r + (1 − f) × w) + k × T × r.
  7. In the recorded sessions, turn 1 read 1,463 tokens on its first request. The usage counters do not record read/write timing. That is f = 19% (8,778 of 46,978 tokens, 6 sessions). Its origin was not isolated. The dollar figures use f = 0, the whole prefix new, which is the stricter case. The break-even rows show both.
  8. T is the rounded mean turn-1 input (uncached input + cache reads + cache writes) of a model’s recorded sessions. For Sonnet 5.5 it is 7,831 tokens (range 7,831 to 7,832, n = 3). For Opus 5.5 it is 7,828 tokens (the same in every session, n = 3). A session of N turns is N requests that send the prefix, so k = N − 1. The cost tables give 1,000 sessions of 1, 2, 3, 5, 10 and 20 turns.
  9. Split case: 10 requests as one 10-turn session (one write, 9 reads) against 10 one-turn sessions (10 writes). Each recorded session ran in a fresh process and a new temporary working folder. Of 4 later sessions, 0 met the reuse proxy. The split case assumes a wholly new prefix with no shared reuse; the receipts do not test that exact case. The reuse proxy follows the caching study’s rule: turn 1 reads more than it writes.
  10. Opus 5.5 price question: the price list gives a cache read of $0.2 per million (5% of the $4 input price). Another table of the product lists $0.4 (10%). This study did not check the vendor price. Every Opus 5.5 row and the second Opus line show both values.
  11. Check against the recorded sessions, on the input side only (output excluded). The recorded 5-turn sessions saved 55.4% on Sonnet 5.5 (n = 3; session range 54.3% to 57.1%) and 59.2% on Opus 5.5 (n = 3; range 58.5% to 59.7%). These ranges are not confidence intervals. The formula gives 52.0% and 56.0% with a new prefix, and 59.1% and 63.3% with 19% already cached. The two formula cases bracket the recorded figure for both models. The recorded sessions also wrote the tokens that each turn added.
  12. With output included, the recorded saving is 50.0% for Sonnet 5.5 (n = 3; session range 47.5% to 54.1%) and 53.1% for Opus 5.5 (n = 3; session range 51.4% to 54.3%). These are list-price calculations. The ranges are not confidence intervals.
  13. Everything is a list-price calculation, not a bill. The sample behind T is n = 3 sessions per model.

Caveats

  • The retained Claude protocol file was created at 14:43:55 UTC, after the first counted session at 14:35:39 UTC on 2026-10-06. Its declaration says 14:25 UTC, but file times do not verify that claim. Treat this as a retrospective protocol record.
  • The source has 30 attempted Claude turns, 0 failed turns and 0 turns without token usage. Failed turns with usage remain in cost totals. Missing usage cannot be priced. No quality rate or cache-caused speed effect is claimed.
  • 7 counted turns record a provider rate-limit warning status. The retained outcome log says there was no warning. The protocol calls for a stop on limit text in an error or warning. These receipts do not establish whether that text appeared; this discrepancy remains unresolved.
  • One shared Mac and one synthetic ledger supplied these tokens. The ledger was shortened once after a two-call probe and before the counted sessions. All 30 counted answers passed, so this task set hits a quality ceiling. This calculation does not rank model quality or speed.
  • A calculation, not a run and not a bill. The recorded calls used flat subscriptions. The prices are list prices dated 2026-09-21 (Anthropic) and 2026-10-03 (OpenAI); they can change.
  • The formula prices the shared prefix only. A real turn also writes new tokens to the cache at the write price. In the recorded sessions that was 58 to 1,117 tokens per turn after turn 1. A real turn also produces output, which costs the same with or without the cache. The check against the recorded sessions shows the size of this effect.
  • The formula assumes that every reuse arrives before the cache entry expires. An entry lasts 5 minutes or 1 hour, by write type. When the gap between requests is longer, the write repeats and the saving shrinks.
  • The recorded sessions wrote 1-hour entries only (0 of 30 turns wrote a 5-minute entry). The 5-minute rows use an assumed 1.25 multiplier that this study did not test.
  • Only 0 of 4 later sessions met the read-more-than-write proxy (95% Wilson 0% to 49%). They still read some cached tokens. This proxy does not establish cache provenance. Each recorded session ran in a fresh process and a new temporary working folder. The caching study protocol names that folder as one untested hypothesis for the miss, and this study did not test it. The split case assumes no shared reuse and a wholly new prefix. It does not price the partly cached recorded first turns. A script that keeps one working folder may read an earlier session’s cache: check your own receipts before you use the split case.
  • The 19% already-cached share is not a constant. We recorded it on Sonnet 5.5 and Opus 5.5 sessions. For Haiku 4.5 and Fable 5.1 it is a what-if. The one Haiku 4.5 probe session (outside every cell) read 0 tokens from the cache on turn 1.
  • The Opus 5.5 cache-read price is open: $0.2 or $0.4 per million. This study shows both. Resolve the price before you quote an Opus figure.
  • The dollar figures depend on T: one synthetic ledger in one CLI version, 3 sessions per model. A different prefix changes the dollars. It does not change the break-even reuses, which depend on the price ratios only.
  • The GPT rows follow the price list, which has no write surcharge. Another table of the product lists a write price of 1.25× input for both models. At that price they need 1 reuse (exact break-even 0.26 and 0.28), and a prefix used once costs 1.25 times as much. We did not check the vendor price. The Codex app-server reports no cache writes, and this study did not test reuse for GPT. The caching study measured Codex reads on a larger context.

Sources

  • Cost with and without the prompt cache (calculation)

    Calculation ·

    Recorded tokens per turn × Anthropic list prices. With the cache: uncached input at the input price, cache reads at the cache-read price, 1-hour cache writes at twice the input price, 5-minute writes at 1.25 times (an assumption; none occurred). Without a cache: every input token at the input price. Output is priced the same in both. Not a bill.

    Raw data: caching-consistency/caching.json

  • Caching sessions and repeated prompts (Claude Code and Codex CLI)

    Our recorded runs ·

    Part 1: 5-turn CLI sessions over a fixed synthetic ledger, with the cache counters each provider reports per turn. Part 2: three prompts with deterministic validators, 10 repetitions per model. Declared protocols, validator controls before inference, every attempt kept; answers are published as ordinal ids, never as text.

    Raw data: caching-consistency/caching.json, caching-consistency/consistency.json

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • OpenAI list prices

    Vendor price list ·

    Token prices as listed by the vendor on 2026-10-03.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Prompt cache break-even: after how many reuses does a cached prefix cost less?”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/prompt-cache-break-even.

Models and comparisons in this study

More studies

All benchmarks
Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Live story
  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.