Does Claude Code reuse the prompt cache across sessions? A fixed-folder test
Claude Code showed near-full cache reads in later fixed-folder sessions (4/4; 95% interval 51–100%). New folders: 0/2 (0–66%). Exploratory, 30 calls.
TL;DR
- Yes, Claude Code showed near-full turn-1 cache reads in later sessions with a fixed folder. Across setups B and C, 4 of 4 later sessions did so (pooled calculation; 95% Wilson interval 51% to 100%). With a new folder per session, 0 of 2 met the 50% read-share rule (0% to 66%). Those intervals overlap. This test does not show that the folder caused the difference.
- What it costs. Mean later-session turn-1 cost at list price was $0.0259 in A (n = 2; range $0.025871 to $0.025879). Across B and C it was $0.0016 (n = 4; range $0.001597 to $0.001601). The ratio is 16.2× (calculation, not a bill).
- Both ledger placements showed near-full reads with a fixed folder. Setup B put the ledger in the first message; C used
--append-system-prompt. Each had 2 of 2 later sessions meet the rule (95% interval 34% to 100%). This does not prove equal reuse rates. - No isolated speed effect. Turn-1 medians were 1.56 s in A, 1.76 s in B and 1.77 s in C. Each has n = 3; ranges appear below. Folder, ledger seed and call order differed.
- Codex CLI did not reach the post-hoc 90% rule in any later session: 0 of 4 (95% interval 0% to 49%). The 50% rule in the surviving protocol counts 2 of 4 (15% to 85%). Neither rule identifies which tokens were read.
- For a script that starts a new session per task: test a fixed folder and check your receipts. These observations do not guarantee reuse on another workload.
Study: /benchmarks/prompt-cache-across-sessions. Earlier study: /benchmarks/caching-consistency.
Prompt caching and consistency: what the cache saves, and how much answers vary
Calculation at list price: the cache cut a 5-question session 50% on Sonnet 5.5 and 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every repetition.
Transcript
- Caching and consistency · 135 calls. What the cache saves, and how much answers vary. Five-question sessions on a fixed context. Then the same prompt, 10 times.
- 135 calls: 45 cache turns, 90 repeated prompts. Every call counted. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- Claude Code turn 1 reads 19% from the cache, the CLI’s own prefix. Turns 2–5 read 90%–99%. Codex CLI: 98%–99%. Chart: Share of input read from the cache · mean of 3 sessions per turn (n = 3 each). Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- A calculation: Sonnet 5.5 $0.1350 with the cache vs $0.2698 without, 50% less. Opus 5.5 $0.2551 vs $0.5442, 53% less. Chart: List-price cost of the recorded sessions · calculation (n = 15 each). Calculation, not a run. Caveat: Costs are list-price calculations; the calls used flat subscriptions.
- No clear speed effect: Sonnet 5.5 1.6 s on turn 1 vs 1.6 s later; Opus 5.5 1.9 s vs 2.4 s. The ranges overlap. Sonnet 5.5: median turn 1 vs turns 2–5: 1.6 s vs 1.6 s (ranges 1.6–1.8 s and 1.4–5.6 s · n = 3 and 12). Opus 5.5: median turn 1 vs turns 2–5: 1.9 s vs 2.4 s (ranges 1.8–4.4 s and 1.6–12.7 s · n = 3 and 12). Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- 7 of 9 cells passed 10/10. Haiku 4.5: 0/10 on the exact number (10 wrong, 1 distinct answer), 1/10 on JSON (9 format misses). Chart: Same prompt, 10 times · strict passes · a 10/10 is 72%–100% at 95% (n = 10 each). Caveat: Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
- Consistent is not correct: Haiku 4.5 gave the same wrong number all 10 times. Code fix, distinct correct bodies: Haiku 4.5 6, Sonnet 5.5 3, GPT-6.1 Sol (medium) 6. Chart: Same prompt, 10 times · distinct answers (n = 10 each). Caveat: The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
- Cache the fixed context; check answers, not agreement. Every call online.
The question
Our earlier caching study found that 0 of 4 later Claude Code sessions met the ledger-reuse rule on turn 1 (95% interval 0% to 49%). We wrote that we did not know why (How much does prompt caching actually save?). Many scripts start a new session for every task. This post tests whether a fixed working folder changes the observed cache counts. Background on the mechanism: how prompt caching works.
The chart is from the earlier study. In the Claude sessions, turn 1 read about 19% of its input from the cache (calculation; n = 3 sessions per model). We interpret that part as the CLI's own prefix. No token-level trace tests this reading. Every session had run in a new temporary folder.
What we ran
We used Claude Sonnet 5.5 in Claude Code 2.1.286 at default effort. Tools were off and the working folder was empty. A session is one CLI process with two turns. Turn 1 sends a synthetic stock ledger and a question. The ledger has 100 lines; total turn-1 input is about 7.8k tokens. Turn 2 sends a second question. We ran three sessions per setup, back to back. Between one session's last call and the next session's first call, 508 to 720 ms passed (calculation from recorded times).
| Setup | Working folder | Where the ledger goes |
|---|---|---|
| A | a new temporary folder for every session (the earlier setup) | the first user message |
| B | one fixed folder | the first user message |
| C | the same fixed folder | --append-system-prompt |
Each setup had its own ledger, from its own random seed. Different seeds prevent full ledger-prefix reuse between setups. Shared CLI-prefix reads remain possible. A and B differ in folder, seed and call order. We made 18 counted Claude calls; all passed the normalized-answer checker (95% interval 82% to 100%). One probe call was not counted.
The surviving protocol claims a declaration at 00:32 UTC. Its file birth time is 00:49:27 UTC, after all counted calls. The first counted call started at 00:32:33 UTC. We cannot verify a pre-call protocol. Treat the analysis as exploratory. We kept every counted attempt and retried nothing.
Result: near-full reads in later fixed-folder sessions
Each bar is turn 1 of one session. Session 1 was the first use of its setup's ledger. Sessions 2 and 3 are the later sessions. The table shows individual receipts; read share is a calculation from the reported input counters.
| Setup | Later session | Read from cache | Written to cache | Read share (calculation) | Turn-1 list cost (calculation) |
|---|---|---|---|---|---|
| A, new folder | 2 | 1,463 | 6,386 | 19% | $0.0259 |
| A, new folder | 3 | 1,463 | 6,388 | 19% | $0.0259 |
| B, fixed folder | 2 | 7,833 | 0 | 99.97% | $0.0016 |
| B, fixed folder | 3 | 7,833 | 0 | 99.97% | $0.0016 |
| C, fixed folder, ledger in system prompt | 2 | 7,815 | 0 | 99.97% | $0.0016 |
| C, fixed folder, ledger in system prompt | 3 | 7,815 | 0 | 99.97% | $0.0016 |
In A, both later sessions read 1,463 tokens and wrote 6,386 to 6,388 tokens. In B and C, later sessions read nearly all input and wrote nothing. Two uncached input tokens remained in each call. We interpret A's reads as a shared prefix; the receipts do not identify those tokens.
The surviving protocol labels at least 50% read share as reuse. Each setup has n = 2 later sessions. B and C each met that rule 2 of 2 times (95% Wilson interval 34% to 100%). A met it 0 of 2 times (0% to 66%). These intervals overlap.
We also pooled sessions after seeing the counts. This is a post-hoc calculation. New-folder later sessions from both studies: 0 of 6 (95% interval 0% to 39%). Fixed-folder later sessions, B and C together: 4 of 4 (51% to 100%). These intervals do not overlap.
The four earlier sessions are not a match for the new ones. They differ in model, ledger and number of turns, and ran under another Claude login. Sonnet and Opus sessions ran interleaved. Same-model sessions were 13 to 27 seconds apart (calculation), rather than under one second. B and C also differ in ledger placement. The pooled counts are a rough check, not a controlled comparison.
What the folder costs you
This is a list-price calculation from recorded tokens, not a bill. The calls ran on a subscription. At the study's Sonnet 5.5 prices, a cache read costs a tenth of the input price. A 1-hour cache write costs twice the input price. Every write in this run used that 1-hour tier.
- Later session, fixed folder: mean $0.0016 for turn 1 across B and C (n = 4; range $0.001597 to $0.001601).
- Later session, new folder: mean $0.0259 for turn 1 in A (n = 2; range $0.025871 to $0.025879). That is 16.2 times as much (calculation). It is about 1.6 times the mean $0.0157 without a cache (calculation on the same two A receipts; range $0.015732 to $0.015736).
Our total turn-1 input is about 7.8k tokens. At these list prices, writing more tokens increases the dollar cost (calculation). This does not predict cache matches or token use on another workload.
Did the cache make turn 1 faster?
This run does not isolate a speed effect. Median turn-1 time was 1.56 s for A (range 1.34 to 1.62 s), 1.76 s for B (1.71 to 1.86 s) and 1.77 s for C (1.69 to 1.78 s). Each setup has n = 3 sessions. Ranges are fastest to slowest, not confidence intervals. These medians include session 1, when each setup wrote its ledger.
Turn 2 read from the cache in all nine Claude sessions. Its medians were 1.02 s in A (range 0.88 to 1.14 s), 1.33 s in B (1.31 to 1.37 s) and 1.13 s in C (1.09 to 1.29 s). Each has n = 3. The Mac also ran other agent work. Folder, ledger and order changed together, so these times cannot show that caching has no speed benefit.
Codex CLI: no near-full turn-1 reads in this sample
We ran A and B with GPT-6.1 Sol (medium) in Codex CLI 0.160.0. All 12 counted calls passed the normalized-answer checker (95% interval 76% to 100%). One probe call was not counted. Turn-1 input ranged from 16,033 to 16,053 tokens across six sessions. Codex adds a large system prompt of its own.
Across all six turn-1 calls, cached input ranged from 0 to 8,960 tokens. The highest read share was 55.9% (calculation). The 50% threshold in the surviving protocol counts 2 of 4 later sessions (95% interval 15% to 85%). A session 2 cached 8,960 of 16,051 tokens (55.8%, calculation). B session 2 cached 8,960 of 16,033 (55.9%, calculation).
The probe also showed 8,960 cached tokens. We interpret that count as Codex's system prefix, but did not test which tokens it held. The post-hoc 90% rule counts 0 of 4 later sessions (95% interval 0% to 49%). We cannot verify that the surviving protocol's 50% rule was recorded before the calls. No token-level trace identifies the ledger under either rule.
Turn 2 cached 15,872 tokens in all six sessions. Its read share ranged from 98.7% to 98.8% (calculation). Thus near-full reads appeared inside sessions, but not on turn 1 under the 90% rule. This small sample cannot show that folders never matter. Codex reports no cache-write count, so we calculate no cost for it. Do not rank Codex against Claude here: the prefixes and counters differ. More on the two CLIs: Claude Code vs Codex CLI.
What the conditions show, and what they do not
- Observed: later fixed-folder Claude sessions had near-full cache reads. Later new-folder sessions had partial reads and wrote most input again. A and B differ in folder, ledger seed and call order. The observations do not prove a folder cause or a requirement to use one fixed folder.
- Observed: both ledger placements had near-full reads with a fixed folder. Two later sessions per setup cannot establish equal reuse rates.
- Not shown: whether the folder path, name or another property breaks the match. We did not inspect the outgoing request. The run notes identify
--exclude-dynamic-system-prompt-sectionsas an option for moving per-machine sections into the first user message. We did not test it or verify its effect on these requests. - Not shown: whether a ledger in the system prompt survives a changing folder. Setup C ran only with a fixed folder.
- Not shown: reuse after a long gap. The sessions ran less than a second apart, and the cache entries were 1-hour writes.
What this means for a script that starts a new session per task
- Test one fixed working folder when you can. In this sample, later fixed-folder sessions had near-full cache reads. That does not guarantee the same result for your script.
- Budget for possible cache writes with fresh folders. A's later sessions wrote most input again. Their partial reads did not make turn 1 cheaper than the no-cache calculation.
- Check your own receipts. Claude Code stream output reports cache read and cache write tokens. Use those counts to calculate list-price cost; subscription calls are not per-token bills.
- Test the CLI option before you rely on it. The run notes name a possible option to investigate. We have not tested it, so we do not claim it fixes reuse.
- Do not generalize past this setup. Tools were off, the folder was empty, and the shared prefix was small. A real repository adds text that can differ between sessions.
Caveats
- Small n. There are two later sessions per setup and route. Per-setup Claude intervals overlap. Both pooled counts are post-hoc calculations.
- Exploratory protocol. The retained protocol file was created after all counted calls. Separate pre-edit copies do not verify the claimed amendment times. The summary was regenerated later.
- Correctness ceiling. All 30 answers passed after text normalization (95% Wilson interval 89% to 100%). The checker trims spaces, quotes, backticks, a final period and currency units. These short synthetic lookups cannot establish coding quality or general task success. Checker controls ran after the counted calls; cache measures do not depend on them.
- One CLI version and one model per route: Claude Code 2.1.286 with Sonnet 5.5, and Codex CLI 0.160.0 with GPT-6.1 Sol.
- Costs are calculations. They are list price times recorded tokens. The calls used flat subscriptions.
- Timing has confounders. The Mac ran other agent work. Folder, ledger seed and call order differed.
- The Codex criterion changed after inspection. The 90% rule gives 0 of 4 (95% interval 0% to 49%). The surviving protocol's 50% rule gives 2 of 4 (15% to 85%). Neither identifies the cached tokens.
- The pooled new-folder count is not a matched sample. Earlier sessions used another login and ledger, with five turns and interleaved Sonnet and Opus calls. Their same-model gaps were 13 to 27 seconds (calculation).
- The Codex allowance gate is not verified. The protocol lists conflicting entry thresholds, and no retained gate receipt proves the allowance at run time. The call caps were kept.
- I build Agent, so I want agent runs to be cheap. The study page and sanitized receipts are public. We kept every counted attempt and retried nothing.
What to read next
- How much does prompt caching actually save?
- How prompt caching works
- Claude Code vs Codex CLI: the hidden context tax
- How to estimate your AI coding bill
- All the numbers: /benchmarks/prompt-cache-across-sessions
- Sanitized cache-session receipts
See your own cache share
Agent records cache reads and writes for every model call. You can check whether new sessions write most input again. Try Agent and check your own numbers.