• Prompt Caching
  • Cache Reuse
  • Claude Code
  • Codex CLI

Does Claude Code reuse the prompt cache across sessions? A fixed-folder test

Claude Code showed near-full cache reads in later fixed-folder sessions (4/4; 95% interval 51–100%). New folders: 0/2 (0–66%). Exploratory, 30 calls.

TL;DR

  • Yes, Claude Code showed near-full turn-1 cache reads in later sessions with a fixed folder. Across setups B and C, 4 of 4 later sessions did so (pooled calculation; 95% Wilson interval 51% to 100%). With a new folder per session, 0 of 2 met the 50% read-share rule (0% to 66%). Those intervals overlap. This test does not show that the folder caused the difference.
  • What it costs. Mean later-session turn-1 cost at list price was $0.0259 in A (n = 2; range $0.025871 to $0.025879). Across B and C it was $0.0016 (n = 4; range $0.001597 to $0.001601). The ratio is 16.2× (calculation, not a bill).
  • Both ledger placements showed near-full reads with a fixed folder. Setup B put the ledger in the first message; C used --append-system-prompt. Each had 2 of 2 later sessions meet the rule (95% interval 34% to 100%). This does not prove equal reuse rates.
  • No isolated speed effect. Turn-1 medians were 1.56 s in A, 1.76 s in B and 1.77 s in C. Each has n = 3; ranges appear below. Folder, ledger seed and call order differed.
  • Codex CLI did not reach the post-hoc 90% rule in any later session: 0 of 4 (95% interval 0% to 49%). The 50% rule in the surviving protocol counts 2 of 4 (15% to 85%). Neither rule identifies which tokens were read.
  • For a script that starts a new session per task: test a fixed folder and check your receipts. These observations do not guarantee reuse on another workload.

Study: /benchmarks/prompt-cache-across-sessions. Earlier study: /benchmarks/caching-consistency.

Live story · 49 sPrompt caching and consistency: what the cache saves, and how much answers vary

Prompt caching and consistency: what the cache saves, and how much answers vary

Calculation at list price: the cache cut a 5-question session 50% on Sonnet 5.5 and 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every repetition.

Transcript
  1. Caching and consistency · 135 calls. What the cache saves, and how much answers vary. Five-question sessions on a fixed context. Then the same prompt, 10 times.
  2. 135 calls: 45 cache turns, 90 repeated prompts. Every call counted. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
  3. Claude Code turn 1 reads 19% from the cache, the CLI’s own prefix. Turns 2–5 read 90%–99%. Codex CLI: 98%–99%. Chart: Share of input read from the cache · mean of 3 sessions per turn (n = 3 each). Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  4. A calculation: Sonnet 5.5 $0.1350 with the cache vs $0.2698 without, 50% less. Opus 5.5 $0.2551 vs $0.5442, 53% less. Chart: List-price cost of the recorded sessions · calculation (n = 15 each). Calculation, not a run. Caveat: Costs are list-price calculations; the calls used flat subscriptions.
  5. No clear speed effect: Sonnet 5.5 1.6 s on turn 1 vs 1.6 s later; Opus 5.5 1.9 s vs 2.4 s. The ranges overlap. Sonnet 5.5: median turn 1 vs turns 2–5: 1.6 s vs 1.6 s (ranges 1.6–1.8 s and 1.4–5.6 s · n = 3 and 12). Opus 5.5: median turn 1 vs turns 2–5: 1.9 s vs 2.4 s (ranges 1.8–4.4 s and 1.6–12.7 s · n = 3 and 12). Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
  6. 7 of 9 cells passed 10/10. Haiku 4.5: 0/10 on the exact number (10 wrong, 1 distinct answer), 1/10 on JSON (9 format misses). Chart: Same prompt, 10 times · strict passes · a 10/10 is 72%–100% at 95% (n = 10 each). Caveat: Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
  7. Consistent is not correct: Haiku 4.5 gave the same wrong number all 10 times. Code fix, distinct correct bodies: Haiku 4.5 6, Sonnet 5.5 3, GPT-6.1 Sol (medium) 6. Chart: Same prompt, 10 times · distinct answers (n = 10 each). Caveat: The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
  8. Cache the fixed context; check answers, not agreement. Every call online.

The question

Our earlier caching study found that 0 of 4 later Claude Code sessions met the ledger-reuse rule on turn 1 (95% interval 0% to 49%). We wrote that we did not know why (How much does prompt caching actually save?). Many scripts start a new session for every task. This post tests whether a fixed working folder changes the observed cache counts. Background on the mechanism: how prompt caching works.

  • Claude Sonnet 5.5 · Claude Code
  • Claude Opus 5.5 · Claude Code
  • GPT-6.1 Sol (medium) · Codex CLI*

* The note below the chart says what this route does not report.

5 turn in the sessions, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (medium) · Codex CLI. Claude Sonnet 5.5 · Claude Code: highest Turn 2 99% (n 3). Lowest Turn 1 19% (n 3). Claude Opus 5.5 · Claude Code: highest Turn 2 99% (n 3). Lowest Turn 1 19% (n 3).

Notesn = 3 per row

Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each

Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

The chart is from the earlier study. In the Claude sessions, turn 1 read about 19% of its input from the cache (calculation; n = 3 sessions per model). We interpret that part as the CLI's own prefix. No token-level trace tests this reading. Every session had run in a new temporary folder.

What we ran

We used Claude Sonnet 5.5 in Claude Code 2.1.286 at default effort. Tools were off and the working folder was empty. A session is one CLI process with two turns. Turn 1 sends a synthetic stock ledger and a question. The ledger has 100 lines; total turn-1 input is about 7.8k tokens. Turn 2 sends a second question. We ran three sessions per setup, back to back. Between one session's last call and the next session's first call, 508 to 720 ms passed (calculation from recorded times).

SetupWorking folderWhere the ledger goes
Aa new temporary folder for every session (the earlier setup)the first user message
Bone fixed folderthe first user message
Cthe same fixed folder--append-system-prompt

Each setup had its own ledger, from its own random seed. Different seeds prevent full ledger-prefix reuse between setups. Shared CLI-prefix reads remain possible. A and B differ in folder, seed and call order. We made 18 counted Claude calls; all passed the normalized-answer checker (95% interval 82% to 100%). One probe call was not counted.

The surviving protocol claims a declaration at 00:32 UTC. Its file birth time is 00:49:27 UTC, after all counted calls. The first counted call started at 00:32:33 UTC. We cannot verify a pre-call protocol. Treat the analysis as exploratory. We kept every counted attempt and retried nothing.

Result: near-full reads in later fixed-folder sessions

Calculation

Session 1 (first in its setup)

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

Session 2

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

Session 3

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

One panel per series, all on the same axis.

List-price calculation, not a run. 3 setups, 3 series: Session 1 (first in its setup), Session 2, Session 3. Session 1 (first in its setup): highest B: fixed folder 19% (n 1). Lowest A: new folder each time 6.8% (n 1). Session 2: highest B: fixed folder 100% (n 1). Lowest A: new folder each time 19% (n 1).

Notesn = 1 per row

One bar per call: turn 1 of one session; 3 sessions per setup, run back to back; read-share calculation

Session 1 was the first session to use its setup’s ledger. Sessions 2 and 3 used that same ledger. Different seeds prevent full ledger-prefix reuse between setups; shared CLI-prefix reads remain possible. We interpret session-1 reads as a shared CLI prefix; no token-level trace proves this. Read share (calculation) = cache reads ÷ (uncached input + cache reads + cache writes), as the provider reports them. Each bar is one call, not a rate; the counts per setup are in the table. 2 later sessions per setup is a small number.

Source: Prompt cache across sessions

Each bar is turn 1 of one session. Session 1 was the first use of its setup's ledger. Sessions 2 and 3 are the later sessions. The table shows individual receipts; read share is a calculation from the reported input counters.

SetupLater sessionRead from cacheWritten to cacheRead share (calculation)Turn-1 list cost (calculation)
A, new folder21,4636,38619%$0.0259
A, new folder31,4636,38819%$0.0259
B, fixed folder27,833099.97%$0.0016
B, fixed folder37,833099.97%$0.0016
C, fixed folder, ledger in system prompt27,815099.97%$0.0016
C, fixed folder, ledger in system prompt37,815099.97%$0.0016

In A, both later sessions read 1,463 tokens and wrote 6,386 to 6,388 tokens. In B and C, later sessions read nearly all input and wrote nothing. Two uncached input tokens remained in each call. We interpret A's reads as a shared prefix; the receipts do not identify those tokens.

The surviving protocol labels at least 50% read share as reuse. Each setup has n = 2 later sessions. B and C each met that rule 2 of 2 times (95% Wilson interval 34% to 100%). A met it 0 of 2 times (0% to 66%). These intervals overlap.

We also pooled sessions after seeing the counts. This is a post-hoc calculation. New-folder later sessions from both studies: 0 of 6 (95% interval 0% to 39%). Fixed-folder later sessions, B and C together: 4 of 4 (51% to 100%). These intervals do not overlap.

The four earlier sessions are not a match for the new ones. They differ in model, ledger and number of turns, and ran under another Claude login. Sonnet and Opus sessions ran interleaved. Same-model sessions were 13 to 27 seconds apart (calculation), rather than under one second. B and C also differ in ledger placement. The pooled counts are a rough check, not a controlled comparison.

What the folder costs you

Calculation

Session 1 (first in its setup)

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

Session 2

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

Session 3

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

Same call without a cache (every input token at the input price)

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

One panel per series, all on the same axis.

List-price calculation, not a run. 3 setups, 4 series: Session 1 (first in its setup), Session 2, Session 3, Same call without a cache (every input token at the input price). Session 1 (first in its setup): highest A: new folder each time $0.029 (n 1). Lowest B: fixed folder $0.026 (n 1). Session 2: highest A: new folder each time $0.026 (n 1). Lowest C: fixed folder, ledger in system prompt $0.0016 (n 1).

Notesn 1–3 per row

The recorded tokens of each call priced at list price; the last bar shows the mean of the 3 calls priced without any cache

Calculation, not a bill: the calls ran on a subscription. Uncached input at the input price, cache reads at the cache-read price, 1-hour cache writes at 2× the input price (every write in this run was a 1-hour write); Sonnet 5.5 list prices, effective 2026-09-21. A write costs more than plain input, so a session that writes the ledger again costs more than no cache at all.

Sources: Prompt cache across sessions, Cost with and without the prompt cache (calculation), Anthropic list prices (Claude models)

This is a list-price calculation from recorded tokens, not a bill. The calls ran on a subscription. At the study's Sonnet 5.5 prices, a cache read costs a tenth of the input price. A 1-hour cache write costs twice the input price. Every write in this run used that 1-hour tier.

  • Later session, fixed folder: mean $0.0016 for turn 1 across B and C (n = 4; range $0.001597 to $0.001601).
  • Later session, new folder: mean $0.0259 for turn 1 in A (n = 2; range $0.025871 to $0.025879). That is 16.2 times as much (calculation). It is about 1.6 times the mean $0.0157 without a cache (calculation on the same two A receipts; range $0.015732 to $0.015736).

Our total turn-1 input is about 7.8k tokens. At these list prices, writing more tokens increases the dollar cost (calculation). This does not predict cache matches or token use on another workload.

Did the cache make turn 1 faster?

  • Turn 1 (the ledger and question 1)
  • Turn 2 (question 2)
Entrance: medians race at 1.3× real timeMotion reduced: press Replay to animateThe slowest median is 1.8 s. The clock runs at the recorded speed.
A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

3 rows, 2 series: Turn 1 (the ledger and question 1), Turn 2 (question 2). Turn 1 (the ledger and question 1): slowest C: fixed folder, ledger in system prompt 1.8 s (range 1.7 s–1.8 s, n 3). Fastest A: new folder each time 1.6 s (range 1.3 s–1.6 s, n 3). Not all run ranges overlap. Turn 2 (question 2): slowest B: fixed folder 1.3 s (range 1.3 s–1.4 s, n 3). Fastest A: new folder each time 1 s (range 0.9 s–1.1 s, n 3). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Median of 3 sessions; whiskers = fastest and slowest of the 3

Whiskers are a range (fastest and slowest of 3 sessions), not a confidence interval. Turn 2 read the cache in every setup. This design does not isolate a cache effect on speed. Ledger seed and call order also differ. The Mac also ran other agent work during these calls.

Source: Prompt cache across sessions

This run does not isolate a speed effect. Median turn-1 time was 1.56 s for A (range 1.34 to 1.62 s), 1.76 s for B (1.71 to 1.86 s) and 1.77 s for C (1.69 to 1.78 s). Each setup has n = 3 sessions. Ranges are fastest to slowest, not confidence intervals. These medians include session 1, when each setup wrote its ledger.

Turn 2 read from the cache in all nine Claude sessions. Its medians were 1.02 s in A (range 0.88 to 1.14 s), 1.33 s in B (1.31 to 1.37 s) and 1.13 s in C (1.09 to 1.29 s). Each has n = 3. The Mac also ran other agent work. Folder, ledger and order changed together, so these times cannot show that caching has no speed benefit.

Codex CLI: no near-full turn-1 reads in this sample

  • Session 1 (first in its setup)
  • Session 2
  • Session 3
A: new folder each time
B: fixed folder

2 setups, 3 series: Session 1 (first in its setup), Session 2, Session 3. Session 1 (first in its setup): highest A: new folder each time 4,864 (n 1). Lowest B: fixed folder 0 (n 1). Session 2: all at 8,960.

Notesn = 1 per row

Turn-1 input was about 16,042 tokens in every call

Cached input as the Codex app-server reports it (input includes the cached tokens; it reports no cache writes). Codex sends a large system prompt of its own, so a cached count of about 9,000 can come from that prefix alone. Each bar is one call. Not comparable with the Claude Code bars: different prefix, different cache.

Source: Prompt cache across sessions

We ran A and B with GPT-6.1 Sol (medium) in Codex CLI 0.160.0. All 12 counted calls passed the normalized-answer checker (95% interval 76% to 100%). One probe call was not counted. Turn-1 input ranged from 16,033 to 16,053 tokens across six sessions. Codex adds a large system prompt of its own.

Across all six turn-1 calls, cached input ranged from 0 to 8,960 tokens. The highest read share was 55.9% (calculation). The 50% threshold in the surviving protocol counts 2 of 4 later sessions (95% interval 15% to 85%). A session 2 cached 8,960 of 16,051 tokens (55.8%, calculation). B session 2 cached 8,960 of 16,033 (55.9%, calculation).

The probe also showed 8,960 cached tokens. We interpret that count as Codex's system prefix, but did not test which tokens it held. The post-hoc 90% rule counts 0 of 4 later sessions (95% interval 0% to 49%). We cannot verify that the surviving protocol's 50% rule was recorded before the calls. No token-level trace identifies the ledger under either rule.

Turn 2 cached 15,872 tokens in all six sessions. Its read share ranged from 98.7% to 98.8% (calculation). Thus near-full reads appeared inside sessions, but not on turn 1 under the 90% rule. This small sample cannot show that folders never matter. Codex reports no cache-write count, so we calculate no cost for it. Do not rank Codex against Claude here: the prefixes and counters differ. More on the two CLIs: Claude Code vs Codex CLI.

What the conditions show, and what they do not

  • Observed: later fixed-folder Claude sessions had near-full cache reads. Later new-folder sessions had partial reads and wrote most input again. A and B differ in folder, ledger seed and call order. The observations do not prove a folder cause or a requirement to use one fixed folder.
  • Observed: both ledger placements had near-full reads with a fixed folder. Two later sessions per setup cannot establish equal reuse rates.
  • Not shown: whether the folder path, name or another property breaks the match. We did not inspect the outgoing request. The run notes identify --exclude-dynamic-system-prompt-sections as an option for moving per-machine sections into the first user message. We did not test it or verify its effect on these requests.
  • Not shown: whether a ledger in the system prompt survives a changing folder. Setup C ran only with a fixed folder.
  • Not shown: reuse after a long gap. The sessions ran less than a second apart, and the cache entries were 1-hour writes.

What this means for a script that starts a new session per task

  1. Test one fixed working folder when you can. In this sample, later fixed-folder sessions had near-full cache reads. That does not guarantee the same result for your script.
  2. Budget for possible cache writes with fresh folders. A's later sessions wrote most input again. Their partial reads did not make turn 1 cheaper than the no-cache calculation.
  3. Check your own receipts. Claude Code stream output reports cache read and cache write tokens. Use those counts to calculate list-price cost; subscription calls are not per-token bills.
  4. Test the CLI option before you rely on it. The run notes name a possible option to investigate. We have not tested it, so we do not claim it fixes reuse.
  5. Do not generalize past this setup. Tools were off, the folder was empty, and the shared prefix was small. A real repository adds text that can differ between sessions.

Caveats

  • Small n. There are two later sessions per setup and route. Per-setup Claude intervals overlap. Both pooled counts are post-hoc calculations.
  • Exploratory protocol. The retained protocol file was created after all counted calls. Separate pre-edit copies do not verify the claimed amendment times. The summary was regenerated later.
  • Correctness ceiling. All 30 answers passed after text normalization (95% Wilson interval 89% to 100%). The checker trims spaces, quotes, backticks, a final period and currency units. These short synthetic lookups cannot establish coding quality or general task success. Checker controls ran after the counted calls; cache measures do not depend on them.
  • One CLI version and one model per route: Claude Code 2.1.286 with Sonnet 5.5, and Codex CLI 0.160.0 with GPT-6.1 Sol.
  • Costs are calculations. They are list price times recorded tokens. The calls used flat subscriptions.
  • Timing has confounders. The Mac ran other agent work. Folder, ledger seed and call order differed.
  • The Codex criterion changed after inspection. The 90% rule gives 0 of 4 (95% interval 0% to 49%). The surviving protocol's 50% rule gives 2 of 4 (15% to 85%). Neither identifies the cached tokens.
  • The pooled new-folder count is not a matched sample. Earlier sessions used another login and ledger, with five turns and interleaved Sonnet and Opus calls. Their same-model gaps were 13 to 27 seconds (calculation).
  • The Codex allowance gate is not verified. The protocol lists conflicting entry thresholds, and no retained gate receipt proves the allowance at run time. The call caps were kept.
  • I build Agent, so I want agent runs to be cheap. The study page and sanitized receipts are public. We kept every counted attempt and retried nothing.

See your own cache share

Agent records cache reads and writes for every model call. You can check whether new sessions write most input again. Try Agent and check your own numbers.

The data behind this post

  • Prompt Caching
  • Cache Reuse

Does a new Claude Code session reuse the prompt cache of an earlier one?

30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.

0of 2 (95% interval 0% to 66%) · Later sessions with at least 50% of turn-1 input cached, A: new folder each time · n = 2

4 chartsUpdated October 7, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.