• Prompt Caching
  • Cache Reuse
  • Claude Code
  • Codex CLI
  • Claude Sonnet
  • GPT-6.1 Sol
  • Working Folder
  • Calculation

Does a new Claude Code session reuse the prompt cache of an earlier one?

In the earlier caching study, a new Claude Code session did not read the cache that an earlier session wrote. Does a fixed working folder change that, and does putting the ledger in the system prompt help?

Published · 4 charts · Download the data or a carousel

0

95% CI 0%–66% · n = 2

Later sessions with at least 50% of turn-1 input cached, A: new folder each time · of 2 (95% interval 0% to 66%)

2

95% CI 34%–100% · n = 2

Later sessions with at least 50% of turn-1 input cached, B: fixed folder · of 2 (95% interval 34% to 100%)

The answer

Later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder. Setup A (new folder): 0 of 2 met the 50% rule, with a 95% interval of 0% to 66%. Setup B (fixed folder): 2 of 2, with a 95% interval of 34% to 100%. Setup C (fixed folder): 2 of 2, with a 95% interval of 34% to 100%. The per-setup intervals overlap. Later A sessions read 1,463 tokens and wrote 6,386 to 6,388 tokens. Later B sessions read 7,833 tokens and wrote 0 tokens. Later C sessions read 7,815 tokens and wrote 0 tokens. We interpret A’s reads as a shared CLI prefix. No token-level trace tests this reading. A post-hoc calculation pools both studies. New folder: 0 of 6, with a 95% interval of 0% to 39%. Fixed folder: 4 of 4, with a 95% interval of 51% to 100%. These intervals do not overlap. The earlier sessions used another login, ledger and turn count, with interleaved models. This pool is a rough check, not a controlled comparison. Mean later-session turn-1 cost at list price was $0.0259 in A and $0.0016 across B and C (calculation). The ratio is 16.2 (calculation). The cost stats show n and ranges. These are subscription calls, not bills. Turn-1 times do not isolate a cache effect. Each setup has n = 3 sessions. A: median 1.56 s, range 1.34 s to 1.62 s. B: median 1.76 s, range 1.71 s to 1.86 s. C: median 1.77 s, range 1.69 s to 1.78 s. Ranges are not confidence intervals. Codex turn-1 cached input ranged from 0 to 8,960 tokens across 6 calls. Its highest read share was 56% (calculation). The 50% rule counts 2 of 4 later sessions, with a 95% interval of 15% to 85%. The post-hoc 90% rule counts 0 of 4, with a 95% interval of 0% to 49%. Neither rule proves which tokens came from the ledger. A and B differ in folder, ledger seed and call order. These observations do not show that a fixed folder is necessary or that the folder caused the difference. The surviving protocol file dates from after the calls. Treat this analysis as exploratory.

Key numbers

2

Later sessions with at least 50% of turn-1 input cached, C: fixed folder, ledger in system prompt

of 2 (95% interval 34% to 100%) · 95% CI 34%–100% · n = 2

0

Later sessions with at least 50% of turn-1 input cached, new folders, pooled across studies

of 6 (95% interval 0% to 39%) · 95% CI 0%–39% · n = 6

4

Later sessions with at least 50% of turn-1 input cached, fixed folder, B and C pooled

of 4 (95% interval 51% to 100%) · 95% CI 51%–100% · n = 4

$0.0259

Turn-1 list-price cost of a later session, new folder each time (calculation)

n = 2

$0.0016

Turn-1 list-price cost of a later session, fixed folder (calculation)

n = 4

16.2×

Turn-1 cost of a later session: new folder as a multiple of fixed folder (calculation)

n = 6

2

Later Codex CLI sessions with at least 50% of turn-1 input cached (protocol threshold)

of 4 (95% interval 15% to 85%) · 95% CI 15%–85% · n = 4

0

Later Codex CLI sessions with at least 90% of turn-1 input cached (post-hoc rule) (setups A and B)

of 4 (95% interval 0% to 49%) · 95% CI 0%–49% · n = 4

30

Counted calls in this study (every one counted)

(18 Claude Code, 12 Codex CLI), plus 2 uncounted probe calls

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Calculation

Session 1 (first in its setup)

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

Session 2

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

Session 3

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

One panel per series, all on the same axis.

List-price calculation, not a run. 3 setups, 3 series: Session 1 (first in its setup), Session 2, Session 3. Session 1 (first in its setup): highest B: fixed folder 19% (n 1). Lowest A: new folder each time 6.8% (n 1). Session 2: highest B: fixed folder 100% (n 1). Lowest A: new folder each time 19% (n 1).

Notesn = 1 per row

One bar per call: turn 1 of one session; 3 sessions per setup, run back to back; read-share calculation

Session 1 was the first session to use its setup’s ledger. Sessions 2 and 3 used that same ledger. Different seeds prevent full ledger-prefix reuse between setups; shared CLI-prefix reads remain possible. We interpret session-1 reads as a shared CLI prefix; no token-level trace proves this. Read share (calculation) = cache reads ÷ (uncached input + cache reads + cache writes), as the provider reports them. Each bar is one call, not a rate; the counts per setup are in the table. 2 later sessions per setup is a small number.

Source: Prompt cache across sessions

Share card (PNG)
Calculation

Session 1 (first in its setup)

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

Session 2

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

Session 3

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

Same call without a cache (every input token at the input price)

A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

One panel per series, all on the same axis.

List-price calculation, not a run. 3 setups, 4 series: Session 1 (first in its setup), Session 2, Session 3, Same call without a cache (every input token at the input price). Session 1 (first in its setup): highest A: new folder each time $0.029 (n 1). Lowest B: fixed folder $0.026 (n 1). Session 2: highest A: new folder each time $0.026 (n 1). Lowest C: fixed folder, ledger in system prompt $0.0016 (n 1).

Notesn 1–3 per row

The recorded tokens of each call priced at list price; the last bar shows the mean of the 3 calls priced without any cache

Calculation, not a bill: the calls ran on a subscription. Uncached input at the input price, cache reads at the cache-read price, 1-hour cache writes at 2× the input price (every write in this run was a 1-hour write); Sonnet 5.5 list prices, effective 2026-09-21. A write costs more than plain input, so a session that writes the ledger again costs more than no cache at all.

Sources: Prompt cache across sessions, Cost with and without the prompt cache (calculation), Anthropic list prices (Claude models)

Share card (PNG)
  • Turn 1 (the ledger and question 1)
  • Turn 2 (question 2)
Entrance: medians race at 1.3× real timeMotion reduced: press Replay to animateThe slowest median is 1.8 s. The clock runs at the recorded speed.
A: new folder each time
B: fixed folder
C: fixed folder, ledger in system prompt

3 rows, 2 series: Turn 1 (the ledger and question 1), Turn 2 (question 2). Turn 1 (the ledger and question 1): slowest C: fixed folder, ledger in system prompt 1.8 s (range 1.7 s–1.8 s, n 3). Fastest A: new folder each time 1.6 s (range 1.3 s–1.6 s, n 3). Not all run ranges overlap. Turn 2 (question 2): slowest B: fixed folder 1.3 s (range 1.3 s–1.4 s, n 3). Fastest A: new folder each time 1 s (range 0.9 s–1.1 s, n 3). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Median of 3 sessions; whiskers = fastest and slowest of the 3

Whiskers are a range (fastest and slowest of 3 sessions), not a confidence interval. Turn 2 read the cache in every setup. This design does not isolate a cache effect on speed. Ledger seed and call order also differ. The Mac also ran other agent work during these calls.

Source: Prompt cache across sessions

Share card (PNG)
  • Session 1 (first in its setup)
  • Session 2
  • Session 3
A: new folder each time
B: fixed folder

2 setups, 3 series: Session 1 (first in its setup), Session 2, Session 3. Session 1 (first in its setup): highest A: new folder each time 4,864 (n 1). Lowest B: fixed folder 0 (n 1). Session 2: all at 8,960.

Notesn = 1 per row

Turn-1 input was about 16,042 tokens in every call

Cached input as the Codex app-server reports it (input includes the cached tokens; it reports no cache writes). Codex sends a large system prompt of its own, so a cached count of about 9,000 can come from that prefix alone. Each bar is one call. Not comparable with the Claude Code bars: different prefix, different cache.

Source: Prompt cache across sessions

Share card (PNG)

Tables

Every counted turn

Route and modelSetupSessionTurnSeconds since the previous session’s last callInput tokens (all)Read from cacheWritten to cacheRead shareTime (s)USD with cache (calculation)Answer passes after normalization
Claude Sonnet 5.5 · Claude CodeA: new folder each time11—7,85353173206.8%1.3 s$0.029yes
Claude Sonnet 5.5 · Claude CodeA: new folder each time12—7,9117,8515899%0.9 s$0.0019yes
Claude Sonnet 5.5 · Claude CodeA: new folder each time210.7 s7,8511,463638619%1.6 s$0.026yes
Claude Sonnet 5.5 · Claude CodeA: new folder each time22—7,9097,8495899%1.1 s$0.0019yes
Claude Sonnet 5.5 · Claude CodeA: new folder each time310.5 s7,8531,463638819%1.6 s$0.026yes
Claude Sonnet 5.5 · Claude CodeA: new folder each time32—7,9117,8515899%1 s$0.0019yes
Claude Sonnet 5.5 · Claude CodeB: fixed folder11—7,8351,463637019%1.9 s$0.026yes
Claude Sonnet 5.5 · Claude CodeB: fixed folder12—7,8937,8335899%1.3 s$0.0019yes
Claude Sonnet 5.5 · Claude CodeB: fixed folder210.5 s7,8357,8330100%1.7 s$0.0016yes
Claude Sonnet 5.5 · Claude CodeB: fixed folder22—7,8937,8910100%1.4 s$0.0016yes
Claude Sonnet 5.5 · Claude CodeB: fixed folder310.5 s7,8357,8330100%1.8 s$0.0016yes
Claude Sonnet 5.5 · Claude CodeB: fixed folder32—7,8937,8910100%1.3 s$0.0016yes
Claude Sonnet 5.5 · Claude CodeC: fixed folder, ledger in system prompt11—7,81754072756.9%1.7 s$0.029yes
Claude Sonnet 5.5 · Claude CodeC: fixed folder, ledger in system prompt12—7,8757,8155899%1.3 s$0.0018yes
Claude Sonnet 5.5 · Claude CodeC: fixed folder, ledger in system prompt210.6 s7,8177,8150100%1.8 s$0.0016yes
Claude Sonnet 5.5 · Claude CodeC: fixed folder, ledger in system prompt22—7,8757,8730100%1.1 s$0.0016yes
Claude Sonnet 5.5 · Claude CodeC: fixed folder, ledger in system prompt310.7 s7,8177,8150100%1.8 s$0.0016yes
Claude Sonnet 5.5 · Claude CodeC: fixed folder, ledger in system prompt32—7,8757,8730100%1.1 s$0.0016yes
GPT-6.1 Sol (medium) · Codex CLIA: new folder each time11—16,0494,864not reported30%2.8 s—yes
GPT-6.1 Sol (medium) · Codex CLIA: new folder each time12—16,07815,872not reported99%1.6 s—yes
GPT-6.1 Sol (medium) · Codex CLIA: new folder each time210.5 s16,0518,960not reported56%4 s—yes
GPT-6.1 Sol (medium) · Codex CLIA: new folder each time22—16,08015,872not reported99%1.7 s—yes
GPT-6.1 Sol (medium) · Codex CLIA: new folder each time310.6 s16,0530not reported0%4.1 s—yes
GPT-6.1 Sol (medium) · Codex CLIA: new folder each time32—16,08215,872not reported99%1.7 s—yes
GPT-6.1 Sol (medium) · Codex CLIB: fixed folder11—16,0330not reported0%2.6 s—yes
GPT-6.1 Sol (medium) · Codex CLIB: fixed folder12—16,06215,872not reported99%1.5 s—yes
GPT-6.1 Sol (medium) · Codex CLIB: fixed folder210.6 s16,0338,960not reported56%2.8 s—yes
GPT-6.1 Sol (medium) · Codex CLIB: fixed folder22—16,06215,872not reported99%1.6 s—yes
GPT-6.1 Sol (medium) · Codex CLIB: fixed folder310.6 s16,0330not reported0%3.9 s—yes
GPT-6.1 Sol (medium) · Codex CLIB: fixed folder32—16,06215,872not reported99%2.9 s—yes

Method

  1. The protocol claims a declaration at 00:32 UTC. Its surviving file has a birth time of 00:49:27 UTC. The first counted call started at 00:32:33 UTC. We cannot verify a pre-call protocol.
  2. The run kept every try and made no retries. The run log lists three amendments. They cover the post-hoc Codex rule, a post-hoc pooled calculation and text corrections after a check.
  3. Claude Code 2.1.286 ran Sonnet 5.5 at its default effort. Tools were off. No MCP servers. No session persistence across processes. Caching was at the provider default.
  4. A session is one CLI process with 2 turns. Turn 1 asks a quantity lookup over a seeded synthetic stock ledger. The ledger has 100 lines and about 9,600 characters. Turn 2 asks a warehouse lookup. Each question has one reference answer. The checker trims spaces, quotes, backticks, a final period and currency units before comparison.
  5. Each setup had 3 sessions with 2 turns each (6 calls). A used a new folder each time. B used a fixed folder. Both put the ledger in the first user message. C used a fixed folder and put the ledger in the system prompt.
  6. Each setup has its own ledger seed. Different seeds prevent full ledger-prefix reuse between setups. Shared CLI-prefix reads remain possible. Sessions of a setup ran back to back. The gap between one session's last call and the next session's first call was 508 to 720 ms.
  7. A later session is session 2 or 3. The protocol labels it reuse when turn 1 read at least half of its input from the cache. This is a counter-based rule; no token-level trace identifies the ledger. The cache counters are the provider’s: uncached input, cache reads and cache writes (with the 5-minute and 1-hour split).
  8. List-price cost is a calculation. It prices uncached input at the input price and reads at the cache-read price. It prices 1-hour writes at 2× the input price. The calls ran on a subscription, so nothing here is a bill.
  9. Codex CLI ran GPT-6.1 Sol at medium effort in setups A and B. Each setup had 3 sessions with 2 turns each. Tools, apps, plugins and web search were off. Each session used an ephemeral read-only thread. The app-server reports input and cached input, but no cache writes. Its own large system prompt could explain a cached count.
  10. The surviving protocol records a 50% threshold but dates from after the calls. This rule counts 2 of 4 later Codex sessions. The post-hoc 90% rule counts 0 of 4. The respective Wilson 95% intervals are 15% to 85% and 0% to 49%. The highest turn-1 read share was 56% (calculation). No token-level trace identifies the ledger.
  11. One uncounted probe call per route, with its own ledger seed, checked the driver before the counted calls. The answer checker is the one of the earlier study. Its control test (every reference answer passes, every planted wrong answer fails) ran after the counted calls, because no cache measure depends on it.
  12. No batch stopped early. Nothing was trimmed. The Claude CLI reported a rate-limit status of "allowed_warning" on 9 of 18 Claude turns. No call was refused and the run did not stop.
  13. The normalized-answer checker passed 18 of 18 Claude answers (95% interval 82% to 100%).
  14. The normalized-answer checker passed 12 of 12 Codex answers (95% interval 76% to 100%).

Caveats

  • The surviving protocol file was created after all counted calls. Its claimed 00:32 UTC declaration is not supported by its file birth time. Amendment 1 and 2 state 00:36 and 00:37 UTC, but separate pre-edit copies do not verify those times. The current summary was regenerated at 07:41 UTC. Treat the analysis as exploratory.
  • The synthetic ledger and short lookup questions reuse a designed task shape from the earlier study. All 30 answers passed after text normalization (95% interval 89% to 100%), so correctness hits a ceiling. This is a cache-counter probe, not evidence about coding quality or general task success.
  • The protocol gives conflicting Codex entry gates: 30% weekly allowance remaining for the optional half, but 20% before the batch. The gate script enforces 20%. No retained gate receipt proves the allowance at run time. Call caps were kept: 18 Claude calls and 12 Codex calls, plus one probe each.
  • The sample is small: 2 later sessions per setup and route. A count of 2 of 2 has a 95% interval of 34% to 100%, so the per-setup counts alone do not separate the setups. The token counts were stable: every later session in A read the same 1,463 tokens, and later sessions in B read 7,833 tokens each and those in C read 7,815 each. The pooled comparison across both studies is post hoc.
  • The 4 earlier sessions in the pooled new-folder count differ from this study's sessions. They ran under a different Claude login, with another ledger and 5 turns per session. Sonnet and Opus sessions ran interleaved, so same-model sessions were 13 to 27 seconds apart (calculation from the recorded call times), not under 1 second. The pooled counts are a rough check, not a controlled comparison.
  • Consecutive sessions of the same setup ran less than one second apart. The cache entries were 1-hour writes. This run says nothing about reuse after a longer gap or after an entry expires.
  • A and B differ in working folder, ledger seed and call order. We did not test whether the folder path, its name or another property of a new empty folder breaks the match. The CLI documents an option that moves its per-machine system-prompt sections (the working folder is one) into the first user message. We did not test it.
  • Setup C ran only with a fixed folder. We did not test whether a ledger in the system prompt protects the cache against a changing folder.
  • Tools were off and the working folder was empty. A real repository adds other session-specific text (for example git state). We did not test that.
  • Total turn-1 Claude input, including the ledger, is about 7.8k tokens. We interpret the 1,463 reads in later A sessions as the shared CLI prefix; this was not tested. We tested one CLI version and one model per route: Claude Code 2.1.286 with Sonnet 5.5, and Codex CLI 0.160.0 with GPT-6.1 Sol. Larger rewritten inputs cost more at these list prices (calculation); this does not predict another workload’s cache use.
  • Codex CLI reports no cache-write count and has a large system prompt of its own. Its turn-1 cached counts varied from 0 to 8,960 in each setup. Neither setup reached the post-hoc 90% threshold. The small sample cannot show that folders never matter. We did not test why. Do not rank Codex against Claude on these numbers.
  • The Codex criterion for a full read (90% or more of its input) was set after we saw the Codex counts (Amendment 1). Under the 50% threshold we recorded in the protocol, 2 of 4 later Codex sessions would count (8,960 tokens cached in each, the count the Codex probe call showed, which we read as the Codex system prefix). The 90% rule, set after we saw the counts, gives 0 of 4. Their Wilson 95% intervals are 15% to 85% and 0% to 49%, respectively.
  • Costs are list-price calculations. The calls used flat subscriptions. Turn times come from a Mac that also ran other agent work, so contention can add noise.

Sources

  • Prompt cache across sessions

    Our recorded runs ·

    Sanitized cache-session receipts: one CLI process per session, 2 turns each, a seeded synthetic ledger (a different seed per condition) and 2 short questions with exact answers. Conditions: A a new temporary working folder per session, B one fixed folder, C a fixed folder with the ledger in the system prompt (Claude Code only). The Codex app-server reports cached input only, so its rows have no cache-write count. Probes are single uncounted calls with their own ledger seed. The answers and the ledger itself are not published. Session 1 was the first use of each setup’s ledger; sessions 2 and 3 reused that ledger. Different seeds prevent full ledger-prefix reuse between setups, but shared CLI-prefix reads remain possible. Correctness uses the reference checker after it trims spaces, surrounding quotes and backticks, a final period and currency units. The surviving protocol file dates from after the counted calls; pre-call declaration is not verified.

    Raw data: cache-sessions/sessions.json

  • Cost with and without the prompt cache (calculation)

    Calculation ·

    Recorded tokens per turn × Anthropic list prices. With the cache: uncached input at the input price, cache reads at the cache-read price, 1-hour cache writes at twice the input price, 5-minute writes at 1.25 times (an assumption; none occurred). Without a cache: every input token at the input price. Output is priced the same in both. Not a bill.

    Raw data: caching-consistency/caching.json

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • Caching sessions and repeated prompts (Claude Code and Codex CLI)

    Our recorded runs ·

    Part 1: 5-turn CLI sessions over a fixed synthetic ledger, with the cache counters each provider reports per turn. Part 2: three prompts with deterministic validators, 10 repetitions per model. Declared protocols, validator controls before inference, every attempt kept; answers are published as ordinal ids, never as text.

    Raw data: caching-consistency/caching.json, caching-consistency/consistency.json

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Does a new Claude Code session reuse the prompt cache of an earlier one?”, updated October 7, 2026, https://agent.sasid.ai/benchmarks/prompt-cache-across-sessions.

More studies

All benchmarks
Live story
  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.