{"i":7,"study":{"slug":"caching-consistency","title":"Prompt caching and run-to-run consistency in Claude Code and Codex CLI","seoTitle":"Prompt caching savings and LLM consistency, measured","description":"135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.","question":"When a CLI session reuses a fixed context, how much input comes from the cache, what does that save at list price, and does it change latency? When the same prompt runs 10 times, how much do the pass rate, the answer and the time vary?","answer":"Caching: inside one Claude Code session, turns 2-5 read 97% of their input from the cache on average (turn 1: 19%, the CLI's own prefix). At list price, a calculation, all recorded turns cost Sonnet $0.1350 vs $0.2698 (50% less) and Opus $0.2551 vs $0.5442 (53% less) without the cache. Turn 1 costs more with the cache, because a 1-hour cache write costs twice the input price. A new session did not reuse the cache of an earlier one: on turn 1, all 4 later sessions wrote the ledger to the cache again. The cause was not tested. The cache showed no clear speed effect: median turn time was Sonnet 1.6 s on turn 1 vs 1.6 s on turns 2-5 and Opus 1.9 s on turn 1 vs 2.4 s on turns 2-5, and the fastest-to-slowest ranges overlap. Codex CLI (GPT-6.1 Sol) read 99% of later-turn input from its cache on a larger context; its app-server reports no cache writes, so no cost is calculated for it. Consistency: 7 of 9 model-and-prompt cells passed all 10 repetitions (95% interval 72% to 100%). Haiku passed 0/10 on the exact-number prompt; Haiku passed 1/10 on the JSON prompt (9 more were correct but in the wrong format). Haiku gave the same wrong answer every time (289; expected 282): consistent is not the same as correct. The code-fix prompt gave 6 different code bodies for Haiku, 3 different code bodies for Sonnet and 6 different code bodies for GPT-6.1 Sol (medium).","date":"2026-10-06","updated":"2026-10-06","tags":["prompt-caching","consistency","variance","claude-code","codex-cli","claude-haiku","claude-sonnet","claude-opus","gpt-6-1-sol","calculation"],"caveats":["Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.","Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.","Costs are list-price calculations; the calls used flat subscriptions.","Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.","The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.","Claude rows ran at default effort and GPT-6.1 Sol at medium; Codex CLI adds its own system prompt and tool schemas. Rows across routes compare route + model pairs."],"sourceIds":["agent-caching-consistency","calc-cache-pricing","price-anthropic"],"stats":{"$k":["id","label","value","unit","display","note","n"],"$r":[["caching-saving-sonnet","List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)",0.4996,"rate","50% ($0.1350 vs $0.2698)","Calculation from recorded tokens and list prices; not a bill.","\u0001"],["caching-saving-opus","List-price saving from the cache over 15 turns, Claude Opus 5.5 · Claude Code (calculation)",0.5313,"rate","53% ($0.2551 vs $0.5442)","Calculation from recorded tokens and list prices; not a bill.","\u0001"],["caching-cross-session-reuse","Later sessions whose first turn read the ledger from an earlier session’s cache",0,"count","0 of 4","\u0001",4],["consistency-perfect-cells","Model-and-prompt cells that passed all 10 repetitions",7,"count","7 of 9","\u0001",9],["caching-consistency-calls","Calls in this study (every one counted)",135,"calls","135 (45 cache turns, 90 repeated prompts)","\u0001","\u0001"]]},"charts":{"$k":["id","title","subtitle","kind","unit","xLabel","yLabel","series","note","sourceIds","whisker"],"$r":[["caching-read-share-by-turn","Share of input read from the cache, by turn in a session","Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each","line","rate","Turn in the session","Input tokens read from cache",{"$k":["name","points"],"$r":[["Claude Sonnet 5.5 · Claude Code",{"$k":["label","value","n"],"$r":[["Turn 1",0.1868,3],["Turn 2",0.9924,3],["Turn 3",0.9896,3],["Turn 4",0.992,3],["Turn 5",0.9074,3]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","value","n"],"$r":[["Turn 1",0.1869,3],["Turn 2",0.9924,3],["Turn 3",0.9866,3],["Turn 4",0.9893,3],["Turn 5",0.9031,3]]}],["GPT-6.1 Sol (medium) · Codex CLI",{"$k":["label","value","n"],"$r":[["Turn 1",0.5545,3],["Turn 2",0.9871,3],["Turn 3",0.9892,3],["Turn 4",0.9898,3],["Turn 5",0.9786,3]]}]]},"Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.",["agent-caching-consistency"],"\u0001"],["caching-cost-with-without","List-price cost of 5-question sessions with and without the cache (calculation)","All recorded turns per model; the same reported tokens priced two ways","grouped-bar","usd","\u0001","USD (list price)",[{"name":"With the cache, as recorded","points":[{"label":"Claude Sonnet 5.5 · Claude Code","value":0.135003,"n":15},{"label":"Claude Opus 5.5 · Claude Code","value":0.255057,"n":15}]},{"name":"Without a cache: every input token at the input price","points":[{"label":"Claude Sonnet 5.5 · Claude Code","value":0.269788,"n":15},{"label":"Claude Opus 5.5 · Claude Code","value":0.544228,"n":15}]}],"Calculation, not a bill: the calls ran on a subscription. Cache reads at the cache-read price, 1-hour cache writes at 2× the input price (every write in this run was a 1-hour write). Codex CLI is not priced here: it reports no cache-write count.",["agent-caching-consistency","calc-cache-pricing","price-anthropic"],"\u0001"],["caching-latency-first-vs-later","Time per turn: first turn vs later turns in a cached session","Median; whiskers = fastest and slowest turn","dot-range","seconds","\u0001","Seconds",[{"name":"Turn 1 (writes the ledger to the cache)","points":[{"label":"Claude Sonnet 5.5 · Claude Code","value":1.64,"lo":1.58,"hi":1.79,"n":3},{"label":"Claude Opus 5.5 · Claude Code","value":1.9,"lo":1.78,"hi":4.36,"n":3}]},{"name":"Turns 2-5 (read the ledger from the cache)","points":[{"label":"Claude Sonnet 5.5 · Claude Code","value":1.61,"lo":1.35,"hi":5.63,"n":12},{"label":"Claude Opus 5.5 · Claude Code","value":2.4,"lo":1.63,"hi":12.67,"n":12}]}],"Whiskers are a range (fastest and slowest turn), not a confidence interval. Turns ask different questions: the slow later turns are the counting question, which produced the most output.",["agent-caching-consistency"],"minmax"],["consistency-pass-rate","Same prompt, 10 times: strict pass rate","One series per prompt; whiskers are 95% Wilson intervals","dot-range","rate","\u0001","Passed",{"$k":["name","points"],"$r":[["Exact number",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",0,0,0.2775,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}],["JSON object",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",0.1,0.0179,0.4042,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}],["Code fix",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,0.7225,1,10],["Claude Sonnet 5.5 · Claude Code",1,0.7225,1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,0.7225,1,10]]}]]},"Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.",["agent-caching-consistency"],"ci95"],["consistency-distinct-answers","Same prompt, 10 times: how many different answers","Distinct normalized answers over 10 repetitions (1 = the same answer every time)","grouped-bar","count","\u0001","Distinct answers",{"$k":["name","points"],"$r":[["Exact number",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,10],["Claude Sonnet 5.5 · Claude Code",1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,10]]}],["JSON object",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,10],["Claude Sonnet 5.5 · Claude Code",1,10],["GPT-6.1 Sol (medium) · Codex CLI",1,10]]}],["Code fix",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 · Claude Code",6,10],["Claude Sonnet 5.5 · Claude Code",3,10],["GPT-6.1 Sol (medium) · Codex CLI",6,10]]}]]},"Normalized: the last number for the exact-number prompt; key-sorted JSON for the JSON prompt; for the code fix, the code without fences, comments, whitespace, quote style and line-end semicolons. One distinct answer is not the same as a correct answer: an answer can be the same and wrong every time. Different code can be equally correct.",["agent-caching-consistency"],"\u0001"],["consistency-latency-spread","Same prompt, 10 times: time per call","Median; whiskers = fastest and slowest of 10 calls","dot-range","seconds","\u0001","Seconds",{"$k":["name","points"],"$r":[["Exact number",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",5.06,4.42,6.2,10],["Claude Sonnet 5.5 · Claude Code",6.89,5.81,7.81,10],["GPT-6.1 Sol (medium) · Codex CLI",13.38,12.29,17.97,10]]}],["JSON object",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",7.03,5.28,12.27,10],["Claude Sonnet 5.5 · Claude Code",2.89,2.68,5.3,10],["GPT-6.1 Sol (medium) · Codex CLI",6.42,5.25,8.26,10]]}],["Code fix",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",5.95,4.89,7.33,10],["Claude Sonnet 5.5 · Claude Code",2.67,2.32,4.34,10],["GPT-6.1 Sol (medium) · Codex CLI",11.29,9.08,14.85,10]]}]]},"Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.",["agent-caching-consistency"],"minmax"]]},"related":["effort-ladder","cli-model-latency-tokens"]}}