{"method":["The protocol claims a declaration at 00:32 UTC. Its surviving file has a birth time of 00:49:27 UTC. The first counted call started at 00:32:33 UTC. We cannot verify a pre-call protocol.","The run kept every try and made no retries. The run log lists three amendments. They cover the post-hoc Codex rule, a post-hoc pooled calculation and text corrections after a check.","Claude Code 2.1.286 ran Sonnet 5.5 at its default effort. Tools were off. No MCP servers. No session persistence across processes. Caching was at the provider default.","A session is one CLI process with 2 turns. Turn 1 asks a quantity lookup over a seeded synthetic stock ledger. The ledger has 100 lines and about 9,600 characters. Turn 2 asks a warehouse lookup. Each question has one reference answer. The checker trims spaces, quotes, backticks, a final period and currency units before comparison.","Each setup had 3 sessions with 2 turns each (6 calls). A used a new folder each time. B used a fixed folder. Both put the ledger in the first user message. C used a fixed folder and put the ledger in the system prompt.","Each setup has its own ledger seed. Different seeds prevent full ledger-prefix reuse between setups. Shared CLI-prefix reads remain possible. Sessions of a setup ran back to back. The gap between one session's last call and the next session's first call was 508 to 720 ms.","A later session is session 2 or 3. The protocol labels it reuse when turn 1 read at least half of its input from the cache. This is a counter-based rule; no token-level trace identifies the ledger. The cache counters are the provider’s: uncached input, cache reads and cache writes (with the 5-minute and 1-hour split).","List-price cost is a calculation. It prices uncached input at the input price and reads at the cache-read price. It prices 1-hour writes at 2× the input price. The calls ran on a subscription, so nothing here is a bill.","Codex CLI ran GPT-6.1 Sol at medium effort in setups A and B. Each setup had 3 sessions with 2 turns each. Tools, apps, plugins and web search were off. Each session used an ephemeral read-only thread. The app-server reports input and cached input, but no cache writes. Its own large system prompt could explain a cached count.","The surviving protocol records a 50% threshold but dates from after the calls. This rule counts 2 of 4 later Codex sessions. The post-hoc 90% rule counts 0 of 4. The respective Wilson 95% intervals are 15% to 85% and 0% to 49%. The highest turn-1 read share was 56% (calculation). No token-level trace identifies the ledger.","One uncounted probe call per route, with its own ledger seed, checked the driver before the counted calls. The answer checker is the one of the earlier study. Its control test (every reference answer passes, every planted wrong answer fails) ran after the counted calls, because no cache measure depends on it.","No batch stopped early. Nothing was trimmed. The Claude CLI reported a rate-limit status of \"allowed_warning\" on 9 of 18 Claude turns. No call was refused and the run did not stop.","The normalized-answer checker passed 18 of 18 Claude answers (95% interval 82% to 100%).","The normalized-answer checker passed 12 of 12 Codex answers (95% interval 76% to 100%)."]}