{
  "schema": "agent-public-bench@1",
  "generatedAt": "2026-10-07T00:00:00.000Z",
  "url": "https://agent.sasid.ai/benchmarks/caching-consistency",
  "study": {
    "slug": "caching-consistency",
    "title": "Prompt caching and run-to-run consistency in Claude Code and Codex CLI",
    "seoTitle": "Prompt caching savings and LLM consistency, measured",
    "description": "135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.",
    "question": "When a CLI session reuses a fixed context, how much input comes from the cache, what does that save at list price, and does it change latency? When the same prompt runs 10 times, how much do the pass rate, the answer and the time vary?",
    "answer": "Caching: inside one Claude Code session, turns 2-5 read 97% of their input from the cache on average (turn 1: 19%, the CLI's own prefix). At list price, a calculation, all recorded turns cost Sonnet $0.1350 vs $0.2698 (50% less) and Opus $0.2551 vs $0.5442 (53% less) without the cache. Turn 1 costs more with the cache, because a 1-hour cache write costs twice the input price. A new session did not reuse the cache of an earlier one: on turn 1, all 4 later sessions wrote the ledger to the cache again. The cause was not tested. The cache showed no clear speed effect: median turn time was Sonnet 1.6 s on turn 1 vs 1.6 s on turns 2-5 and Opus 1.9 s on turn 1 vs 2.4 s on turns 2-5, and the fastest-to-slowest ranges overlap. Codex CLI (GPT-6.1 Sol) read 99% of later-turn input from its cache on a larger context; its app-server reports no cache writes, so no cost is calculated for it. Consistency: 7 of 9 model-and-prompt cells passed all 10 repetitions (95% interval 72% to 100%). Haiku passed 0/10 on the exact-number prompt; Haiku passed 1/10 on the JSON prompt (9 more were correct but in the wrong format). Haiku gave the same wrong answer every time (289; expected 282): consistent is not the same as correct. The code-fix prompt gave 6 different code bodies for Haiku, 3 different code bodies for Sonnet and 6 different code bodies for GPT-6.1 Sol (medium).",
    "date": "2026-10-06",
    "updated": "2026-10-06",
    "tags": [
      "prompt-caching",
      "consistency",
      "variance",
      "claude-code",
      "codex-cli",
      "claude-haiku",
      "claude-sonnet",
      "claude-opus",
      "gpt-6-1-sol",
      "calculation"
    ],
    "method": [
      "Protocols declared before the first call, one per route. Every attempt is kept; nothing was retried.",
      "Caching: 9 sessions (3 × Claude Sonnet 5.5 · Claude Code, 3 × Claude Opus 5.5 · Claude Code and 3 × GPT-6.1 Sol (medium) · Codex CLI), 5 turns each. A session is one CLI process. Turn 1 sends a seeded synthetic stock ledger plus question 1; turns 2-5 send one short question each (lookups, a count, an arg-max), each with one exact answer. Each turn is one model request.",
      "A 2-call probe sized the context before the run (not part of any cell). The ledger was larger than declared, so it was cut once, from 170 to 100 lines, and the answers were recomputed, as the protocol allowed.",
      "The Codex sessions ran before that cut, on the 170-line ledger (16,197 characters vs 9,651). The two routes are reported side by side, never as a like-for-like pair.",
      "Cache counters as the provider reports them: Claude Code gives uncached input, cache reads and cache writes (with the 5-minute and 1-hour split); the Codex app-server gives input (cached included) and cached input, and no cache writes.",
      "Cost with and without the cache (Claude only): see the calculation source. Every write in this run was a 1-hour write, so the 5-minute multiplier (an assumption) was not used.",
      "Consistency: 3 prompts (an exact number, a JSON object with exact keys, a small code fix), each with a deterministic validator; 10 repetitions per prompt for Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI. One-shot calls, one at a time per account. Answer diversity counts distinct normalized answers; the raw extract holds ordinal answer ids, never the text.",
      "Validator controls ran before inference on each route: every reference answer passes, every plausible wrong answer fails, and a wrapped reference is flagged as a format miss.",
      "Isolation as in the hard head-to-head: fresh empty working folder, tools off, no MCP servers, no session persistence across processes, provider-default caching. Claude at its default effort; GPT-6.1 Sol at medium.",
      "No batch stopped early and nothing was trimmed. The Claude CLI reported a rate-limit status of \"allowed_warning\" on 7 of 30 cache turns; no call was refused and the run did not stop."
    ],
    "caveats": [
      "Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.",
      "Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.",
      "Costs are list-price calculations; the calls used flat subscriptions.",
      "Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.",
      "The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.",
      "Claude rows ran at default effort and GPT-6.1 Sol at medium; Codex CLI adds its own system prompt and tool schemas. Rows across routes compare route + model pairs."
    ],
    "sourceIds": [
      "agent-caching-consistency",
      "calc-cache-pricing",
      "price-anthropic"
    ],
    "stats": [
      {
        "id": "caching-saving-sonnet",
        "label": "List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)",
        "value": 0.4996,
        "unit": "rate",
        "display": "50% ($0.1350 vs $0.2698)",
        "note": "Calculation from recorded tokens and list prices; not a bill."
      },
      {
        "id": "caching-saving-opus",
        "label": "List-price saving from the cache over 15 turns, Claude Opus 5.5 · Claude Code (calculation)",
        "value": 0.5313,
        "unit": "rate",
        "display": "53% ($0.2551 vs $0.5442)",
        "note": "Calculation from recorded tokens and list prices; not a bill."
      },
      {
        "id": "caching-cross-session-reuse",
        "label": "Later sessions whose first turn read the ledger from an earlier session’s cache",
        "value": 0,
        "unit": "count",
        "display": "0 of 4",
        "n": 4
      },
      {
        "id": "consistency-perfect-cells",
        "label": "Model-and-prompt cells that passed all 10 repetitions",
        "value": 7,
        "unit": "count",
        "display": "7 of 9",
        "n": 9
      },
      {
        "id": "caching-consistency-calls",
        "label": "Calls in this study (every one counted)",
        "value": 135,
        "unit": "calls",
        "display": "135 (45 cache turns, 90 repeated prompts)"
      }
    ],
    "charts": [
      {
        "id": "caching-read-share-by-turn",
        "title": "Share of input read from the cache, by turn in a session",
        "subtitle": "Mean over sessions; turn 1 sends the ledger, turns 2-5 send one short question each",
        "kind": "line",
        "unit": "rate",
        "xLabel": "Turn in the session",
        "yLabel": "Input tokens read from cache",
        "series": [
          {
            "name": "Claude Sonnet 5.5 · Claude Code",
            "points": [
              {
                "label": "Turn 1",
                "value": 0.1868,
                "n": 3
              },
              {
                "label": "Turn 2",
                "value": 0.9924,
                "n": 3
              },
              {
                "label": "Turn 3",
                "value": 0.9896,
                "n": 3
              },
              {
                "label": "Turn 4",
                "value": 0.992,
                "n": 3
              },
              {
                "label": "Turn 5",
                "value": 0.9074,
                "n": 3
              }
            ]
          },
          {
            "name": "Claude Opus 5.5 · Claude Code",
            "points": [
              {
                "label": "Turn 1",
                "value": 0.1869,
                "n": 3
              },
              {
                "label": "Turn 2",
                "value": 0.9924,
                "n": 3
              },
              {
                "label": "Turn 3",
                "value": 0.9866,
                "n": 3
              },
              {
                "label": "Turn 4",
                "value": 0.9893,
                "n": 3
              },
              {
                "label": "Turn 5",
                "value": 0.9031,
                "n": 3
              }
            ]
          },
          {
            "name": "GPT-6.1 Sol (medium) · Codex CLI",
            "points": [
              {
                "label": "Turn 1",
                "value": 0.5545,
                "n": 3
              },
              {
                "label": "Turn 2",
                "value": 0.9871,
                "n": 3
              },
              {
                "label": "Turn 3",
                "value": 0.9892,
                "n": 3
              },
              {
                "label": "Turn 4",
                "value": 0.9898,
                "n": 3
              },
              {
                "label": "Turn 5",
                "value": 0.9786,
                "n": 3
              }
            ]
          }
        ],
        "note": "Claude Code: cache reads ÷ (uncached input + cache reads + cache writes). Codex CLI: cached input ÷ input, as its app-server reports them; it reports no cache writes, and its input includes its own system prompt and tool schemas. The Codex sessions used the ledger before it was cut (16,197 characters vs 9,651), so the two routes are not a like-for-like pair. Measured shares, not pass rates.",
        "sourceIds": [
          "agent-caching-consistency"
        ]
      },
      {
        "id": "caching-cost-with-without",
        "title": "List-price cost of 5-question sessions with and without the cache (calculation)",
        "subtitle": "All recorded turns per model; the same reported tokens priced two ways",
        "kind": "grouped-bar",
        "unit": "usd",
        "yLabel": "USD (list price)",
        "series": [
          {
            "name": "With the cache, as recorded",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.135003,
                "n": 15
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.255057,
                "n": 15
              }
            ]
          },
          {
            "name": "Without a cache: every input token at the input price",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 0.269788,
                "n": 15
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.544228,
                "n": 15
              }
            ]
          }
        ],
        "note": "Calculation, not a bill: the calls ran on a subscription. Cache reads at the cache-read price, 1-hour cache writes at 2× the input price (every write in this run was a 1-hour write). Codex CLI is not priced here: it reports no cache-write count.",
        "sourceIds": [
          "agent-caching-consistency",
          "calc-cache-pricing",
          "price-anthropic"
        ]
      },
      {
        "id": "caching-latency-first-vs-later",
        "title": "Time per turn: first turn vs later turns in a cached session",
        "subtitle": "Median; whiskers = fastest and slowest turn",
        "kind": "dot-range",
        "unit": "seconds",
        "yLabel": "Seconds",
        "series": [
          {
            "name": "Turn 1 (writes the ledger to the cache)",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1.64,
                "lo": 1.58,
                "hi": 1.79,
                "n": 3
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 1.9,
                "lo": 1.78,
                "hi": 4.36,
                "n": 3
              }
            ]
          },
          {
            "name": "Turns 2-5 (read the ledger from the cache)",
            "points": [
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1.61,
                "lo": 1.35,
                "hi": 5.63,
                "n": 12
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 2.4,
                "lo": 1.63,
                "hi": 12.67,
                "n": 12
              }
            ]
          }
        ],
        "note": "Whiskers are a range (fastest and slowest turn), not a confidence interval. Turns ask different questions: the slow later turns are the counting question, which produced the most output.",
        "whisker": "minmax",
        "sourceIds": [
          "agent-caching-consistency"
        ]
      },
      {
        "id": "consistency-pass-rate",
        "title": "Same prompt, 10 times: strict pass rate",
        "subtitle": "One series per prompt; whiskers are 95% Wilson intervals",
        "kind": "dot-range",
        "unit": "rate",
        "yLabel": "Passed",
        "series": [
          {
            "name": "Exact number",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 0,
                "lo": 0,
                "hi": 0.2775,
                "n": 10
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1,
                "lo": 0.7225,
                "hi": 1,
                "n": 10
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 1,
                "lo": 0.7225,
                "hi": 1,
                "n": 10
              }
            ]
          },
          {
            "name": "JSON object",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 0.1,
                "lo": 0.0179,
                "hi": 0.4042,
                "n": 10
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1,
                "lo": 0.7225,
                "hi": 1,
                "n": 10
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 1,
                "lo": 0.7225,
                "hi": 1,
                "n": 10
              }
            ]
          },
          {
            "name": "Code fix",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 1,
                "lo": 0.7225,
                "hi": 1,
                "n": 10
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1,
                "lo": 0.7225,
                "hi": 1,
                "n": 10
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 1,
                "lo": 0.7225,
                "hi": 1,
                "n": 10
              }
            ]
          }
        ],
        "note": "Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.",
        "whisker": "ci95",
        "sourceIds": [
          "agent-caching-consistency"
        ]
      },
      {
        "id": "consistency-distinct-answers",
        "title": "Same prompt, 10 times: how many different answers",
        "subtitle": "Distinct normalized answers over 10 repetitions (1 = the same answer every time)",
        "kind": "grouped-bar",
        "unit": "count",
        "yLabel": "Distinct answers",
        "series": [
          {
            "name": "Exact number",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 1,
                "n": 10
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1,
                "n": 10
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 1,
                "n": 10
              }
            ]
          },
          {
            "name": "JSON object",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 1,
                "n": 10
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1,
                "n": 10
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 1,
                "n": 10
              }
            ]
          },
          {
            "name": "Code fix",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 6,
                "n": 10
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 3,
                "n": 10
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 6,
                "n": 10
              }
            ]
          }
        ],
        "note": "Normalized: the last number for the exact-number prompt; key-sorted JSON for the JSON prompt; for the code fix, the code without fences, comments, whitespace, quote style and line-end semicolons. One distinct answer is not the same as a correct answer: an answer can be the same and wrong every time. Different code can be equally correct.",
        "sourceIds": [
          "agent-caching-consistency"
        ]
      },
      {
        "id": "consistency-latency-spread",
        "title": "Same prompt, 10 times: time per call",
        "subtitle": "Median; whiskers = fastest and slowest of 10 calls",
        "kind": "dot-range",
        "unit": "seconds",
        "yLabel": "Seconds",
        "series": [
          {
            "name": "Exact number",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 5.06,
                "lo": 4.42,
                "hi": 6.2,
                "n": 10
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 6.89,
                "lo": 5.81,
                "hi": 7.81,
                "n": 10
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 13.38,
                "lo": 12.29,
                "hi": 17.97,
                "n": 10
              }
            ]
          },
          {
            "name": "JSON object",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 7.03,
                "lo": 5.28,
                "hi": 12.27,
                "n": 10
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 2.89,
                "lo": 2.68,
                "hi": 5.3,
                "n": 10
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 6.42,
                "lo": 5.25,
                "hi": 8.26,
                "n": 10
              }
            ]
          },
          {
            "name": "Code fix",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 5.95,
                "lo": 4.89,
                "hi": 7.33,
                "n": 10
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 2.67,
                "lo": 2.32,
                "hi": 4.34,
                "n": 10
              },
              {
                "label": "GPT-6.1 Sol (medium) · Codex CLI",
                "value": 11.29,
                "lo": 9.08,
                "hi": 14.85,
                "n": 10
              }
            ]
          }
        ],
        "note": "Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.",
        "whisker": "minmax",
        "sourceIds": [
          "agent-caching-consistency"
        ]
      }
    ],
    "tables": [
      {
        "id": "caching-turns",
        "title": "Cache counters per turn (mean over sessions)",
        "columns": [
          {
            "key": "config",
            "label": "Configuration",
            "unit": "text"
          },
          {
            "key": "turn",
            "label": "Turn",
            "unit": "count"
          },
          {
            "key": "sessions",
            "label": "Sessions",
            "unit": "count"
          },
          {
            "key": "correct",
            "label": "Exact answers",
            "unit": "text"
          },
          {
            "key": "input",
            "label": "Input tokens (all)",
            "unit": "tokens"
          },
          {
            "key": "read",
            "label": "Read from cache",
            "unit": "tokens"
          },
          {
            "key": "written",
            "label": "Written to cache",
            "unit": "text"
          },
          {
            "key": "share",
            "label": "Read share",
            "unit": "rate"
          },
          {
            "key": "medianTotal",
            "label": "Median time (s)",
            "unit": "seconds"
          },
          {
            "key": "withCache",
            "label": "USD with cache (calculation)",
            "unit": "usd"
          },
          {
            "key": "without",
            "label": "USD without cache (calculation)",
            "unit": "usd"
          }
        ],
        "rows": [
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "turn": 1,
            "sessions": 3,
            "correct": "3/3",
            "input": 7831,
            "read": 1463,
            "written": "6,366",
            "share": 0.1868,
            "medianTotal": 1.64,
            "withCache": 0.025792,
            "without": 0.015693
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "turn": 2,
            "sessions": 3,
            "correct": "3/3",
            "input": 7889,
            "read": 7829,
            "written": "58",
            "share": 0.9924,
            "medianTotal": 1.54,
            "withCache": 0.001862,
            "without": 0.015839
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "turn": 3,
            "sessions": 3,
            "correct": "3/3",
            "input": 7970,
            "read": 7887,
            "written": "81",
            "share": 0.9896,
            "medianTotal": 1.53,
            "withCache": 0.001955,
            "without": 0.015991
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "turn": 4,
            "sessions": 3,
            "correct": "3/3",
            "input": 8032,
            "read": 7968,
            "written": "62",
            "share": 0.992,
            "medianTotal": 5.2,
            "withCache": 0.009576,
            "without": 0.023795
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "turn": 5,
            "sessions": 3,
            "correct": "3/3",
            "input": 8861,
            "read": 8030,
            "written": "829",
            "share": 0.9074,
            "medianTotal": 1.77,
            "withCache": 0.005816,
            "without": 0.018613
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "turn": 1,
            "sessions": 3,
            "correct": "3/3",
            "input": 7828,
            "read": 1463,
            "written": "6,363",
            "share": 0.1869,
            "medianTotal": 1.9,
            "withCache": 0.051265,
            "without": 0.031372
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "turn": 2,
            "sessions": 3,
            "correct": "3/3",
            "input": 7886,
            "read": 7826,
            "written": "58",
            "share": 0.9924,
            "medianTotal": 1.63,
            "withCache": 0.002637,
            "without": 0.032144
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "turn": 3,
            "sessions": 3,
            "correct": "3/3",
            "input": 7991,
            "read": 7884,
            "written": "105",
            "share": 0.9866,
            "medianTotal": 1.81,
            "withCache": 0.002965,
            "without": 0.032504
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "turn": 4,
            "sessions": 3,
            "correct": "3/3",
            "input": 8075,
            "read": 7989,
            "written": "84",
            "share": 0.9893,
            "medianTotal": 6.09,
            "withCache": 0.018471,
            "without": 0.048493
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "turn": 5,
            "sessions": 3,
            "correct": "3/3",
            "input": 8941,
            "read": 8073,
            "written": "866",
            "share": 0.9031,
            "medianTotal": 2.09,
            "withCache": 0.009681,
            "without": 0.036896
          },
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "turn": 1,
            "sessions": 3,
            "correct": "3/3",
            "input": 18774,
            "read": 10411,
            "written": "not reported",
            "share": 0.5545,
            "medianTotal": 3.56,
            "withCache": null,
            "without": null
          },
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "turn": 2,
            "sessions": 3,
            "correct": "3/3",
            "input": 18803,
            "read": 18560,
            "written": "not reported",
            "share": 0.9871,
            "medianTotal": 2.39,
            "withCache": null,
            "without": null
          },
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "turn": 3,
            "sessions": 3,
            "correct": "3/3",
            "input": 18848,
            "read": 18645,
            "written": "not reported",
            "share": 0.9892,
            "medianTotal": 3.11,
            "withCache": null,
            "without": null
          },
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "turn": 4,
            "sessions": 3,
            "correct": "3/3",
            "input": 18880,
            "read": 18688,
            "written": "not reported",
            "share": 0.9898,
            "medianTotal": 5.97,
            "withCache": null,
            "without": null
          },
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "turn": 5,
            "sessions": 3,
            "correct": "3/3",
            "input": 19097,
            "read": 18688,
            "written": "not reported",
            "share": 0.9786,
            "medianTotal": 2.38,
            "withCache": null,
            "without": null
          }
        ]
      },
      {
        "id": "consistency-cells",
        "title": "Every consistency cell (10 repetitions each)",
        "columns": [
          {
            "key": "config",
            "label": "Configuration",
            "unit": "text"
          },
          {
            "key": "prompt",
            "label": "Prompt",
            "unit": "text"
          },
          {
            "key": "strict",
            "label": "Strict passes",
            "unit": "text"
          },
          {
            "key": "ci",
            "label": "95% interval",
            "unit": "text"
          },
          {
            "key": "formatMisses",
            "label": "Format misses",
            "unit": "count"
          },
          {
            "key": "wrong",
            "label": "Wrong answers",
            "unit": "count"
          },
          {
            "key": "distinct",
            "label": "Distinct answers",
            "unit": "count"
          },
          {
            "key": "distinctRaw",
            "label": "Distinct raw replies",
            "unit": "count"
          },
          {
            "key": "medianTotal",
            "label": "Median time (s)",
            "unit": "seconds"
          },
          {
            "key": "iqr",
            "label": "Interquartile range (s)",
            "unit": "seconds"
          },
          {
            "key": "cv",
            "label": "Coefficient of variation",
            "unit": "score"
          }
        ],
        "rows": [
          {
            "config": "Claude Haiku 4.5 · Claude Code",
            "prompt": "Exact number",
            "strict": "0/10",
            "ci": "0% to 28%",
            "formatMisses": 0,
            "wrong": 10,
            "distinct": 1,
            "distinctRaw": 10,
            "medianTotal": 5.06,
            "iqr": 1.12,
            "cv": 0.13
          },
          {
            "config": "Claude Haiku 4.5 · Claude Code",
            "prompt": "JSON object",
            "strict": "1/10",
            "ci": "2% to 40%",
            "formatMisses": 9,
            "wrong": 0,
            "distinct": 1,
            "distinctRaw": 3,
            "medianTotal": 7.03,
            "iqr": 2.21,
            "cv": 0.28
          },
          {
            "config": "Claude Haiku 4.5 · Claude Code",
            "prompt": "Code fix",
            "strict": "10/10",
            "ci": "72% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "distinct": 6,
            "distinctRaw": 6,
            "medianTotal": 5.95,
            "iqr": 1.09,
            "cv": 0.13
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "prompt": "Exact number",
            "strict": "10/10",
            "ci": "72% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "distinct": 1,
            "distinctRaw": 1,
            "medianTotal": 6.89,
            "iqr": 0.95,
            "cv": 0.09
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "prompt": "JSON object",
            "strict": "10/10",
            "ci": "72% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "distinct": 1,
            "distinctRaw": 1,
            "medianTotal": 2.89,
            "iqr": 0.54,
            "cv": 0.28
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "prompt": "Code fix",
            "strict": "10/10",
            "ci": "72% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "distinct": 3,
            "distinctRaw": 3,
            "medianTotal": 2.67,
            "iqr": 1.17,
            "cv": 0.26
          },
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "prompt": "Exact number",
            "strict": "10/10",
            "ci": "72% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "distinct": 1,
            "distinctRaw": 1,
            "medianTotal": 13.38,
            "iqr": 1.53,
            "cv": 0.14
          },
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "prompt": "JSON object",
            "strict": "10/10",
            "ci": "72% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "distinct": 1,
            "distinctRaw": 1,
            "medianTotal": 6.42,
            "iqr": 1.8,
            "cv": 0.17
          },
          {
            "config": "GPT-6.1 Sol (medium) · Codex CLI",
            "prompt": "Code fix",
            "strict": "10/10",
            "ci": "72% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "distinct": 6,
            "distinctRaw": 6,
            "medianTotal": 11.29,
            "iqr": 2.25,
            "cv": 0.15
          }
        ]
      }
    ],
    "related": [
      "effort-ladder",
      "cli-model-latency-tokens"
    ]
  },
  "sources": [
    {
      "id": "agent-caching-consistency",
      "title": "Caching sessions and repeated prompts (Claude Code and Codex CLI)",
      "kind": "run",
      "date": "2026-10-06",
      "note": "Part 1: 5-turn CLI sessions over a fixed synthetic ledger, with the cache counters each provider reports per turn. Part 2: three prompts with deterministic validators, 10 repetitions per model. Declared protocols, validator controls before inference, every attempt kept; answers are published as ordinal ids, never as text.",
      "data": [
        "/benchmarks/raw/caching-consistency/caching.json",
        "/benchmarks/raw/caching-consistency/consistency.json"
      ]
    },
    {
      "id": "price-anthropic",
      "title": "Anthropic list prices (Claude models)",
      "kind": "price-list",
      "date": "2026-09-21",
      "url": "https://platform.claude.com/docs/en/about-claude/pricing",
      "note": "Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price."
    },
    {
      "id": "calc-cache-pricing",
      "title": "Cost with and without the prompt cache (calculation)",
      "kind": "calculation",
      "date": "2026-10-06",
      "note": "Recorded tokens per turn × Anthropic list prices. With the cache: uncached input at the input price, cache reads at the cache-read price, 1-hour cache writes at twice the input price, 5-minute writes at 1.25 times (an assumption; none occurred). Without a cache: every input token at the input price. Output is priced the same in both. Not a bill.",
      "data": [
        "/benchmarks/raw/caching-consistency/caching.json"
      ]
    }
  ]
}
