{
  "schema": "agent-public-bench@1",
  "generatedAt": "2026-10-07T00:00:00.000Z",
  "url": "https://agent.sasid.ai/benchmarks/llm-speed-anatomy",
  "study": {
    "slug": "llm-speed-anatomy",
    "title": "Where the seconds go: first text, output speed and prompt size for 6 LLMs",
    "seoTitle": "LLM speed: first text and tokens per second through CLIs",
    "description": "60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.",
    "question": "For 6 models run through their own coding CLIs (Haiku, Sonnet, Opus, Fable, Sol (low) and Luna (low)), how long until the first text, how fast does text stream after it, and what does a longer prompt add?",
    "answer": "We measured CLI start-up, first text and total time on one 250-line answer and a ledger lookup at three prompt sizes. There were 60 counted calls: 4 per model in part A and 3 per size in part B. The receipts do not separate provider queueing, prompt processing and reasoning time. Time to first text, median (range; n calls): Sonnet 2.0 s (0.9 to 4.1; n = 4). Opus 2.0 s (1.7 to 2.4; n = 4). Luna (low) 3.3 s (3.2 to 3.5; n = 4). Sol (low) 3.5 s (2.7 to 4.4; n = 4). Haiku 4.0 s (2.8 to 6.4; n = 4). Fable 4.4 s (2.3 to 4.6; n = 4). Opus’s slowest call (2.4 s) was faster than the fastest call of Luna (low) (3.2 s), Sol (low) (2.7 s) and Haiku (2.8 s). Sonnet and Fable overlap it. There are 4 calls per model. The samples are below the protocol’s five-call minimum for naming a speed winner. Characters per second is a calculation: correct reply length ÷ time from first text to call end. Median (range; n exact replies): Haiku 547 (546 to 548; n = 3). Luna (low) 524 (225 to 1,052; n = 4). Sonnet 517 (513 to 519; n = 4). Opus 347 (345 to 349; n = 4). Sol (low) 323 (291 to 327; n = 4). Fable 273 (270 to 293; n = 4). Haiku and Fable have non-overlapping observed speed ranges (3 and 4 exact replies). Both samples are below the protocol’s five-call minimum for naming a speed winner. Luna (low)’s 4 exact replies ran from 225 to 1,052 characters per second, so its median hides this large spread. These few calls cannot establish two speed groups. Token rates use different token units across models. The same exact reply has 4,327 characters. Median visible tokens: 1,066 for Sol (low) (n = 4) and Luna (low) (n = 4). 1,210 for Haiku (n = 3). 1,941 for Sonnet (n = 4), Opus (n = 4) and Fable (n = 4). Visible tokens per second (calculation), median (range; n calls): Sonnet 232 (230 to 233; n = 4). Opus 156 (155 to 156; n = 4). Haiku 153 (153 to 216; n = 4). Luna (low) 129 (55 to 259; n = 4). Fable 123 (121 to 131; n = 4). Sol (low) 80 (72 to 80; n = 4). These ranges show individual calls, not confidence intervals. Of the 24 part A replies, 23 matched all 250 lines exactly; 1 was wrong (Haiku, 288 lines), and every call stays in the timings. First-text medians differ by prompt size (calculation: difference of medians, 64k minus 1k): Haiku +0.9 s (1.9 s to 2.8 s). Sonnet +1.6 s (1.4 s to 3.1 s, ranges overlap). Opus +0.3 s (1.5 s to 1.8 s, ranges overlap). Sol (low) +0.6 s (3.4 s to 3.9 s, ranges overlap). Exact lookup answers (95% Wilson intervals): Haiku 9/9 (70% to 100%). Sonnet 9/9 (70% to 100%). Opus 5/9 (27% to 81%). Sol (low) 9/9 (70% to 100%). All 4 models’ intervals overlap; this sample cannot rank their lookup rates. 4 Opus replies echoed the ledger line and then gave the right number last. The declared scoring counts these as wrong answers. The table shows the post-hoc reading. Cache reads did not grow with the ledger for any model. This fits no ledger reuse, but the counts do not prove which text the provider cached.",
    "date": "2026-10-07",
    "updated": "2026-10-07",
    "tags": [
      "latency",
      "time-to-first-token",
      "tokens-per-second",
      "output-speed",
      "prompt-size",
      "claude-code",
      "codex-cli",
      "claude-haiku",
      "claude-sonnet",
      "claude-opus",
      "claude-fable",
      "gpt-6-1-sol",
      "gpt-6-luna"
    ],
    "method": [
      "The protocol file birth precedes the first probe and counted call. Later edits have no frozen versions. Two probes checked the routes and helped calibrate ledger sizes. They stay outside all cells and charts.",
      "Part A used 24 calls: 4 per model. The prompt asked for 250 numbers in words, one per line. The check compares the reply with all 250 expected lines. Models: Haiku, Sonnet, Opus, Fable, Sol (low) and Luna (low).\n\nControls ran before the first counted call. The reference passes. All 8/8 wrapped references are format misses. All 11/11 planted wrong answers fail.",
      "Part B used 36 calls: 3 per size for each model. Models: Haiku, Sonnet, Opus and Sol (low). Each call asks one exact lookup question about a synthetic stock ledger. Ledger sizes: 18 lines for 1k, 327 lines for 16k and 1,315 lines for 64k.\n\nThe 1k, 16k and 64k labels are approximate targets from probe calibration calculations, using estimated Haiku prefix counts. A new seed changes each ledger. Some fixed text repeats. Cache counts do not identify cached text.",
      "First text means the first non-empty text after CLI start. The time includes CLI start-up and work before the first word. The table also subtracts the CLI ready time (calculation).",
      "Output speed is a calculation. Visible tokens equal output tokens minus reported reasoning tokens. When the CLI reports no reasoning count, the calculation uses zero. That does not prove the model did no reasoning. Divide visible tokens by the time from first text to call end. \n\nThe plain form divides all output tokens by that time. It can overstate visible speed when reasoning comes first.",
      "The harness uses the hard head-to-head machinery. It disables tools, MCP servers and saved sessions. Claude Code uses default effort; Codex CLI uses low effort. Calls run one at a time on one Mac.\n\nThe timeout was 300 s. Claude Code had a 16,000-token output cap. Codex had no output-token cap. CLI versions: 2.1.286 (Claude Code) and codex-cli 0.160.0.",
      "Stop rules require a stop at the first usage-limit text. The gate checks other study markers before each batch. No stored batch stopped or trimmed a cell. No counted call duplicates another. \n\nA sequence process ended at the gate. Codex part B ran later. Every counted call completed.",
      "Cache-read share is a calculation: cache-read tokens divided by reported input tokens for each call. The cache table shows the median, observed range and n for each cell."
    ],
    "caveats": [
      "4 calls per model in part A and 3 per size in part B. Medians of so few calls move with one slow call, and the ranges are not confidence intervals. The 95% Wilson intervals on the lookup rates are wide.",
      "First text includes CLI start-up and reasoning time. It is not the API’s time to first token; a direct API call would skip the CLI start-up (the CLI reported ready after a median 0.33 s to 0.84 s per cell).",
      "Claude Code ran models at their default effort; Codex CLI ran at low effort. Effort levels are not one scale, and a CLI adds its own system prompt, so a row mixes a model and its CLI.",
      "Output speed is a calculation from reported token counts, the length of the reply and the clock. It depends on how each CLI counts reasoning, and part A has one prompt: other text, other days or an API route can differ.",
      "All 36 counted ledger prompts differ. Their input counts do not compare tokenizers on the same text. The size calibration uses prefix estimates from earlier short prompts. Those calculations do not measure the prefix in each counted call.",
      "The tasks hit a ceiling. Part A passed 23/24. In part B, 3 of 4 models passed every lookup. See the rate stats and chart for 95% intervals. This small task set cannot rank general capability.",
      "One shared Mac: the gate checked other study markers before each batch. No full host-load record proves that all other work stopped.",
      "Prompt sizes were calibrated after two probes. The cases are synthetic and tuned to this measurement, not a sample of real workloads.",
      "Amendment 3 says 06:05 UTC and claims to precede Codex part B. Those calls ran at 06:03:43 to 06:04:19 UTC. The same-file amendment is retrospective; its date does not prove advance declaration.",
      "The controls receipt was overwritten. The surviving checks ran after the probes and before counted calls; an earlier control run cannot be audited.",
      "Each size used different ledger text. Prompt-size differences also include content and provider-load changes; they do not isolate a cause.",
      "The failed Haiku reply contained tool-shaped text and extra output. No structured tool event was recorded; it stays a strict failure and remains in timings.",
      "One Mac, one network, two sessions on one night. Provider load changes timings from hour to hour. The Codex CLI part B batch ran about 4 hours after the batches before it (a run script stopped while it waited for another run to finish timing), so its timings come from a later hour."
    ],
    "sourceIds": [
      "agent-speed-anatomy"
    ],
    "hero": {
      "statIds": [
        "speed-anatomy-token-count-ratio"
      ]
    },
    "stats": [
      {
        "id": "speed-anatomy-calls",
        "label": "Counted calls in the speed anatomy study",
        "value": 60,
        "unit": "calls",
        "display": "60 calls (24 in part A, 36 in part B; 43 through Claude Code, 17 through Codex CLI)",
        "n": 60
      },
      {
        "id": "speed-anatomy-part-a-exact",
        "label": "Part A replies that matched all 250 lines exactly (strict)",
        "value": 0.9583,
        "unit": "rate",
        "display": "96% (23/24)",
        "n": 24,
        "ci": [
          0.7976,
          0.9926
        ],
        "note": "0 more were right but wrapped or laid out differently (format misses); 1 was wrong."
      },
      {
        "id": "speed-anatomy-lookup-exact",
        "label": "Part B lookups answered exactly (strict)",
        "value": 0.8889,
        "unit": "rate",
        "display": "89% (32/36)",
        "n": 36,
        "ci": [
          0.7469,
          0.9559
        ],
        "note": "0 more were right with extra words (format misses); 4 were wrong."
      },
      {
        "id": "speed-anatomy-token-count-ratio",
        "label": "The same 4,327-character reply: most tokens ÷ fewest tokens across models (calculation)",
        "value": 1.8,
        "unit": "ratio",
        "display": "1.8x",
        "n": 23,
        "note": "Median visible tokens of exact replies per model (observed token range; n): Haiku 1,210 (1,210 to 1,210; n = 3). Sonnet 1,941 (1,941 to 1,941; n = 4). Opus 1,941 (1,941 to 1,941; n = 4). Fable 1,941 (1,941 to 1,941; n = 4). Sol (low) 1,066 (1,066 to 1,066; n = 4). Luna (low) 1,066 (1,066 to 1,066; n = 4). Each vendor counts the same text differently, so tokens per second do not compare across vendors. A calculation, not a run."
      },
      {
        "id": "speed-anatomy-lookup-last-line",
        "label": "Part B replies whose last line was exactly the right number (post-hoc reading, not a pass)",
        "value": 1,
        "unit": "rate",
        "display": "100% (36/36)",
        "n": 36,
        "ci": [
          0.9036,
          1
        ],
        "note": "A reading declared after the Claude part B batch (protocol Amendment 2), derived from the stored replies. The strict score above stays the headline; this one is never counted as a pass or a format miss."
      },
      {
        "id": "speed-anatomy-cache-reads-grew",
        "label": "Models whose cache reads grew with the ledger size",
        "value": 0,
        "unit": "count",
        "display": "0 of 4",
        "n": 36,
        "note": "A new ledger for every call. The descriptive threshold is the 1k-prompt maximum plus 10% and 128 tokens. Growth does not prove ledger reuse; the cache table lists every cell."
      },
      {
        "id": "speed-anatomy-size-cost-haiku",
        "label": "Haiku: extra time to first text at 64k vs 1k (calculation)",
        "value": 0.86,
        "unit": "seconds",
        "display": "+0.9 s",
        "n": 6,
        "note": "Difference of two medians (2.78 s minus 1.93 s); the ranges are 1.85 to 2.04 s and 2.45 to 2.89 s. A calculation, not a run. Each call used separately seeded text. The table shows reported input counts, not a matched-text tokenizer comparison."
      },
      {
        "id": "speed-anatomy-size-cost-sonnet",
        "label": "Sonnet: extra time to first text at 64k vs 1k (calculation)",
        "value": 1.62,
        "unit": "seconds",
        "display": "+1.6 s",
        "n": 6,
        "note": "Difference of two medians (3.07 s minus 1.45 s); the ranges are 1.23 to 1.72 s and 1.38 to 3.61 s and overlap. A calculation, not a run. Each call used separately seeded text. The table shows reported input counts, not a matched-text tokenizer comparison."
      },
      {
        "id": "speed-anatomy-size-cost-opus",
        "label": "Opus: extra time to first text at 64k vs 1k (calculation)",
        "value": 0.28,
        "unit": "seconds",
        "display": "+0.3 s",
        "n": 6,
        "note": "Difference of two medians (1.79 s minus 1.51 s); the ranges are 1.46 to 2.01 s and 1.72 to 3.72 s and overlap. A calculation, not a run. Each call used separately seeded text. The table shows reported input counts, not a matched-text tokenizer comparison."
      },
      {
        "id": "speed-anatomy-size-cost-sol-low",
        "label": "Sol (low): extra time to first text at 64k vs 1k (calculation)",
        "value": 0.57,
        "unit": "seconds",
        "display": "+0.6 s",
        "n": 6,
        "note": "Difference of two medians (3.93 s minus 3.36 s); the ranges are 3.36 to 4.75 s and 3.42 to 4.38 s and overlap. A calculation, not a run. Each call used separately seeded text. The table shows reported input counts, not a matched-text tokenizer comparison."
      }
    ],
    "charts": [
      {
        "id": "speed-anatomy-first-text",
        "title": "Time to first text: a 250-line answer, six models",
        "subtitle": "Median of 4 calls per model; whiskers = fastest and slowest call",
        "kind": "dot-range",
        "unit": "seconds",
        "yLabel": "Seconds to first text",
        "series": [
          {
            "name": "Time to first text",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 4,
                "lo": 2.84,
                "hi": 6.38,
                "n": 4
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1.96,
                "lo": 0.88,
                "hi": 4.09,
                "n": 4
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 1.97,
                "lo": 1.7,
                "hi": 2.35,
                "n": 4
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 4.43,
                "lo": 2.27,
                "hi": 4.64,
                "n": 4
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 3.52,
                "lo": 2.75,
                "hi": 4.42,
                "n": 4
              },
              {
                "label": "GPT-6 Luna (low) · Codex CLI",
                "value": 3.3,
                "lo": 3.19,
                "hi": 3.47,
                "n": 4
              }
            ]
          }
        ],
        "note": "Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.",
        "whisker": "minmax",
        "sourceIds": [
          "agent-speed-anatomy"
        ]
      },
      {
        "id": "speed-anatomy-output-speed",
        "title": "Output speed after the first text: visible tokens per second (calculation)",
        "subtitle": "Median of 4 calls per model; whiskers = slowest and fastest call",
        "kind": "dot-range",
        "unit": "tokens",
        "yLabel": "Visible tokens per second",
        "series": [
          {
            "name": "Visible tokens per second",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 153.2,
                "lo": 152.6,
                "hi": 216.1,
                "n": 4
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 231.7,
                "lo": 230.3,
                "hi": 233,
                "n": 4
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 155.5,
                "lo": 154.6,
                "hi": 156.4,
                "n": 4
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 122.6,
                "lo": 120.9,
                "hi": 131.4,
                "n": 4
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 79.6,
                "lo": 71.6,
                "hi": 80.5,
                "n": 4
              },
              {
                "label": "GPT-6 Luna (low) · Codex CLI",
                "value": 129.1,
                "lo": 55.5,
                "hi": 259.1,
                "n": 4
              }
            ]
          }
        ],
        "note": "Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.",
        "whisker": "minmax",
        "sourceIds": [
          "agent-speed-anatomy"
        ]
      },
      {
        "id": "speed-anatomy-chars-per-second",
        "title": "Output speed in characters per second after the first text (calculation)",
        "subtitle": "Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call",
        "kind": "dot-range",
        "unit": "count",
        "yLabel": "Characters per second",
        "series": [
          {
            "name": "Characters per second",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 547,
                "lo": 546,
                "hi": 548,
                "n": 3
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 517,
                "lo": 513,
                "hi": 519,
                "n": 4
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 347,
                "lo": 345,
                "hi": 349,
                "n": 4
              },
              {
                "label": "Claude Fable 5.1 · Claude Code",
                "value": 273,
                "lo": 270,
                "hi": 293,
                "n": 4
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 323,
                "lo": 291,
                "hi": 327,
                "n": 4
              },
              {
                "label": "GPT-6 Luna (low) · Codex CLI",
                "value": 524,
                "lo": 225,
                "hi": 1052,
                "n": 4
              }
            ]
          }
        ],
        "note": "Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.",
        "whisker": "minmax",
        "sourceIds": [
          "agent-speed-anatomy"
        ]
      },
      {
        "id": "speed-anatomy-prompt-size",
        "title": "Time to first text as the prompt grows",
        "subtitle": "Median of 3 calls per size; whiskers = fastest and slowest call",
        "kind": "line",
        "unit": "seconds",
        "xLabel": "Prompt-size target (approximate Haiku tokens; calibration calculation)",
        "yLabel": "Seconds to first text",
        "series": [
          {
            "name": "Claude Haiku 4.5 · Claude Code",
            "points": [
              {
                "label": "1k",
                "value": 1.93,
                "lo": 1.85,
                "hi": 2.04,
                "n": 3
              },
              {
                "label": "16k",
                "value": 2.27,
                "lo": 2.22,
                "hi": 2.47,
                "n": 3
              },
              {
                "label": "64k",
                "value": 2.78,
                "lo": 2.45,
                "hi": 2.89,
                "n": 3
              }
            ]
          },
          {
            "name": "Claude Sonnet 5.5 · Claude Code",
            "points": [
              {
                "label": "1k",
                "value": 1.45,
                "lo": 1.23,
                "hi": 1.72,
                "n": 3
              },
              {
                "label": "16k",
                "value": 1.78,
                "lo": 1.64,
                "hi": 2.11,
                "n": 3
              },
              {
                "label": "64k",
                "value": 3.07,
                "lo": 1.38,
                "hi": 3.61,
                "n": 3
              }
            ]
          },
          {
            "name": "Claude Opus 5.5 · Claude Code",
            "points": [
              {
                "label": "1k",
                "value": 1.51,
                "lo": 1.46,
                "hi": 2.01,
                "n": 3
              },
              {
                "label": "16k",
                "value": 1.74,
                "lo": 1.7,
                "hi": 2.97,
                "n": 3
              },
              {
                "label": "64k",
                "value": 1.79,
                "lo": 1.72,
                "hi": 3.72,
                "n": 3
              }
            ]
          },
          {
            "name": "GPT-6.1 Sol (low) · Codex CLI",
            "points": [
              {
                "label": "1k",
                "value": 3.36,
                "lo": 3.36,
                "hi": 4.75,
                "n": 3
              },
              {
                "label": "16k",
                "value": 4.02,
                "lo": 3.3,
                "hi": 4.28,
                "n": 3
              },
              {
                "label": "64k",
                "value": 3.93,
                "lo": 3.42,
                "hi": 4.38,
                "n": 3
              }
            ]
          }
        ],
        "note": "Each call used a new ledger seed. Cache-read counts stayed within the short-prompt baseline (see the cache table). This does not identify which tokens were cached. Sizes name the text we send; each model’s reported input tokens are in the table and include the CLI’s own prefix. The size calibration subtracts estimated prefixes from probe input counts; these are calculations, not measured prefix counts for each call. Whiskers are a range of calls, not a confidence interval.",
        "whisker": "minmax",
        "sourceIds": [
          "agent-speed-anatomy"
        ]
      },
      {
        "id": "speed-anatomy-total-by-size",
        "title": "Total time per call by prompt size",
        "subtitle": "Median of 3 calls per bar; whiskers = fastest and slowest call",
        "kind": "grouped-bar",
        "unit": "seconds",
        "yLabel": "Seconds, whole call",
        "series": [
          {
            "name": "1k prompt",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 2.34,
                "lo": 2.22,
                "hi": 2.46,
                "n": 3
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1.78,
                "lo": 1.57,
                "hi": 2.12,
                "n": 3
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 1.83,
                "lo": 1.82,
                "hi": 2.41,
                "n": 3
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 3.43,
                "lo": 3.43,
                "hi": 4.92,
                "n": 3
              }
            ]
          },
          {
            "name": "16k prompt",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 2.79,
                "lo": 2.58,
                "hi": 2.84,
                "n": 3
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 2.1,
                "lo": 1.98,
                "hi": 2.48,
                "n": 3
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 2.36,
                "lo": 2.11,
                "hi": 3.4,
                "n": 3
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 4.14,
                "lo": 3.96,
                "hi": 4.68,
                "n": 3
              }
            ]
          },
          {
            "name": "64k prompt",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 3.13,
                "lo": 2.84,
                "hi": 3.28,
                "n": 3
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 3.44,
                "lo": 1.74,
                "hi": 4.38,
                "n": 3
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 2.35,
                "lo": 2.26,
                "hi": 4.29,
                "n": 3
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 3.96,
                "lo": 3.47,
                "hi": 4.44,
                "n": 3
              }
            ]
          }
        ],
        "note": "Whole call: CLI start-up, first text and the one-line answer. Whiskers are a range of calls, not a confidence interval. Each call used a new ledger.",
        "whisker": "minmax",
        "sourceIds": [
          "agent-speed-anatomy"
        ]
      },
      {
        "id": "speed-anatomy-lookup-correct",
        "title": "Exact lookup answers at the 1k, 16k and 64k prompt-size targets",
        "subtitle": "All sizes together per model; whiskers = 95% Wilson intervals",
        "kind": "dot-range",
        "unit": "rate",
        "polarity": "higher",
        "yLabel": "Lookups answered exactly",
        "series": [
          {
            "name": "Exact answer",
            "points": [
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 1,
                "lo": 0.7009,
                "hi": 1,
                "n": 9
              },
              {
                "label": "Claude Sonnet 5.5 · Claude Code",
                "value": 1,
                "lo": 0.7009,
                "hi": 1,
                "n": 9
              },
              {
                "label": "Claude Opus 5.5 · Claude Code",
                "value": 0.5556,
                "lo": 0.2667,
                "hi": 0.8112,
                "n": 9
              },
              {
                "label": "GPT-6.1 Sol (low) · Codex CLI",
                "value": 1,
                "lo": 0.7009,
                "hi": 1,
                "n": 9
              }
            ]
          }
        ],
        "note": "Whiskers are 95% Wilson intervals. One lookup question per call; a reply with extra words is a format miss, not a pass. With 9 calls per model, a perfect score still has a wide interval.",
        "whisker": "ci95",
        "sourceIds": [
          "agent-speed-anatomy"
        ]
      }
    ],
    "tables": [
      {
        "id": "speed-anatomy-cells",
        "title": "Every speed-anatomy cell, with reasoning tokens",
        "columns": [
          {
            "key": "part",
            "label": "Part",
            "unit": "text"
          },
          {
            "key": "config",
            "label": "Configuration",
            "unit": "text"
          },
          {
            "key": "size",
            "label": "Prompt size",
            "unit": "text"
          },
          {
            "key": "calls",
            "label": "Calls",
            "unit": "count"
          },
          {
            "key": "medianFirst",
            "label": "Median first text (s)",
            "unit": "seconds"
          },
          {
            "key": "rangeFirst",
            "label": "Fastest to slowest (s)",
            "unit": "text"
          },
          {
            "key": "medianAfterReady",
            "label": "Median first text after the CLI was ready (s, calculation)",
            "unit": "seconds"
          },
          {
            "key": "rangeAfterReady",
            "label": "After ready, fastest to slowest (s)",
            "unit": "text"
          },
          {
            "key": "medianTotal",
            "label": "Median total (s)",
            "unit": "seconds"
          },
          {
            "key": "rangeTotal",
            "label": "Total, fastest to slowest (s)",
            "unit": "text"
          },
          {
            "key": "medianInput",
            "label": "Median input tokens reported",
            "unit": "tokens"
          },
          {
            "key": "rangeInput",
            "label": "Input tokens, lowest to highest",
            "unit": "text"
          },
          {
            "key": "medianOut",
            "label": "Median output tokens",
            "unit": "tokens"
          },
          {
            "key": "rangeOut",
            "label": "Output tokens, lowest to highest",
            "unit": "text"
          },
          {
            "key": "medianReasoning",
            "label": "Median reasoning tokens",
            "unit": "tokens"
          },
          {
            "key": "rangeReasoning",
            "label": "Reasoning tokens, lowest to highest",
            "unit": "text"
          },
          {
            "key": "medianVisible",
            "label": "Median visible tokens",
            "unit": "tokens"
          },
          {
            "key": "rangeVisible",
            "label": "Visible tokens, lowest to highest",
            "unit": "text"
          },
          {
            "key": "speedVisible",
            "label": "Visible tokens per second (calculation)",
            "unit": "tokens"
          },
          {
            "key": "rangeSpeed",
            "label": "Visible tokens/s, lowest to highest",
            "unit": "text"
          },
          {
            "key": "speedPlain",
            "label": "All output tokens per second (calculation; inflated when reasoning comes first)",
            "unit": "tokens"
          },
          {
            "key": "rangePlain",
            "label": "All output tokens/s, lowest to highest",
            "unit": "text"
          },
          {
            "key": "charsPerSecond",
            "label": "Characters per second (calculation; exact replies only)",
            "unit": "count"
          },
          {
            "key": "rangeChars",
            "label": "Characters/s, lowest to highest",
            "unit": "text"
          },
          {
            "key": "charsN",
            "label": "Exact replies used for characters/s",
            "unit": "count"
          },
          {
            "key": "correct",
            "label": "Exactly right",
            "unit": "text"
          },
          {
            "key": "lastLine",
            "label": "Not a pass, but the last line is the right number (post-hoc reading)",
            "unit": "count"
          }
        ],
        "rows": [
          {
            "part": "A: 250 numbers in words",
            "config": "Claude Haiku 4.5 · Claude Code",
            "size": null,
            "calls": 4,
            "medianFirst": 4,
            "rangeFirst": "2.84 to 6.38",
            "medianAfterReady": 3.51,
            "medianTotal": 11.9,
            "rangeTotal": "10.91 to 14.31",
            "medianInput": 3658,
            "rangeInput": "3655 to 3658",
            "medianOut": 1780,
            "rangeOut": "1510 to 1997",
            "medianReasoning": 413,
            "rangeReasoning": "255 to 614",
            "medianVisible": 1210,
            "rangeVisible": "1210 to 1742",
            "speedVisible": 153.2,
            "rangeSpeed": "152.6 to 216.1",
            "rangeChars": "546 to 548",
            "charsN": 3,
            "rangeAfterReady": "2.36 to 5.78",
            "speedPlain": 225,
            "rangePlain": "190.9 to 247.7",
            "charsPerSecond": 547,
            "correct": "3/4",
            "lastLine": null
          },
          {
            "part": "A: 250 numbers in words",
            "config": "Claude Sonnet 5.5 · Claude Code",
            "size": null,
            "calls": 4,
            "medianFirst": 1.96,
            "rangeFirst": "0.88 to 4.09",
            "medianAfterReady": 1.51,
            "medianTotal": 10.33,
            "rangeTotal": "9.31 to 12.42",
            "medianInput": 1913,
            "rangeInput": "1912 to 1914",
            "medianOut": 2024,
            "rangeOut": "1941 to 2339",
            "medianReasoning": 83,
            "rangeReasoning": "0 to 398",
            "medianVisible": 1941,
            "rangeVisible": "1941 to 1941",
            "speedVisible": 231.7,
            "rangeSpeed": "230.3 to 233",
            "rangeChars": "513 to 519",
            "charsN": 4,
            "rangeAfterReady": "0.42 to 3.64",
            "speedPlain": 241.6,
            "rangePlain": "230.3 to 280.8",
            "charsPerSecond": 517,
            "correct": "4/4",
            "lastLine": null
          },
          {
            "part": "A: 250 numbers in words",
            "config": "Claude Opus 5.5 · Claude Code",
            "size": null,
            "calls": 4,
            "medianFirst": 1.97,
            "rangeFirst": "1.7 to 2.35",
            "medianAfterReady": 1.52,
            "medianTotal": 14.43,
            "rangeTotal": "14.25 to 14.8",
            "medianInput": 1909,
            "rangeInput": "1908 to 1909",
            "medianOut": 1990,
            "rangeOut": "1985 to 1994",
            "medianReasoning": 49,
            "rangeReasoning": "44 to 53",
            "medianVisible": 1941,
            "rangeVisible": "1941 to 1941",
            "speedVisible": 155.5,
            "rangeSpeed": "154.6 to 156.4",
            "rangeChars": "345 to 349",
            "charsN": 4,
            "rangeAfterReady": "1.24 to 1.87",
            "speedPlain": 159.4,
            "rangePlain": "158.7 to 160.3",
            "charsPerSecond": 347,
            "correct": "4/4",
            "lastLine": null
          },
          {
            "part": "A: 250 numbers in words",
            "config": "Claude Fable 5.1 · Claude Code",
            "size": null,
            "calls": 4,
            "medianFirst": 4.43,
            "rangeFirst": "2.27 to 4.64",
            "medianAfterReady": 3.95,
            "medianTotal": 19.81,
            "rangeTotal": "18.32 to 20.31",
            "medianInput": 3768,
            "rangeInput": "3766 to 3768",
            "medianOut": 2092,
            "rangeOut": "1941 to 2156",
            "medianReasoning": 151,
            "rangeReasoning": "0 to 215",
            "medianVisible": 1941,
            "rangeVisible": "1941 to 1941",
            "speedVisible": 122.6,
            "rangeSpeed": "120.9 to 131.4",
            "rangeChars": "270 to 293",
            "charsN": 4,
            "rangeAfterReady": "1.65 to 4.21",
            "speedPlain": 130.9,
            "rangePlain": "123.2 to 145.9",
            "charsPerSecond": 273,
            "correct": "4/4",
            "lastLine": null
          },
          {
            "part": "A: 250 numbers in words",
            "config": "GPT-6.1 Sol (low) · Codex CLI",
            "size": null,
            "calls": 4,
            "medianFirst": 3.52,
            "rangeFirst": "2.75 to 4.42",
            "medianAfterReady": 3.07,
            "medianTotal": 17.51,
            "rangeTotal": "16.36 to 17.75",
            "medianInput": 12028,
            "rangeInput": "12026 to 12028",
            "medianOut": 1066,
            "rangeOut": "1066 to 1066",
            "medianReasoning": 0,
            "rangeReasoning": "0 to 0",
            "medianVisible": 1066,
            "rangeVisible": "1066 to 1066",
            "speedVisible": 79.6,
            "rangeSpeed": "71.6 to 80.5",
            "rangeChars": "291 to 327",
            "charsN": 4,
            "rangeAfterReady": "2.34 to 4.01",
            "speedPlain": 79.6,
            "rangePlain": "71.6 to 80.5",
            "charsPerSecond": 323,
            "correct": "4/4",
            "lastLine": null
          },
          {
            "part": "A: 250 numbers in words",
            "config": "GPT-6 Luna (low) · Codex CLI",
            "size": null,
            "calls": 4,
            "medianFirst": 3.3,
            "rangeFirst": "3.19 to 3.47",
            "medianAfterReady": 2.51,
            "medianTotal": 15.5,
            "rangeTotal": "7.39 to 22.69",
            "medianInput": 11335,
            "rangeInput": "11332 to 11336",
            "medianOut": 1066,
            "rangeOut": "1066 to 1066",
            "medianReasoning": 0,
            "rangeReasoning": "0 to 0",
            "medianVisible": 1066,
            "rangeVisible": "1066 to 1066",
            "speedVisible": 129.1,
            "rangeSpeed": "55.5 to 259.1",
            "rangeChars": "225 to 1052",
            "charsN": 4,
            "rangeAfterReady": "2.14 to 2.88",
            "speedPlain": 129.1,
            "rangePlain": "55.5 to 259.1",
            "charsPerSecond": 524,
            "correct": "4/4",
            "lastLine": null
          },
          {
            "part": "B: ledger lookup",
            "config": "Claude Haiku 4.5 · Claude Code",
            "size": "1k",
            "calls": 3,
            "medianFirst": 1.93,
            "rangeFirst": "1.85 to 2.04",
            "medianAfterReady": 1.41,
            "medianTotal": 2.34,
            "rangeTotal": "2.22 to 2.46",
            "medianInput": 4592,
            "rangeInput": "4580 to 4610",
            "medianOut": 130,
            "rangeOut": "126 to 133",
            "medianReasoning": 123,
            "rangeReasoning": "119 to 126",
            "medianVisible": 7,
            "rangeVisible": "7 to 7",
            "speedVisible": 16.9,
            "rangeSpeed": "16.5 to 18.9",
            "rangeChars": "7 to 8",
            "charsN": 3,
            "rangeAfterReady": "1.38 to 1.43",
            "speedPlain": 314,
            "rangePlain": "296.5 to 358.5",
            "charsPerSecond": 7,
            "correct": "3/3",
            "lastLine": 0
          },
          {
            "part": "B: ledger lookup",
            "config": "Claude Haiku 4.5 · Claude Code",
            "size": "16k",
            "calls": 3,
            "medianFirst": 2.27,
            "rangeFirst": "2.22 to 2.47",
            "medianAfterReady": 1.76,
            "medianTotal": 2.79,
            "rangeTotal": "2.58 to 2.84",
            "medianInput": 19618,
            "rangeInput": "19604 to 19632",
            "medianOut": 136,
            "rangeOut": "129 to 139",
            "medianReasoning": 129,
            "rangeReasoning": "122 to 132",
            "medianVisible": 7,
            "rangeVisible": "7 to 7",
            "speedVisible": 19.1,
            "rangeSpeed": "13.5 to 19.4",
            "rangeChars": "6 to 8",
            "charsN": 3,
            "rangeAfterReady": "1.64 to 2.01",
            "speedPlain": 358.3,
            "rangePlain": "268.3 to 371.6",
            "charsPerSecond": 8,
            "correct": "3/3",
            "lastLine": 0
          },
          {
            "part": "B: ledger lookup",
            "config": "Claude Haiku 4.5 · Claude Code",
            "size": "64k",
            "calls": 3,
            "medianFirst": 2.78,
            "rangeFirst": "2.45 to 2.89",
            "medianAfterReady": 2.32,
            "medianTotal": 3.13,
            "rangeTotal": "2.84 to 3.28",
            "medianInput": 67605,
            "rangeInput": "67578 to 67639",
            "medianOut": 157,
            "rangeOut": "132 to 167",
            "medianReasoning": 150,
            "rangeReasoning": "125 to 160",
            "medianVisible": 7,
            "rangeVisible": "7 to 7",
            "speedVisible": 18.2,
            "rangeSpeed": "17.9 to 19.9",
            "rangeChars": "8 to 9",
            "charsN": 3,
            "rangeAfterReady": "1.94 to 2.42",
            "speedPlain": 401.5,
            "rangePlain": "343.8 to 475.8",
            "charsPerSecond": 8,
            "correct": "3/3",
            "lastLine": 0
          },
          {
            "part": "B: ledger lookup",
            "config": "Claude Sonnet 5.5 · Claude Code",
            "size": "1k",
            "calls": 3,
            "medianFirst": 1.45,
            "rangeFirst": "1.23 to 1.72",
            "medianAfterReady": 0.98,
            "medianTotal": 1.78,
            "rangeTotal": "1.57 to 2.12",
            "medianInput": 3051,
            "rangeInput": "3049 to 3057",
            "medianOut": 3,
            "rangeOut": "3 to 3",
            "medianReasoning": 0,
            "rangeReasoning": "0 to 0",
            "medianVisible": 3,
            "rangeVisible": "3 to 3",
            "speedVisible": 8.7,
            "rangeSpeed": "7.4 to 9.1",
            "rangeChars": "7 to 9",
            "charsN": 3,
            "rangeAfterReady": "0.8 to 0.98",
            "speedPlain": 8.7,
            "rangePlain": "7.4 to 9.1",
            "charsPerSecond": 9,
            "correct": "3/3",
            "lastLine": 0
          },
          {
            "part": "B: ledger lookup",
            "config": "Claude Sonnet 5.5 · Claude Code",
            "size": "16k",
            "calls": 3,
            "medianFirst": 1.78,
            "rangeFirst": "1.64 to 2.11",
            "medianAfterReady": 1.28,
            "medianTotal": 2.1,
            "rangeTotal": "1.98 to 2.48",
            "medianInput": 21080,
            "rangeInput": "21050 to 21099",
            "medianOut": 3,
            "rangeOut": "3 to 3",
            "medianReasoning": 0,
            "rangeReasoning": "0 to 0",
            "medianVisible": 3,
            "rangeVisible": "3 to 3",
            "speedVisible": 8.9,
            "rangeSpeed": "8.2 to 9.2",
            "rangeChars": "5 to 9",
            "charsN": 3,
            "rangeAfterReady": "1.15 to 1.32",
            "speedPlain": 8.9,
            "rangePlain": "8.2 to 9.2",
            "charsPerSecond": 9,
            "correct": "3/3",
            "lastLine": 0
          },
          {
            "part": "B: ledger lookup",
            "config": "Claude Sonnet 5.5 · Claude Code",
            "size": "64k",
            "calls": 3,
            "medianFirst": 3.07,
            "rangeFirst": "1.38 to 3.61",
            "medianAfterReady": 2.36,
            "medianTotal": 3.44,
            "rangeTotal": "1.74 to 4.38",
            "medianInput": 78596,
            "rangeInput": "78577 to 78745",
            "medianOut": 3,
            "rangeOut": "3 to 3",
            "medianReasoning": 0,
            "rangeReasoning": "0 to 0",
            "medianVisible": 3,
            "rangeVisible": "3 to 3",
            "speedVisible": 8,
            "rangeSpeed": "3.9 to 8.3",
            "rangeChars": "4 to 8",
            "charsN": 3,
            "rangeAfterReady": "0.94 to 3.13",
            "speedPlain": 8,
            "rangePlain": "3.9 to 8.3",
            "charsPerSecond": 8,
            "correct": "3/3",
            "lastLine": 0
          },
          {
            "part": "B: ledger lookup",
            "config": "Claude Opus 5.5 · Claude Code",
            "size": "1k",
            "calls": 3,
            "medianFirst": 1.51,
            "rangeFirst": "1.46 to 2.01",
            "medianAfterReady": 1.02,
            "medianTotal": 1.83,
            "rangeTotal": "1.82 to 2.41",
            "medianInput": 3040,
            "rangeInput": "3040 to 3044",
            "medianOut": 3,
            "rangeOut": "3 to 3",
            "medianReasoning": 0,
            "rangeReasoning": "0 to 0",
            "medianVisible": 3,
            "rangeVisible": "3 to 3",
            "speedVisible": 8.1,
            "rangeSpeed": "7.5 to 9.4",
            "rangeChars": "8 to 9",
            "charsN": 3,
            "rangeAfterReady": "1 to 1.36",
            "speedPlain": 8.1,
            "rangePlain": "7.5 to 9.4",
            "charsPerSecond": 8,
            "correct": "3/3",
            "lastLine": 0
          },
          {
            "part": "B: ledger lookup",
            "config": "Claude Opus 5.5 · Claude Code",
            "size": "16k",
            "calls": 3,
            "medianFirst": 1.74,
            "rangeFirst": "1.7 to 2.97",
            "medianAfterReady": 1.28,
            "medianTotal": 2.36,
            "rangeTotal": "2.11 to 3.4",
            "medianInput": 21074,
            "rangeInput": "20998 to 21089",
            "medianOut": 3,
            "rangeOut": "3 to 36",
            "medianReasoning": 0,
            "rangeReasoning": "0 to 0",
            "medianVisible": 3,
            "rangeVisible": "3 to 36",
            "speedVisible": 7.3,
            "rangeSpeed": "7 to 58.3",
            "rangeChars": "7 to 7",
            "charsN": 2,
            "rangeAfterReady": "1.19 to 2.53",
            "speedPlain": 7.3,
            "rangePlain": "7 to 58.3",
            "charsPerSecond": 7,
            "correct": "2/3",
            "lastLine": 1
          },
          {
            "part": "B: ledger lookup",
            "config": "Claude Opus 5.5 · Claude Code",
            "size": "64k",
            "calls": 3,
            "medianFirst": 1.79,
            "rangeFirst": "1.72 to 3.72",
            "medianAfterReady": 1.32,
            "medianTotal": 2.35,
            "rangeTotal": "2.26 to 4.29",
            "medianInput": 78674,
            "rangeInput": "78633 to 78684",
            "medianOut": 33,
            "rangeOut": "31 to 34",
            "medianReasoning": 0,
            "rangeReasoning": "0 to 0",
            "medianVisible": 33,
            "rangeVisible": "31 to 34",
            "speedVisible": 57.6,
            "rangeSpeed": "57.5 to 60.4",
            "rangeChars": null,
            "charsN": 0,
            "rangeAfterReady": "1.29 to 3.25",
            "speedPlain": 57.6,
            "rangePlain": "57.5 to 60.4",
            "charsPerSecond": null,
            "correct": "0/3",
            "lastLine": 3
          },
          {
            "part": "B: ledger lookup",
            "config": "GPT-6.1 Sol (low) · Codex CLI",
            "size": "1k",
            "calls": 3,
            "medianFirst": 3.36,
            "rangeFirst": "3.36 to 4.75",
            "medianAfterReady": 3.02,
            "medianTotal": 3.43,
            "rangeTotal": "3.43 to 4.92",
            "medianInput": 12790,
            "rangeInput": "12787 to 12797",
            "medianOut": 5,
            "rangeOut": "5 to 5",
            "medianReasoning": 0,
            "rangeReasoning": "0 to 0",
            "medianVisible": 5,
            "rangeVisible": "5 to 5",
            "speedVisible": 67.6,
            "rangeSpeed": "28.4 to 68.5",
            "rangeChars": "17 to 41",
            "charsN": 3,
            "rangeAfterReady": "2.76 to 3.29",
            "speedPlain": 67.6,
            "rangePlain": "28.4 to 68.5",
            "charsPerSecond": 41,
            "correct": "3/3",
            "lastLine": 0
          },
          {
            "part": "B: ledger lookup",
            "config": "GPT-6.1 Sol (low) · Codex CLI",
            "size": "16k",
            "calls": 3,
            "medianFirst": 4.02,
            "rangeFirst": "3.3 to 4.28",
            "medianAfterReady": 3.68,
            "medianTotal": 4.14,
            "rangeTotal": "3.96 to 4.68",
            "medianInput": 25038,
            "rangeInput": "25030 to 25045",
            "medianOut": 5,
            "rangeOut": "5 to 5",
            "medianReasoning": 0,
            "rangeReasoning": "0 to 0",
            "medianVisible": 5,
            "rangeVisible": "5 to 5",
            "speedVisible": 12.5,
            "rangeSpeed": "7.5 to 39.4",
            "rangeChars": "5 to 24",
            "charsN": 3,
            "rangeAfterReady": "2.88 to 3.95",
            "speedPlain": 12.5,
            "rangePlain": "7.5 to 39.4",
            "charsPerSecond": 7,
            "correct": "3/3",
            "lastLine": 0
          },
          {
            "part": "B: ledger lookup",
            "config": "GPT-6.1 Sol (low) · Codex CLI",
            "size": "64k",
            "calls": 3,
            "medianFirst": 3.93,
            "rangeFirst": "3.42 to 4.38",
            "medianAfterReady": 3.59,
            "medianTotal": 3.96,
            "rangeTotal": "3.47 to 4.44",
            "medianInput": 64172,
            "rangeInput": "64118 to 64198",
            "medianOut": 5,
            "rangeOut": "5 to 5",
            "medianReasoning": 0,
            "rangeReasoning": "0 to 0",
            "medianVisible": 5,
            "rangeVisible": "5 to 5",
            "speedVisible": 86.2,
            "rangeSpeed": "73.5 to 166.7",
            "rangeChars": "44 to 100",
            "charsN": 3,
            "rangeAfterReady": "3.02 to 4.06",
            "speedPlain": 86.2,
            "rangePlain": "73.5 to 166.7",
            "charsPerSecond": 52,
            "correct": "3/3",
            "lastLine": 0
          }
        ]
      },
      {
        "id": "speed-anatomy-cache-reads",
        "title": "Cache reads and writes per call; share = cache-read tokens ÷ reported input tokens (calculation)",
        "columns": [
          {
            "key": "config",
            "label": "Configuration",
            "unit": "text"
          },
          {
            "key": "size",
            "label": "Prompt size",
            "unit": "text"
          },
          {
            "key": "calls",
            "label": "Calls",
            "unit": "count"
          },
          {
            "key": "medianInput",
            "label": "Median input tokens reported",
            "unit": "tokens"
          },
          {
            "key": "rangeInput",
            "label": "Input tokens, lowest to highest",
            "unit": "text"
          },
          {
            "key": "medianRead",
            "label": "Median cache-read tokens",
            "unit": "tokens"
          },
          {
            "key": "rangeRead",
            "label": "Read tokens, lowest to highest",
            "unit": "text"
          },
          {
            "key": "maxRead",
            "label": "Most cache-read tokens in one call",
            "unit": "tokens"
          },
          {
            "key": "medianWrite",
            "label": "Median cache-write tokens",
            "unit": "tokens"
          },
          {
            "key": "rangeWrite",
            "label": "Write tokens, lowest to highest",
            "unit": "text"
          },
          {
            "key": "readShare",
            "label": "Median cache-read share of input (calculation)",
            "unit": "rate"
          },
          {
            "key": "rangeReadShare",
            "label": "Cache-read share, lowest to highest (calculation)",
            "unit": "text"
          },
          {
            "key": "readShareN",
            "label": "Calls used for cache-read share",
            "unit": "count"
          }
        ],
        "rows": [
          {
            "config": "Claude Haiku 4.5 · Claude Code",
            "size": "1k",
            "calls": 3,
            "medianInput": 4592,
            "rangeInput": "4580 to 4610",
            "medianRead": 0,
            "rangeRead": "0 to 0",
            "maxRead": 0,
            "medianWrite": 4582,
            "rangeWrite": "4570 to 4600",
            "readShare": 0,
            "rangeReadShare": "0.0% to 0.0%",
            "readShareN": 3
          },
          {
            "config": "Claude Haiku 4.5 · Claude Code",
            "size": "16k",
            "calls": 3,
            "medianInput": 19618,
            "rangeInput": "19604 to 19632",
            "medianRead": 0,
            "rangeRead": "0 to 0",
            "maxRead": 0,
            "medianWrite": 19608,
            "rangeWrite": "19594 to 19622",
            "readShare": 0,
            "rangeReadShare": "0.0% to 0.0%",
            "readShareN": 3
          },
          {
            "config": "Claude Haiku 4.5 · Claude Code",
            "size": "64k",
            "calls": 3,
            "medianInput": 67605,
            "rangeInput": "67578 to 67639",
            "medianRead": 0,
            "rangeRead": "0 to 0",
            "maxRead": 0,
            "medianWrite": 67595,
            "rangeWrite": "67568 to 67629",
            "readShare": 0,
            "rangeReadShare": "0.0% to 0.0%",
            "readShareN": 3
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "size": "1k",
            "calls": 3,
            "medianInput": 3051,
            "rangeInput": "3049 to 3057",
            "medianRead": 1463,
            "rangeRead": "1463 to 1463",
            "maxRead": 1463,
            "medianWrite": 1586,
            "rangeWrite": "1584 to 1592",
            "readShare": 0.48,
            "rangeReadShare": "47.9% to 48.0%",
            "readShareN": 3
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "size": "16k",
            "calls": 3,
            "medianInput": 21080,
            "rangeInput": "21050 to 21099",
            "medianRead": 1463,
            "rangeRead": "1463 to 1463",
            "maxRead": 1463,
            "medianWrite": 19615,
            "rangeWrite": "19585 to 19634",
            "readShare": 0.069,
            "rangeReadShare": "6.9% to 7.0%",
            "readShareN": 3
          },
          {
            "config": "Claude Sonnet 5.5 · Claude Code",
            "size": "64k",
            "calls": 3,
            "medianInput": 78596,
            "rangeInput": "78577 to 78745",
            "medianRead": 1463,
            "rangeRead": "1463 to 1463",
            "maxRead": 1463,
            "medianWrite": 77131,
            "rangeWrite": "77112 to 77280",
            "readShare": 0.019,
            "rangeReadShare": "1.9% to 1.9%",
            "readShareN": 3
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "size": "1k",
            "calls": 3,
            "medianInput": 3040,
            "rangeInput": "3040 to 3044",
            "medianRead": 1463,
            "rangeRead": "1463 to 1463",
            "maxRead": 1463,
            "medianWrite": 1575,
            "rangeWrite": "1575 to 1579",
            "readShare": 0.481,
            "rangeReadShare": "48.1% to 48.1%",
            "readShareN": 3
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "size": "16k",
            "calls": 3,
            "medianInput": 21074,
            "rangeInput": "20998 to 21089",
            "medianRead": 1463,
            "rangeRead": "1463 to 1463",
            "maxRead": 1463,
            "medianWrite": 19609,
            "rangeWrite": "19533 to 19624",
            "readShare": 0.069,
            "rangeReadShare": "6.9% to 7.0%",
            "readShareN": 3
          },
          {
            "config": "Claude Opus 5.5 · Claude Code",
            "size": "64k",
            "calls": 3,
            "medianInput": 78674,
            "rangeInput": "78633 to 78684",
            "medianRead": 1463,
            "rangeRead": "1463 to 1463",
            "maxRead": 1463,
            "medianWrite": 77209,
            "rangeWrite": "77168 to 77219",
            "readShare": 0.019,
            "rangeReadShare": "1.9% to 1.9%",
            "readShareN": 3
          },
          {
            "config": "GPT-6.1 Sol (low) · Codex CLI",
            "size": "1k",
            "calls": 3,
            "medianInput": 12790,
            "rangeInput": "12787 to 12797",
            "medianRead": 8960,
            "rangeRead": "0 to 8960",
            "maxRead": 8960,
            "medianWrite": 0,
            "rangeWrite": "0 to 0",
            "readShare": 0.7,
            "rangeReadShare": "0.0% to 70.1%",
            "readShareN": 3
          },
          {
            "config": "GPT-6.1 Sol (low) · Codex CLI",
            "size": "16k",
            "calls": 3,
            "medianInput": 25038,
            "rangeInput": "25030 to 25045",
            "medianRead": 8960,
            "rangeRead": "8960 to 8960",
            "maxRead": 8960,
            "medianWrite": 0,
            "rangeWrite": "0 to 0",
            "readShare": 0.358,
            "rangeReadShare": "35.8% to 35.8%",
            "readShareN": 3
          },
          {
            "config": "GPT-6.1 Sol (low) · Codex CLI",
            "size": "64k",
            "calls": 3,
            "medianInput": 64172,
            "rangeInput": "64118 to 64198",
            "medianRead": 8960,
            "rangeRead": "8960 to 8960",
            "maxRead": 8960,
            "medianWrite": 0,
            "rangeWrite": "0 to 0",
            "readShare": 0.14,
            "rangeReadShare": "14.0% to 14.0%",
            "readShareN": 3
          }
        ]
      }
    ],
    "related": [
      "cli-model-latency-tokens",
      "model-head-to-head",
      "routing-overhead"
    ]
  },
  "sources": [
    {
      "id": "agent-speed-anatomy",
      "title": "LLM speed anatomy",
      "kind": "run",
      "date": "2026-10-07",
      "data": [
        "/benchmarks/raw/speed-anatomy/receipts.json"
      ],
      "note": "Receipts of the speed anatomy study: one prompt that asks for 250 numbers in words (output speed, 24 calls) and a seeded synthetic ledger at three sizes with one lookup question (prompt size, 36 calls). Every attempt is kept, failures included. A new seed for every ledger call; no ledger text, prompt or model output is copied."
    }
  ]
}
