Where the seconds go: first text, output speed and prompt size for 6 LLMs
For 6 models run through their own coding CLIs (Haiku, Sonnet, Opus, Fable, Sol (low) and Luna (low)), how long until the first text, how fast does text stream after it, and what does a longer prompt add?
Published · 6 charts · Download the data or a carousel
1.8x
The answer
We measured CLI start-up, first text and total time on one 250-line answer and a ledger lookup at three prompt sizes. There were 60 counted calls: 4 per model in part A and 3 per size in part B. The receipts do not separate provider queueing, prompt processing and reasoning time. Time to first text, median (range; n calls): Sonnet 2.0 s (0.9 to 4.1; n = 4). Opus 2.0 s (1.7 to 2.4; n = 4). Luna (low) 3.3 s (3.2 to 3.5; n = 4). Sol (low) 3.5 s (2.7 to 4.4; n = 4). Haiku 4.0 s (2.8 to 6.4; n = 4). Fable 4.4 s (2.3 to 4.6; n = 4). Opus’s slowest call (2.4 s) was faster than the fastest call of Luna (low) (3.2 s), Sol (low) (2.7 s) and Haiku (2.8 s). Sonnet and Fable overlap it. There are 4 calls per model. The samples are below the protocol’s five-call minimum for naming a speed winner. Characters per second is a calculation: correct reply length ÷ time from first text to call end. Median (range; n exact replies): Haiku 547 (546 to 548; n = 3). Luna (low) 524 (225 to 1,052; n = 4). Sonnet 517 (513 to 519; n = 4). Opus 347 (345 to 349; n = 4). Sol (low) 323 (291 to 327; n = 4). Fable 273 (270 to 293; n = 4). Haiku and Fable have non-overlapping observed speed ranges (3 and 4 exact replies). Both samples are below the protocol’s five-call minimum for naming a speed winner. Luna (low)’s 4 exact replies ran from 225 to 1,052 characters per second, so its median hides this large spread. These few calls cannot establish two speed groups. Token rates use different token units across models. The same exact reply has 4,327 characters. Median visible tokens: 1,066 for Sol (low) (n = 4) and Luna (low) (n = 4). 1,210 for Haiku (n = 3). 1,941 for Sonnet (n = 4), Opus (n = 4) and Fable (n = 4). Visible tokens per second (calculation), median (range; n calls): Sonnet 232 (230 to 233; n = 4). Opus 156 (155 to 156; n = 4). Haiku 153 (153 to 216; n = 4). Luna (low) 129 (55 to 259; n = 4). Fable 123 (121 to 131; n = 4). Sol (low) 80 (72 to 80; n = 4). These ranges show individual calls, not confidence intervals. Of the 24 part A replies, 23 matched all 250 lines exactly; 1 was wrong (Haiku, 288 lines), and every call stays in the timings. First-text medians differ by prompt size (calculation: difference of medians, 64k minus 1k): Haiku +0.9 s (1.9 s to 2.8 s). Sonnet +1.6 s (1.4 s to 3.1 s, ranges overlap). Opus +0.3 s (1.5 s to 1.8 s, ranges overlap). Sol (low) +0.6 s (3.4 s to 3.9 s, ranges overlap). Exact lookup answers (95% Wilson intervals): Haiku 9/9 (70% to 100%). Sonnet 9/9 (70% to 100%). Opus 5/9 (27% to 81%). Sol (low) 9/9 (70% to 100%). All 4 models’ intervals overlap; this sample cannot rank their lookup rates. 4 Opus replies echoed the ledger line and then gave the right number last. The declared scoring counts these as wrong answers. The table shows the post-hoc reading. Cache reads did not grow with the ledger for any model. This fits no ledger reuse, but the counts do not prove which text the provider cached.
Key numbers
60
Counted calls in the speed anatomy study
calls (24 in part A, 36 in part B; 43 through Claude Code, 17 through Codex CLI) · n = 60
96% (23/24)
Part A replies that matched all 250 lines exactly (strict)
95% CI 80%–99% · n = 24
89% (32/36)
Part B lookups answered exactly (strict)
95% CI 75%–96% · n = 36
100% (36/36)
Part B replies whose last line was exactly the right number (post-hoc reading, not a pass)
95% CI 90%–100% · n = 36
0 of 4
Models whose cache reads grew with the ledger size
n = 36
+0.9s
Haiku: extra time to first text at 64k vs 1k (calculation)
n = 6
+1.6s
Sonnet: extra time to first text at 64k vs 1k (calculation)
n = 6
+0.3s
Opus: extra time to first text at 64k vs 1k (calculation)
n = 6
+0.6s
Sol (low): extra time to first text at 64k vs 1k (calculation)
n = 6
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first text | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 4 s | 2.8 s–6.4 s | 4 |
| Claude Sonnet 5.5 · Claude Code | 2 s | 0.9 s–4.1 s | 4 |
| Claude Opus 5.5 · Claude Code | 2 s | 1.7 s–2.4 s | 4 |
| Claude Fable 5.1 · Claude Code | 4.4 s | 2.3 s–4.6 s | 4 |
| GPT-6.1 Sol (low) · Codex CLI | 3.5 s | 2.8 s–4.4 s | 4 |
| GPT-6 Luna (low) · Codex CLI | 3.3 s | 3.2 s–3.5 s | 4 |
6 rows. Slowest Claude Fable 5.1 · Claude Code 4.4 s (range 2.3 s–4.6 s, n 4). Fastest Claude Sonnet 5.5 · Claude Code 2 s (range 0.9 s–4.1 s, n 4). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = fastest and slowest call
Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.
Source: LLM speed anatomy
| Item | Visible tokens per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 153 | 153–216 | 4 |
| Claude Sonnet 5.5 · Claude Code | 232 | 230–233 | 4 |
| Claude Opus 5.5 · Claude Code | 156 | 155–156 | 4 |
| Claude Fable 5.1 · Claude Code | 123 | 121–131 | 4 |
| GPT-6.1 Sol (low) · Codex CLI | 80 | 72–81 | 4 |
| GPT-6 Luna (low) · Codex CLI | 129 | 56–259 | 4 |
List-price calculation, not a run. 6 rows. Highest Claude Sonnet 5.5 · Claude Code 232 (range 230–233, n 4). Lowest GPT-6.1 Sol (low) · Codex CLI 80 (range 72–81, n 4). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = slowest and fastest call
Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.
Source: LLM speed anatomy
| Item | Characters per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 547 | 546–548 | 3 |
| Claude Sonnet 5.5 · Claude Code | 517 | 513–519 | 4 |
| Claude Opus 5.5 · Claude Code | 347 | 345–349 | 4 |
| Claude Fable 5.1 · Claude Code | 273 | 270–293 | 4 |
| GPT-6.1 Sol (low) · Codex CLI | 323 | 291–327 | 4 |
| GPT-6 Luna (low) · Codex CLI | 524 | 225–1,052 | 4 |
List-price calculation, not a run. 6 rows. Highest Claude Haiku 4.5 · Claude Code 547 (range 546–548, n 3). Lowest Claude Fable 5.1 · Claude Code 273 (range 270–293, n 4). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 3–4 per row
Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call
Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
- Claude Haiku 4.5 · Claude Code
- Claude Sonnet 5.5 · Claude Code
- Claude Opus 5.5 · Claude Code
- GPT-6.1 Sol (low) · Codex CLI
| Prompt-size target (approximate Haiku tokens; calibration calculation) | Claude Haiku 4.5 · Claude Code | Claude Sonnet 5.5 · Claude Code | Claude Opus 5.5 · Claude Code | GPT-6.1 Sol (low) · Codex CLI | Range (lowest–highest run) | n |
|---|---|---|---|---|---|---|
| 1k | 1.9 s | 1.5 s | 1.5 s | 3.4 s | Claude Haiku 4.5 · Claude Code: 1.9 s–2 s; Claude Sonnet 5.5 · Claude Code: 1.2 s–1.7 s; Claude Opus 5.5 · Claude Code: 1.5 s–2 s; GPT-6.1 Sol (low) · Codex CLI: 3.4 s–4.8 s | 3 |
| 16k | 2.3 s | 1.8 s | 1.7 s | 4 s | Claude Haiku 4.5 · Claude Code: 2.2 s–2.5 s; Claude Sonnet 5.5 · Claude Code: 1.6 s–2.1 s; Claude Opus 5.5 · Claude Code: 1.7 s–3 s; GPT-6.1 Sol (low) · Codex CLI: 3.3 s–4.3 s | 3 |
| 64k | 2.8 s | 3.1 s | 1.8 s | 3.9 s | Claude Haiku 4.5 · Claude Code: 2.5 s–2.9 s; Claude Sonnet 5.5 · Claude Code: 1.4 s–3.6 s; Claude Opus 5.5 · Claude Code: 1.7 s–3.7 s; GPT-6.1 Sol (low) · Codex CLI: 3.4 s–4.4 s | 3 |
List-price calculation, not a run. 3 rows, 4 series: Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (low) · Codex CLI. Claude Haiku 4.5 · Claude Code: slowest 64k 2.8 s (range 2.5 s–2.9 s, n 3). Fastest 1k 1.9 s (range 1.9 s–2 s, n 3). Not all run ranges overlap. Claude Sonnet 5.5 · Claude Code: slowest 64k 3.1 s (range 1.4 s–3.6 s, n 3). Fastest 1k 1.5 s (range 1.2 s–1.7 s, n 3). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Median of 3 calls per size; whiskers = fastest and slowest call
Each call used a new ledger seed. Cache-read counts stayed within the short-prompt baseline (see the cache table). This does not identify which tokens were cached. Sizes name the text we send; each model’s reported input tokens are in the table and include the CLI’s own prefix. The size calibration subtracts estimated prefixes from probe input counts; these are calculations, not measured prefix counts for each call. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
1k prompt
16k prompt
64k prompt
One panel per series, all on the same axis; whiskers are the fastest–slowest run (not an interval).
| Item | 1k prompt | 16k prompt | 64k prompt | Range (lowest–highest run) | n |
|---|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 2.3 s | 2.8 s | 3.1 s | 1k prompt: 2.2 s–2.5 s; 16k prompt: 2.6 s–2.8 s; 64k prompt: 2.8 s–3.3 s | 3 |
| Claude Sonnet 5.5 · Claude Code | 1.8 s | 2.1 s | 3.4 s | 1k prompt: 1.6 s–2.1 s; 16k prompt: 2 s–2.5 s; 64k prompt: 1.7 s–4.4 s | 3 |
| Claude Opus 5.5 · Claude Code | 1.8 s | 2.4 s | 2.4 s | 1k prompt: 1.8 s–2.4 s; 16k prompt: 2.1 s–3.4 s; 64k prompt: 2.3 s–4.3 s | 3 |
| GPT-6.1 Sol (low) · Codex CLI | 3.4 s | 4.1 s | 4 s | 1k prompt: 3.4 s–4.9 s; 16k prompt: 4 s–4.7 s; 64k prompt: 3.5 s–4.4 s | 3 |
4 rows, 3 series: 1k prompt, 16k prompt, 64k prompt. 1k prompt: slowest GPT-6.1 Sol (low) · Codex CLI 3.4 s (range 3.4 s–4.9 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 1.8 s (range 1.6 s–2.1 s, n 3). Not all run ranges overlap. 16k prompt: slowest GPT-6.1 Sol (low) · Codex CLI 4.1 s (range 4 s–4.7 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 2.1 s (range 2 s–2.5 s, n 3). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Median of 3 calls per bar; whiskers = fastest and slowest call
Whole call: CLI start-up, first text and the one-line answer. Whiskers are a range of calls, not a confidence interval. Each call used a new ledger.
Source: LLM speed anatomy
Every interval overlaps every other: this chart does not order these rows.
| Item | Exact answer | 95% interval | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 100% | 70%–100% | 9 |
| Claude Sonnet 5.5 · Claude Code | 100% | 70%–100% | 9 |
| Claude Opus 5.5 · Claude Code | 56% | 27%–81% | 9 |
| GPT-6.1 Sol (low) · Codex CLI | 100% | 70%–100% | 9 |
4 rows. Highest Claude Haiku 4.5 · Claude Code 100% (95% interval 70%–100%, n 9). Lowest Claude Opus 5.5 · Claude Code 56% (95% interval 27%–81%, n 9). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 9 per row3 of 4 at 100%: this task set cannot separate them.
All sizes together per model; whiskers = 95% Wilson intervals
Whiskers are 95% Wilson intervals. One lookup question per call; a reply with extra words is a format miss, not a pass. With 9 calls per model, a perfect score still has a wide interval.
Source: LLM speed anatomy
Tables
Every speed-anatomy cell, with reasoning tokens
| Part | Configuration | Prompt size | Calls | Median first text (s) | Fastest to slowest (s) | Median first text after the CLI was ready (s, calculation) | After ready, fastest to slowest (s) | Median total (s) | Total, fastest to slowest (s) | Median input tokens reported | Input tokens, lowest to highest | Median output tokens | Output tokens, lowest to highest | Median reasoning tokens | Reasoning tokens, lowest to highest | Median visible tokens | Visible tokens, lowest to highest | Visible tokens per second (calculation) | Visible tokens/s, lowest to highest | All output tokens per second (calculation; inflated when reasoning comes first) | All output tokens/s, lowest to highest | Characters per second (calculation; exact replies only) | Characters/s, lowest to highest | Exact replies used for characters/s | Exactly right | Not a pass, but the last line is the right number (post-hoc reading) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| A: 250 numbers in words | Claude Haiku 4.5 · Claude Code | — | 4 | 4 s | 2.84 to 6.38 | 3.5 s | 2.36 to 5.78 | 11.9 s | 10.91 to 14.31 | 3,658 | 3655 to 3658 | 1,780 | 1510 to 1997 | 413 | 255 to 614 | 1,210 | 1210 to 1742 | 153 | 152.6 to 216.1 | 225 | 190.9 to 247.7 | 547 | 546 to 548 | 3 | 3/4 | — |
| A: 250 numbers in words | Claude Sonnet 5.5 · Claude Code | — | 4 | 2 s | 0.88 to 4.09 | 1.5 s | 0.42 to 3.64 | 10.3 s | 9.31 to 12.42 | 1,913 | 1912 to 1914 | 2,024 | 1941 to 2339 | 83 | 0 to 398 | 1,941 | 1941 to 1941 | 232 | 230.3 to 233 | 242 | 230.3 to 280.8 | 517 | 513 to 519 | 4 | 4/4 | — |
| A: 250 numbers in words | Claude Opus 5.5 · Claude Code | — | 4 | 2 s | 1.7 to 2.35 | 1.5 s | 1.24 to 1.87 | 14.4 s | 14.25 to 14.8 | 1,909 | 1908 to 1909 | 1,990 | 1985 to 1994 | 49 | 44 to 53 | 1,941 | 1941 to 1941 | 156 | 154.6 to 156.4 | 159 | 158.7 to 160.3 | 347 | 345 to 349 | 4 | 4/4 | — |
| A: 250 numbers in words | Claude Fable 5.1 · Claude Code | — | 4 | 4.4 s | 2.27 to 4.64 | 4 s | 1.65 to 4.21 | 19.8 s | 18.32 to 20.31 | 3,768 | 3766 to 3768 | 2,092 | 1941 to 2156 | 151 | 0 to 215 | 1,941 | 1941 to 1941 | 123 | 120.9 to 131.4 | 131 | 123.2 to 145.9 | 273 | 270 to 293 | 4 | 4/4 | — |
| A: 250 numbers in words | GPT-6.1 Sol (low) · Codex CLI | — | 4 | 3.5 s | 2.75 to 4.42 | 3.1 s | 2.34 to 4.01 | 17.5 s | 16.36 to 17.75 | 12,028 | 12026 to 12028 | 1,066 | 1066 to 1066 | 0 | 0 to 0 | 1,066 | 1066 to 1066 | 80 | 71.6 to 80.5 | 80 | 71.6 to 80.5 | 323 | 291 to 327 | 4 | 4/4 | — |
| A: 250 numbers in words | GPT-6 Luna (low) · Codex CLI | — | 4 | 3.3 s | 3.19 to 3.47 | 2.5 s | 2.14 to 2.88 | 15.5 s | 7.39 to 22.69 | 11,335 | 11332 to 11336 | 1,066 | 1066 to 1066 | 0 | 0 to 0 | 1,066 | 1066 to 1066 | 129 | 55.5 to 259.1 | 129 | 55.5 to 259.1 | 524 | 225 to 1052 | 4 | 4/4 | — |
| B: ledger lookup | Claude Haiku 4.5 · Claude Code | 1k | 3 | 1.9 s | 1.85 to 2.04 | 1.4 s | 1.38 to 1.43 | 2.3 s | 2.22 to 2.46 | 4,592 | 4580 to 4610 | 130 | 126 to 133 | 123 | 119 to 126 | 7 | 7 to 7 | 17 | 16.5 to 18.9 | 314 | 296.5 to 358.5 | 7 | 7 to 8 | 3 | 3/3 | 0 |
| B: ledger lookup | Claude Haiku 4.5 · Claude Code | 16k | 3 | 2.3 s | 2.22 to 2.47 | 1.8 s | 1.64 to 2.01 | 2.8 s | 2.58 to 2.84 | 19,618 | 19604 to 19632 | 136 | 129 to 139 | 129 | 122 to 132 | 7 | 7 to 7 | 19 | 13.5 to 19.4 | 358 | 268.3 to 371.6 | 8 | 6 to 8 | 3 | 3/3 | 0 |
| B: ledger lookup | Claude Haiku 4.5 · Claude Code | 64k | 3 | 2.8 s | 2.45 to 2.89 | 2.3 s | 1.94 to 2.42 | 3.1 s | 2.84 to 3.28 | 67,605 | 67578 to 67639 | 157 | 132 to 167 | 150 | 125 to 160 | 7 | 7 to 7 | 18 | 17.9 to 19.9 | 402 | 343.8 to 475.8 | 8 | 8 to 9 | 3 | 3/3 | 0 |
| B: ledger lookup | Claude Sonnet 5.5 · Claude Code | 1k | 3 | 1.5 s | 1.23 to 1.72 | 1 s | 0.8 to 0.98 | 1.8 s | 1.57 to 2.12 | 3,051 | 3049 to 3057 | 3 | 3 to 3 | 0 | 0 to 0 | 3 | 3 to 3 | 9 | 7.4 to 9.1 | 9 | 7.4 to 9.1 | 9 | 7 to 9 | 3 | 3/3 | 0 |
| B: ledger lookup | Claude Sonnet 5.5 · Claude Code | 16k | 3 | 1.8 s | 1.64 to 2.11 | 1.3 s | 1.15 to 1.32 | 2.1 s | 1.98 to 2.48 | 21,080 | 21050 to 21099 | 3 | 3 to 3 | 0 | 0 to 0 | 3 | 3 to 3 | 9 | 8.2 to 9.2 | 9 | 8.2 to 9.2 | 9 | 5 to 9 | 3 | 3/3 | 0 |
| B: ledger lookup | Claude Sonnet 5.5 · Claude Code | 64k | 3 | 3.1 s | 1.38 to 3.61 | 2.4 s | 0.94 to 3.13 | 3.4 s | 1.74 to 4.38 | 78,596 | 78577 to 78745 | 3 | 3 to 3 | 0 | 0 to 0 | 3 | 3 to 3 | 8 | 3.9 to 8.3 | 8 | 3.9 to 8.3 | 8 | 4 to 8 | 3 | 3/3 | 0 |
| B: ledger lookup | Claude Opus 5.5 · Claude Code | 1k | 3 | 1.5 s | 1.46 to 2.01 | 1 s | 1 to 1.36 | 1.8 s | 1.82 to 2.41 | 3,040 | 3040 to 3044 | 3 | 3 to 3 | 0 | 0 to 0 | 3 | 3 to 3 | 8 | 7.5 to 9.4 | 8 | 7.5 to 9.4 | 8 | 8 to 9 | 3 | 3/3 | 0 |
| B: ledger lookup | Claude Opus 5.5 · Claude Code | 16k | 3 | 1.7 s | 1.7 to 2.97 | 1.3 s | 1.19 to 2.53 | 2.4 s | 2.11 to 3.4 | 21,074 | 20998 to 21089 | 3 | 3 to 36 | 0 | 0 to 0 | 3 | 3 to 36 | 7 | 7 to 58.3 | 7 | 7 to 58.3 | 7 | 7 to 7 | 2 | 2/3 | 1 |
| B: ledger lookup | Claude Opus 5.5 · Claude Code | 64k | 3 | 1.8 s | 1.72 to 3.72 | 1.3 s | 1.29 to 3.25 | 2.4 s | 2.26 to 4.29 | 78,674 | 78633 to 78684 | 33 | 31 to 34 | 0 | 0 to 0 | 33 | 31 to 34 | 58 | 57.5 to 60.4 | 58 | 57.5 to 60.4 | — | — | 0 | 0/3 | 3 |
| B: ledger lookup | GPT-6.1 Sol (low) · Codex CLI | 1k | 3 | 3.4 s | 3.36 to 4.75 | 3 s | 2.76 to 3.29 | 3.4 s | 3.43 to 4.92 | 12,790 | 12787 to 12797 | 5 | 5 to 5 | 0 | 0 to 0 | 5 | 5 to 5 | 68 | 28.4 to 68.5 | 68 | 28.4 to 68.5 | 41 | 17 to 41 | 3 | 3/3 | 0 |
| B: ledger lookup | GPT-6.1 Sol (low) · Codex CLI | 16k | 3 | 4 s | 3.3 to 4.28 | 3.7 s | 2.88 to 3.95 | 4.1 s | 3.96 to 4.68 | 25,038 | 25030 to 25045 | 5 | 5 to 5 | 0 | 0 to 0 | 5 | 5 to 5 | 13 | 7.5 to 39.4 | 13 | 7.5 to 39.4 | 7 | 5 to 24 | 3 | 3/3 | 0 |
| B: ledger lookup | GPT-6.1 Sol (low) · Codex CLI | 64k | 3 | 3.9 s | 3.42 to 4.38 | 3.6 s | 3.02 to 4.06 | 4 s | 3.47 to 4.44 | 64,172 | 64118 to 64198 | 5 | 5 to 5 | 0 | 0 to 0 | 5 | 5 to 5 | 86 | 73.5 to 166.7 | 86 | 73.5 to 166.7 | 52 | 44 to 100 | 3 | 3/3 | 0 |
Cache reads and writes per call; share = cache-read tokens ÷ reported input tokens (calculation)
| Configuration | Prompt size | Calls | Median input tokens reported | Input tokens, lowest to highest | Median cache-read tokens | Read tokens, lowest to highest | Most cache-read tokens in one call | Median cache-write tokens | Write tokens, lowest to highest | Median cache-read share of input (calculation) | Cache-read share, lowest to highest (calculation) | Calls used for cache-read share |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 1k | 3 | 4,592 | 4580 to 4610 | 0 | 0 to 0 | 0 | 4,582 | 4570 to 4600 | 0% | 0.0% to 0.0% | 3 |
| Claude Haiku 4.5 · Claude Code | 16k | 3 | 19,618 | 19604 to 19632 | 0 | 0 to 0 | 0 | 19,608 | 19594 to 19622 | 0% | 0.0% to 0.0% | 3 |
| Claude Haiku 4.5 · Claude Code | 64k | 3 | 67,605 | 67578 to 67639 | 0 | 0 to 0 | 0 | 67,595 | 67568 to 67629 | 0% | 0.0% to 0.0% | 3 |
| Claude Sonnet 5.5 · Claude Code | 1k | 3 | 3,051 | 3049 to 3057 | 1,463 | 1463 to 1463 | 1,463 | 1,586 | 1584 to 1592 | 48% | 47.9% to 48.0% | 3 |
| Claude Sonnet 5.5 · Claude Code | 16k | 3 | 21,080 | 21050 to 21099 | 1,463 | 1463 to 1463 | 1,463 | 19,615 | 19585 to 19634 | 6.9% | 6.9% to 7.0% | 3 |
| Claude Sonnet 5.5 · Claude Code | 64k | 3 | 78,596 | 78577 to 78745 | 1,463 | 1463 to 1463 | 1,463 | 77,131 | 77112 to 77280 | 1.9% | 1.9% to 1.9% | 3 |
| Claude Opus 5.5 · Claude Code | 1k | 3 | 3,040 | 3040 to 3044 | 1,463 | 1463 to 1463 | 1,463 | 1,575 | 1575 to 1579 | 48% | 48.1% to 48.1% | 3 |
| Claude Opus 5.5 · Claude Code | 16k | 3 | 21,074 | 20998 to 21089 | 1,463 | 1463 to 1463 | 1,463 | 19,609 | 19533 to 19624 | 6.9% | 6.9% to 7.0% | 3 |
| Claude Opus 5.5 · Claude Code | 64k | 3 | 78,674 | 78633 to 78684 | 1,463 | 1463 to 1463 | 1,463 | 77,209 | 77168 to 77219 | 1.9% | 1.9% to 1.9% | 3 |
| GPT-6.1 Sol (low) · Codex CLI | 1k | 3 | 12,790 | 12787 to 12797 | 8,960 | 0 to 8960 | 8,960 | 0 | 0 to 0 | 70% | 0.0% to 70.1% | 3 |
| GPT-6.1 Sol (low) · Codex CLI | 16k | 3 | 25,038 | 25030 to 25045 | 8,960 | 8960 to 8960 | 8,960 | 0 | 0 to 0 | 36% | 35.8% to 35.8% | 3 |
| GPT-6.1 Sol (low) · Codex CLI | 64k | 3 | 64,172 | 64118 to 64198 | 8,960 | 8960 to 8960 | 8,960 | 0 | 0 to 0 | 14% | 14.0% to 14.0% | 3 |
Method
- The protocol file birth precedes the first probe and counted call. Later edits have no frozen versions. Two probes checked the routes and helped calibrate ledger sizes. They stay outside all cells and charts.
- Part A used 24 calls: 4 per model. The prompt asked for 250 numbers in words, one per line. The check compares the reply with all 250 expected lines. Models: Haiku, Sonnet, Opus, Fable, Sol (low) and Luna (low). Controls ran before the first counted call. The reference passes. All 8/8 wrapped references are format misses. All 11/11 planted wrong answers fail.
- Part B used 36 calls: 3 per size for each model. Models: Haiku, Sonnet, Opus and Sol (low). Each call asks one exact lookup question about a synthetic stock ledger. Ledger sizes: 18 lines for 1k, 327 lines for 16k and 1,315 lines for 64k. The 1k, 16k and 64k labels are approximate targets from probe calibration calculations, using estimated Haiku prefix counts. A new seed changes each ledger. Some fixed text repeats. Cache counts do not identify cached text.
- First text means the first non-empty text after CLI start. The time includes CLI start-up and work before the first word. The table also subtracts the CLI ready time (calculation).
- Output speed is a calculation. Visible tokens equal output tokens minus reported reasoning tokens. When the CLI reports no reasoning count, the calculation uses zero. That does not prove the model did no reasoning. Divide visible tokens by the time from first text to call end. The plain form divides all output tokens by that time. It can overstate visible speed when reasoning comes first.
- The harness uses the hard head-to-head machinery. It disables tools, MCP servers and saved sessions. Claude Code uses default effort; Codex CLI uses low effort. Calls run one at a time on one Mac. The timeout was 300 s. Claude Code had a 16,000-token output cap. Codex had no output-token cap. CLI versions: 2.1.286 (Claude Code) and codex-cli 0.160.0.
- Stop rules require a stop at the first usage-limit text. The gate checks other study markers before each batch. No stored batch stopped or trimmed a cell. No counted call duplicates another. A sequence process ended at the gate. Codex part B ran later. Every counted call completed.
- Cache-read share is a calculation: cache-read tokens divided by reported input tokens for each call. The cache table shows the median, observed range and n for each cell.
Caveats
- 4 calls per model in part A and 3 per size in part B. Medians of so few calls move with one slow call, and the ranges are not confidence intervals. The 95% Wilson intervals on the lookup rates are wide.
- First text includes CLI start-up and reasoning time. It is not the API’s time to first token; a direct API call would skip the CLI start-up (the CLI reported ready after a median 0.33 s to 0.84 s per cell).
- Claude Code ran models at their default effort; Codex CLI ran at low effort. Effort levels are not one scale, and a CLI adds its own system prompt, so a row mixes a model and its CLI.
- Output speed is a calculation from reported token counts, the length of the reply and the clock. It depends on how each CLI counts reasoning, and part A has one prompt: other text, other days or an API route can differ.
- All 36 counted ledger prompts differ. Their input counts do not compare tokenizers on the same text. The size calibration uses prefix estimates from earlier short prompts. Those calculations do not measure the prefix in each counted call.
- The tasks hit a ceiling. Part A passed 23/24. In part B, 3 of 4 models passed every lookup. See the rate stats and chart for 95% intervals. This small task set cannot rank general capability.
- One shared Mac: the gate checked other study markers before each batch. No full host-load record proves that all other work stopped.
- Prompt sizes were calibrated after two probes. The cases are synthetic and tuned to this measurement, not a sample of real workloads.
- Amendment 3 says 06:05 UTC and claims to precede Codex part B. Those calls ran at 06:03:43 to 06:04:19 UTC. The same-file amendment is retrospective; its date does not prove advance declaration.
- The controls receipt was overwritten. The surviving checks ran after the probes and before counted calls; an earlier control run cannot be audited.
- Each size used different ledger text. Prompt-size differences also include content and provider-load changes; they do not isolate a cause.
- The failed Haiku reply contained tool-shaped text and extra output. No structured tool event was recorded; it stays a strict failure and remains in timings.
- One Mac, one network, two sessions on one night. Provider load changes timings from hour to hour. The Codex CLI part B batch ran about 4 hours after the batches before it (a run script stopped while it waited for another run to finish timing), so its timings come from a later hour.
Sources
LLM speed anatomy
Receipts of the speed anatomy study: one prompt that asks for 250 numbers in words (output speed, 24 calls) and a seeded synthetic ledger at three sizes with one lookup question (prompt size, 36 calls). Every attempt is kept, failures included. A new seed for every ledger call; no ledger text, prompt or model output is copied.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “Where the seconds go: first text, output speed and prompt size for 6 LLMs”, updated October 7, 2026, https://agent.sasid.ai/benchmarks/llm-speed-anatomy.
More comparisons based on this study (6)
These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.
Models and comparisons in this study
- Claude Sonnet 5.5
- Claude Opus 5.5
- Claude Haiku 4.5
- Claude Fable 5.1
- GPT-6.1 Sol (Codex CLI)
- Claude Code
- Codex CLI
- GPT-6 Luna (Codex CLI)
- Claude Sonnet 5.5 vs Claude Opus 5.5
- Claude Haiku 4.5 vs Claude Sonnet 5.5
- Claude Opus 5.5 vs Claude Fable 5.1
- Claude Sonnet 5.5 vs Claude Fable 5.1
- Claude Haiku 4.5 vs Claude Opus 5.5
- Claude Code vs Codex CLI
- Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI)
- Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI)
- All comparisons
Write-ups on this study
Claude tokens per second and time to first token: six models timed, and what a longer prompt adds
60 timed calls: time to first text for Claude Haiku, Sonnet, Opus, Fable, GPT-6.1 Sol and Luna, their output speed, and what a 64k prompt adds.
More studies
All benchmarksHow much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.