• Latency
  • Time To First Token
  • Tokens Per Second
  • Output Speed
  • Prompt Size
  • Claude Code
  • Codex CLI
  • Claude Haiku
  • Claude Sonnet
  • Claude Opus
  • Claude Fable
  • GPT-6.1 Sol
  • Gpt 6 Luna

Where the seconds go: first text, output speed and prompt size for 6 LLMs

For 6 models run through their own coding CLIs (Haiku, Sonnet, Opus, Fable, Sol (low) and Luna (low)), how long until the first text, how fast does text stream after it, and what does a longer prompt add?

Published · 6 charts · Download the data or a carousel

1.8x

Calculation

n = 23

The same 4,327-character reply: most tokens ÷ fewest tokens across models (calculation)

Median visible tokens of exact replies per model (observed token range; n): Haiku 1,210 (1,210 to 1,210; n = 3). Sonnet 1,941 (1,941 to 1,941; n = 4). Opus 1,941 (1,941 to 1,941; n = 4). Fable 1,941 (1,941 to 1,941; n = 4). Sol (low) 1,066 (1,066 to 1,066; n = 4). Luna (low) 1,066 (1,066 to 1,066; n = 4). Each vendor counts the same text differently, so tokens per second do not compare across vendors. A calculation, not a run.

The answer

We measured CLI start-up, first text and total time on one 250-line answer and a ledger lookup at three prompt sizes. There were 60 counted calls: 4 per model in part A and 3 per size in part B. The receipts do not separate provider queueing, prompt processing and reasoning time. Time to first text, median (range; n calls): Sonnet 2.0 s (0.9 to 4.1; n = 4). Opus 2.0 s (1.7 to 2.4; n = 4). Luna (low) 3.3 s (3.2 to 3.5; n = 4). Sol (low) 3.5 s (2.7 to 4.4; n = 4). Haiku 4.0 s (2.8 to 6.4; n = 4). Fable 4.4 s (2.3 to 4.6; n = 4). Opus’s slowest call (2.4 s) was faster than the fastest call of Luna (low) (3.2 s), Sol (low) (2.7 s) and Haiku (2.8 s). Sonnet and Fable overlap it. There are 4 calls per model. The samples are below the protocol’s five-call minimum for naming a speed winner. Characters per second is a calculation: correct reply length ÷ time from first text to call end. Median (range; n exact replies): Haiku 547 (546 to 548; n = 3). Luna (low) 524 (225 to 1,052; n = 4). Sonnet 517 (513 to 519; n = 4). Opus 347 (345 to 349; n = 4). Sol (low) 323 (291 to 327; n = 4). Fable 273 (270 to 293; n = 4). Haiku and Fable have non-overlapping observed speed ranges (3 and 4 exact replies). Both samples are below the protocol’s five-call minimum for naming a speed winner. Luna (low)’s 4 exact replies ran from 225 to 1,052 characters per second, so its median hides this large spread. These few calls cannot establish two speed groups. Token rates use different token units across models. The same exact reply has 4,327 characters. Median visible tokens: 1,066 for Sol (low) (n = 4) and Luna (low) (n = 4). 1,210 for Haiku (n = 3). 1,941 for Sonnet (n = 4), Opus (n = 4) and Fable (n = 4). Visible tokens per second (calculation), median (range; n calls): Sonnet 232 (230 to 233; n = 4). Opus 156 (155 to 156; n = 4). Haiku 153 (153 to 216; n = 4). Luna (low) 129 (55 to 259; n = 4). Fable 123 (121 to 131; n = 4). Sol (low) 80 (72 to 80; n = 4). These ranges show individual calls, not confidence intervals. Of the 24 part A replies, 23 matched all 250 lines exactly; 1 was wrong (Haiku, 288 lines), and every call stays in the timings. First-text medians differ by prompt size (calculation: difference of medians, 64k minus 1k): Haiku +0.9 s (1.9 s to 2.8 s). Sonnet +1.6 s (1.4 s to 3.1 s, ranges overlap). Opus +0.3 s (1.5 s to 1.8 s, ranges overlap). Sol (low) +0.6 s (3.4 s to 3.9 s, ranges overlap). Exact lookup answers (95% Wilson intervals): Haiku 9/9 (70% to 100%). Sonnet 9/9 (70% to 100%). Opus 5/9 (27% to 81%). Sol (low) 9/9 (70% to 100%). All 4 models’ intervals overlap; this sample cannot rank their lookup rates. 4 Opus replies echoed the ledger line and then gave the right number last. The declared scoring counts these as wrong answers. The table shows the post-hoc reading. Cache reads did not grow with the ledger for any model. This fits no ledger reuse, but the counts do not prove which text the provider cached.

Key numbers

60

Counted calls in the speed anatomy study

calls (24 in part A, 36 in part B; 43 through Claude Code, 17 through Codex CLI) · n = 60

96% (23/24)

Part A replies that matched all 250 lines exactly (strict)

95% CI 80%–99% · n = 24

89% (32/36)

Part B lookups answered exactly (strict)

95% CI 75%–96% · n = 36

100% (36/36)

Part B replies whose last line was exactly the right number (post-hoc reading, not a pass)

95% CI 90%–100% · n = 36

0 of 4

Models whose cache reads grew with the ledger size

n = 36

+0.9s

Haiku: extra time to first text at 64k vs 1k (calculation)

n = 6

+1.6s

Sonnet: extra time to first text at 64k vs 1k (calculation)

n = 6

+0.3s

Opus: extra time to first text at 64k vs 1k (calculation)

n = 6

+0.6s

Sol (low): extra time to first text at 64k vs 1k (calculation)

n = 6

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Entrance: medians race at 3.2× real timeMotion reduced: press Replay to animateThe slowest median is 4.4 s. The clock runs at the recorded speed.
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6 Luna (low) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

6 rows. Slowest Claude Fable 5.1 · Claude Code 4.4 s (range 2.3 s–4.6 s, n 4). Fastest Claude Sonnet 5.5 · Claude Code 2 s (range 0.9 s–4.1 s, n 4). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 4 per row

Median of 4 calls per model; whiskers = fastest and slowest call

Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.

Source: LLM speed anatomy

Share card (PNG)
Calculation
Claude Haiku 4.5
Claude Sonnet 5.5
Claude Opus 5.5
Claude Fable 5.1
GPT-6.1 Sol (low)
GPT-6 Luna (low)

List-price calculation, not a run. 6 rows. Highest Claude Sonnet 5.5 · Claude Code 232 (range 230–233, n 4). Lowest GPT-6.1 Sol (low) · Codex CLI 80 (range 72–81, n 4). Not all run ranges overlap.

NotesLines: lowest–highest run (not an interval)n = 4 per row

Median of 4 calls per model; whiskers = slowest and fastest call

Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.

Source: LLM speed anatomy

Share card (PNG)
Calculation
Claude Haiku 4.5
Claude Sonnet 5.5
Claude Opus 5.5
Claude Fable 5.1
GPT-6.1 Sol (low)
GPT-6 Luna (low)

List-price calculation, not a run. 6 rows. Highest Claude Haiku 4.5 · Claude Code 547 (range 546–548, n 3). Lowest Claude Fable 5.1 · Claude Code 273 (range 270–293, n 4). Not all run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 3–4 per row

Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call

Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.

Source: LLM speed anatomy

Share card (PNG)
Calculation
  • Claude Haiku 4.5 · Claude Code
  • Claude Sonnet 5.5 · Claude Code
  • Claude Opus 5.5 · Claude Code
  • GPT-6.1 Sol (low) · Codex CLI

List-price calculation, not a run. 3 rows, 4 series: Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (low) · Codex CLI. Claude Haiku 4.5 · Claude Code: slowest 64k 2.8 s (range 2.5 s–2.9 s, n 3). Fastest 1k 1.9 s (range 1.9 s–2 s, n 3). Not all run ranges overlap. Claude Sonnet 5.5 · Claude Code: slowest 64k 3.1 s (range 1.4 s–3.6 s, n 3). Fastest 1k 1.5 s (range 1.2 s–1.7 s, n 3). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Median of 3 calls per size; whiskers = fastest and slowest call

Each call used a new ledger seed. Cache-read counts stayed within the short-prompt baseline (see the cache table). This does not identify which tokens were cached. Sizes name the text we send; each model’s reported input tokens are in the table and include the CLI’s own prefix. The size calibration subtracts estimated prefixes from probe input counts; these are calculations, not measured prefix counts for each call. Whiskers are a range of calls, not a confidence interval.

Source: LLM speed anatomy

Share card (PNG)

1k prompt

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI

16k prompt

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI

64k prompt

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (low) · Codex CLI

One panel per series, all on the same axis; whiskers are the fastest–slowest run (not an interval).

4 rows, 3 series: 1k prompt, 16k prompt, 64k prompt. 1k prompt: slowest GPT-6.1 Sol (low) · Codex CLI 3.4 s (range 3.4 s–4.9 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 1.8 s (range 1.6 s–2.1 s, n 3). Not all run ranges overlap. 16k prompt: slowest GPT-6.1 Sol (low) · Codex CLI 4.1 s (range 4 s–4.7 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 2.1 s (range 2 s–2.5 s, n 3). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Median of 3 calls per bar; whiskers = fastest and slowest call

Whole call: CLI start-up, first text and the one-line answer. Whiskers are a range of calls, not a confidence interval. Each call used a new ledger.

Source: LLM speed anatomy

Share card (PNG)
Claude Haiku 4.5
Claude Sonnet 5.5
Claude Opus 5.5
GPT-6.1 Sol (low)

Every interval overlaps every other: this chart does not order these rows.

4 rows. Highest Claude Haiku 4.5 · Claude Code 100% (95% interval 70%–100%, n 9). Lowest Claude Opus 5.5 · Claude Code 56% (95% interval 27%–81%, n 9). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 9 per row3 of 4 at 100%: this task set cannot separate them.

All sizes together per model; whiskers = 95% Wilson intervals

Whiskers are 95% Wilson intervals. One lookup question per call; a reply with extra words is a format miss, not a pass. With 9 calls per model, a perfect score still has a wide interval.

Source: LLM speed anatomy

Share card (PNG)

Tables

Every speed-anatomy cell, with reasoning tokens

PartConfigurationPrompt sizeCallsMedian first text (s)Fastest to slowest (s)Median first text after the CLI was ready (s, calculation)After ready, fastest to slowest (s)Median total (s)Total, fastest to slowest (s)Median input tokens reportedInput tokens, lowest to highestMedian output tokensOutput tokens, lowest to highestMedian reasoning tokensReasoning tokens, lowest to highestMedian visible tokensVisible tokens, lowest to highestVisible tokens per second (calculation)Visible tokens/s, lowest to highestAll output tokens per second (calculation; inflated when reasoning comes first)All output tokens/s, lowest to highestCharacters per second (calculation; exact replies only)Characters/s, lowest to highestExact replies used for characters/sExactly rightNot a pass, but the last line is the right number (post-hoc reading)
A: 250 numbers in wordsClaude Haiku 4.5 · Claude Code—44 s2.84 to 6.383.5 s2.36 to 5.7811.9 s10.91 to 14.313,6583655 to 36581,7801510 to 1997413255 to 6141,2101210 to 1742153152.6 to 216.1225190.9 to 247.7547546 to 54833/4—
A: 250 numbers in wordsClaude Sonnet 5.5 · Claude Code—42 s0.88 to 4.091.5 s0.42 to 3.6410.3 s9.31 to 12.421,9131912 to 19142,0241941 to 2339830 to 3981,9411941 to 1941232230.3 to 233242230.3 to 280.8517513 to 51944/4—
A: 250 numbers in wordsClaude Opus 5.5 · Claude Code—42 s1.7 to 2.351.5 s1.24 to 1.8714.4 s14.25 to 14.81,9091908 to 19091,9901985 to 19944944 to 531,9411941 to 1941156154.6 to 156.4159158.7 to 160.3347345 to 34944/4—
A: 250 numbers in wordsClaude Fable 5.1 · Claude Code—44.4 s2.27 to 4.644 s1.65 to 4.2119.8 s18.32 to 20.313,7683766 to 37682,0921941 to 21561510 to 2151,9411941 to 1941123120.9 to 131.4131123.2 to 145.9273270 to 29344/4—
A: 250 numbers in wordsGPT-6.1 Sol (low) · Codex CLI—43.5 s2.75 to 4.423.1 s2.34 to 4.0117.5 s16.36 to 17.7512,02812026 to 120281,0661066 to 106600 to 01,0661066 to 10668071.6 to 80.58071.6 to 80.5323291 to 32744/4—
A: 250 numbers in wordsGPT-6 Luna (low) · Codex CLI—43.3 s3.19 to 3.472.5 s2.14 to 2.8815.5 s7.39 to 22.6911,33511332 to 113361,0661066 to 106600 to 01,0661066 to 106612955.5 to 259.112955.5 to 259.1524225 to 105244/4—
B: ledger lookupClaude Haiku 4.5 · Claude Code1k31.9 s1.85 to 2.041.4 s1.38 to 1.432.3 s2.22 to 2.464,5924580 to 4610130126 to 133123119 to 12677 to 71716.5 to 18.9314296.5 to 358.577 to 833/30
B: ledger lookupClaude Haiku 4.5 · Claude Code16k32.3 s2.22 to 2.471.8 s1.64 to 2.012.8 s2.58 to 2.8419,61819604 to 19632136129 to 139129122 to 13277 to 71913.5 to 19.4358268.3 to 371.686 to 833/30
B: ledger lookupClaude Haiku 4.5 · Claude Code64k32.8 s2.45 to 2.892.3 s1.94 to 2.423.1 s2.84 to 3.2867,60567578 to 67639157132 to 167150125 to 16077 to 71817.9 to 19.9402343.8 to 475.888 to 933/30
B: ledger lookupClaude Sonnet 5.5 · Claude Code1k31.5 s1.23 to 1.721 s0.8 to 0.981.8 s1.57 to 2.123,0513049 to 305733 to 300 to 033 to 397.4 to 9.197.4 to 9.197 to 933/30
B: ledger lookupClaude Sonnet 5.5 · Claude Code16k31.8 s1.64 to 2.111.3 s1.15 to 1.322.1 s1.98 to 2.4821,08021050 to 2109933 to 300 to 033 to 398.2 to 9.298.2 to 9.295 to 933/30
B: ledger lookupClaude Sonnet 5.5 · Claude Code64k33.1 s1.38 to 3.612.4 s0.94 to 3.133.4 s1.74 to 4.3878,59678577 to 7874533 to 300 to 033 to 383.9 to 8.383.9 to 8.384 to 833/30
B: ledger lookupClaude Opus 5.5 · Claude Code1k31.5 s1.46 to 2.011 s1 to 1.361.8 s1.82 to 2.413,0403040 to 304433 to 300 to 033 to 387.5 to 9.487.5 to 9.488 to 933/30
B: ledger lookupClaude Opus 5.5 · Claude Code16k31.7 s1.7 to 2.971.3 s1.19 to 2.532.4 s2.11 to 3.421,07420998 to 2108933 to 3600 to 033 to 3677 to 58.377 to 58.377 to 722/31
B: ledger lookupClaude Opus 5.5 · Claude Code64k31.8 s1.72 to 3.721.3 s1.29 to 3.252.4 s2.26 to 4.2978,67478633 to 786843331 to 3400 to 03331 to 345857.5 to 60.45857.5 to 60.4——00/33
B: ledger lookupGPT-6.1 Sol (low) · Codex CLI1k33.4 s3.36 to 4.753 s2.76 to 3.293.4 s3.43 to 4.9212,79012787 to 1279755 to 500 to 055 to 56828.4 to 68.56828.4 to 68.54117 to 4133/30
B: ledger lookupGPT-6.1 Sol (low) · Codex CLI16k34 s3.3 to 4.283.7 s2.88 to 3.954.1 s3.96 to 4.6825,03825030 to 2504555 to 500 to 055 to 5137.5 to 39.4137.5 to 39.475 to 2433/30
B: ledger lookupGPT-6.1 Sol (low) · Codex CLI64k33.9 s3.42 to 4.383.6 s3.02 to 4.064 s3.47 to 4.4464,17264118 to 6419855 to 500 to 055 to 58673.5 to 166.78673.5 to 166.75244 to 10033/30

Cache reads and writes per call; share = cache-read tokens ÷ reported input tokens (calculation)

ConfigurationPrompt sizeCallsMedian input tokens reportedInput tokens, lowest to highestMedian cache-read tokensRead tokens, lowest to highestMost cache-read tokens in one callMedian cache-write tokensWrite tokens, lowest to highestMedian cache-read share of input (calculation)Cache-read share, lowest to highest (calculation)Calls used for cache-read share
Claude Haiku 4.5 · Claude Code1k34,5924580 to 461000 to 004,5824570 to 46000%0.0% to 0.0%3
Claude Haiku 4.5 · Claude Code16k319,61819604 to 1963200 to 0019,60819594 to 196220%0.0% to 0.0%3
Claude Haiku 4.5 · Claude Code64k367,60567578 to 6763900 to 0067,59567568 to 676290%0.0% to 0.0%3
Claude Sonnet 5.5 · Claude Code1k33,0513049 to 30571,4631463 to 14631,4631,5861584 to 159248%47.9% to 48.0%3
Claude Sonnet 5.5 · Claude Code16k321,08021050 to 210991,4631463 to 14631,46319,61519585 to 196346.9%6.9% to 7.0%3
Claude Sonnet 5.5 · Claude Code64k378,59678577 to 787451,4631463 to 14631,46377,13177112 to 772801.9%1.9% to 1.9%3
Claude Opus 5.5 · Claude Code1k33,0403040 to 30441,4631463 to 14631,4631,5751575 to 157948%48.1% to 48.1%3
Claude Opus 5.5 · Claude Code16k321,07420998 to 210891,4631463 to 14631,46319,60919533 to 196246.9%6.9% to 7.0%3
Claude Opus 5.5 · Claude Code64k378,67478633 to 786841,4631463 to 14631,46377,20977168 to 772191.9%1.9% to 1.9%3
GPT-6.1 Sol (low) · Codex CLI1k312,79012787 to 127978,9600 to 89608,96000 to 070%0.0% to 70.1%3
GPT-6.1 Sol (low) · Codex CLI16k325,03825030 to 250458,9608960 to 89608,96000 to 036%35.8% to 35.8%3
GPT-6.1 Sol (low) · Codex CLI64k364,17264118 to 641988,9608960 to 89608,96000 to 014%14.0% to 14.0%3

Method

  1. The protocol file birth precedes the first probe and counted call. Later edits have no frozen versions. Two probes checked the routes and helped calibrate ledger sizes. They stay outside all cells and charts.
  2. Part A used 24 calls: 4 per model. The prompt asked for 250 numbers in words, one per line. The check compares the reply with all 250 expected lines. Models: Haiku, Sonnet, Opus, Fable, Sol (low) and Luna (low). Controls ran before the first counted call. The reference passes. All 8/8 wrapped references are format misses. All 11/11 planted wrong answers fail.
  3. Part B used 36 calls: 3 per size for each model. Models: Haiku, Sonnet, Opus and Sol (low). Each call asks one exact lookup question about a synthetic stock ledger. Ledger sizes: 18 lines for 1k, 327 lines for 16k and 1,315 lines for 64k. The 1k, 16k and 64k labels are approximate targets from probe calibration calculations, using estimated Haiku prefix counts. A new seed changes each ledger. Some fixed text repeats. Cache counts do not identify cached text.
  4. First text means the first non-empty text after CLI start. The time includes CLI start-up and work before the first word. The table also subtracts the CLI ready time (calculation).
  5. Output speed is a calculation. Visible tokens equal output tokens minus reported reasoning tokens. When the CLI reports no reasoning count, the calculation uses zero. That does not prove the model did no reasoning. Divide visible tokens by the time from first text to call end. The plain form divides all output tokens by that time. It can overstate visible speed when reasoning comes first.
  6. The harness uses the hard head-to-head machinery. It disables tools, MCP servers and saved sessions. Claude Code uses default effort; Codex CLI uses low effort. Calls run one at a time on one Mac. The timeout was 300 s. Claude Code had a 16,000-token output cap. Codex had no output-token cap. CLI versions: 2.1.286 (Claude Code) and codex-cli 0.160.0.
  7. Stop rules require a stop at the first usage-limit text. The gate checks other study markers before each batch. No stored batch stopped or trimmed a cell. No counted call duplicates another. A sequence process ended at the gate. Codex part B ran later. Every counted call completed.
  8. Cache-read share is a calculation: cache-read tokens divided by reported input tokens for each call. The cache table shows the median, observed range and n for each cell.

Caveats

  • 4 calls per model in part A and 3 per size in part B. Medians of so few calls move with one slow call, and the ranges are not confidence intervals. The 95% Wilson intervals on the lookup rates are wide.
  • First text includes CLI start-up and reasoning time. It is not the API’s time to first token; a direct API call would skip the CLI start-up (the CLI reported ready after a median 0.33 s to 0.84 s per cell).
  • Claude Code ran models at their default effort; Codex CLI ran at low effort. Effort levels are not one scale, and a CLI adds its own system prompt, so a row mixes a model and its CLI.
  • Output speed is a calculation from reported token counts, the length of the reply and the clock. It depends on how each CLI counts reasoning, and part A has one prompt: other text, other days or an API route can differ.
  • All 36 counted ledger prompts differ. Their input counts do not compare tokenizers on the same text. The size calibration uses prefix estimates from earlier short prompts. Those calculations do not measure the prefix in each counted call.
  • The tasks hit a ceiling. Part A passed 23/24. In part B, 3 of 4 models passed every lookup. See the rate stats and chart for 95% intervals. This small task set cannot rank general capability.
  • One shared Mac: the gate checked other study markers before each batch. No full host-load record proves that all other work stopped.
  • Prompt sizes were calibrated after two probes. The cases are synthetic and tuned to this measurement, not a sample of real workloads.
  • Amendment 3 says 06:05 UTC and claims to precede Codex part B. Those calls ran at 06:03:43 to 06:04:19 UTC. The same-file amendment is retrospective; its date does not prove advance declaration.
  • The controls receipt was overwritten. The surviving checks ran after the probes and before counted calls; an earlier control run cannot be audited.
  • Each size used different ledger text. Prompt-size differences also include content and provider-load changes; they do not isolate a cause.
  • The failed Haiku reply contained tool-shaped text and extra output. No structured tool event was recorded; it stays a strict failure and remains in timings.
  • One Mac, one network, two sessions on one night. Provider load changes timings from hour to hour. The Codex CLI part B batch ran about 4 hours after the batches before it (a run script stopped while it waited for another run to finish timing), so its timings come from a later hour.

Sources

  • LLM speed anatomy

    Our recorded runs ·

    Receipts of the speed anatomy study: one prompt that asks for 250 numbers in words (output speed, 24 calls) and a seeded synthetic ledger at three sizes with one lookup question (prompt size, 36 calls). Every attempt is kept, failures included. A new seed for every ledger call; no ledger text, prompt or model output is copied.

    Raw data: speed-anatomy/receipts.json

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Where the seconds go: first text, output speed and prompt size for 6 LLMs”, updated October 7, 2026, https://agent.sasid.ai/benchmarks/llm-speed-anatomy.

More comparisons based on this study (6)

These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.

Models and comparisons in this study

More studies

All benchmarks
Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Live story
  • Head to head
  • Hard Tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks

152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.

91% (139/152)Calls that passed strictly (hard set) · n = 152

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.