{"i":23,"study":{"slug":"llm-speed-anatomy","title":"Where the seconds go: first text, output speed and prompt size for 6 LLMs","seoTitle":"LLM speed: first text and tokens per second through CLIs","description":"60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.","question":"For 6 models run through their own coding CLIs (Haiku, Sonnet, Opus, Fable, Sol (low) and Luna (low)), how long until the first text, how fast does text stream after it, and what does a longer prompt add?","answer":"We measured CLI start-up, first text and total time on one 250-line answer and a ledger lookup at three prompt sizes. There were 60 counted calls: 4 per model in part A and 3 per size in part B. The receipts do not separate provider queueing, prompt processing and reasoning time. Time to first text, median (range; n calls): Sonnet 2.0 s (0.9 to 4.1; n = 4). Opus 2.0 s (1.7 to 2.4; n = 4). Luna (low) 3.3 s (3.2 to 3.5; n = 4). Sol (low) 3.5 s (2.7 to 4.4; n = 4). Haiku 4.0 s (2.8 to 6.4; n = 4). Fable 4.4 s (2.3 to 4.6; n = 4). Opus’s slowest call (2.4 s) was faster than the fastest call of Luna (low) (3.2 s), Sol (low) (2.7 s) and Haiku (2.8 s). Sonnet and Fable overlap it. There are 4 calls per model. The samples are below the protocol’s five-call minimum for naming a speed winner. Characters per second is a calculation: correct reply length ÷ time from first text to call end. Median (range; n exact replies): Haiku 547 (546 to 548; n = 3). Luna (low) 524 (225 to 1,052; n = 4). Sonnet 517 (513 to 519; n = 4). Opus 347 (345 to 349; n = 4). Sol (low) 323 (291 to 327; n = 4). Fable 273 (270 to 293; n = 4). Haiku and Fable have non-overlapping observed speed ranges (3 and 4 exact replies). Both samples are below the protocol’s five-call minimum for naming a speed winner. Luna (low)’s 4 exact replies ran from 225 to 1,052 characters per second, so its median hides this large spread. These few calls cannot establish two speed groups. Token rates use different token units across models. The same exact reply has 4,327 characters. Median visible tokens: 1,066 for Sol (low) (n = 4) and Luna (low) (n = 4). 1,210 for Haiku (n = 3). 1,941 for Sonnet (n = 4), Opus (n = 4) and Fable (n = 4). Visible tokens per second (calculation), median (range; n calls): Sonnet 232 (230 to 233; n = 4). Opus 156 (155 to 156; n = 4). Haiku 153 (153 to 216; n = 4). Luna (low) 129 (55 to 259; n = 4). Fable 123 (121 to 131; n = 4). Sol (low) 80 (72 to 80; n = 4). These ranges show individual calls, not confidence intervals. Of the 24 part A replies, 23 matched all 250 lines exactly; 1 was wrong (Haiku, 288 lines), and every call stays in the timings. First-text medians differ by prompt size (calculation: difference of medians, 64k minus 1k): Haiku +0.9 s (1.9 s to 2.8 s). Sonnet +1.6 s (1.4 s to 3.1 s, ranges overlap). Opus +0.3 s (1.5 s to 1.8 s, ranges overlap). Sol (low) +0.6 s (3.4 s to 3.9 s, ranges overlap). Exact lookup answers (95% Wilson intervals): Haiku 9/9 (70% to 100%). Sonnet 9/9 (70% to 100%). Opus 5/9 (27% to 81%). Sol (low) 9/9 (70% to 100%). All 4 models’ intervals overlap; this sample cannot rank their lookup rates. 4 Opus replies echoed the ledger line and then gave the right number last. The declared scoring counts these as wrong answers. The table shows the post-hoc reading. Cache reads did not grow with the ledger for any model. This fits no ledger reuse, but the counts do not prove which text the provider cached.","date":"2026-10-07","updated":"2026-10-07","tags":["latency","time-to-first-token","tokens-per-second","output-speed","prompt-size","claude-code","codex-cli","claude-haiku","claude-sonnet","claude-opus","claude-fable","gpt-6-1-sol","gpt-6-luna"],"caveats":["4 calls per model in part A and 3 per size in part B. Medians of so few calls move with one slow call, and the ranges are not confidence intervals. The 95% Wilson intervals on the lookup rates are wide.","First text includes CLI start-up and reasoning time. It is not the API’s time to first token; a direct API call would skip the CLI start-up (the CLI reported ready after a median 0.33 s to 0.84 s per cell).","Claude Code ran models at their default effort; Codex CLI ran at low effort. Effort levels are not one scale, and a CLI adds its own system prompt, so a row mixes a model and its CLI.","Output speed is a calculation from reported token counts, the length of the reply and the clock. It depends on how each CLI counts reasoning, and part A has one prompt: other text, other days or an API route can differ.","All 36 counted ledger prompts differ. Their input counts do not compare tokenizers on the same text. The size calibration uses prefix estimates from earlier short prompts. Those calculations do not measure the prefix in each counted call.","The tasks hit a ceiling. Part A passed 23/24. In part B, 3 of 4 models passed every lookup. See the rate stats and chart for 95% intervals. This small task set cannot rank general capability.","One shared Mac: the gate checked other study markers before each batch. No full host-load record proves that all other work stopped.","Prompt sizes were calibrated after two probes. The cases are synthetic and tuned to this measurement, not a sample of real workloads.","Amendment 3 says 06:05 UTC and claims to precede Codex part B. Those calls ran at 06:03:43 to 06:04:19 UTC. The same-file amendment is retrospective; its date does not prove advance declaration.","The controls receipt was overwritten. The surviving checks ran after the probes and before counted calls; an earlier control run cannot be audited.","Each size used different ledger text. Prompt-size differences also include content and provider-load changes; they do not isolate a cause.","The failed Haiku reply contained tool-shaped text and extra output. No structured tool event was recorded; it stays a strict failure and remains in timings.","One Mac, one network, two sessions on one night. Provider load changes timings from hour to hour. The Codex CLI part B batch ran about 4 hours after the batches before it (a run script stopped while it waited for another run to finish timing), so its timings come from a later hour."],"sourceIds":["agent-speed-anatomy"],"hero":{"statIds":["speed-anatomy-token-count-ratio"]},"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["speed-anatomy-calls","Counted calls in the speed anatomy study",60,"calls","60 calls (24 in part A, 36 in part B; 43 through Claude Code, 17 through Codex CLI)",60,"\u0001","\u0001"],["speed-anatomy-part-a-exact","Part A replies that matched all 250 lines exactly (strict)",0.9583,"rate","96% (23/24)",24,[0.7976,0.9926],"0 more were right but wrapped or laid out differently (format misses); 1 was wrong."],["speed-anatomy-lookup-exact","Part B lookups answered exactly (strict)",0.8889,"rate","89% (32/36)",36,[0.7469,0.9559],"0 more were right with extra words (format misses); 4 were wrong."],["speed-anatomy-token-count-ratio","The same 4,327-character reply: most tokens ÷ fewest tokens across models (calculation)",1.8,"ratio","1.8x",23,"\u0001","Median visible tokens of exact replies per model (observed token range; n): Haiku 1,210 (1,210 to 1,210; n = 3). Sonnet 1,941 (1,941 to 1,941; n = 4). Opus 1,941 (1,941 to 1,941; n = 4). Fable 1,941 (1,941 to 1,941; n = 4). Sol (low) 1,066 (1,066 to 1,066; n = 4). Luna (low) 1,066 (1,066 to 1,066; n = 4). Each vendor counts the same text differently, so tokens per second do not compare across vendors. A calculation, not a run."],["speed-anatomy-lookup-last-line","Part B replies whose last line was exactly the right number (post-hoc reading, not a pass)",1,"rate","100% (36/36)",36,[0.9036,1],"A reading declared after the Claude part B batch (protocol Amendment 2), derived from the stored replies. The strict score above stays the headline; this one is never counted as a pass or a format miss."],["speed-anatomy-cache-reads-grew","Models whose cache reads grew with the ledger size",0,"count","0 of 4",36,"\u0001","A new ledger for every call. The descriptive threshold is the 1k-prompt maximum plus 10% and 128 tokens. Growth does not prove ledger reuse; the cache table lists every cell."],["speed-anatomy-size-cost-haiku","Haiku: extra time to first text at 64k vs 1k (calculation)",0.86,"seconds","+0.9 s",6,"\u0001","Difference of two medians (2.78 s minus 1.93 s); the ranges are 1.85 to 2.04 s and 2.45 to 2.89 s. A calculation, not a run. Each call used separately seeded text. The table shows reported input counts, not a matched-text tokenizer comparison."],["speed-anatomy-size-cost-sonnet","Sonnet: extra time to first text at 64k vs 1k (calculation)",1.62,"seconds","+1.6 s",6,"\u0001","Difference of two medians (3.07 s minus 1.45 s); the ranges are 1.23 to 1.72 s and 1.38 to 3.61 s and overlap. A calculation, not a run. Each call used separately seeded text. The table shows reported input counts, not a matched-text tokenizer comparison."],["speed-anatomy-size-cost-opus","Opus: extra time to first text at 64k vs 1k (calculation)",0.28,"seconds","+0.3 s",6,"\u0001","Difference of two medians (1.79 s minus 1.51 s); the ranges are 1.46 to 2.01 s and 1.72 to 3.72 s and overlap. A calculation, not a run. Each call used separately seeded text. The table shows reported input counts, not a matched-text tokenizer comparison."],["speed-anatomy-size-cost-sol-low","Sol (low): extra time to first text at 64k vs 1k (calculation)",0.57,"seconds","+0.6 s",6,"\u0001","Difference of two medians (3.93 s minus 3.36 s); the ranges are 3.36 to 4.75 s and 3.42 to 4.38 s and overlap. A calculation, not a run. Each call used separately seeded text. The table shows reported input counts, not a matched-text tokenizer comparison."]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","whisker","sourceIds","xLabel","polarity"],"$r":[["speed-anatomy-first-text","Time to first text: a 250-line answer, six models","Median of 4 calls per model; whiskers = fastest and slowest call","dot-range","seconds","Seconds to first text",[{"name":"Time to first text","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",4,2.84,6.38,4],["Claude Sonnet 5.5 · Claude Code",1.96,0.88,4.09,4],["Claude Opus 5.5 · Claude Code",1.97,1.7,2.35,4],["Claude Fable 5.1 · Claude Code",4.43,2.27,4.64,4],["GPT-6.1 Sol (low) · Codex CLI",3.52,2.75,4.42,4],["GPT-6 Luna (low) · Codex CLI",3.3,3.19,3.47,4]]}}],"Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.","minmax",["agent-speed-anatomy"],"\u0001","\u0001"],["speed-anatomy-output-speed","Output speed after the first text: visible tokens per second (calculation)","Median of 4 calls per model; whiskers = slowest and fastest call","dot-range","tokens","Visible tokens per second",[{"name":"Visible tokens per second","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",153.2,152.6,216.1,4],["Claude Sonnet 5.5 · Claude Code",231.7,230.3,233,4],["Claude Opus 5.5 · Claude Code",155.5,154.6,156.4,4],["Claude Fable 5.1 · Claude Code",122.6,120.9,131.4,4],["GPT-6.1 Sol (low) · Codex CLI",79.6,71.6,80.5,4],["GPT-6 Luna (low) · Codex CLI",129.1,55.5,259.1,4]]}}],"Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.","minmax",["agent-speed-anatomy"],"\u0001","\u0001"],["speed-anatomy-chars-per-second","Output speed in characters per second after the first text (calculation)","Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call","dot-range","count","Characters per second",[{"name":"Characters per second","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",547,546,548,3],["Claude Sonnet 5.5 · Claude Code",517,513,519,4],["Claude Opus 5.5 · Claude Code",347,345,349,4],["Claude Fable 5.1 · Claude Code",273,270,293,4],["GPT-6.1 Sol (low) · Codex CLI",323,291,327,4],["GPT-6 Luna (low) · Codex CLI",524,225,1052,4]]}}],"Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.","minmax",["agent-speed-anatomy"],"\u0001","\u0001"],["speed-anatomy-prompt-size","Time to first text as the prompt grows","Median of 3 calls per size; whiskers = fastest and slowest call","line","seconds","Seconds to first text",{"$k":["name","points"],"$r":[["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["1k",1.93,1.85,2.04,3],["16k",2.27,2.22,2.47,3],["64k",2.78,2.45,2.89,3]]}],["Claude Sonnet 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["1k",1.45,1.23,1.72,3],["16k",1.78,1.64,2.11,3],["64k",3.07,1.38,3.61,3]]}],["Claude Opus 5.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["1k",1.51,1.46,2.01,3],["16k",1.74,1.7,2.97,3],["64k",1.79,1.72,3.72,3]]}],["GPT-6.1 Sol (low) · Codex CLI",{"$k":["label","value","lo","hi","n"],"$r":[["1k",3.36,3.36,4.75,3],["16k",4.02,3.3,4.28,3],["64k",3.93,3.42,4.38,3]]}]]},"Each call used a new ledger seed. Cache-read counts stayed within the short-prompt baseline (see the cache table). This does not identify which tokens were cached. Sizes name the text we send; each model’s reported input tokens are in the table and include the CLI’s own prefix. The size calibration subtracts estimated prefixes from probe input counts; these are calculations, not measured prefix counts for each call. Whiskers are a range of calls, not a confidence interval.","minmax",["agent-speed-anatomy"],"Prompt-size target (approximate Haiku tokens; calibration calculation)","\u0001"],["speed-anatomy-total-by-size","Total time per call by prompt size","Median of 3 calls per bar; whiskers = fastest and slowest call","grouped-bar","seconds","Seconds, whole call",{"$k":["name","points"],"$r":[["1k prompt",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",2.34,2.22,2.46,3],["Claude Sonnet 5.5 · Claude Code",1.78,1.57,2.12,3],["Claude Opus 5.5 · Claude Code",1.83,1.82,2.41,3],["GPT-6.1 Sol (low) · Codex CLI",3.43,3.43,4.92,3]]}],["16k prompt",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",2.79,2.58,2.84,3],["Claude Sonnet 5.5 · Claude Code",2.1,1.98,2.48,3],["Claude Opus 5.5 · Claude Code",2.36,2.11,3.4,3],["GPT-6.1 Sol (low) · Codex CLI",4.14,3.96,4.68,3]]}],["64k prompt",{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",3.13,2.84,3.28,3],["Claude Sonnet 5.5 · Claude Code",3.44,1.74,4.38,3],["Claude Opus 5.5 · Claude Code",2.35,2.26,4.29,3],["GPT-6.1 Sol (low) · Codex CLI",3.96,3.47,4.44,3]]}]]},"Whole call: CLI start-up, first text and the one-line answer. Whiskers are a range of calls, not a confidence interval. Each call used a new ledger.","minmax",["agent-speed-anatomy"],"\u0001","\u0001"],["speed-anatomy-lookup-correct","Exact lookup answers at the 1k, 16k and 64k prompt-size targets","All sizes together per model; whiskers = 95% Wilson intervals","dot-range","rate","Lookups answered exactly",[{"name":"Exact answer","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 · Claude Code",1,0.7009,1,9],["Claude Sonnet 5.5 · Claude Code",1,0.7009,1,9],["Claude Opus 5.5 · Claude Code",0.5556,0.2667,0.8112,9],["GPT-6.1 Sol (low) · Codex CLI",1,0.7009,1,9]]}}],"Whiskers are 95% Wilson intervals. One lookup question per call; a reply with extra words is a format miss, not a pass. With 9 calls per model, a perfect score still has a wide interval.","ci95",["agent-speed-anatomy"],"\u0001","higher"]]},"related":["cli-model-latency-tokens","model-head-to-head","routing-overhead"]}}