Claude tokens per second and time to first token: six models timed, and what a longer prompt adds
60 timed calls: time to first text for Claude Haiku, Sonnet, Opus, Fable, GPT-6.1 Sol and Luna, their output speed, and what a 64k prompt adds.
TL;DR
- Time to first text on a 250-line answer, median of 4 calls per model: Sonnet 5.5 2.0 s (0.9–4.1), Opus 5.5 2.0 s (1.7–2.4), GPT-6 Luna 3.3 s (3.2–3.5), GPT-6.1 Sol 3.5 s (2.7–4.4), Haiku 4.5 4.0 s (2.8–6.4), Fable 5.1 4.4 s (2.3–4.6). Parentheses show observed ranges in seconds. Opus's slowest call (2.4 s) was faster than the fastest call of Haiku, Sol and Luna. With 4 calls each, that is a pattern, not a ranking.
- Output speed (a calculation: characters per second after the first text): Haiku 547 (546–548; n = 3 exact replies), Sonnet 517 (513–519), Opus 347 (345–349), Sol 323 (291–327), Fable 273 (270–293). The other models each have 4 exact replies; parentheses show ranges. Luna's 4 calls ran from 225 to 1,052.
- Do not compare tokens per second across vendors. The same 4,327-character reply counted as 1,066 tokens for GPT-6.1 Sol and Luna, 1,210 for Haiku and 1,941 for Sonnet, Opus and Fable.
- Prompt-size medians differed by little time. From about 1k to about 64k tokens of text, the median time to first text differed by +0.3 s to +1.6 s (a calculation; 3 calls per size). Only Haiku's ranges separate (+0.9 s). For the other three, the ranges at 1k and 64k overlap.
- Test output length first. In these runs, 1,000 output tokens took 4.3 s to 12.6 s to stream (a calculation on medians). The prompt-size differences imply at most about 0.02 s per 1,000 reported input tokens (a calculation). Different ledger text and provider load can affect those differences.
- Cache counts fit no ledger reuse. Cache reads did not grow with prompt size in 36 calls. The counts do not identify cached text.
The protocol requires at least 5 calls per side before naming a speed winner. These samples fall below that minimum.
Study and data: Where the seconds go: speed anatomy, also at https://agent.sasid.ai/benchmarks/llm-speed-anatomy. It holds 60 timed calls. Public receipts hold the per-call counts and times. Earlier CLI timings: CLI vs API latency.
What "time to first token" means here
We did not call an API. We ran each model through its own coding CLI: Claude Code for Haiku 4.5, Sonnet 5.5, Opus 5.5 and Fable 5.1, and the Codex CLI for GPT-6.1 Sol and GPT-6 Luna. So our clock measures time to first text. It starts when the CLI process starts and stops at the first non-empty piece of text.
That time can include several steps. These receipts do not separate provider queueing, prompt processing and reasoning time:
- CLI start-up. The CLI ready-time medians were 0.33 s to 0.84 s across cells (3 or 4 calls each). Individual ready times ranged from 0.31 s to 1.46 s.
- Thinking before the first word. Some models write reasoning tokens first.
- The model's own first token.
The earlier start-up test shows the same effect on a one-word answer:
A direct API call would skip item 1. We did not measure an API call in this study.
1. Time to first text: six models, one 250-line answer
The prompt was the same for every call: write the whole numbers from 1 to 250 in English words, one per line, and nothing else. A strict check compared each reply with the 250 expected lines. We ran 4 calls per model, one at a time.
| Model and CLI | First text, median (fastest to slowest) | Total time, median (fastest to slowest) | Visible tokens (median, exact replies) | Reasoning tokens (median) | Exact replies (95% Wilson interval) |
|---|---|---|---|---|---|
| Haiku 4.5, Claude Code | 4.00 s (2.84 to 6.38) | 11.90 s (10.91 to 14.31) | 1,210 | 413 | 3/4 (30%–95%) |
| Sonnet 5.5, Claude Code | 1.96 s (0.88 to 4.09) | 10.33 s (9.31 to 12.42) | 1,941 | 83 | 4/4 (51%–100%) |
| Opus 5.5, Claude Code | 1.97 s (1.70 to 2.35) | 14.43 s (14.25 to 14.80) | 1,941 | 49 | 4/4 (51%–100%) |
| Fable 5.1, Claude Code | 4.43 s (2.27 to 4.64) | 19.81 s (18.32 to 20.31) | 1,941 | 151 | 4/4 (51%–100%) |
| GPT-6.1 Sol (low), Codex CLI | 3.52 s (2.75 to 4.42) | 17.51 s (16.36 to 17.75) | 1,066 | 0 | 4/4 (51%–100%) |
| GPT-6 Luna (low), Codex CLI | 3.30 s (3.19 to 3.47) | 15.50 s (7.39 to 22.69) | 1,066 | 0 | 4/4 (51%–100%) |
Claude Code ran at each model's default effort. The Codex CLI ran at low effort. The first-text ranges of Sonnet (0.88 to 4.09 s) overlap every other range, so Sonnet's place in the order is not separated. Opus's range (1.70 to 2.35 s) sits below the ranges of Haiku, Sol and Luna.
Reasoning counts and waiting varied together in some calls. Haiku's four calls wrote 255, 300, 525 and 614 reasoning tokens. Their first text came at 2.8 s, 3.4 s, 4.6 s and 6.4 s. Sonnet's call with 398 reasoning tokens had its slowest first text (4.1 s). Its two calls with no reasoning tokens had its fastest (0.9 s and 1.3 s). Fable does not follow this: its call with no reasoning tokens took 4.5 s. With 4 calls per model, we call this an observation, not a rule.
The one wrong reply in part A came from Haiku: 288 lines instead of 250. It stays in every timing.
2. Output speed: tokens per second, and why they do not compare
Output speed here is a calculation, not a measurement. We divide the reply size by the time from the first text to the end of the call. That end includes the CLI's exit overhead. We did not measure its size, so this is not pure model throughput.
First, tokens. We subtract reported reasoning tokens. The receipts do not locate all reasoning in time. A missing count becomes zero in the calculation; it does not prove no reasoning:
In each vendor's own visible tokens, the medians are Sonnet 232 (230–233), Opus 156 (155–156), Haiku 153 (153–216), Luna 129 (55–259), Fable 123 (121–131) and Sol 80 (71.6–80.5) tokens per second. Parentheses show observed ranges; n = 4 each. But those tokens are not the same size. The 23 exact replies had the same 4,327-character text. Haiku's failed reply had extra output. Haiku counted the exact text as 1,210 tokens, Sonnet, Opus and Fable as 1,941, and Sol and Luna as 1,066. The most tokens divided by the fewest is 1.8x (a calculation). A model with small tokens needs more of them for the same text, so its tokens per second look high.
So we also give characters per second, which is the same unit for every model. This uses only replies that matched all 250 lines:
The medians are Haiku 547 (3 exact replies, range 546 to 548), Sonnet 517 (513 to 519), Opus 347 (345 to 349), Sol 323 (291 to 327) and Fable 273 (270 to 293). The ranges of Haiku, Sonnet, Opus and Sol do not overlap one another. Fable's overlaps Sol's.
Luna had a large spread. Its total time was 7.4 s and 8.6 s in two calls and 22.4 s and 22.7 s in the other two. First text was 3.2 s to 3.5 s in all four, and every reply had 1,066 tokens. The observed gap sits after the first text: 203 and 259 tokens per second in the fast calls, about 55 in both slow ones (a calculation). We do not know why. Four calls cannot establish two speed groups. Its median of 524 characters per second describes none of the four calls. Sol was steady: total time 16.4 s to 17.7 s.
3. What does a longer prompt cost?
Part B asked one lookup question over a synthetic stock ledger at three sizes: about 1k, 16k and 64k tokens of text (18, 327 and 1,315 lines; approximate targets from probe calibration calculations and estimated Haiku prefix counts). The answer was one number. We used a new seed for every call, so each ledger differed. Some fixed text repeated. Cache counts do not prove which text was cached. We ran 3 calls per size for Haiku, Sonnet, Opus and Sol: 36 calls.
| Model | First text at 1k, 16k, 64k (median) | 64k minus 1k (calculation) | First-text ranges at 1k; 16k; 64k (s), n = 3 each |
|---|---|---|---|
| Haiku 4.5 | 1.93 s, 2.27 s, 2.78 s | +0.9 s | 1.85–2.04; 2.22–2.47; 2.45–2.89 (1k and 64k separate) |
| Sonnet 5.5 | 1.45 s, 1.78 s, 3.07 s | +1.6 s | 1.23–1.72; 1.64–2.11; 1.38–3.61 (1k and 64k overlap) |
| Opus 5.5 | 1.51 s, 1.74 s, 1.79 s | +0.3 s | 1.46–2.01; 1.70–2.97; 1.72–3.72 (1k and 64k overlap) |
| GPT-6.1 Sol (low) | 3.36 s, 4.02 s, 3.93 s | +0.6 s | 3.36–4.75; 3.30–4.28; 3.42–4.38 (1k and 64k overlap) |
Only Haiku's ranges at 1k and 64k do not overlap. Sonnet's 64k calls ran from 1.38 s to 3.61 s, so one 64k call was faster than its 1k median. Total time tells the same story:
The reported input tokens include each CLI's own prefix. The 1k cells reported a median 3,040 to 4,592 input tokens for Claude Code and 12,790 for the Codex CLI, because the prefix is part of the call. The 64k cells reported a median 67,605 to 78,674 for Claude Code and 64,172 for the Codex CLI. Every counted ledger differed. These input counts cannot compare tokenizers on the same text. The prefix size per counted call was not measured separately.
Per-token time calculations differ greatly. Take the largest difference of size medians, Sonnet's +1.6 s over 75,545 extra input tokens: about 0.02 s per 1,000 prompt tokens. Haiku's separated effect is about 0.014 s per 1,000. Opus is 0.004 s and Sol 0.011 s. Writing is slower: 1,000 visible output tokens took 4.3 s for Sonnet, 6.4 s for Opus, 6.5 s for Haiku, 8.2 s for Fable and 12.6 s for Sol (all calculations on medians, from part A). These per-token figures are rough: three of the four prompt effects are inside the noise.
4. Did the cache help?
Cache reads did not grow with ledger size. That fits no ledger reuse, but counts do not identify cached text:
- Sonnet and Opus read 1,463 tokens from the cache on every call, at every size. That fits reuse of fixed prefix text, but the counters do not identify it.
- Haiku read 0 tokens and wrote nearly all of its input to the cache each time.
- Sol read 8,960 tokens on nearly every call, at every size. That fits reuse of fixed prefix text, but the counters do not identify it. The first call read 0.
The study's cache table lists every cell. The size differences do not isolate prompt-processing time or a cache effect. A repeated prompt can behave differently (see prompt caching).
5. A side result: exact lookups
Haiku, Sonnet and Sol answered 9 of 9 lookups exactly (95% interval 70% to 100%). Opus answered 5 of 9 (27% to 81%). Four Opus replies (3 at 64k and 1 at 16k) echoed the matching ledger line and then gave the right number as the last line. The prompt said "reply with only the number", so the strict check marks them wrong, and we kept that scoring. All four models' intervals overlap, so these lookup rates do not separate them. The declared strict and lenient checks count the four Opus replies as wrong, not format misses. A post-hoc reading shows all 36 replies ended on the right number (95% interval 90% to 100%); this is not a pass rate. The table in the study keeps both. Those four calls took 2.3 s to 4.3 s in total.
Quality hit a ceiling in this study: 23 of 24 part A replies were exact (95% interval 80% to 99%). These runs measure speed. They cannot rank accuracy.
What to change (our reading of the data)
We did not test these changes before and after. Each rests on a number above.
- Test a shorter answer before a shorter prompt. The calculations above suggest this test. They do not prove the effect of changing your task.
- Test context changes on your task. The 64k-minus-1k median differences were +0.3 s to +1.6 s (calculation; ranges above). Context still costs money, and it can change quality.
- Check default thinking before a fast step. Haiku's first text rose with its reasoning tokens (2.8 s at 255 tokens, 6.4 s at 614). Test your step with thinking limited.
- Compare speed in characters or in seconds for a fixed task. Tokens per second depend on the vendor's tokenizer.
- Time your API route directly. It skips the CLI process. The CLI ready-time medians were 0.33 s to 0.84 s per cell (3 or 4 calls). This does not predict API latency.
- Plan for spread. Luna's 4 calls ranged from 7.4 s to 22.7 s on one prompt. Budget for the slow calls, not the median (see tail latency).
How we measured
- Protocol timing. The file was created at 01:09:29 UTC, before the first probe at 01:10:04 UTC. Earlier contents were not frozen. Later edits cannot prove advance declaration. Amendment 3 says 06:05 UTC and claims to precede Codex part B, which ran at 06:03:43–06:04:19 UTC. It is retrospective. The surviving controls ran at 01:10:57 UTC, after the probes and before counted calls. Earlier controls were overwritten. The surviving reference, wrapped-answer and planted-error checks pass.
- Part A: 24 calls. Claude Haiku 4.5, Sonnet 5.5, Opus 5.5 and Fable 5.1 at default effort (16 calls), GPT-6.1 Sol and GPT-6 Luna at low effort (8 calls). 4 calls per model.
- Part B: 36 calls. Haiku, Sonnet and Opus (27 calls) and Sol (9 calls), 3 calls per size, a new ledger seed per call.
- Isolation: a fresh folder, no tools, no MCP servers, no session, one turn, one call at a time on one Mac. Two probe calls ran outside every cell and are not in any chart. Nothing was retried. Every call completed.
- Calculations: output speed, per-1,000-token times and the 64k-minus-1k differences are our arithmetic on reported counts and medians. A range (fastest to slowest) is not a confidence interval. The lookup rates carry 95% Wilson intervals.
Caveats
- Few calls. 4 calls per model in part A and 3 per size in part B. A median of so few calls moves with one slow call.
- A CLI is not an API. First text includes CLI start-up and each CLI's own system prompt. A model and its CLI are one row.
- Effort is not one scale. Claude Code ran at default effort and the Codex CLI at low effort.
- One prompt for speed. Other text, other days and other hours can differ. Output speed subtracts reasoning tokens as the CLIs report them.
- Two sessions on one night. The Claude batches and the Codex part A ran within 8 minutes. The Codex part B batch (Sol, prompt size) ran about 4 hours later, after a run script stopped. Sol's first-text median was 3.52 s in part A (2.75–4.42 s; n = 4). Part B medians were 3.36 s to 4.02 s (cell ranges above; n = 3 each). The observed ranges overlap. Different tasks and so few calls cannot isolate an hour effect.
- Synthetic ledger. Each call used different text. Size differences also include content and provider-load changes; they do not isolate a cause.
- Shared host. No full host-load record proves that all other work stopped.
- Haiku failure. Its failed reply included tool-shaped text. No structured tool event was recorded; the failure stays in every timing.
- We build Agent. I am Shahrukh Siddiqui, and I build Agent. Agent was not one of the models timed here. All models ran through public CLIs.
What to read next
- Claude Code vs Codex CLI vs the API: latency, time to first token and hidden prompts
- Why is Claude Code slow? Where the seconds go
- Time to first token (TTFT), explained
- Claude Code vs Codex CLI: every measured row
- Tokens per call and the CLI context tax
Time your own calls
Agent records the model, the tokens and the time of every step, so you can see which calls earn their seconds. Try Agent and measure your own work.