Explainer · Time to first token
Time to first token (TTFT), explained with CLI and API timings
Definition
Time to first token (TTFT) is the wait from sending a model request to receiving its first reply token. Start-up, processing, queues and the network can add delay. We measure first useful output, a different clock from raw TTFT.
Agent team · · 5 min read · Every number is from the public studies
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first text | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 4 s | 2.8 s–6.4 s | 4 |
| Claude Sonnet 5.5 · Claude Code | 2 s | 0.9 s–4.1 s | 4 |
| Claude Opus 5.5 · Claude Code | 2 s | 1.7 s–2.4 s | 4 |
| Claude Fable 5.1 · Claude Code | 4.4 s | 2.3 s–4.6 s | 4 |
| GPT-6.1 Sol (low) · Codex CLI | 3.5 s | 2.8 s–4.4 s | 4 |
| GPT-6 Luna (low) · Codex CLI | 3.3 s | 3.2 s–3.5 s | 4 |
6 rows. Slowest Claude Fable 5.1 · Claude Code 4.4 s (range 2.3 s–4.6 s, n 4). Fastest Claude Sonnet 5.5 · Claude Code 2 s (range 0.9 s–4.1 s, n 4). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = fastest and slowest call
Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.
Source: LLM speed anatomy
Three clocks, not one
- Time to first token: request to first model token.
- First useful output: first streamed answer text, including CLI start-up and its prompt. Our head-to-head and CLI studies use this clock.
- Total time: launch to exit, including start-up.
The routing overhead study marks first output event and first model output. Claude Code records system:init, then assistant; Codex records thread.started, then item.completed:agent_message. The first-event ranges end before the model-output ranges start in both CLIs. Neither event measures the raw first token on the wire.
What adds to the wait
- CLI start-up and context. Coding CLIs start a process and add prompt/tool context. Codex sent median 19,551 input tokens versus the API's 17 for a one-line request (GPT-6.1 Sol, low, n = 5). Claude Code's one-word overhead was median 1,690 ms (1,533 to 1,811 ms, n = 5): a calculation, wall time minus CLI-reported API time. Some falls after the answer. See CLI context tax.
- Thinking. In the hard-task study, Haiku 4.5 at CLI default reported median 4,556 reasoning tokens. First useful output: 35.54 s (12.88 to 70.31 s), versus Sonnet 5.5's 5.95 s (0.86 to 30.57 s); n = 24 each. Ranges overlap: no ranking or tested cause.
- Network. Included in every timing; we did not measure it alone.
Our numbers
One host, one network: these sessions give no general guarantee. Timings are medians with run ranges, not 95% intervals. Small samples give directional evidence.
CLI start-up, one-word answer, 5 runs each. Claude Code used Haiku 4.5 with tools, MCP servers and sessions off; Codex used its default model and read-only sandbox.
- First event: Claude Code 563 ms (519 to 726 ms); Codex 489 ms (354 to 1,304 ms). Ranges overlap: neither ahead.
- First model output: 1,461 ms (1,206 to 2,308 ms) versus 5,059 ms (4,391 to 5,478 ms). Ranges do not overlap: Claude Code ahead here. The models differ, so the study cannot separate CLI and model effects.
Same model, API versus CLI, one-line answer. Each cell has n = 5. First useful output is in seconds: median (fastest to slowest).
| Model and effort | OpenAI API | Codex CLI |
|---|---|---|
| GPT-6 Luna, none | 0.82 (0.51 to 1.37) | 2.79 (2.46 to 3.42) |
| GPT-6.1 Sol, low | 0.87 (0.84 to 1.74) | 3.75 (3.44 to 4.10) |
| GPT-6.1 Sol, high | 1.34 (1.26 to 2.12) | 3.79 (3.37 to 4.30) |
Ranges do not overlap in each pair: the API is ahead on this task (comparison).
Across models, five short tasks, 10 to 15 calls each.
Claude Code medians (n = 15 each): 1.2 s for Fable 5.1 (0.95 to 7.9 s) to 3.63 s for Haiku 4.5 (2.78 to 22.27 s). GPT-6.1 Sol/Codex medians: medium 5.05 s (3.36 to 17.82 s, n = 15); low 5.14 s (4.02 to 8.50 s, n = 10); high 5.32 s (3.64 to 16.37 s, n = 15). Neighbouring ranges overlap: no ranking.
How to measure it
- Stream: an all-at-once reply gives only total time.
- Stamp request, first event, model content, answer text and end.
- Repeat one prompt, one call at a time, on one host. Five calls are descriptive, not a general minimum. Pair calls back to back; vendor latency changes during the day.
- Report median and range; a range is not an interval.
- Change one thing at a time: model, effort or route.
How to cut it
- Direct API for short tasks. Ahead in every one-line pair (n = 5 per cell).
- Less context. API: 0.87 s (0.84 to 1.74 s), 17 input tokens, n = 5. Wire-matched batch: 1.58 s (1.15 to 1.96 s), median 8,297 tokens, n = 3. Both GPT-6.1 Sol, low. Overlapping ranges and different batches do not isolate input size.
- Lower effort where quality holds. All 11 effort-ladder configurations passed 16 of 16 hard-task calls (95% interval 81% to 100%). The set hits a ceiling; test your tasks.
- Cache: no clear wait reduction in our sessions.
Frequently asked questions
What is a good time to first token?
We did not test what delay people accept, so we give no target. One-line medians span 0.82 to 1.34 s (API), 2.79 to 3.79 s (Codex), n = 5 per cell. These are spans of medians; the table above gives each run range.
Why is the Codex CLI slower to first output than the API?
The CLI adds prompt/tool context: 19,551 versus 17 input tokens (n = 5). Total-time calculation: 3.5 times the API median. Pooled across the three model/effort cells: CLI 3.9 s (2.88 to 4.69 s), API 1.1 s (0.65 to 2.23 s), n = 15 each.
Input size may not explain everything. Wire-matched first useful output: API 1.58 s (1.15 to 1.96 s), CLI 2.53 s (1.44 to 2.57 s). Both sent median 8,297 tokens, GPT-6.1 Sol, low, n = 3 each. Overlapping ranges establish no speed difference.
Does prompt caching reduce time to first token?
No clear effect in our Claude Code sessions. We timed whole turns, not first tokens; no cache-off control. Cache reads: mean 97% of input on turns 2 to 5 (calculation: 24 turns, 6 sessions), range 88% to 99%. Median turn time, turn 1 versus turns 2 to 5 (n = 3 versus 12 per model): Sonnet 1.64 s (1.58 to 1.79 s) versus 1.61 s (1.35 to 5.63 s); Opus 1.90 s (1.78 to 4.36 s) versus 2.40 s (1.63 to 12.67 s). The ranges overlap; the turns ask different questions. See cache-session timings and prompt caching for cost.
Does reasoning effort change TTFT?
No clear one-line change for GPT-6.1 Sol, low versus high. API: 0.87 s (0.84 to 1.74 s) versus 1.34 s (1.26 to 2.12 s); Codex: 3.75 s (3.44 to 4.10 s) versus 3.79 s (3.37 to 4.30 s), n = 5 each. Ranges overlap. Small coding task: API 1.05 s (0.97 to 1.40 s) versus 5.31 s (4.99 to 6.42 s); Codex 13.60 s (12.52 to 13.83 s) versus 17.27 s (17.13 to 21.86 s). Ranges do not overlap, but n = 3 each is too few to separate. See reasoning effort.
Watch the data
Claude Code CLI vs Codex CLI vs the API: a latency race
For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.
Transcript
- Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.
- A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
- It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
- A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
- Open benchmarks: intervals, sources and every failure kept.