Explainer · Time to first token

Time to first token (TTFT), explained with CLI and API timings

Definition

Time to first token (TTFT) is the wait from sending a model request to receiving its first reply token. Start-up, processing, queues and the network can add delay. We measure first useful output, a different clock from raw TTFT.

Agent team · · 5 min read · Every number is from the public studies

Entrance: medians race at 3.2× real timeMotion reduced: press Replay to animateThe slowest median is 4.4 s. The clock runs at the recorded speed.
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (low) · Codex CLI
GPT-6 Luna (low) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

6 rows. Slowest Claude Fable 5.1 · Claude Code 4.4 s (range 2.3 s–4.6 s, n 4). Fastest Claude Sonnet 5.5 · Claude Code 2 s (range 0.9 s–4.1 s, n 4). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 4 per row

Median of 4 calls per model; whiskers = fastest and slowest call

Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.

Source: LLM speed anatomy

Three clocks, not one

  • Time to first token: request to first model token.
  • First useful output: first streamed answer text, including CLI start-up and its prompt. Our head-to-head and CLI studies use this clock.
  • Total time: launch to exit, including start-up.

The routing overhead study marks first output event and first model output. Claude Code records system:init, then assistant; Codex records thread.started, then item.completed:agent_message. The first-event ranges end before the model-output ranges start in both CLIs. Neither event measures the raw first token on the wire.

What adds to the wait

  • CLI start-up and context. Coding CLIs start a process and add prompt/tool context. Codex sent median 19,551 input tokens versus the API's 17 for a one-line request (GPT-6.1 Sol, low, n = 5). Claude Code's one-word overhead was median 1,690 ms (1,533 to 1,811 ms, n = 5): a calculation, wall time minus CLI-reported API time. Some falls after the answer. See CLI context tax.
  • Thinking. In the hard-task study, Haiku 4.5 at CLI default reported median 4,556 reasoning tokens. First useful output: 35.54 s (12.88 to 70.31 s), versus Sonnet 5.5's 5.95 s (0.86 to 30.57 s); n = 24 each. Ranges overlap: no ranking or tested cause.
  • Network. Included in every timing; we did not measure it alone.

Our numbers

One host, one network: these sessions give no general guarantee. Timings are medians with run ranges, not 95% intervals. Small samples give directional evidence.

CLI start-up, one-word answer, 5 runs each. Claude Code used Haiku 4.5 with tools, MCP servers and sessions off; Codex used its default model and read-only sandbox.

  • First output event
  • First model output
  • Total wall time
Entrance: medians race at 4.3× real timeMotion reduced: press Replay to animateThe slowest median is 6 s. The clock runs at the recorded speed.
Claude Code · Claude Haiku 4.5
Codex CLI (default model)

2 rows, 3 series: First output event, First model output, Total wall time. First output event: slowest Claude Code · Claude Haiku 4.5 563 ms (range 519 ms–726 ms, n 5). Fastest Codex CLI (default model) 489 ms (range 354 ms–1.3 s, n 5). All run ranges overlap. First model output: slowest Codex CLI (default model) 5.06 s (range 4.39 s–5.48 s, n 5). Fastest Claude Code · Claude Haiku 4.5 1.46 s (range 1.21 s–2.31 s, n 5). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 5 per row

Median of 5 runs; whiskers = fastest and slowest run

Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.

Source: Routing overhead runs: policy microbenchmark and CLI start-up

  • First event: Claude Code 563 ms (519 to 726 ms); Codex 489 ms (354 to 1,304 ms). Ranges overlap: neither ahead.
  • First model output: 1,461 ms (1,206 to 2,308 ms) versus 5,059 ms (4,391 to 5,478 ms). Ranges do not overlap: Claude Code ahead here. The models differ, so the study cannot separate CLI and model effects.

Same model, API versus CLI, one-line answer. Each cell has n = 5. First useful output is in seconds: median (fastest to slowest).

  • Total time
  • First useful output
Entrance: medians race at 3× real timeMotion reduced: press Replay to animateThe slowest median is 4.2 s. The clock runs at the recorded speed.
OpenAI API · GPT-6 Luna · none
OpenAI API · GPT-6.1 Sol · low
OpenAI API · GPT-6.1 Sol · high
Codex CLI · GPT-6 Luna · none
Codex CLI · GPT-6.1 Sol · low
Codex CLI · GPT-6.1 Sol · high

6 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · high 4.2 s (range 3.8 s–4.7 s, n 5). Fastest OpenAI API · GPT-6 Luna · none 1 s (range 0.7 s–1.5 s, n 5). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · high 3.8 s (range 3.4 s–4.3 s, n 5). Fastest OpenAI API · GPT-6 Luna · none 0.8 s (range 0.5 s–1.4 s, n 5). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 5 per row

Matched cohort, fixed exact reply, 5 runs per configuration

Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.

Source: Provider explorer receipts: CLI vs API

Model and effortOpenAI APICodex CLI
GPT-6 Luna, none0.82 (0.51 to 1.37)2.79 (2.46 to 3.42)
GPT-6.1 Sol, low0.87 (0.84 to 1.74)3.75 (3.44 to 4.10)
GPT-6.1 Sol, high1.34 (1.26 to 2.12)3.79 (3.37 to 4.30)

Ranges do not overlap in each pair: the API is ahead on this task (comparison).

Across models, five short tasks, 10 to 15 calls each.

Entrance: medians race at 3.8× real timeMotion reduced: press Replay to animateThe slowest median is 5.3 s. The clock runs at the recorded speed.
Claude Fable 5.1 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
GPT-6.1 Sol (low) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

9 rows. Slowest GPT-6.1 Sol (high) · Codex CLI 5.3 s (range 3.6 s–16.4 s, n 15). Fastest Claude Fable 5.1 · Claude Code 1.2 s (range 1 s–7.9 s, n 15). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 10–15 per row

Median per configuration; whiskers = fastest and slowest call

One host, one network, one day. Whiskers are a range, not a confidence interval.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Claude Code medians (n = 15 each): 1.2 s for Fable 5.1 (0.95 to 7.9 s) to 3.63 s for Haiku 4.5 (2.78 to 22.27 s). GPT-6.1 Sol/Codex medians: medium 5.05 s (3.36 to 17.82 s, n = 15); low 5.14 s (4.02 to 8.50 s, n = 10); high 5.32 s (3.64 to 16.37 s, n = 15). Neighbouring ranges overlap: no ranking.

How to measure it

  1. Stream: an all-at-once reply gives only total time.
  2. Stamp request, first event, model content, answer text and end.
  3. Repeat one prompt, one call at a time, on one host. Five calls are descriptive, not a general minimum. Pair calls back to back; vendor latency changes during the day.
  4. Report median and range; a range is not an interval.
  5. Change one thing at a time: model, effort or route.

How to cut it

  • Direct API for short tasks. Ahead in every one-line pair (n = 5 per cell).
  • Less context. API: 0.87 s (0.84 to 1.74 s), 17 input tokens, n = 5. Wire-matched batch: 1.58 s (1.15 to 1.96 s), median 8,297 tokens, n = 3. Both GPT-6.1 Sol, low. Overlapping ranges and different batches do not isolate input size.
  • Lower effort where quality holds. All 11 effort-ladder configurations passed 16 of 16 hard-task calls (95% interval 81% to 100%). The set hits a ceiling; test your tasks.
  • Cache: no clear wait reduction in our sessions.

Frequently asked questions

What is a good time to first token?

We did not test what delay people accept, so we give no target. One-line medians span 0.82 to 1.34 s (API), 2.79 to 3.79 s (Codex), n = 5 per cell. These are spans of medians; the table above gives each run range.

Why is the Codex CLI slower to first output than the API?

The CLI adds prompt/tool context: 19,551 versus 17 input tokens (n = 5). Total-time calculation: 3.5 times the API median. Pooled across the three model/effort cells: CLI 3.9 s (2.88 to 4.69 s), API 1.1 s (0.65 to 2.23 s), n = 15 each.

Input size may not explain everything. Wire-matched first useful output: API 1.58 s (1.15 to 1.96 s), CLI 2.53 s (1.44 to 2.57 s). Both sent median 8,297 tokens, GPT-6.1 Sol, low, n = 3 each. Overlapping ranges establish no speed difference.

Does prompt caching reduce time to first token?

No clear effect in our Claude Code sessions. We timed whole turns, not first tokens; no cache-off control. Cache reads: mean 97% of input on turns 2 to 5 (calculation: 24 turns, 6 sessions), range 88% to 99%. Median turn time, turn 1 versus turns 2 to 5 (n = 3 versus 12 per model): Sonnet 1.64 s (1.58 to 1.79 s) versus 1.61 s (1.35 to 5.63 s); Opus 1.90 s (1.78 to 4.36 s) versus 2.40 s (1.63 to 12.67 s). The ranges overlap; the turns ask different questions. See cache-session timings and prompt caching for cost.

Does reasoning effort change TTFT?

No clear one-line change for GPT-6.1 Sol, low versus high. API: 0.87 s (0.84 to 1.74 s) versus 1.34 s (1.26 to 2.12 s); Codex: 3.75 s (3.44 to 4.10 s) versus 3.79 s (3.37 to 4.30 s), n = 5 each. Ranges overlap. Small coding task: API 1.05 s (0.97 to 1.40 s) versus 5.31 s (4.99 to 6.42 s); Codex 13.60 s (12.52 to 13.83 s) versus 17.27 s (17.13 to 21.86 s). Ranges do not overlap, but n = 3 each is too few to separate. See reasoning effort.

Watch the data

Live story · 33 sClaude Code CLI vs Codex CLI vs the API: a latency race

Claude Code CLI vs Codex CLI vs the API: a latency race

For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.

Transcript
  1. Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.
  2. A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
  3. It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
  4. A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
  5. Open benchmarks: intervals, sources and every failure kept.

The data behind this explainer

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

More explainers

  • CLI context tax

    Tokens per call and the CLI context tax

    A coding CLI wraps every request in its own system prompt and tools. How many input tokens that adds per call, what it costs in time, and when it matters.

  • Reasoning effort

    What is reasoning effort?

    Reasoning effort sets how much a model thinks before it answers. What low, medium and high change in quality, time, tokens and cost, measured on hard tasks.

  • Agent harness

    What is an agent harness? The code around the model, measured

    An agent harness is the code around a model: the loop, tools, prompts, memory and checks. What it changes in time and tokens, measured on real runs.

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.