A latency budget for voice agents: which LLM steps fit in one turn?
136.5 ms for Jev, 0.82 s for a small-model API, 2.79 s to 3.79 s for Codex CLI: which steps fit a voice agent latency budget? A thought experiment.
TL;DR
- A thought experiment. We set measured step times against three turn budgets: 0.5 s, 1 s and 2 s. They are our assumptions, not data.
- Decision step: rules took a median 1.42 µs. Jev 1.13 took 136.5 ms (p95 195.7 ms, n = 246) in a separate keyed run on 2026-10-06. Both medians and p95 values fit every assumed budget. Sonnet 5.5 through Claude Code took 2.60 s and fits none.
- Model step: GPT-6 Luna through the OpenAI API gave first useful output in a median 0.82 s (0.51 s to 1.37 s, n = 5). The median fits 1 s. The slowest run does not.
- Coding CLIs: no median fit 1 s. Codex CLI: 2.79 s to 3.79 s.
- Sample pipeline (a calculation): Jev plus one Luna call: 0.96 s on medians, 1.57 s on p95 plus slowest run. It fits 1 s on medians only. Sonnet 5.5 plus Luna totals 3.42 s on medians.
Claude Code CLI vs Codex CLI vs the API: a latency race
For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.
Transcript
- Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.
- A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
- It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
- A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
- Open benchmarks: intervals, sources and every failure kept.
What is the experiment?
A voice turn runs from the end of the user's speech to the start of the spoken reply. We chose the three budgets. They are assumptions, and we cite no outside study. They cover the LLM steps only.
We judge each step alone:
- yes: the median and the listed p95 or slowest run are under the budget.
- median: only the median is.
- no: the median is over.
These labels are calculations, not service guarantees. A p95 is not the worst case; some calls take longer.
A real turn shares one budget between all steps, so this is the easy test.
Which measured steps fit?
| Step (route, n) | Median | Observed range | p95 or slowest | 0.5 s | 1 s | 2 s |
|---|---|---|---|---|---|---|
| Rules (in process, 20,000) | 1.42 µs | Maximum 2,538.21 µs; minimum not reported | p95 2.33 µs | yes | yes | yes |
| Jev 1.13 (HTTPS, 246) | 136.5 ms | 100.9 ms to 297.3 ms | p95 195.7 ms | yes | yes | yes |
| GPT-6 Luna, first useful (API, 5) | 0.82 s | 0.51 s to 1.37 s | Slowest 1.37 s | no | median | yes |
| GPT-6.1 Sol low, first useful (API, 5) | 0.87 s | 0.84 s to 1.74 s | Slowest 1.74 s | no | median | yes |
| Fable 5.1, first useful (Claude Code, 15) | 1.20 s | 0.95 s to 7.90 s | Slowest 7.90 s | no | no | median |
| Haiku 4.5, first model output (Claude Code, 5) | 1,461 ms | 1,206 ms to 2,308 ms | Slowest 2,308 ms | no | no | median |
| Sonnet 5.5 as router (Claude Code, 82) | 2,597 ms | 1,993 ms to 5,583 ms | p95 4,298 ms | no | no | no |
| Codex CLI, first useful (3 setups, 5 each) | 2.79 s to 3.79 s | See the three ranges below | Slowest across setups 4.30 s | no | no | no |
| Haiku 4.5 as router (Claude Code, 82) | 12,543 ms | 5,857 ms to 51,278 ms | p95 34,481 ms | no | no | no |
The listed tail figure is p95 where n is 82 or more, and the slowest run elsewhere. Neither a range nor p95 is a confidence interval. The Jev row comes from a separate keyed run on 2026-10-06 over the network from one Mac. The Fable row comes from our five-task study.
Decision step: rules, a decision model or an LLM router?
The 1.42 µs is the pure decision. With its database write, the recorded rule arm took a median 1 ms (p95 2 ms, maximum 3 ms, n = 419). Jev 1.13, a small routing model, took a median 136.5 ms (p95 195.7 ms) over 246 calls. The cold first call took 224.7 ms.
Jev answered 221 of 246 calls exactly (89.8%). These repeat the same 82 decisions, so they are not 246 independent samples. The dataset's 95% Wilson interval is 81.9% to 95.0%, using the rounded count of 74/82 on the case scale. Both charts now include the live Jev timings.
Sonnet 5.5 through Claude Code took a median 2,597 ms (p95 4,298 ms, n = 82 recorded calls). Haiku 4.5 took 12,543 ms. This chart reads 2.60 s and 12.67 s (the routing study's own summary). The table uses per-call medians from the overhead study; the summary medians differ slightly.
The CLI also split out its own time. For Sonnet, model time was a median 1.596 s (p95 2.583 s; range 1.060 s to 4.757 s, n = 82). CLI time was a median 973 ms (p95 1,277 ms; range 826 ms to 3,673 ms). Model time alone was over 1 s. We did not time a direct Claude API call (no key).
Model step: API or CLI?
Through the OpenAI API, GPT-6 Luna gave first useful output in a median 0.82 s (0.51 s to 1.37 s). It finished in 0.97 s (0.65 s to 1.50 s). GPT-6.1 Sol at low effort took 0.87 s (0.84 s to 1.74 s). At high effort it took 1.34 s (1.26 s to 2.12 s), which misses 1 s. Each setup has n = 5; these are ranges, not confidence intervals.
A decision must finish before the next step starts, so use finish time for a decision. For this calculation, first useful text ends the reply step. Spoken playback still needs speech-ready text and speech synthesis.
Through the Codex CLI, the same models took a median 2.79 s (Luna; range 2.46 s to 3.42 s), 3.75 s (Sol, low; 3.44 s to 4.10 s) and 3.79 s (Sol, high; 3.37 s to 4.30 s). Each setup has n = 5. For each model and effort, the API range and the CLI range do not overlap, so the API is ahead on first useful output for this task in this small sample.
Does a coding CLI fit a voice turn?
Claude Code with Haiku 4.5 got a one-word prompt with tools off (n = 5). Its first output event came at a median 563 ms (519 ms to 726 ms), but that is not when the model starts to answer. First model output came at a median 1,461 ms (1,206 ms to 2,308 ms).
In our five-task study, Fable 5.1 in Claude Code had the lowest observed median first useful output of 9 setups. Its median was 1.20 s (range 0.95 s to 7.90 s, n = 15). Ranges overlap across setups, so this does not establish a speed ranking.
Sample pipeline: a decision plus one model call
| Pipeline | Sum of medians | p95 plus slowest | 0.5 s | 1 s | 2 s |
|---|---|---|---|---|---|
| Rules + Luna | 0.82 s | 1.37 s | no | median | yes |
| Jev 1.13 + Luna | 0.96 s | 1.57 s | no | median | yes |
| Sonnet 5.5 router + Luna | 3.42 s | 5.67 s | no | no | no |
Each sum is a calculation: decision time plus first useful output. The median of a sum is not exactly the sum of the medians. The sums use unrounded receipts where available. A p95 plus the slowest of 5 runs is a tail scenario, not a measured pipeline p95 or an upper bound. No run timed a whole pipeline. At 1 s, Jev plus Luna leaves about 41 ms (calculation) for the rest of the turn.
What we read from this
This is our reading of small samples, not a rule.
- Among these measured decision steps, rules and Jev fit the assumed budgets on median and p95. The LLM routers through a CLI miss even 2 s on their medians (Sonnet: 2,597 ms).
- A small model through its API can fit 1 s on the median. Luna's slowest run (1.37 s) did not. At 0.5 s no measured reply median fit: Luna's fastest run took 0.51 s.
- None of the measured coding CLI setups fits on its median at 1 s or under. At 2 s it fits some medians, never the slowest call.
Our advice, not a measured result: decide with rules first. Add a small decision model only where rules cannot decide. Call the model API directly and stream the reply. Keep coding CLIs out of the turn loop.
How we measured
- No new model calls. Every figure comes from our benchmark dataset or its source receipts and run summaries. The sums and fit labels are our calculations.
- Rules: one Apple M3 Ultra Mac, 20,000 decisions after 5,000 warm-up calls. Jev 1.13: 246 calls (3 reps of 82 typed decisions), one at a time, over HTTPS from a home network.
- API and Codex rows: one fixed short reply, 5 runs per setup, on 2026-10-03. All 30 calls passed (30/30; 95% Wilson interval 88.6% to 100%). This short task hits a quality ceiling; it cannot rank the routes on quality.
Public receipts: Jev calls, routing overhead, API and Codex calls, and five-task calls.
Caveats
- One host, one network. Network and vendor latency change over the day.
- Small samples. The API and matched Codex rows have n = 5 per setup; Fable has n = 15 across five tasks. Five runs say little about the tail.
- Different tasks. A fixed reply, a one-word prompt and five short tasks are not real voice conversations. First useful text is not necessarily speech-ready text.
- Different routes and settings. Jev uses HTTPS; Claude routers use Claude Code. Sonnet runs at low effort; Haiku uses default thinking. These timings cannot isolate model compute time.
- Different days and routes. The Jev run and the API runs took place on different days. We timed no direct Claude API call (no key).
- Assumed budgets, text only. We did not time speech-to-text, speech synthesis or audio transport.
- Home advantage. We revised the 82 decision cases against Jev's answers.
What to read next
- The routing hub
- Claude Code vs Codex CLI vs the API: latency
- What is an LLM router?
- What does a router cost you?
- Studies: routing overhead, Jev vs Claude, five tasks.
Time your own routes
Disclosure: I build Agent, the product behind these benchmarks.
Agent keeps a receipt for each task: model, route, tokens, time and cost. Try Agent and time your own routes.