A voice agent latency budget, with measured times: what fits in one turn?
Rules and Jev 1.13 fit every budget we assumed; a Claude router through a CLI fits none. 14 measured steps vs 300, 800 and 1,500 ms. A thought experiment.
Routing overhead: a 1.42 µs policy vs LLM routers
A deterministic routing policy decides in 1.42 µs (median, n = 20,000); Sonnet 5.5 as a router takes 2.60 s through the CLI. Per-task costs and delays are calculations.
Transcript
- Routing overhead · policy vs LLM routers. What a routing decision costs before the work starts. Time and money per decision, then per task. Jev is timed over its API, a different route from the CLI routers.
- The latency ladder: the in-process policy decides in 1.42 µs. Jev, a direct API call, takes 137 ms. Sonnet 5.5 as a router takes 2.60 s through the CLI, Haiku 4.5 12.5 s. Chart: Time per decision · log scale · median to p95 (n = 82–20000 each). Caveat: Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
- The LLM router takes about 1.8 million times longer per decision. Jev costs $0.0337 per 1,000 decisions, provider-reported. Policy decision, median (20,000 timed): 1.42 µs (n = 20000, p95 2.33 µs). Sonnet 5.5 router ÷ policy, medians: 1.8 million× (n = 82). Jev 1.13 per 1,000 decisions (provider-reported): $0.0337 (n = 82, 137 ms median per call over its API). Caveat: OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
- Of Sonnet 5.5’s 2.60 s per decision, a median 973 ms is CLI and harness time, not the model. Chart: Where an LLM router’s time goes · median per call (n = 82 each). Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
- Route all 49.5 model calls per task: $1.67 per 1,000 tasks with Jev, $247.30 with Sonnet 5.5. Route 7 decisions: $34.97. Chart: Added routing cost per 1,000 tasks. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
- If each decision waits in line, Sonnet 5.5 adds up to 129 s per task, or 18.2 s for 7 decisions. The policy adds 70.3 µs. Chart: Added routing delay per task · upper bound. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
- Start-up tax for a one-word answer: Claude Code 2.53 s and 6,761 input tokens; Codex CLI 6.00 s and 17,051 input tokens, 13,184 of them read from the cache. Chart: CLI start-up tax · one-word answer · median of 5 runs (n = 5 each). Caveat: CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
- Route in process where a rule is enough. Every timing, cost and gap online.
TL;DR
- A thought experiment on measured times. We set 14 measured steps against three budgets per step: 300 ms, 800 ms and 1,500 ms. The budgets are our assumptions, not data. We made no new model call.
- Decision step. Rules decided in a median 1.42 µs (p95 2.33 µs, n = 20,000; maximum 2.5 ms, minimum not retained). Jev 1.13, a small routing model, took a median 136.5 ms over HTTPS (p95 195.7 ms, n = 246). Both fit all three budgets, at the median and at p95. One Jev decision uses 46% of a 300 ms budget at the median and 65% at p95 (a calculation).
- A Claude router through the CLI fits none. Sonnet 5.5 through Claude Code took a median 2,598 ms (p95 4,298 ms, n = 82). That is 173% of a 1,500 ms budget at the median (a calculation).
- Model step. GPT-6 Luna through the OpenAI API gave its first useful output in a median 823 ms (slowest of 5 runs: 1,371 ms). The median misses an 800 ms budget by 23 ms (a calculation). It fits 1,500 ms.
- The count (a calculation). At the slow end, 3 of 14 steps fit 300 ms, 3 fit 800 ms and 4 fit 1,500 ms. At the median, 3, 3 and 8 fit.
- A component scenario (a calculation). Jev plus one Luna call is 959.5 ms on medians. That leaves 540.5 ms of 1,500 ms for speech-to-text, speech synthesis, turn detection and the user-to-agent network. We did not time those. API times already include the client-to-provider network. On the slow ends the sum is 1,566.7 ms, which is over the budget.
The full study, with every table: voice agent latency budget study.
What is a latency budget for a voice agent?
A voice turn runs from the moment the user stops speaking to the moment the agent starts to speak. In that gap the agent runs speech-to-text, one or more decisions, a model call, speech synthesis and the network. A budget says how long each part may take.
This calculation compares two kinds of step. The policy row times only policy logic and excludes database work. A decision is a whole call, from request to answer. A first useful output is the time until the harness records the first model text. It may not be enough text to start speech. We do not time speech or audio.
We chose three budgets per step: 300 ms, 800 ms and 1,500 ms. They are assumptions: budgets a voice team might set. Nothing in our data says what listeners accept.
Which steps fit?
Every time below comes from a run that another study published. yes means the median and the slow end are at or under the budget. median means only the median is. no means the median is over. The slow end is the p95 for steps with 30 or more runs. For steps with 5 or 15 runs it is the slowest observed run. The cutoff is a display rule, not proof of a reliable tail estimate. Observed ranges below are not confidence intervals.
| Step (route) | Runs | Median | Slow end | Observed range | 300 ms | 800 ms | 1,500 ms |
|---|---|---|---|---|---|---|---|
| Rules: Agent's routing policy (in process) | 20,000 | 1.42 µs | p95 2.33 µs | Minimum not retained to 2.5 ms (n = 20,000) | yes | yes | yes |
| Rule-based System One decision, record write excluded (in process) | 419 | 1 ms | p95 2 ms | Minimum not retained to 3 ms (n = 419) | yes | yes | yes |
| Jev 1.13 decision (HTTPS) | 246 | 136.5 ms | p95 195.7 ms | 100.9 ms to 297.3 ms (n = 246) | yes | yes | yes |
| GPT-6 Luna, effort none, first output (OpenAI API) | 5 | 823 ms | slowest 1,371 ms | 506 ms to 1,371 ms (n = 5) | no | no | yes |
| GPT-6.1 Sol, low, first output (OpenAI API) | 5 | 873 ms | slowest 1,742 ms | 835 ms to 1,742 ms (n = 5) | no | no | median |
| Claude Fable 5.1, first output on five short tasks (Claude Code) | 15 | 1,196 ms | slowest 7,902 ms | 947 ms to 7,902 ms (n = 15) | no | no | median |
| GPT-6.1 Sol, high, first output (OpenAI API) | 5 | 1,341 ms | slowest 2,115 ms | 1,259 ms to 2,115 ms (n = 5) | no | no | median |
| Claude Haiku 4.5, CLI start-up, first model output (Claude Code) | 5 | 1,461 ms | slowest 2,308 ms | 1,206 ms to 2,308 ms (n = 5) | no | no | median |
| Sonnet 5.5 as a router, effort low (Claude Code) | 82 | 2,598 ms | p95 4,298 ms | 1,993 ms to 5,583 ms (n = 82) | no | no | no |
| GPT-6 Luna, effort none, first output (Codex CLI) | 5 | 2,787 ms | slowest 3,418 ms | 2,461 ms to 3,418 ms (n = 5) | no | no | no |
| GPT-6.1 Sol, low, first output (Codex CLI) | 5 | 3,753 ms | slowest 4,103 ms | 3,435 ms to 4,103 ms (n = 5) | no | no | no |
| GPT-6.1 Sol, high, first output (Codex CLI) | 5 | 3,786 ms | slowest 4,296 ms | 3,366 ms to 4,296 ms (n = 5) | no | no | no |
| Codex CLI start-up, first model output | 5 | 5,059 ms | slowest 5,478 ms | 4,391 ms to 5,478 ms (n = 5) | no | no | no |
| Haiku 4.5 as a router, thinking on (Claude Code) | 82 | 12,674 ms | p95 34,413 ms | 5,857 ms to 51.28 s (n = 82) | no | no | no |
The steps differ in route and in what they time, so read the count as a summary of this table, not as a ranking.
Decision steps: rules, a small model or an LLM?
The chart above is from the routing overhead study. It puts decision times on a log axis, so a microsecond rule and a multi-second router share one scale. The next chart adds the budgets as dots.
- Rules. The routing policy took a median 1.42 µs (p95 2.33 µs, n = 20,000). A recorded rule-based decision, with its database record write excluded, took a median 1 ms (p95 2 ms, n = 419, millisecond resolution).
- Jev 1.13. A median 136.5 ms and a p95 of 195.7 ms over 246 calls: 3 repeats of 82 typed decisions, one call at a time, over HTTPS from one Mac. The first call used a fresh connection and took 224.7 ms (n = 1). The later-call median was 136.4 ms (n = 245). The difference is 88.3 ms (a calculation). It does not isolate connection cost or server warmth.
- How many fit in a row (a calculation). At p95, 4 Jev decisions fit back to back in 800 ms (5 at the median) and 7 fit in 1,500 ms (10 at the median). At 300 ms only 1 fits at p95 (2 at the median). These counts assume a fixed duration each time. They do not give a probability that a sequence fits.
Jev is a small routing model. It decides a typed choice, such as an intent or a failure class. It does not write the reply. This study times it and does not score it. Its accuracy is in the Jev vs LLM routers study, where the intervals of the three routers overlap.
Sonnet 5.5 as a router took a median 2,598 ms (p95 4,298 ms, n = 82). Haiku 4.5, with the CLI's default thinking, took a median 12,674 ms (p95 34,413 ms, n = 82). Both ran through the Claude Code CLI, which adds time: the overhead study reports CLI and harness p50 of 973 ms (nearest-rank quantile; p95 1,277 ms, range 826 to 3,673 ms, n = 82) (the router cost post). We did not time a direct Claude API call, so we claim nothing about one.
Model steps: how fast is the first useful output?
These charts hold the steps with a median of one second or more. The first chart shows the routers from median to p95. The second shows observed minimum-to-maximum ranges for the smaller samples. Neither chart shows a confidence interval.
For a one-line answer, GPT-6 Luna through the OpenAI API had a median first useful output of 823 ms and a slowest run of 1,371 ms (n = 5). GPT-6.1 Sol at low effort had 873 ms and 1,742 ms (n = 5). At high effort it had 1,341 ms and 2,115 ms (n = 5). API ranges overlap, so this is not a speed ranking. Only Luna fits 1,500 ms at both the median and the slowest run. None fits 800 ms at the slowest run, and none fits 300 ms at the median.
The same models through the Codex CLI had medians of 2,787 ms (Luna), 3,753 ms (Sol, low) and 3,786 ms (Sol, high). For each model, the API's slowest run is faster than the CLI's median. The CLI-minus-API difference in median first output was 1,964 to 2,880 ms for the same model (a calculation, 3 setups, 5 runs each). No Codex CLI step has a median under 1,500 ms. The CLI start-up probe, with a one-word prompt, took a median 5,059 ms to its first model output (n = 5). It includes a model call.
Two Claude Code cells had medians below 1,500 ms. Claude Fable 5.1 had a median first useful output of 1,196 ms over five short tasks (n = 15), but its slowest call took 7,902 ms. Claude Haiku 4.5 through Claude Code took a median 1,461 ms to the first model output of a one-word prompt (n = 5; slowest 2,308 ms).
A router in front of a model (a calculation)
A voice agent often decides first and then calls a model. Each step adds its time. We added the medians and the slow ends of runs that were not made together. Neither sum is a measured pair percentile or a guaranteed upper bound. Medians do not add to the median of a pair. Sums of p95 values do not guarantee a pair p95. Jev has n = 246; each API cell has n = 5; Sonnet has n = 82.
| Router in front | Model first output after it | Sum of medians | Sum of slow ends | Fits 1,500 ms |
|---|---|---|---|---|
| Rules (routing policy) | GPT-6 Luna, API | 823 ms | 1,371 ms | median and slow end |
| Jev 1.13 | GPT-6 Luna, API | 959.5 ms | 1,566.7 ms | median only |
| Jev 1.13 | GPT-6.1 Sol (low), API | 1,009.5 ms | 1,937.7 ms | median only |
| Sonnet 5.5 router, Claude Code | GPT-6 Luna, API | 3,421 ms | 5,669 ms | neither |
None of these pairs fits 300 ms or 800 ms, at the median or at the slow end. With Jev and Luna, 540.5 ms of 1,500 ms is left at the median. At the slow ends the pair is 66.7 ms over. This arithmetic remainder is not measured audio latency. We did not time speech recognition, synthesis, turn detection or the user-to-agent network. API timings include the client-to-provider network.
What we take from this
This is our reading of small samples, not a rule.
- In these rows, rules and Jev fit the assumed decision budgets. Both stay under 300 ms at the slow end in our data. No LLM step does.
- An LLM router through a coding CLI fits none. The Sonnet median alone is 173% of a 1,500 ms budget (calculation).
- API first output can fit 1,500 ms, but not 800 ms at its slowest run. The sample is 5 runs, so the tail is not known.
- The median hides the tail. 8 of 14 steps fit 1,500 ms at the median, but only 4 fit at the slow end. These samples do not establish a production tail.
What we would do
These are our recommendations. We did not test them as a package.
- Set the budget on the slow end. Read p95, not the median. Ask whether the step fits at its slowest runs.
- Decide with rules first. Add a small decision model only where rules cannot decide. Check its time on your own route.
- Keep a coding CLI out of the turn loop. Call the model API directly. The Codex CLI-minus-API median difference was 1,964 to 2,880 ms (calculation; n = 5 per setup). We did not time a direct Claude API call, so this is a Codex result.
- Test when speech can start. Streaming text may let speech start before the reply is complete. The first recorded output may be too short for speech. We timed text, not audio.
- Time the whole turn on your own route. Add speech-to-text, speech synthesis, turn detection and the user-to-agent network. The API times already include the provider network. Then count how many decisions you can afford in the budget.
- Try your own budget. The study page lists the share of each budget a step uses. Divide a step's time by your budget. A share of 100% or less fits.
How we measured
- Rules: the policy timing of the routing overhead run (p50 and p95 in µs), in process on one Mac. The System One decision time is the recorded decisions of the same run.
- Jev 1.13: the live run of 2026-10-06 (latency per call: n, median, p95, and the cold first call), over HTTPS from one Mac on a home network.
- Claude routers: the recorded routing runs of 2026-10-05 (wall time per call), through Claude Code, one call at a time. The budget table uses conventional medians and interpolated p95. The overhead chart uses nearest-rank quantiles, so its values differ slightly.
- CLI start-up: 5 runs per CLI, a one-word prompt (first model output).
- Model first output: the provider explorer receipts of 2026-10-03: the matched cohort on one fixed short reply, 5 runs per setup. Plus the Claude Code setup with the lowest observed first-output median in the five-task head-to-head (2026-10-05), 15 runs. We selected it after the run; overlapping ranges prevent a speed ranking.
- Calculations: share of a budget = time ÷ budget. A step fits when its time is at or under the budget. A sum adds two medians or two slow ends.
Caveats
- The budgets are assumptions. A product may need less or more.
- Different routes, days and samples. Routers are whole decisions. Models are the time to first text. Claude ran through a CLI. Runs are from 2026-10-03 to 2026-10-06.
- One Mac. Jev was timed over a home network in one run of 35 seconds. Vendor latency changes over a day, and a server near the API would see other times.
- Small samples. n = 5 for most model and start-up rows. A range from 5 runs says little about the tail.
- No speech or user-to-agent network round trip is timed. API times already include the client-to-provider network.
- Timing limits. Policy times exclude database reads, policy merge and record insertion. The rule-arm timing stops before record insertion. Individual policy and rule timings are not retained, so their summaries cannot be rebuilt.
- Ceiling and selection. The short Claude head-to-head tasks hit a pass-rate ceiling. We selected Fable by its observed median after the run. Ranges overlap. Accounts could run concurrently on the shared Mac.
- Protocol limit. The imported explorer receipts have no frozen pre-run protocol. They do not show that controls or call caps were set before the run.
- Tail limit. Calls can exceed p95. Summed slow ends and repeated-step counts are scenarios, not guarantees.
- Home advantage. We revised the routing decision cases against Jev's answers. This study uses Jev's time only, not its accuracy.
What to read next
- The voice agent latency budget study: every step, table and share.
- The routing hub: /routing
- Routing overhead study: rules vs Claude routers, with CLI start-up.
- CLI vs API latency study: first output and total time per route.
- What does a router cost you?
Time your own routes
Disclosure: I build Agent, the product behind these benchmarks. Agent uses the rules policy and Jev 1.13 in production, so two of the fast steps are our own choices.
Agent keeps a receipt for each task: the model, the route, the tokens, the time, the cost and the validation result. Try Agent and time your own routes.