Voice agent latency budget: component calculations, not a measured turn
Which decision steps fit inside a voice agent’s turn budget, using only measured times?
Published · 5 charts · Download the data or a carousel
14
The answer
Rules and Jev 1.13 fit all three assumed budgets at the median and p95. LLM routers through the Claude Code CLI fit none at the median. Calculation: the budgets are 300 ms, 800 ms and 1,500 ms per step. At the slow end, 3, 3 and 4 of 14 measured steps fit these budgets, respectively. At the median, 3, 3 and 8 fit. The deterministic policy took a median 1.42 µs (p95 2.33 µs, n = 20,000). Its maximum was 2.5 ms; its minimum was not retained. Jev 1.13 took a median 136.5 ms over HTTPS (p95 195.7 ms, n = 246). Its range was 100.9 ms to 297.3 ms. Calculation: this uses 46% of a 300 ms budget at the median and 65% at p95. Sonnet 5.5 routing through Claude Code took a median 2,598 ms (p95 4,298 ms, n = 82). Its range was 1,993 ms to 5,583 ms. Its median exceeds a 1,500 ms budget. GPT-6 Luna had the lowest observed API median for first output: 823 ms (n = 5). Its range was 506 ms to 1,371 ms. Calculation: its median is over an 800 ms budget. Its slowest run is over that budget. API ranges overlap, so this is not a speed ranking. These ranges are not confidence intervals. All times come from other studies on one Mac, over different routes. The source runs did not measure a voice turn or audio.
Key numbers
3
Steps that fit a 300 ms budget at the slow end (calculation)
of 14 (median: 3 of 14) · n = 14
3
Steps that fit an 800 ms budget at the slow end (calculation)
of 14 (median: 3 of 14) · n = 14
4
Steps that fit a 1,500 ms budget at the slow end (calculation)
of 14 (median: 8 of 14) · n = 14
46%
Share of a 300 ms budget that one Jev 1.13 decision uses at the median (calculation)
at the median, 65% at p95 · n = 246
4
Back-to-back Jev 1.13 decisions that fit an 800 ms budget at the slow end (calculation)
(5 at the median) · n = 246
173%
Share of a 1,500 ms budget that one Sonnet 5.5 router decision uses at the median (calculation)
at the median, 287% at p95 · n = 82
823ms
Lowest observed API median time to the first output of a one-line answer
(slowest run 1,371 ms) · n = 5
+1,964 to +2,880ms
Difference in median first-output time across 3 Codex model setups: Codex CLI minus API (calculation)
n = 5
540.5ms
Time left of a 1,500 ms budget after Jev 1.13 and the API cell with the lowest observed median, at the median (calculation)
left (sum of medians 959.5 ms) · n = 5
88.3ms
First Jev call minus the later-call median (calculation)
once · n = 1
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
- Time per step
- Budget (assumption: a budget a voice team might set)
Time per step · log scale: each gridline is 10 times the one before
| Item | Time per step | Budget (assumption: a budget a voice team might set) | Median to p95 | n |
|---|---|---|---|---|
| Deterministic routing policy (Agent, in process) | 1.42 µs | — | Time per step: 1.42 µs–2.33 µs | 20000 |
| Rule-based System One decision (record write excluded) | 1 ms | — | Time per step: 1 ms–2 ms | 419 |
| Jev 1.13 (TypeSafe, direct HTTPS) | 137 ms | — | Time per step: 137 ms–196 ms | 246 |
| Budget 300 ms | — | 300 ms | — | — |
| Budget 800 ms | — | 800 ms | — | — |
| Budget 1,500 ms | — | 1.5 s | — | — |
6 rows, 2 series: Time per step, Budget (assumption: a budget a voice team might set). Time per step: slowest Jev 1.13 (TypeSafe, direct HTTPS) 137 ms (median to p95 137 ms–196 ms, n 246). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap. Budget (assumption: a budget a voice team might set): slowest Budget 1,500 ms 1.5 s. Fastest Budget 300 ms 300 ms. Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n 246–20000 per row
Median per step; whiskers = median to p95; budget dots are assumptions
These 5 steps have a median under 1,000 ms. Decisions are whole calls; model steps are the time to the first useful output of a one-line answer. The dot is the median. The whisker runs from the median to the 95th percentile (30 or more runs per step). Observed limits stay in the measured-step table; some minima were not retained. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Provider explorer receipts: CLI vs API, Provider head-to-head: Claude Code models vs Codex efforts, Jev live run: 246 timed calls on the 82 routing decisions
- Time per step
- Budget (assumption: a budget a voice team might set)
| Item | Time per step | Budget (assumption: a budget a voice team might set) | Range (lowest–highest run) | n |
|---|---|---|---|---|
| OpenAI API · GPT-6 Luna · none (first output, one-line answer) | 823 ms | — | Time per step: 506 ms–1.37 s | 5 |
| OpenAI API · GPT-6.1 Sol · low (first output, one-line answer) | 873 ms | — | Time per step: 835 ms–1.74 s | 5 |
| Budget 300 ms | — | 300 ms | — | — |
| Budget 800 ms | — | 800 ms | — | — |
| Budget 1,500 ms | — | 1.5 s | — | — |
5 rows, 2 series: Time per step, Budget (assumption: a budget a voice team might set). Time per step: slowest OpenAI API · GPT-6.1 Sol · low (first output, one-line answer) 873 ms (range 835 ms–1.74 s, n 5). Fastest OpenAI API · GPT-6 Luna · none (first output, one-line answer) 823 ms (range 506 ms–1.37 s, n 5). All run ranges overlap. Budget (assumption: a budget a voice team might set): slowest Budget 1,500 ms 1.5 s. Fastest Budget 300 ms 300 ms. Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 5 per row
Median per step; whiskers = fastest to slowest run; budget dots are assumptions
These 5 steps have a median under 1,000 ms. Decisions are whole calls; model steps are the time to the first useful output of a one-line answer. The dot is the median. The whisker runs from the fastest to the slowest observed run (fewer than 30 runs per step). The maximum is not a p95 estimate. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Provider explorer receipts: CLI vs API, Provider head-to-head: Claude Code models vs Codex efforts, Jev live run: 246 timed calls on the 82 routing decisions
- Time per step
- Budget (assumption: a budget a voice team might set)
Time per step · log scale: each gridline is 10 times the one before
| Item | Time per step | Budget (assumption: a budget a voice team might set) | Median to p95 | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 (router, effort low, via Claude Code) | 2.6 s | — | Time per step: 2.6 s–4.3 s | 82 |
| Claude Haiku 4.5 (router, thinking on, via Claude Code) | 12.67 s | — | Time per step: 12.67 s–34.41 s | 82 |
| Budget 300 ms | — | 300 ms | — | — |
| Budget 800 ms | — | 800 ms | — | — |
| Budget 1,500 ms | — | 1.5 s | — | — |
5 rows, 2 series: Time per step, Budget (assumption: a budget a voice team might set). Time per step: slowest Claude Haiku 4.5 (router, thinking on, via Claude Code) 12.67 s (median to p95 12.67 s–34.41 s, n 82). Fastest Claude Sonnet 5.5 (router, effort low, via Claude Code) 2.6 s (median to p95 2.6 s–4.3 s, n 82). Not all run ranges overlap. Budget (assumption: a budget a voice team might set): slowest Budget 1,500 ms 1.5 s. Fastest Budget 300 ms 300 ms. Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n = 82 per row
Median per step; whiskers = median to p95; budget dots are assumptions
These 9 steps have a median of 1,000 ms or more. Routers are whole calls through the Claude Code CLI; the rest are the time to the first output of a model through a CLI or the API. The dot is the median. The whisker runs from the median to the 95th percentile (30 or more runs per step). Observed limits stay in the measured-step table; some minima were not retained. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Provider explorer receipts: CLI vs API, Provider head-to-head: Claude Code models vs Codex efforts, Jev live run: 246 timed calls on the 82 routing decisions
- Time per step
- Budget (assumption: a budget a voice team might set)
| Item | Time per step | Budget (assumption: a budget a voice team might set) | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Code · Claude Fable 5.1 (first output, 5 short tasks) | 1.2 s | — | Time per step: 947 ms–7.9 s | 15 |
| OpenAI API · GPT-6.1 Sol · high (first output, one-line answer) | 1.34 s | — | Time per step: 1.26 s–2.12 s | 5 |
| Claude Code · Claude Haiku 4.5 (start-up, first model output) | 1.46 s | — | Time per step: 1.21 s–2.31 s | 5 |
| Codex CLI · GPT-6 Luna · none (first output, one-line answer) | 2.79 s | — | Time per step: 2.46 s–3.42 s | 5 |
| Codex CLI · GPT-6.1 Sol · low (first output, one-line answer) | 3.75 s | — | Time per step: 3.44 s–4.1 s | 5 |
| Codex CLI · GPT-6.1 Sol · high (first output, one-line answer) | 3.79 s | — | Time per step: 3.37 s–4.3 s | 5 |
| Codex CLI (start-up, first model output) | 5.06 s | — | Time per step: 4.39 s–5.48 s | 5 |
| Budget 300 ms | — | 300 ms | — | — |
| Budget 800 ms | — | 800 ms | — | — |
| Budget 1,500 ms | — | 1.5 s | — | — |
10 rows, 2 series: Time per step, Budget (assumption: a budget a voice team might set). Time per step: slowest Codex CLI (start-up, first model output) 5.06 s (range 4.39 s–5.48 s, n 5). Fastest Claude Code · Claude Fable 5.1 (first output, 5 short tasks) 1.2 s (range 947 ms–7.9 s, n 15). Not all run ranges overlap. Budget (assumption: a budget a voice team might set): slowest Budget 1,500 ms 1.5 s. Fastest Budget 300 ms 300 ms. Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 5–15 per row
Median per step; whiskers = fastest to slowest run; budget dots are assumptions
These 9 steps have a median of 1,000 ms or more. Routers are whole calls through the Claude Code CLI; the rest are the time to the first output of a model through a CLI or the API. The dot is the median. The whisker runs from the fastest to the slowest observed run (fewer than 30 runs per step). The maximum is not a p95 estimate. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Provider explorer receipts: CLI vs API, Provider head-to-head: Claude Code models vs Codex efforts, Jev live run: 246 timed calls on the 82 routing decisions
- Fits at the median
- Fits at the slow end (square)
Gap labels, Fits at the slow end vs Fits at the median: Fits at the slow end is x% higher (+) or lower (−) than Fits at the median, calculated from the two values shown (the change counted from Fits at the median’s value).
| Item | Fits at the median | Fits at the slow end | n |
|---|---|---|---|
| 300 ms budget | 3 | 3 | 14 |
| 800 ms budget | 3 | 3 | 14 |
| 1,500 ms budget | 8 | 4 | 14 |
List-price calculation, not a run. 3 rows, 2 series: Fits at the median, Fits at the slow end. Fits at the median: highest 1,500 ms budget 8 (n 14). Lowest 800 ms budget 3 (n 14). Fits at the slow end: highest 1,500 ms budget 4 (n 14). Lowest 800 ms budget 3 (n 14).
Notesn = 14 per row
Count of 14 measured steps; n = steps
A calculation on measured times, with budgets of 300 ms, 800 ms and 1,500 ms (assumption: a budget a voice team might set). A step fits when its median, or its slow end, is at or below the budget. The slow end is the 95th percentile for steps with 30 or more runs and the slowest run for the rest, which is stricter. The steps differ in route and in what they time, so read the count as a summary of the table below, not as a ranking.
Sources: Voice-agent latency budget calculation, Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Provider explorer receipts: CLI vs API, Provider head-to-head: Claude Code models vs Codex efforts, Jev live run: 246 timed calls on the 82 routing decisions
Tables
Every step and the time it was measured to take
| Step | What the time covers | Route | Runs | Median | Observed range (not an interval) | Slow end | Slow end is | Where the time comes from |
|---|---|---|---|---|---|---|---|---|
| Deterministic routing policy (Agent, in process) | pure in-process policy decision; no database work | in process, no model call | 20,000 | 1.42 µs | Minimum not retained to 2.5 ms (n = 20,000) | 2.33 µs | 95th percentile | Routing overhead run: in-process policy timing, p50 and p95 (µs) |
| Rule-based System One decision (record write excluded) | whole decision, request to answer | in process, record write excluded | 419 | 1 ms | Minimum not retained to 3 ms (n = 419) | 2 ms | 95th percentile | Routing overhead run: recorded System One decisions (millisecond resolution) |
| Jev 1.13 (TypeSafe, direct HTTPS) | whole decision, request to answer | direct HTTPS, home network | 246 | 137 ms | 100.9 ms to 297.3 ms (n = 246) | 196 ms | 95th percentile | Jev live run: client wall time per call |
| OpenAI API · GPT-6 Luna · none (first output, one-line answer) | time to first useful output | OpenAI API | 5 | 823 ms | 506 ms to 1,371 ms (n = 5) | 1.37 s | slowest run | Provider explorer receipts: first useful output, matched cohort, fixed exact reply |
| OpenAI API · GPT-6.1 Sol · low (first output, one-line answer) | time to first useful output | OpenAI API | 5 | 873 ms | 835 ms to 1,742 ms (n = 5) | 1.74 s | slowest run | Provider explorer receipts: first useful output, matched cohort, fixed exact reply |
| Claude Code · Claude Fable 5.1 (first output, 5 short tasks) | time to first useful output | Claude Code CLI | 15 | 1.2 s | 947 ms to 7,902 ms (n = 15) | 7.9 s | slowest run | Provider head-to-head receipts: first useful output, configuration with the lowest observed median |
| OpenAI API · GPT-6.1 Sol · high (first output, one-line answer) | time to first useful output | OpenAI API | 5 | 1.34 s | 1,259 ms to 2,115 ms (n = 5) | 2.12 s | slowest run | Provider explorer receipts: first useful output, matched cohort, fixed exact reply |
| Claude Code · Claude Haiku 4.5 (start-up, first model output) | time to first useful output | Claude Code CLI, one-word prompt | 5 | 1.46 s | 1,206 ms to 2,308 ms (n = 5) | 2.31 s | slowest run | Routing overhead run: CLI start-up, time to the first model output |
| Claude Sonnet 5.5 (router, effort low, via Claude Code) | whole decision, request to answer | Claude Code CLI | 82 | 2.6 s | 1,993 ms to 5,583 ms (n = 82) | 4.3 s | 95th percentile | Routing run: median and interpolated p95; overhead extract: n and observed limits |
| Codex CLI · GPT-6 Luna · none (first output, one-line answer) | time to first useful output | Codex CLI | 5 | 2.79 s | 2,461 ms to 3,418 ms (n = 5) | 3.42 s | slowest run | Provider explorer receipts: first useful output, matched cohort, fixed exact reply |
| Codex CLI · GPT-6.1 Sol · low (first output, one-line answer) | time to first useful output | Codex CLI | 5 | 3.75 s | 3,435 ms to 4,103 ms (n = 5) | 4.1 s | slowest run | Provider explorer receipts: first useful output, matched cohort, fixed exact reply |
| Codex CLI · GPT-6.1 Sol · high (first output, one-line answer) | time to first useful output | Codex CLI | 5 | 3.79 s | 3,366 ms to 4,296 ms (n = 5) | 4.3 s | slowest run | Provider explorer receipts: first useful output, matched cohort, fixed exact reply |
| Codex CLI (start-up, first model output) | time to first useful output | Codex CLI, one-word prompt | 5 | 5.06 s | 4,391 ms to 5,478 ms (n = 5) | 5.48 s | slowest run | Routing overhead run: CLI start-up, time to the first model output |
| Claude Haiku 4.5 (router, thinking on, via Claude Code) | whole decision, request to answer | Claude Code CLI | 82 | 12.67 s | 5,857 ms to 51.28 s (n = 82) | 34.41 s | 95th percentile | Routing run: median and interpolated p95; overhead extract: n and observed limits |
Which steps fit each assumed budget (calculation)
| Step | Fits 300 ms | Fits 800 ms | Fits 1,500 ms |
|---|---|---|---|
| Deterministic routing policy (Agent, in process) | median and slow end | median and slow end | median and slow end |
| Rule-based System One decision (record write excluded) | median and slow end | median and slow end | median and slow end |
| Jev 1.13 (TypeSafe, direct HTTPS) | median and slow end | median and slow end | median and slow end |
| OpenAI API · GPT-6 Luna · none (first output, one-line answer) | neither | neither | median and slow end |
| OpenAI API · GPT-6.1 Sol · low (first output, one-line answer) | neither | neither | median only |
| Claude Code · Claude Fable 5.1 (first output, 5 short tasks) | neither | neither | median only |
| OpenAI API · GPT-6.1 Sol · high (first output, one-line answer) | neither | neither | median only |
| Claude Code · Claude Haiku 4.5 (start-up, first model output) | neither | neither | median only |
| Claude Sonnet 5.5 (router, effort low, via Claude Code) | neither | neither | neither |
| Codex CLI · GPT-6 Luna · none (first output, one-line answer) | neither | neither | neither |
| Codex CLI · GPT-6.1 Sol · low (first output, one-line answer) | neither | neither | neither |
| Codex CLI · GPT-6.1 Sol · high (first output, one-line answer) | neither | neither | neither |
| Codex CLI (start-up, first model output) | neither | neither | neither |
| Claude Haiku 4.5 (router, thinking on, via Claude Code) | neither | neither | neither |
How many runs of one step fit back to back in each budget (calculation, counts stop at 1,000)
| Step | 300 ms: at the median | 300 ms: at the slow end | 800 ms: at the median | 800 ms: at the slow end | 1,500 ms: at the median | 1,500 ms: at the slow end |
|---|---|---|---|---|---|---|
| Deterministic routing policy (Agent, in process) | 1,000 | 1,000 | 1,000 | 1,000 | 1,000 | 1,000 |
| Rule-based System One decision (record write excluded) | 300 | 150 | 800 | 400 | 1,000 | 750 |
| Jev 1.13 (TypeSafe, direct HTTPS) | 2 | 1 | 5 | 4 | 10 | 7 |
| OpenAI API · GPT-6 Luna · none (first output, one-line answer) | 0 | 0 | 0 | 0 | 1 | 1 |
| OpenAI API · GPT-6.1 Sol · low (first output, one-line answer) | 0 | 0 | 0 | 0 | 1 | 0 |
| Claude Code · Claude Fable 5.1 (first output, 5 short tasks) | 0 | 0 | 0 | 0 | 1 | 0 |
| OpenAI API · GPT-6.1 Sol · high (first output, one-line answer) | 0 | 0 | 0 | 0 | 1 | 0 |
| Claude Code · Claude Haiku 4.5 (start-up, first model output) | 0 | 0 | 0 | 0 | 1 | 0 |
| Claude Sonnet 5.5 (router, effort low, via Claude Code) | 0 | 0 | 0 | 0 | 0 | 0 |
| Codex CLI · GPT-6 Luna · none (first output, one-line answer) | 0 | 0 | 0 | 0 | 0 | 0 |
| Codex CLI · GPT-6.1 Sol · low (first output, one-line answer) | 0 | 0 | 0 | 0 | 0 | 0 |
| Codex CLI · GPT-6.1 Sol · high (first output, one-line answer) | 0 | 0 | 0 | 0 | 0 | 0 |
| Codex CLI (start-up, first model output) | 0 | 0 | 0 | 0 | 0 | 0 |
| Claude Haiku 4.5 (router, thinking on, via Claude Code) | 0 | 0 | 0 | 0 | 0 | 0 |
A router in front of a model: the two times added (calculation)
| Router (decision) | Model (first output) | Sum of medians | Sum of slow ends | Fits 300 ms | Fits 800 ms | Fits 1,500 ms | Left of 1,500 ms at the median (negative: over) | Left of 1,500 ms at the slow end (negative: over) |
|---|---|---|---|---|---|---|---|---|
| Deterministic routing policy (Agent, in process) | OpenAI API · GPT-6 Luna · none (first output, one-line answer) | 823 ms | 1.37 s | neither | neither | median and slow end | 677 ms | 129 ms |
| Deterministic routing policy (Agent, in process) | OpenAI API · GPT-6.1 Sol · low (first output, one-line answer) | 873 ms | 1.74 s | neither | neither | median only | 627 ms | -242 ms |
| Jev 1.13 (TypeSafe, direct HTTPS) | OpenAI API · GPT-6 Luna · none (first output, one-line answer) | 960 ms | 1.57 s | neither | neither | median only | 541 ms | -67 ms |
| Jev 1.13 (TypeSafe, direct HTTPS) | OpenAI API · GPT-6.1 Sol · low (first output, one-line answer) | 1.01 s | 1.94 s | neither | neither | median only | 491 ms | -438 ms |
| Claude Sonnet 5.5 (router, effort low, via Claude Code) | OpenAI API · GPT-6 Luna · none (first output, one-line answer) | 3.42 s | 5.67 s | neither | neither | neither | -1.92 s | -4.17 s |
| Claude Sonnet 5.5 (router, effort low, via Claude Code) | OpenAI API · GPT-6.1 Sol · low (first output, one-line answer) | 3.47 s | 6.04 s | neither | neither | neither | -1.97 s | -4.54 s |
Method
- Thought experiment on recorded times. This study makes no new model calls. Every time comes from a run that another study published.
- Budgets: 300 ms, 800 ms, 1,500 ms per decision step (assumption: a budget a voice team might set). Nothing in the data says what a product needs; change the budget and the table gives the share and the count for it.
- Deterministic routing policy: the policy timing of the routing overhead run (latencyUs p50 and p95, converted to milliseconds). Rule-based System One decision: the recorded decision times of that run (millisecond resolution, record write excluded).
- Jev 1.13: this study checks the per-call rows against the live-run summary. The rows record client wall time over HTTPS. They give n, median and p95 (latencyMs calls). The first call uses a fresh connection (coldFirstCall). This calculation subtracts the later-call median (callsWithoutColdFirst) from the first-call time. The run starts at run.startedAt and ends at run.endedAt. The client uses one Apple Silicon Mac on a home network.
- Claude Sonnet 5.5 and Claude Haiku 4.5 as routers: median and interpolated p95 from the routing extract (latencyMs wallP50 and wallP95). The overhead extract supplies n and the observed limits, through the Claude Code CLI, one call at a time. CLI start-up: first model output on a one-word prompt, from the routing overhead run (firstModelEventMs; 5 runs per CLI).
- Model first output: this study uses the matched provider explorer cohort for the fixed exact reply. Each configuration has 5 runs with firstUsefulMs timings. Routes are the OpenAI API and the Codex CLI. This study also selects the Claude Code configuration with the lowest observed first-output median in the 5-task head-to-head. Its timing sample includes completed validator failures. Failed or untimed calls have no successful first-output time.
- Slow end: the 95th percentile where a step has 30 or more runs, and the slowest run where it has fewer. The slowest run is stricter than a 95th percentile.
- Calculations: share of a budget = time ÷ budget, at the median and at the slow end. A step fits when its time is at or below the budget. Runs that fit back to back = the budget ÷ the time, rounded down, counted up to 1,000. A router in front of a model adds the two medians and the two slow ends, as a scenario. Neither sum is a measured percentile or a guaranteed upper bound for the pair.
- The charts separate steps with a median under 1,000 ms from the rest so each axis stays readable. Separate charts show median-to-p95 bands and observed min-to-max ranges. A range maximum is not a p95.
Caveats
- Timing samples omit 0 failed or untimed matched explorer calls and 0 failed or untimed Claude head-to-head calls. A failure is not a fast successful step. No pass rate is estimated here.
- The overhead extract uses nearest-rank quantiles. For LLM routers this study instead uses the conventional median and interpolated p95 from the routing extract. Per-iteration policy times and individual rule-record times are not retained, so those summaries cannot be independently rebuilt.
- The rule-arm latency is captured before the decision record is written. It excludes that write, despite the source summary label.
- The policy timing excludes database reads, policy merge and decision-record insertion. Its minimum was not retained. The rule-record timing also lacks a retained minimum.
- The Claude head-to-head tasks hit a pass-rate ceiling. This study selects the lowest observed Claude median after the run; ranges overlap and this selection is not a speed ranking. The accounts could run concurrently on the shared Mac.
- The 30-run cutoff is a display rule. It does not make a sample p95 a reliable production tail estimate. Some recorded calls exceed the p95; the table keeps their maxima visible.
- The budgets are assumptions. A voice product may need less or more; this study does not say what listeners accept.
- The steps differ in route and in what they time. Routers are whole decisions; models are the time to the first useful text, as the harness records it. First output is not necessarily enough text to start speech. A router through a CLI is a deployment choice, not a model limit.
- All times were recorded on one Mac, not on a server near the vendor, and on different days. Vendor latency changes over a day, and a server near the API would see different times.
- Small samples: 5 runs per model configuration, 5 per CLI start-up, 15 for the Claude Code configuration. Their slow end is the slowest run, a range and not a confidence interval.
- The imported explorer evidence has no frozen pre-run protocol. Those receipts do not establish that its controls and call caps were set before the run.
- The explorer uses a one-line answer; CLI start-up uses a one-word prompt. The selected Claude Code cell uses five short tasks, including code. Longer prompts and replies can change first-output time. No speech or audio was measured.
- No speech recognition, speech synthesis, turn detection or network round trip to a user is timed here. API wall times include the client-to-provider network. The steps are only the decision and the first model output.
- Calculations add medians and add slow ends across runs that were not made together. Medians do not add up exactly, and sums of p95 values do not guarantee a pair p95 or an upper bound. Repeated-step counts assume the same duration each time; they do not give a probability of fitting.
- Jev was timed in one run of 35 seconds, from one client machine on one network path, with calls one at a time. A hosted API can be faster or slower at another hour or from another region. The Jev decision cases were revised against Jev’s own answers, so its accuracy has a home advantage; this study uses only its time, not its accuracy.
Sources
Voice-agent latency budget calculation
Recorded decision and first-output times compared with assumed 300, 800 and 1,500 ms budgets. No complete voice turn ran.
Routing overhead runs: policy microbenchmark and CLI start-up
In-process timing of the deterministic routing policy (20,000 timed decisions), CLI start-up with a one-word prompt (5 runs per CLI), and decision counts read from recorded bench runs. LLM router timings are reused from the routing runs.
Routing runs: Jev router vs LLM routing
Routing decisions recorded per case and arm.
Provider explorer receipts: CLI vs API
230 imported receipts for short fixed tasks over Claude Code CLI, Codex CLI and the OpenAI API, with time to first useful output, total time, tokens and validation.
Provider head-to-head: Claude Code models vs Codex efforts
Five short tasks with deterministic validators, declared protocol, every attempt kept.
Jev live run: 246 timed calls on the 82 routing decisions
Jev 1.13 called over HTTPS, 3 repeats of the same 82 typed decisions, one call at a time, from one Mac over a home network: client wall time, with the network inside it. The API reports no server time. Cost per 1,000 decisions is a calculation from the reported input tokens and the published price. The case sets were revised against Jev answers, so Jev has a home advantage.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “Voice agent latency budget: component calculations, not a measured turn”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/voice-agent-latency-budget.
Write-ups on this study
A voice agent latency budget, with measured times: what fits in one turn?
Rules and Jev 1.13 fit every budget we assumed; a Claude router through a CLI fits none. 14 measured steps vs 300, 800 and 1,500 ms. A thought experiment.
More studies
All benchmarksWhat does routing a million AI requests a day cost? A calculation from measured runs
A calculation from measured runs: what 10,000 to 10 million routing decisions a day cost with rules, Jev and Claude, with median/p95 time scenarios.
Routing overhead: deterministic policy vs LLM routers vs Jev
How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.
How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.