• Voice Agent
  • Latency
  • Latency Budget
  • Routing
  • Jev
  • Thought experiment
  • Calculation

Voice agent latency budget: component calculations, not a measured turn

Which decision steps fit inside a voice agent’s turn budget, using only measured times?

Published · 5 charts · Download the data or a carousel

14

n = 14

Decision and first-output steps compared · steps

5 under 1,000 ms at the median, 9 at 1,000 ms or more. Every time is a measured median from another study.

The answer

Rules and Jev 1.13 fit all three assumed budgets at the median and p95. LLM routers through the Claude Code CLI fit none at the median. Calculation: the budgets are 300 ms, 800 ms and 1,500 ms per step. At the slow end, 3, 3 and 4 of 14 measured steps fit these budgets, respectively. At the median, 3, 3 and 8 fit. The deterministic policy took a median 1.42 µs (p95 2.33 µs, n = 20,000). Its maximum was 2.5 ms; its minimum was not retained. Jev 1.13 took a median 136.5 ms over HTTPS (p95 195.7 ms, n = 246). Its range was 100.9 ms to 297.3 ms. Calculation: this uses 46% of a 300 ms budget at the median and 65% at p95. Sonnet 5.5 routing through Claude Code took a median 2,598 ms (p95 4,298 ms, n = 82). Its range was 1,993 ms to 5,583 ms. Its median exceeds a 1,500 ms budget. GPT-6 Luna had the lowest observed API median for first output: 823 ms (n = 5). Its range was 506 ms to 1,371 ms. Calculation: its median is over an 800 ms budget. Its slowest run is over that budget. API ranges overlap, so this is not a speed ranking. These ranges are not confidence intervals. All times come from other studies on one Mac, over different routes. The source runs did not measure a voice turn or audio.

Key numbers

3

Steps that fit a 300 ms budget at the slow end (calculation)

of 14 (median: 3 of 14) · n = 14

3

Steps that fit an 800 ms budget at the slow end (calculation)

of 14 (median: 3 of 14) · n = 14

4

Steps that fit a 1,500 ms budget at the slow end (calculation)

of 14 (median: 8 of 14) · n = 14

46%

Share of a 300 ms budget that one Jev 1.13 decision uses at the median (calculation)

at the median, 65% at p95 · n = 246

4

Back-to-back Jev 1.13 decisions that fit an 800 ms budget at the slow end (calculation)

(5 at the median) · n = 246

173%

Share of a 1,500 ms budget that one Sonnet 5.5 router decision uses at the median (calculation)

at the median, 287% at p95 · n = 82

823ms

Lowest observed API median time to the first output of a one-line answer

(slowest run 1,371 ms) · n = 5

+1,964 to +2,880ms

Difference in median first-output time across 3 Codex model setups: Codex CLI minus API (calculation)

n = 5

540.5ms

Time left of a 1,500 ms budget after Jev 1.13 and the API cell with the lowest observed median, at the median (calculation)

left (sum of medians 959.5 ms) · n = 5

88.3ms

First Jev call minus the later-call median (calculation)

once · n = 1

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Thought experiment: not a run. These values reprice recorded tokens at list prices. No model was called again.

  • Time per step
  • Budget (assumption: a budget a voice team might set)
Deterministic routing policy (Agent, in process)
Rule-based System One decision (record write excluded)
Jev 1.13 (TypeSafe, direct HTTPS)
Budget 300 ms
Budget 800 ms
Budget 1,500 ms

Time per step · log scale: each gridline is 10 times the one before

6 rows, 2 series: Time per step, Budget (assumption: a budget a voice team might set). Time per step: slowest Jev 1.13 (TypeSafe, direct HTTPS) 137 ms (median to p95 137 ms–196 ms, n 246). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap. Budget (assumption: a budget a voice team might set): slowest Budget 1,500 ms 1.5 s. Fastest Budget 300 ms 300 ms. Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 246–20000 per row

Median per step; whiskers = median to p95; budget dots are assumptions

These 5 steps have a median under 1,000 ms. Decisions are whole calls; model steps are the time to the first useful output of a one-line answer. The dot is the median. The whisker runs from the median to the 95th percentile (30 or more runs per step). Observed limits stay in the measured-step table; some minima were not retained. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Provider explorer receipts: CLI vs API, Provider head-to-head: Claude Code models vs Codex efforts, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
  • Time per step
  • Budget (assumption: a budget a voice team might set)
Entrance: medians race at 1.1× real timeMotion reduced: press Replay to animateThe slowest median is 1.5 s. The clock runs at the recorded speed.
OpenAI API · GPT-6 Luna · none (first output, one-line answer)
OpenAI API · GPT-6.1 Sol · low (first output, one-line answer)
Budget 300 ms
Budget 800 ms
Budget 1,500 ms

5 rows, 2 series: Time per step, Budget (assumption: a budget a voice team might set). Time per step: slowest OpenAI API · GPT-6.1 Sol · low (first output, one-line answer) 873 ms (range 835 ms–1.74 s, n 5). Fastest OpenAI API · GPT-6 Luna · none (first output, one-line answer) 823 ms (range 506 ms–1.37 s, n 5). All run ranges overlap. Budget (assumption: a budget a voice team might set): slowest Budget 1,500 ms 1.5 s. Fastest Budget 300 ms 300 ms. Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 5 per row

Median per step; whiskers = fastest to slowest run; budget dots are assumptions

These 5 steps have a median under 1,000 ms. Decisions are whole calls; model steps are the time to the first useful output of a one-line answer. The dot is the median. The whisker runs from the fastest to the slowest observed run (fewer than 30 runs per step). The maximum is not a p95 estimate. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Provider explorer receipts: CLI vs API, Provider head-to-head: Claude Code models vs Codex efforts, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
  • Time per step
  • Budget (assumption: a budget a voice team might set)
Claude Sonnet 5.5 (router, effort low, via Claude Code)
Claude Haiku 4.5 (router, thinking on, via Claude Code)
Budget 300 ms
Budget 800 ms
Budget 1,500 ms

Time per step · log scale: each gridline is 10 times the one before

5 rows, 2 series: Time per step, Budget (assumption: a budget a voice team might set). Time per step: slowest Claude Haiku 4.5 (router, thinking on, via Claude Code) 12.67 s (median to p95 12.67 s–34.41 s, n 82). Fastest Claude Sonnet 5.5 (router, effort low, via Claude Code) 2.6 s (median to p95 2.6 s–4.3 s, n 82). Not all run ranges overlap. Budget (assumption: a budget a voice team might set): slowest Budget 1,500 ms 1.5 s. Fastest Budget 300 ms 300 ms. Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n = 82 per row

Median per step; whiskers = median to p95; budget dots are assumptions

These 9 steps have a median of 1,000 ms or more. Routers are whole calls through the Claude Code CLI; the rest are the time to the first output of a model through a CLI or the API. The dot is the median. The whisker runs from the median to the 95th percentile (30 or more runs per step). Observed limits stay in the measured-step table; some minima were not retained. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Provider explorer receipts: CLI vs API, Provider head-to-head: Claude Code models vs Codex efforts, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
  • Time per step
  • Budget (assumption: a budget a voice team might set)
Entrance: medians race at 3.6× real timeMotion reduced: press Replay to animateThe slowest median is 5.06 s. The clock runs at the recorded speed.
Claude Code · Claude Fable 5.1 (first output, 5 short tasks)
OpenAI API · GPT-6.1 Sol · high (first output, one-line answer)
Claude Code · Claude Haiku 4.5 (start-up, first model output)
Codex CLI · GPT-6 Luna · none (first output, one-line answer)
Codex CLI · GPT-6.1 Sol · low (first output, one-line answer)
Codex CLI · GPT-6.1 Sol · high (first output, one-line answer)
Codex CLI (start-up, first model output)
Budget 300 ms
Budget 800 ms
Budget 1,500 ms

10 rows, 2 series: Time per step, Budget (assumption: a budget a voice team might set). Time per step: slowest Codex CLI (start-up, first model output) 5.06 s (range 4.39 s–5.48 s, n 5). Fastest Claude Code · Claude Fable 5.1 (first output, 5 short tasks) 1.2 s (range 947 ms–7.9 s, n 15). Not all run ranges overlap. Budget (assumption: a budget a voice team might set): slowest Budget 1,500 ms 1.5 s. Fastest Budget 300 ms 300 ms. Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 5–15 per row

Median per step; whiskers = fastest to slowest run; budget dots are assumptions

These 9 steps have a median of 1,000 ms or more. Routers are whole calls through the Claude Code CLI; the rest are the time to the first output of a model through a CLI or the API. The dot is the median. The whisker runs from the fastest to the slowest observed run (fewer than 30 runs per step). The maximum is not a p95 estimate. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Provider explorer receipts: CLI vs API, Provider head-to-head: Claude Code models vs Codex efforts, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
Calculation
  • Fits at the median
  • Fits at the slow end (square)
In chart order.
300 ms budget
800 ms budget
1,500 ms budget

Gap labels, Fits at the slow end vs Fits at the median: Fits at the slow end is x% higher (+) or lower (−) than Fits at the median, calculated from the two values shown (the change counted from Fits at the median’s value).

List-price calculation, not a run. 3 rows, 2 series: Fits at the median, Fits at the slow end. Fits at the median: highest 1,500 ms budget 8 (n 14). Lowest 800 ms budget 3 (n 14). Fits at the slow end: highest 1,500 ms budget 4 (n 14). Lowest 800 ms budget 3 (n 14).

Notesn = 14 per row

Count of 14 measured steps; n = steps

A calculation on measured times, with budgets of 300 ms, 800 ms and 1,500 ms (assumption: a budget a voice team might set). A step fits when its median, or its slow end, is at or below the budget. The slow end is the 95th percentile for steps with 30 or more runs and the slowest run for the rest, which is stricter. The steps differ in route and in what they time, so read the count as a summary of the table below, not as a ranking.

Sources: Voice-agent latency budget calculation, Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Provider explorer receipts: CLI vs API, Provider head-to-head: Claude Code models vs Codex efforts, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)

Tables

Every step and the time it was measured to take

StepWhat the time coversRouteRunsMedianObserved range (not an interval)Slow endSlow end isWhere the time comes from
Deterministic routing policy (Agent, in process)pure in-process policy decision; no database workin process, no model call20,0001.42 µsMinimum not retained to 2.5 ms (n = 20,000)2.33 µs95th percentileRouting overhead run: in-process policy timing, p50 and p95 (µs)
Rule-based System One decision (record write excluded)whole decision, request to answerin process, record write excluded4191 msMinimum not retained to 3 ms (n = 419)2 ms95th percentileRouting overhead run: recorded System One decisions (millisecond resolution)
Jev 1.13 (TypeSafe, direct HTTPS)whole decision, request to answerdirect HTTPS, home network246137 ms100.9 ms to 297.3 ms (n = 246)196 ms95th percentileJev live run: client wall time per call
OpenAI API · GPT-6 Luna · none (first output, one-line answer)time to first useful outputOpenAI API5823 ms506 ms to 1,371 ms (n = 5)1.37 sslowest runProvider explorer receipts: first useful output, matched cohort, fixed exact reply
OpenAI API · GPT-6.1 Sol · low (first output, one-line answer)time to first useful outputOpenAI API5873 ms835 ms to 1,742 ms (n = 5)1.74 sslowest runProvider explorer receipts: first useful output, matched cohort, fixed exact reply
Claude Code · Claude Fable 5.1 (first output, 5 short tasks)time to first useful outputClaude Code CLI151.2 s947 ms to 7,902 ms (n = 15)7.9 sslowest runProvider head-to-head receipts: first useful output, configuration with the lowest observed median
OpenAI API · GPT-6.1 Sol · high (first output, one-line answer)time to first useful outputOpenAI API51.34 s1,259 ms to 2,115 ms (n = 5)2.12 sslowest runProvider explorer receipts: first useful output, matched cohort, fixed exact reply
Claude Code · Claude Haiku 4.5 (start-up, first model output)time to first useful outputClaude Code CLI, one-word prompt51.46 s1,206 ms to 2,308 ms (n = 5)2.31 sslowest runRouting overhead run: CLI start-up, time to the first model output
Claude Sonnet 5.5 (router, effort low, via Claude Code)whole decision, request to answerClaude Code CLI822.6 s1,993 ms to 5,583 ms (n = 82)4.3 s95th percentileRouting run: median and interpolated p95; overhead extract: n and observed limits
Codex CLI · GPT-6 Luna · none (first output, one-line answer)time to first useful outputCodex CLI52.79 s2,461 ms to 3,418 ms (n = 5)3.42 sslowest runProvider explorer receipts: first useful output, matched cohort, fixed exact reply
Codex CLI · GPT-6.1 Sol · low (first output, one-line answer)time to first useful outputCodex CLI53.75 s3,435 ms to 4,103 ms (n = 5)4.1 sslowest runProvider explorer receipts: first useful output, matched cohort, fixed exact reply
Codex CLI · GPT-6.1 Sol · high (first output, one-line answer)time to first useful outputCodex CLI53.79 s3,366 ms to 4,296 ms (n = 5)4.3 sslowest runProvider explorer receipts: first useful output, matched cohort, fixed exact reply
Codex CLI (start-up, first model output)time to first useful outputCodex CLI, one-word prompt55.06 s4,391 ms to 5,478 ms (n = 5)5.48 sslowest runRouting overhead run: CLI start-up, time to the first model output
Claude Haiku 4.5 (router, thinking on, via Claude Code)whole decision, request to answerClaude Code CLI8212.67 s5,857 ms to 51.28 s (n = 82)34.41 s95th percentileRouting run: median and interpolated p95; overhead extract: n and observed limits

Which steps fit each assumed budget (calculation)

StepFits 300 msFits 800 msFits 1,500 ms
Deterministic routing policy (Agent, in process)median and slow endmedian and slow endmedian and slow end
Rule-based System One decision (record write excluded)median and slow endmedian and slow endmedian and slow end
Jev 1.13 (TypeSafe, direct HTTPS)median and slow endmedian and slow endmedian and slow end
OpenAI API · GPT-6 Luna · none (first output, one-line answer)neitherneithermedian and slow end
OpenAI API · GPT-6.1 Sol · low (first output, one-line answer)neitherneithermedian only
Claude Code · Claude Fable 5.1 (first output, 5 short tasks)neitherneithermedian only
OpenAI API · GPT-6.1 Sol · high (first output, one-line answer)neitherneithermedian only
Claude Code · Claude Haiku 4.5 (start-up, first model output)neitherneithermedian only
Claude Sonnet 5.5 (router, effort low, via Claude Code)neitherneitherneither
Codex CLI · GPT-6 Luna · none (first output, one-line answer)neitherneitherneither
Codex CLI · GPT-6.1 Sol · low (first output, one-line answer)neitherneitherneither
Codex CLI · GPT-6.1 Sol · high (first output, one-line answer)neitherneitherneither
Codex CLI (start-up, first model output)neitherneitherneither
Claude Haiku 4.5 (router, thinking on, via Claude Code)neitherneitherneither

Share of each assumed budget that one step uses (calculation)

Step300 ms: median300 ms: slow end800 ms: median800 ms: slow end1,500 ms: median1,500 ms: slow end
Deterministic routing policy (Agent, in process)under 0.1%under 0.1%under 0.1%under 0.1%under 0.1%under 0.1%
Rule-based System One decision (record write excluded)0.3%0.7%0.1%0.3%under 0.1%0.1%
Jev 1.13 (TypeSafe, direct HTTPS)46%65%17%24%9.1%13%
OpenAI API · GPT-6 Luna · none (first output, one-line answer)274%457%103%171%55%91%
OpenAI API · GPT-6.1 Sol · low (first output, one-line answer)291%581%109%218%58%116%
Claude Code · Claude Fable 5.1 (first output, 5 short tasks)399%2634%150%988%80%527%
OpenAI API · GPT-6.1 Sol · high (first output, one-line answer)447%705%168%264%89%141%
Claude Code · Claude Haiku 4.5 (start-up, first model output)487%769%183%289%97%154%
Claude Sonnet 5.5 (router, effort low, via Claude Code)866%1433%325%537%173%287%
Codex CLI · GPT-6 Luna · none (first output, one-line answer)929%1139%348%427%186%228%
Codex CLI · GPT-6.1 Sol · low (first output, one-line answer)1251%1368%469%513%250%274%
Codex CLI · GPT-6.1 Sol · high (first output, one-line answer)1262%1432%473%537%252%286%
Codex CLI (start-up, first model output)1686%1826%632%685%337%365%
Claude Haiku 4.5 (router, thinking on, via Claude Code)4225%11471%1584%4302%845%2294%

How many runs of one step fit back to back in each budget (calculation, counts stop at 1,000)

Step300 ms: at the median300 ms: at the slow end800 ms: at the median800 ms: at the slow end1,500 ms: at the median1,500 ms: at the slow end
Deterministic routing policy (Agent, in process)1,0001,0001,0001,0001,0001,000
Rule-based System One decision (record write excluded)3001508004001,000750
Jev 1.13 (TypeSafe, direct HTTPS)2154107
OpenAI API · GPT-6 Luna · none (first output, one-line answer)000011
OpenAI API · GPT-6.1 Sol · low (first output, one-line answer)000010
Claude Code · Claude Fable 5.1 (first output, 5 short tasks)000010
OpenAI API · GPT-6.1 Sol · high (first output, one-line answer)000010
Claude Code · Claude Haiku 4.5 (start-up, first model output)000010
Claude Sonnet 5.5 (router, effort low, via Claude Code)000000
Codex CLI · GPT-6 Luna · none (first output, one-line answer)000000
Codex CLI · GPT-6.1 Sol · low (first output, one-line answer)000000
Codex CLI · GPT-6.1 Sol · high (first output, one-line answer)000000
Codex CLI (start-up, first model output)000000
Claude Haiku 4.5 (router, thinking on, via Claude Code)000000

A router in front of a model: the two times added (calculation)

Router (decision)Model (first output)Sum of mediansSum of slow endsFits 300 msFits 800 msFits 1,500 msLeft of 1,500 ms at the median (negative: over)Left of 1,500 ms at the slow end (negative: over)
Deterministic routing policy (Agent, in process)OpenAI API · GPT-6 Luna · none (first output, one-line answer)823 ms1.37 sneitherneithermedian and slow end677 ms129 ms
Deterministic routing policy (Agent, in process)OpenAI API · GPT-6.1 Sol · low (first output, one-line answer)873 ms1.74 sneitherneithermedian only627 ms-242 ms
Jev 1.13 (TypeSafe, direct HTTPS)OpenAI API · GPT-6 Luna · none (first output, one-line answer)960 ms1.57 sneitherneithermedian only541 ms-67 ms
Jev 1.13 (TypeSafe, direct HTTPS)OpenAI API · GPT-6.1 Sol · low (first output, one-line answer)1.01 s1.94 sneitherneithermedian only491 ms-438 ms
Claude Sonnet 5.5 (router, effort low, via Claude Code)OpenAI API · GPT-6 Luna · none (first output, one-line answer)3.42 s5.67 sneitherneitherneither-1.92 s-4.17 s
Claude Sonnet 5.5 (router, effort low, via Claude Code)OpenAI API · GPT-6.1 Sol · low (first output, one-line answer)3.47 s6.04 sneitherneitherneither-1.97 s-4.54 s

Method

  1. Thought experiment on recorded times. This study makes no new model calls. Every time comes from a run that another study published.
  2. Budgets: 300 ms, 800 ms, 1,500 ms per decision step (assumption: a budget a voice team might set). Nothing in the data says what a product needs; change the budget and the table gives the share and the count for it.
  3. Deterministic routing policy: the policy timing of the routing overhead run (latencyUs p50 and p95, converted to milliseconds). Rule-based System One decision: the recorded decision times of that run (millisecond resolution, record write excluded).
  4. Jev 1.13: this study checks the per-call rows against the live-run summary. The rows record client wall time over HTTPS. They give n, median and p95 (latencyMs calls). The first call uses a fresh connection (coldFirstCall). This calculation subtracts the later-call median (callsWithoutColdFirst) from the first-call time. The run starts at run.startedAt and ends at run.endedAt. The client uses one Apple Silicon Mac on a home network.
  5. Claude Sonnet 5.5 and Claude Haiku 4.5 as routers: median and interpolated p95 from the routing extract (latencyMs wallP50 and wallP95). The overhead extract supplies n and the observed limits, through the Claude Code CLI, one call at a time. CLI start-up: first model output on a one-word prompt, from the routing overhead run (firstModelEventMs; 5 runs per CLI).
  6. Model first output: this study uses the matched provider explorer cohort for the fixed exact reply. Each configuration has 5 runs with firstUsefulMs timings. Routes are the OpenAI API and the Codex CLI. This study also selects the Claude Code configuration with the lowest observed first-output median in the 5-task head-to-head. Its timing sample includes completed validator failures. Failed or untimed calls have no successful first-output time.
  7. Slow end: the 95th percentile where a step has 30 or more runs, and the slowest run where it has fewer. The slowest run is stricter than a 95th percentile.
  8. Calculations: share of a budget = time ÷ budget, at the median and at the slow end. A step fits when its time is at or below the budget. Runs that fit back to back = the budget ÷ the time, rounded down, counted up to 1,000. A router in front of a model adds the two medians and the two slow ends, as a scenario. Neither sum is a measured percentile or a guaranteed upper bound for the pair.
  9. The charts separate steps with a median under 1,000 ms from the rest so each axis stays readable. Separate charts show median-to-p95 bands and observed min-to-max ranges. A range maximum is not a p95.

Caveats

  • Timing samples omit 0 failed or untimed matched explorer calls and 0 failed or untimed Claude head-to-head calls. A failure is not a fast successful step. No pass rate is estimated here.
  • The overhead extract uses nearest-rank quantiles. For LLM routers this study instead uses the conventional median and interpolated p95 from the routing extract. Per-iteration policy times and individual rule-record times are not retained, so those summaries cannot be independently rebuilt.
  • The rule-arm latency is captured before the decision record is written. It excludes that write, despite the source summary label.
  • The policy timing excludes database reads, policy merge and decision-record insertion. Its minimum was not retained. The rule-record timing also lacks a retained minimum.
  • The Claude head-to-head tasks hit a pass-rate ceiling. This study selects the lowest observed Claude median after the run; ranges overlap and this selection is not a speed ranking. The accounts could run concurrently on the shared Mac.
  • The 30-run cutoff is a display rule. It does not make a sample p95 a reliable production tail estimate. Some recorded calls exceed the p95; the table keeps their maxima visible.
  • The budgets are assumptions. A voice product may need less or more; this study does not say what listeners accept.
  • The steps differ in route and in what they time. Routers are whole decisions; models are the time to the first useful text, as the harness records it. First output is not necessarily enough text to start speech. A router through a CLI is a deployment choice, not a model limit.
  • All times were recorded on one Mac, not on a server near the vendor, and on different days. Vendor latency changes over a day, and a server near the API would see different times.
  • Small samples: 5 runs per model configuration, 5 per CLI start-up, 15 for the Claude Code configuration. Their slow end is the slowest run, a range and not a confidence interval.
  • The imported explorer evidence has no frozen pre-run protocol. Those receipts do not establish that its controls and call caps were set before the run.
  • The explorer uses a one-line answer; CLI start-up uses a one-word prompt. The selected Claude Code cell uses five short tasks, including code. Longer prompts and replies can change first-output time. No speech or audio was measured.
  • No speech recognition, speech synthesis, turn detection or network round trip to a user is timed here. API wall times include the client-to-provider network. The steps are only the decision and the first model output.
  • Calculations add medians and add slow ends across runs that were not made together. Medians do not add up exactly, and sums of p95 values do not guarantee a pair p95 or an upper bound. Repeated-step counts assume the same duration each time; they do not give a probability of fitting.
  • Jev was timed in one run of 35 seconds, from one client machine on one network path, with calls one at a time. A hosted API can be faster or slower at another hour or from another region. The Jev decision cases were revised against Jev’s own answers, so its accuracy has a home advantage; this study uses only its time, not its accuracy.

Sources

  • Voice-agent latency budget calculation

    Calculation ·

    Recorded decision and first-output times compared with assumed 300, 800 and 1,500 ms budgets. No complete voice turn ran.

  • Routing overhead runs: policy microbenchmark and CLI start-up

    Our recorded runs ·

    In-process timing of the deterministic routing policy (20,000 timed decisions), CLI start-up with a one-word prompt (5 runs per CLI), and decision counts read from recorded bench runs. LLM router timings are reused from the routing runs.

    Raw data: routing-overhead/results.json

  • Routing runs: Jev router vs LLM routing

    Our recorded runs ·

    Routing decisions recorded per case and arm.

    Raw data: routing/receipts.json

  • Provider explorer receipts: CLI vs API

    Our recorded runs ·

    230 imported receipts for short fixed tasks over Claude Code CLI, Codex CLI and the OpenAI API, with time to first useful output, total time, tokens and validation.

    Raw data: provider-explorer/runs.json

  • Provider head-to-head: Claude Code models vs Codex efforts

    Our recorded runs ·

    Five short tasks with deterministic validators, declared protocol, every attempt kept.

    Raw data: provider-h2h/receipts.json

  • Jev live run: 246 timed calls on the 82 routing decisions

    Our recorded runs ·

    Jev 1.13 called over HTTPS, 3 repeats of the same 82 typed decisions, one call at a time, from one Mac over a home network: client wall time, with the network inside it. The API reports no server time. Cost per 1,000 decisions is a calculation from the reported input tokens and the published price. The case sets were revised against Jev answers, so Jev has a home advantage.

    Raw data: jev-live/summary.json, jev-live/calls.json

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Voice agent latency budget: component calculations, not a measured turn”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/voice-agent-latency-budget.

More studies

All benchmarks
Live story
  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.