{"i":26,"study":{"slug":"voice-agent-latency-budget","title":"Voice agent latency budget: component calculations, not a measured turn","seoTitle":"Voice agent latency budget: calculated component times","description":"Calculation: component times against assumed 300 ms, 800 ms and 1,500 ms budgets. No voice turn or audio was measured.","question":"Which decision steps fit inside a voice agent’s turn budget, using only measured times?","answer":"Rules and Jev 1.13 fit all three assumed budgets at the median and p95. LLM routers through the Claude Code CLI fit none at the median. Calculation: the budgets are 300 ms, 800 ms and 1,500 ms per step. At the slow end, 3, 3 and 4 of 14 measured steps fit these budgets, respectively. At the median, 3, 3 and 8 fit. The deterministic policy took a median 1.42 µs (p95 2.33 µs, n = 20,000). Its maximum was 2.5 ms; its minimum was not retained. Jev 1.13 took a median 136.5 ms over HTTPS (p95 195.7 ms, n = 246). Its range was 100.9 ms to 297.3 ms. Calculation: this uses 46% of a 300 ms budget at the median and 65% at p95. Sonnet 5.5 routing through Claude Code took a median 2,598 ms (p95 4,298 ms, n = 82). Its range was 1,993 ms to 5,583 ms. Its median exceeds a 1,500 ms budget. GPT-6 Luna had the lowest observed API median for first output: 823 ms (n = 5). Its range was 506 ms to 1,371 ms. Calculation: its median is over an 800 ms budget. Its slowest run is over that budget. API ranges overlap, so this is not a speed ranking. These ranges are not confidence intervals. All times come from other studies on one Mac, over different routes. The source runs did not measure a voice turn or audio.","date":"2026-10-06","updated":"2026-10-06","tags":["voice-agent","latency","latency-budget","routing","jev","thought-experiment","calculation"],"caveats":["Timing samples omit 0 failed or untimed matched explorer calls and 0 failed or untimed Claude head-to-head calls. A failure is not a fast successful step. No pass rate is estimated here.","The overhead extract uses nearest-rank quantiles. For LLM routers this study instead uses the conventional median and interpolated p95 from the routing extract. Per-iteration policy times and individual rule-record times are not retained, so those summaries cannot be independently rebuilt.","The rule-arm latency is captured before the decision record is written. It excludes that write, despite the source summary label.","The policy timing excludes database reads, policy merge and decision-record insertion. Its minimum was not retained. The rule-record timing also lacks a retained minimum.","The Claude head-to-head tasks hit a pass-rate ceiling. This study selects the lowest observed Claude median after the run; ranges overlap and this selection is not a speed ranking. The accounts could run concurrently on the shared Mac.","The 30-run cutoff is a display rule. It does not make a sample p95 a reliable production tail estimate. Some recorded calls exceed the p95; the table keeps their maxima visible.","The budgets are assumptions. A voice product may need less or more; this study does not say what listeners accept.","The steps differ in route and in what they time. Routers are whole decisions; models are the time to the first useful text, as the harness records it. First output is not necessarily enough text to start speech. A router through a CLI is a deployment choice, not a model limit.","All times were recorded on one Mac, not on a server near the vendor, and on different days. Vendor latency changes over a day, and a server near the API would see different times.","Small samples: 5 runs per model configuration, 5 per CLI start-up, 15 for the Claude Code configuration. Their slow end is the slowest run, a range and not a confidence interval.","The imported explorer evidence has no frozen pre-run protocol. Those receipts do not establish that its controls and call caps were set before the run.","The explorer uses a one-line answer; CLI start-up uses a one-word prompt. The selected Claude Code cell uses five short tasks, including code. Longer prompts and replies can change first-output time. No speech or audio was measured.","No speech recognition, speech synthesis, turn detection or network round trip to a user is timed here. API wall times include the client-to-provider network. The steps are only the decision and the first model output.","Calculations add medians and add slow ends across runs that were not made together. Medians do not add up exactly, and sums of p95 values do not guarantee a pair p95 or an upper bound. Repeated-step counts assume the same duration each time; they do not give a probability of fitting.","Jev was timed in one run of 35 seconds, from one client machine on one network path, with calls one at a time. A hosted API can be faster or slower at another hour or from another region. The Jev decision cases were revised against Jev’s own answers, so its accuracy has a home advantage; this study uses only its time, not its accuracy."],"sourceIds":["calc-latency-budget","agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"],"stats":{"$k":["id","label","value","unit","display","n","note"],"$r":[["latency-budget-steps-counted","Decision and first-output steps compared",14,"count","14 steps",14,"5 under 1,000 ms at the median, 9 at 1,000 ms or more. Every time is a measured median from another study."],["latency-budget-fit-300-slow","Steps that fit a 300 ms budget at the slow end (calculation)",3,"count","3 of 14 (median: 3 of 14)",14,"Budget (assumption: a budget a voice team might set). Slow end: the 95th percentile for steps with 30 or more runs, the slowest run for the rest."],["latency-budget-fit-800-slow","Steps that fit an 800 ms budget at the slow end (calculation)",3,"count","3 of 14 (median: 3 of 14)",14,"Budget (assumption: a budget a voice team might set). Slow end: the 95th percentile for steps with 30 or more runs, the slowest run for the rest."],["latency-budget-fit-1500-slow","Steps that fit a 1,500 ms budget at the slow end (calculation)",4,"count","4 of 14 (median: 8 of 14)",14,"Budget (assumption: a budget a voice team might set). Slow end: the 95th percentile for steps with 30 or more runs, the slowest run for the rest."],["latency-budget-jev-share-300","Share of a 300 ms budget that one Jev 1.13 decision uses at the median (calculation)",45.5,"percent","46% at the median, 65% at p95",246,"Median 136.5 ms, p95 195.7 ms over HTTPS from one Mac on a home network. Budget (assumption: a budget a voice team might set)."],["latency-budget-jev-in-sequence-800","Back-to-back Jev 1.13 decisions that fit an 800 ms budget at the slow end (calculation)",4,"count","4 (5 at the median)",246,"Budget (assumption: a budget a voice team might set)."],["latency-budget-sonnet-router-share-1500","Share of a 1,500 ms budget that one Sonnet 5.5 router decision uses at the median (calculation)",173,"percent","173% at the median, 287% at p95",82,"Median 2,598 ms, p95 4,298 ms through the Claude Code CLI. Budget (assumption: a budget a voice team might set)."],["latency-budget-fastest-model-first-output","Lowest observed API median time to the first output of a one-line answer",823,"ms","823 ms (slowest run 1,371 ms)",5,"GPT-6 Luna through the OpenAI API. Range 506 ms to 1,371 ms, n = 5; not a confidence interval. API ranges overlap, so this is not a ranking."],["latency-budget-codex-cli-added","Difference in median first-output time across 3 Codex model setups: Codex CLI minus API (calculation)",1964,"ms","+1,964 to +2,880 ms",5,"Median through the Codex CLI minus median through the OpenAI API, for the same model, effort and one-line prompt (GPT-6 Luna: +1,964 ms; GPT-6.1 Sol (effort low): +2,880 ms; GPT-6.1 Sol (effort high): +2,445 ms). For each setup the slowest API run is below the CLI median. A calculation across runs of 5 per setup; the samples are small."],["latency-budget-jev-model-left-1500","Time left of a 1,500 ms budget after Jev 1.13 and the API cell with the lowest observed median, at the median (calculation)",540.5,"ms","540.5 ms left (sum of medians 959.5 ms)",5,"Separate samples: Jev n = 246; API n = 5. Jev median 136.5 ms plus GPT-6 Luna median 823 ms through the OpenAI API. At the slow ends the sum is 1,566.7 ms, 66.7 ms over. This arithmetic remainder is not measured audio latency. API timings already include the client-to-provider network. Budget (assumption: a budget a voice team might set). A calculation across runs that were not made together."],["latency-budget-jev-cold-call","First Jev call minus the later-call median (calculation)",88.3,"ms","88.3 ms once",1,"The first call used a fresh connection and took 224.7 ms; the later-call median was 136.4 ms (n = 245). One first call, so no range. This difference does not isolate connection cost or server warmth."]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","whisker","series","note","sourceIds"],"$r":[["latency-budget-fast-steps","Steps that take under a second, against three assumed budgets (median to p95)","Median per step; whiskers = median to p95; budget dots are assumptions","dot-range","ms","Time per step","p50-p95",[{"name":"Time per step","points":{"$k":["label","value","lo","hi","n"],"$r":[["Deterministic routing policy (Agent, in process)",0.00142,0.00142,0.00233,20000],["Rule-based System One decision (record write excluded)",1,1,2,419],["Jev 1.13 (TypeSafe, direct HTTPS)",136.5,136.5,195.7,246]]}},{"name":"Budget (assumption: a budget a voice team might set)","points":{"$k":["label","value"],"$r":[["Budget 300 ms",300],["Budget 800 ms",800],["Budget 1,500 ms",1500]]}}],"These 5 steps have a median under 1,000 ms. Decisions are whole calls; model steps are the time to the first useful output of a one-line answer. The dot is the median. The whisker runs from the median to the 95th percentile (30 or more runs per step). Observed limits stay in the measured-step table; some minima were not retained. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.",["agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"]],["latency-budget-fast-steps-ranges","Steps that take under a second, against three assumed budgets (observed ranges)","Median per step; whiskers = fastest to slowest run; budget dots are assumptions","dot-range","ms","Time per step","minmax",[{"name":"Time per step","points":[{"label":"OpenAI API · GPT-6 Luna · none (first output, one-line answer)","value":823,"lo":506,"hi":1371,"n":5},{"label":"OpenAI API · GPT-6.1 Sol · low (first output, one-line answer)","value":873,"lo":835,"hi":1742,"n":5}]},{"name":"Budget (assumption: a budget a voice team might set)","points":{"$k":["label","value"],"$r":[["Budget 300 ms",300],["Budget 800 ms",800],["Budget 1,500 ms",1500]]}}],"These 5 steps have a median under 1,000 ms. Decisions are whole calls; model steps are the time to the first useful output of a one-line answer. The dot is the median. The whisker runs from the fastest to the slowest observed run (fewer than 30 runs per step). The maximum is not a p95 estimate. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.",["agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"]],["latency-budget-slow-steps","Steps that take a second or more, against three assumed budgets (median to p95)","Median per step; whiskers = median to p95; budget dots are assumptions","dot-range","ms","Time per step","p50-p95",[{"name":"Time per step","points":[{"label":"Claude Sonnet 5.5 (router, effort low, via Claude Code)","value":2598,"lo":2598,"hi":4298,"n":82},{"label":"Claude Haiku 4.5 (router, thinking on, via Claude Code)","value":12674,"lo":12674,"hi":34413,"n":82}]},{"name":"Budget (assumption: a budget a voice team might set)","points":{"$k":["label","value"],"$r":[["Budget 300 ms",300],["Budget 800 ms",800],["Budget 1,500 ms",1500]]}}],"These 9 steps have a median of 1,000 ms or more. Routers are whole calls through the Claude Code CLI; the rest are the time to the first output of a model through a CLI or the API. The dot is the median. The whisker runs from the median to the 95th percentile (30 or more runs per step). Observed limits stay in the measured-step table; some minima were not retained. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.",["agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"]],["latency-budget-slow-steps-ranges","Steps that take a second or more, against three assumed budgets (observed ranges)","Median per step; whiskers = fastest to slowest run; budget dots are assumptions","dot-range","ms","Time per step","minmax",[{"name":"Time per step","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Code · Claude Fable 5.1 (first output, 5 short tasks)",1196,947,7902,15],["OpenAI API · GPT-6.1 Sol · high (first output, one-line answer)",1341,1259,2115,5],["Claude Code · Claude Haiku 4.5 (start-up, first model output)",1461,1206,2308,5],["Codex CLI · GPT-6 Luna · none (first output, one-line answer)",2787,2461,3418,5],["Codex CLI · GPT-6.1 Sol · low (first output, one-line answer)",3753,3435,4103,5],["Codex CLI · GPT-6.1 Sol · high (first output, one-line answer)",3786,3366,4296,5],["Codex CLI (start-up, first model output)",5059,4391,5478,5]]}},{"name":"Budget (assumption: a budget a voice team might set)","points":{"$k":["label","value"],"$r":[["Budget 300 ms",300],["Budget 800 ms",800],["Budget 1,500 ms",1500]]}}],"These 9 steps have a median of 1,000 ms or more. Routers are whole calls through the Claude Code CLI; the rest are the time to the first output of a model through a CLI or the API. The dot is the median. The whisker runs from the fastest to the slowest observed run (fewer than 30 runs per step). The maximum is not a p95 estimate. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.",["agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"]],["latency-budget-fit","How many of the steps fit each assumed budget (calculation)","Count of 14 measured steps; n = steps","grouped-bar","count","Steps that fit (of 14)","\u0001",[{"name":"Fits at the median","points":{"$k":["label","value","n"],"$r":[["300 ms budget",3,14],["800 ms budget",3,14],["1,500 ms budget",8,14]]}},{"name":"Fits at the slow end","points":{"$k":["label","value","n"],"$r":[["300 ms budget",3,14],["800 ms budget",3,14],["1,500 ms budget",4,14]]}}],"A calculation on measured times, with budgets of 300 ms, 800 ms and 1,500 ms (assumption: a budget a voice team might set). A step fits when its median, or its slow end, is at or below the budget. The slow end is the 95th percentile for steps with 30 or more runs and the slowest run for the rest, which is stricter. The steps differ in route and in what they time, so read the count as a summary of the table below, not as a ranking.",["calc-latency-budget","agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"]]]},"related":["routing-overhead","cli-model-latency-tokens"]}}