{"i":26,"slug":"voice-agent-latency-budget","chart":{"id":"latency-budget-slow-steps-ranges","title":"Steps that take a second or more, against three assumed budgets (observed ranges)","subtitle":"Median per step; whiskers = fastest to slowest run; budget dots are assumptions","kind":"dot-range","unit":"ms","yLabel":"Time per step","whisker":"minmax","series":[{"name":"Time per step","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Code · Claude Fable 5.1 (first output, 5 short tasks)",1196,947,7902,15],["OpenAI API · GPT-6.1 Sol · high (first output, one-line answer)",1341,1259,2115,5],["Claude Code · Claude Haiku 4.5 (start-up, first model output)",1461,1206,2308,5],["Codex CLI · GPT-6 Luna · none (first output, one-line answer)",2787,2461,3418,5],["Codex CLI · GPT-6.1 Sol · low (first output, one-line answer)",3753,3435,4103,5],["Codex CLI · GPT-6.1 Sol · high (first output, one-line answer)",3786,3366,4296,5],["Codex CLI (start-up, first model output)",5059,4391,5478,5]]}},{"name":"Budget (assumption: a budget a voice team might set)","points":{"$k":["label","value"],"$r":[["Budget 300 ms",300],["Budget 800 ms",800],["Budget 1,500 ms",1500]]}}],"note":"These 9 steps have a median of 1,000 ms or more. Routers are whole calls through the Claude Code CLI; the rest are the time to the first output of a model through a CLI or the API. The dot is the median. The whisker runs from the fastest to the slowest observed run (fewer than 30 runs per step). The maximum is not a p95 estimate. A whisker is not a confidence interval. The budget dots are assumptions (assumption: a budget a voice team might set); they are not measured.","sourceIds":["agent-routing-overhead","agent-routing","agent-provider-explorer","agent-provider-h2h","agent-jev-live"]}}