{"method":["Thought experiment on recorded times. This study makes no new model calls. Every time comes from a run that another study published.","Budgets: 300 ms, 800 ms, 1,500 ms per decision step (assumption: a budget a voice team might set). Nothing in the data says what a product needs; change the budget and the table gives the share and the count for it.","Deterministic routing policy: the policy timing of the routing overhead run (latencyUs p50 and p95, converted to milliseconds). Rule-based System One decision: the recorded decision times of that run (millisecond resolution, record write excluded).","Jev 1.13: this study checks the per-call rows against the live-run summary. The rows record client wall time over HTTPS. They give n, median and p95 (latencyMs calls). The first call uses a fresh connection (coldFirstCall). This calculation subtracts the later-call median (callsWithoutColdFirst) from the first-call time. The run starts at run.startedAt and ends at run.endedAt. The client uses one Apple Silicon Mac on a home network.","Claude Sonnet 5.5 and Claude Haiku 4.5 as routers: median and interpolated p95 from the routing extract (latencyMs wallP50 and wallP95). The overhead extract supplies n and the observed limits, through the Claude Code CLI, one call at a time. CLI start-up: first model output on a one-word prompt, from the routing overhead run (firstModelEventMs; 5 runs per CLI).","Model first output: this study uses the matched provider explorer cohort for the fixed exact reply. Each configuration has 5 runs with firstUsefulMs timings. Routes are the OpenAI API and the Codex CLI. This study also selects the Claude Code configuration with the lowest observed first-output median in the 5-task head-to-head. Its timing sample includes completed validator failures. Failed or untimed calls have no successful first-output time.","Slow end: the 95th percentile where a step has 30 or more runs, and the slowest run where it has fewer. The slowest run is stricter than a 95th percentile.","Calculations: share of a budget = time ÷ budget, at the median and at the slow end. A step fits when its time is at or below the budget. Runs that fit back to back = the budget ÷ the time, rounded down, counted up to 1,000. A router in front of a model adds the two medians and the two slow ends, as a scenario. Neither sum is a measured percentile or a guaranteed upper bound for the pair.","The charts separate steps with a median under 1,000 ms from the rest so each axis stays readable. Separate charts show median-to-p95 bands and observed min-to-max ranges. A range maximum is not a p95."]}