• Thought experiment
  • Latency
  • Voice Agents
  • Routing

A latency budget for voice agents: which LLM steps fit in one turn?

136.5 ms for Jev, 0.82 s for a small-model API, 2.79 s to 3.79 s for Codex CLI: which steps fit a voice agent latency budget? A thought experiment.

TL;DR

  • A thought experiment. We set measured step times against three turn budgets: 0.5 s, 1 s and 2 s. They are our assumptions, not data.
  • Decision step: rules took a median 1.42 µs. Jev 1.13 took 136.5 ms (p95 195.7 ms, n = 246) in a separate keyed run on 2026-10-06. Both medians and p95 values fit every assumed budget. Sonnet 5.5 through Claude Code took 2.60 s and fits none.
  • Model step: GPT-6 Luna through the OpenAI API gave first useful output in a median 0.82 s (0.51 s to 1.37 s, n = 5). The median fits 1 s. The slowest run does not.
  • Coding CLIs: no median fit 1 s. Codex CLI: 2.79 s to 3.79 s.
  • Sample pipeline (a calculation): Jev plus one Luna call: 0.96 s on medians, 1.57 s on p95 plus slowest run. It fits 1 s on medians only. Sonnet 5.5 plus Luna totals 3.42 s on medians.
Live story · 33 sClaude Code CLI vs Codex CLI vs the API: a latency race

Claude Code CLI vs Codex CLI vs the API: a latency race

For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.

Transcript
  1. Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.
  2. A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.
  3. It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
  4. A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.
  5. Open benchmarks: intervals, sources and every failure kept.

What is the experiment?

A voice turn runs from the end of the user's speech to the start of the spoken reply. We chose the three budgets. They are assumptions, and we cite no outside study. They cover the LLM steps only.

We judge each step alone:

  • yes: the median and the listed p95 or slowest run are under the budget.
  • median: only the median is.
  • no: the median is over.

These labels are calculations, not service guarantees. A p95 is not the worst case; some calls take longer.

A real turn shares one budget between all steps, so this is the easy test.

Which measured steps fit?

Step (route, n)MedianObserved rangep95 or slowest0.5 s1 s2 s
Rules (in process, 20,000)1.42 µsMaximum 2,538.21 µs; minimum not reportedp95 2.33 µsyesyesyes
Jev 1.13 (HTTPS, 246)136.5 ms100.9 ms to 297.3 msp95 195.7 msyesyesyes
GPT-6 Luna, first useful (API, 5)0.82 s0.51 s to 1.37 sSlowest 1.37 snomedianyes
GPT-6.1 Sol low, first useful (API, 5)0.87 s0.84 s to 1.74 sSlowest 1.74 snomedianyes
Fable 5.1, first useful (Claude Code, 15)1.20 s0.95 s to 7.90 sSlowest 7.90 snonomedian
Haiku 4.5, first model output (Claude Code, 5)1,461 ms1,206 ms to 2,308 msSlowest 2,308 msnonomedian
Sonnet 5.5 as router (Claude Code, 82)2,597 ms1,993 ms to 5,583 msp95 4,298 msnonono
Codex CLI, first useful (3 setups, 5 each)2.79 s to 3.79 sSee the three ranges belowSlowest across setups 4.30 snonono
Haiku 4.5 as router (Claude Code, 82)12,543 ms5,857 ms to 51,278 msp95 34,481 msnonono

The listed tail figure is p95 where n is 82 or more, and the slowest run elsewhere. Neither a range nor p95 is a confidence interval. The Jev row comes from a separate keyed run on 2026-10-06 over the network from one Mac. The Fable row comes from our five-task study.

Decision step: rules, a decision model or an LLM router?

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

Time per decision · log scale: each gridline is 10 times the one before

4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–20000 per row

Median; whiskers = median to 95th percentile

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

The 1.42 µs is the pure decision. With its database write, the recorded rule arm took a median 1 ms (p95 2 ms, maximum 3 ms, n = 419). Jev 1.13, a small routing model, took a median 136.5 ms (p95 195.7 ms) over 246 calls. The cold first call took 224.7 ms.

Jev answered 221 of 246 calls exactly (89.8%). These repeat the same 82 decisions, so they are not 246 independent samples. The dataset's 95% Wilson interval is 81.9% to 95.0%, using the rounded count of 74/82 on the case scale. Both charts now include the live Jev timings.

  • Wall time (CLI)
  • Wall time (direct API call)
  • Model time (API)
Claude Haiku 4.5
Claude Sonnet 5.5
Jev 1.13 (TypeSafe)

Time per decision · log scale: each gridline is 10 times the one before

3 rows, 3 series: Wall time (CLI), Wall time (direct API call), Model time (API). Wall time (CLI): slowest Claude Haiku 4.5 12.67 s (median to p95 12.67 s–34.41 s, n 82). Fastest Claude Sonnet 5.5 2.6 s (median to p95 2.6 s–4.3 s, n 82). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–246 per row

Median wall time, whisker to the 95th percentile

Whiskers run from p50 to p95. The Claude routers ran through the Claude Code CLI, so their wall time includes CLI start-up and the tool schema; one pass of 82 decisions each. Jev was called directly over HTTPS from one Mac on a home network: 246 calls in a 35-second window, client wall time with the network inside it. Its API reports no server time, so Jev has no model-time point. These are different routes: the chart shows what a caller waits per decision, not model compute time.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Sonnet 5.5 through Claude Code took a median 2,597 ms (p95 4,298 ms, n = 82 recorded calls). Haiku 4.5 took 12,543 ms. This chart reads 2.60 s and 12.67 s (the routing study's own summary). The table uses per-call medians from the overhead study; the summary medians differ slightly.

The CLI also split out its own time. For Sonnet, model time was a median 1.596 s (p95 2.583 s; range 1.060 s to 4.757 s, n = 82). CLI time was a median 973 ms (p95 1,277 ms; range 826 ms to 3,673 ms). Model time alone was over 1 s. We did not time a direct Claude API call (no key).

Model step: API or CLI?

  • Total time
  • First useful output
Entrance: medians race at 3× real timeMotion reduced: press Replay to animateThe slowest median is 4.2 s. The clock runs at the recorded speed.
OpenAI API · GPT-6 Luna · none
OpenAI API · GPT-6.1 Sol · low
OpenAI API · GPT-6.1 Sol · high
Codex CLI · GPT-6 Luna · none
Codex CLI · GPT-6.1 Sol · low
Codex CLI · GPT-6.1 Sol · high

6 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · high 4.2 s (range 3.8 s–4.7 s, n 5). Fastest OpenAI API · GPT-6 Luna · none 1 s (range 0.7 s–1.5 s, n 5). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · high 3.8 s (range 3.4 s–4.3 s, n 5). Fastest OpenAI API · GPT-6 Luna · none 0.8 s (range 0.5 s–1.4 s, n 5). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 5 per row

Matched cohort, fixed exact reply, 5 runs per configuration

Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.

Source: Provider explorer receipts: CLI vs API

Through the OpenAI API, GPT-6 Luna gave first useful output in a median 0.82 s (0.51 s to 1.37 s). It finished in 0.97 s (0.65 s to 1.50 s). GPT-6.1 Sol at low effort took 0.87 s (0.84 s to 1.74 s). At high effort it took 1.34 s (1.26 s to 2.12 s), which misses 1 s. Each setup has n = 5; these are ranges, not confidence intervals.

A decision must finish before the next step starts, so use finish time for a decision. For this calculation, first useful text ends the reply step. Spoken playback still needs speech-ready text and speech synthesis.

Through the Codex CLI, the same models took a median 2.79 s (Luna; range 2.46 s to 3.42 s), 3.75 s (Sol, low; 3.44 s to 4.10 s) and 3.79 s (Sol, high; 3.37 s to 4.30 s). Each setup has n = 5. For each model and effort, the API range and the CLI range do not overlap, so the API is ahead on first useful output for this task in this small sample.

Does a coding CLI fit a voice turn?

  • First output event
  • First model output
  • Total wall time
Entrance: medians race at 4.3× real timeMotion reduced: press Replay to animateThe slowest median is 6 s. The clock runs at the recorded speed.
Claude Code · Claude Haiku 4.5
Codex CLI (default model)

2 rows, 3 series: First output event, First model output, Total wall time. First output event: slowest Claude Code · Claude Haiku 4.5 563 ms (range 519 ms–726 ms, n 5). Fastest Codex CLI (default model) 489 ms (range 354 ms–1.3 s, n 5). All run ranges overlap. First model output: slowest Codex CLI (default model) 5.06 s (range 4.39 s–5.48 s, n 5). Fastest Claude Code · Claude Haiku 4.5 1.46 s (range 1.21 s–2.31 s, n 5). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 5 per row

Median of 5 runs; whiskers = fastest and slowest run

Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.

Source: Routing overhead runs: policy microbenchmark and CLI start-up

Claude Code with Haiku 4.5 got a one-word prompt with tools off (n = 5). Its first output event came at a median 563 ms (519 ms to 726 ms), but that is not when the model starts to answer. First model output came at a median 1,461 ms (1,206 ms to 2,308 ms).

In our five-task study, Fable 5.1 in Claude Code had the lowest observed median first useful output of 9 setups. Its median was 1.20 s (range 0.95 s to 7.90 s, n = 15). Ranges overlap across setups, so this does not establish a speed ranking.

Sample pipeline: a decision plus one model call

PipelineSum of mediansp95 plus slowest0.5 s1 s2 s
Rules + Luna0.82 s1.37 snomedianyes
Jev 1.13 + Luna0.96 s1.57 snomedianyes
Sonnet 5.5 router + Luna3.42 s5.67 snonono

Each sum is a calculation: decision time plus first useful output. The median of a sum is not exactly the sum of the medians. The sums use unrounded receipts where available. A p95 plus the slowest of 5 runs is a tail scenario, not a measured pipeline p95 or an upper bound. No run timed a whole pipeline. At 1 s, Jev plus Luna leaves about 41 ms (calculation) for the rest of the turn.

What we read from this

This is our reading of small samples, not a rule.

  1. Among these measured decision steps, rules and Jev fit the assumed budgets on median and p95. The LLM routers through a CLI miss even 2 s on their medians (Sonnet: 2,597 ms).
  2. A small model through its API can fit 1 s on the median. Luna's slowest run (1.37 s) did not. At 0.5 s no measured reply median fit: Luna's fastest run took 0.51 s.
  3. None of the measured coding CLI setups fits on its median at 1 s or under. At 2 s it fits some medians, never the slowest call.

Our advice, not a measured result: decide with rules first. Add a small decision model only where rules cannot decide. Call the model API directly and stream the reply. Keep coding CLIs out of the turn loop.

How we measured

  • No new model calls. Every figure comes from our benchmark dataset or its source receipts and run summaries. The sums and fit labels are our calculations.
  • Rules: one Apple M3 Ultra Mac, 20,000 decisions after 5,000 warm-up calls. Jev 1.13: 246 calls (3 reps of 82 typed decisions), one at a time, over HTTPS from a home network.
  • API and Codex rows: one fixed short reply, 5 runs per setup, on 2026-10-03. All 30 calls passed (30/30; 95% Wilson interval 88.6% to 100%). This short task hits a quality ceiling; it cannot rank the routes on quality.

Public receipts: Jev calls, routing overhead, API and Codex calls, and five-task calls.

Caveats

  • One host, one network. Network and vendor latency change over the day.
  • Small samples. The API and matched Codex rows have n = 5 per setup; Fable has n = 15 across five tasks. Five runs say little about the tail.
  • Different tasks. A fixed reply, a one-word prompt and five short tasks are not real voice conversations. First useful text is not necessarily speech-ready text.
  • Different routes and settings. Jev uses HTTPS; Claude routers use Claude Code. Sonnet runs at low effort; Haiku uses default thinking. These timings cannot isolate model compute time.
  • Different days and routes. The Jev run and the API runs took place on different days. We timed no direct Claude API call (no key).
  • Assumed budgets, text only. We did not time speech-to-text, speech synthesis or audio transport.
  • Home advantage. We revised the 82 decision cases against Jev's answers.

Time your own routes

Disclosure: I build Agent, the product behind these benchmarks.

Agent keeps a receipt for each task: model, route, tokens, time and cost. Try Agent and time your own routes.

The data behind this post

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.