Plan for p95, not the median: LLM tail latency in our runs
4.30 s p95 against a 2.60 s median for Claude Sonnet 5.5; 34.5 s against 12.5 s for Haiku 4.5. Measured LLM tail latency and what to do about it.
TL;DR
- The median hides slow calls. Claude Sonnet 5.5 through Claude Code took a median 2,597 ms per routing decision and 4,298 ms at p95 (n = 82). Haiku 4.5 took 12,543 ms and 34,481 ms (n = 82). Ranges appear below. The p95 is 1.7x and 2.7x the median (calculations).
- On hard tasks, the slowest call took 1.9x to 5.7x the median across 7 configurations (calculation).
- Tails add up. Across 48 recorded tasks, the median was 49.5 model calls (range 13 to 73). At about 50 calls, the chance of at least one call above a population p95 is about 92%. This calculation assumes independent calls with the same latency distribution; it is not a measured task rate.
- Routes differ by orders of magnitude. At p95, rules took 2.33 µs, Jev 1.13 over HTTPS took 195.7 ms (a separate keyed run on 2026-10-06) and Sonnet took 4,298 ms.
- What to do: set a timeout for each route and retry inside a deadline. Run independent calls in parallel. Keep latency-critical steps out of a CLI. Route with rules or a small model.
Why does the median mislead?
The median is the middle call. A person who waits feels the slow call, not the middle one.
The population p95 is the time at or below which 95% of calls finish. About one call in 20 takes longer. A sample p95 only estimates that threshold.
An agent task makes many calls: a median 49.5 in 48 recorded bench tasks (13 to 73). The calculated chance of at least one call above a population p95 is 49% for 13 calls, about 92% for 50 and 98% for 73. Calculation: 1 − 0.95^calls. It assumes independent calls with the same latency distribution. We did not test those assumptions; 49.5 is a task-count summary, not a possible call count.
In a queue, waits add up. A Sonnet router before every call projects 128.6 s per task (calculation: 49.5 × 2.597 s, all decisions sequential). This multiplies two typical values; it is not an upper bound or a measured task median. Routing was off in those tasks. Two empty runs were excluded from the 48-task summary.
What tails did we measure?
Routing decisions
Each row is one way to make a routing decision, so the rows compare deployments, not models. The chart includes Jev from the separate live HTTPS run. Its timings include the network; the Claude timings include the CLI.
| Route | n | Median (p50) | p95 | Fastest to slowest | p95 ÷ p50 (calculation) |
|---|---|---|---|---|---|
| Rules policy, in process | 20,000 | 1.42 µs | 2.33 µs | Minimum not stored; maximum 2,538.21 µs | 1.6x |
| Jev 1.13, direct HTTPS | 246 | 136.5 ms | 195.7 ms | 100.9 to 297.3 ms | 1.4x |
| Claude Sonnet 5.5 (low effort), Claude Code | 82 | 2,597 ms | 4,298 ms | 1,993 to 5,583 ms | 1.7x |
| Claude Haiku 4.5 (thinking on), Claude Code | 82 | 12,543 ms | 34,481 ms | 5,857 to 51,278 ms | 2.7x |
These are observed ranges, not 95% intervals. The rules summary stores no minimum; we show its recorded maximum. The Claude rows use the routing-overhead compiler’s nearest-rank p50 and p95, as shown in the dataset.
Haiku's median (12,543 ms) is above Sonnet's p95 (4,298 ms). In the Jev run, the slowest call took 297.3 ms, 2.2x its median (calculation).
Hard tasks
A p95 from 16 or 24 calls is too noisy. We show ranges.
| Configuration | n | Median | Fastest to slowest | Slowest ÷ median (calculation) |
|---|---|---|---|---|
| Claude Sonnet 5.5, Claude Code | 24 | 7.75 s | 2.26 to 34.79 s | 4.5x |
| Claude Opus 5.5, Claude Code | 24 | 9.18 s | 4.24 to 27.21 s | 3.0x |
| Claude Opus 5.5 (high), Claude Code | 24 | 11.03 s | 3.63 to 63 s | 5.7x |
| GPT-6.1 Sol (medium), Codex CLI | 16 | 13.11 s | 8.54 to 61.6 s | 4.7x |
| Claude Fable 5.1, Claude Code | 24 | 16.13 s | 4.46 to 90 s | 5.6x |
| GPT-6.1 Sol (high), Codex CLI | 16 | 18.12 s | 11.67 to 92.21 s | 5.1x |
| Claude Haiku 4.5, Claude Code | 24 | 39.01 s | 15.27 to 75.13 s | 1.9x |
Haiku passed 11 of 24 calls (46%; 95% Wilson interval 28% to 65%). Its times include wrong answers and format misses. Each of the other six configurations passed every call: 24/24 for each Claude configuration (86% to 100%) and 16/16 for each Codex configuration (81% to 100%). Those six hit this task set’s pass-rate ceiling. All bounds here are 95% Wilson intervals. All seven latency ranges overlap, so we rank none of them on speed.
Repeated prompts
We sent each of 3 prompts 10 times to 3 configurations. In all 9 cells, the slowest call took 1.1x to 1.8x the median (calculation). The Sonnet code-fix cell had a median 2.67 s (range 2.32 to 4.34 s, n = 10). Its interquartile range was 1.17 s; its sample coefficient of variation was 0.26 (calculations). These tails were tighter, but the prompts and call counts differ from the hard set.
Where do tails come from?
We did not test causes. We show only what the data separates.
Thinking tokens. Haiku reported a median 4,556 reasoning tokens per call on the hard set (n = 24; range 1,452 to 8,569). In the routing runs, its model API p50 was 10,508 ms (p95 32,132 ms; range 4,328 to 49,487 ms, n = 82). Sonnet at low effort had p50 1,596 ms (p95 2,583 ms; range 1,060 to 4,757 ms, n = 82). In the effort ladder, the slowest call was longer at high than at low effort in all three families (16 calls per cell, overlapping ranges):
| Family | Low: median; range | High: median; range | High maximum ÷ low maximum (calculation) |
|---|---|---|---|
| Claude Sonnet 5.5 | 5.82 s; 2.78 to 19.96 s | 8.81 s; 2.93 to 35.81 s | 1.8x |
| Claude Opus 5.5 | 7.50 s; 3.34 to 15.82 s | 10.11 s; 3.63 to 63.00 s | 4.0x |
| GPT-6.1 Sol | 13.62 s; 7.94 to 44.29 s | 18.12 s; 11.67 to 92.21 s | 2.1x |
Each cell has n = 16. These observed maxima do not establish a cause for slow calls.
CLI start-up. The CLI adds time outside the model; these runs do not establish a fixed start-up floor. On a one-word answer, Claude Code spent a median 1,690 ms outside the model (Haiku 4.5, 5 runs, 1,533 to 1,811 ms). For Sonnet routing, CLI and harness time had p50 973 ms and p95 1,277 ms (range 826 to 3,673 ms, n = 82). Its p95 ÷ p50 was 1.3x, against 1.6x for model time (calculations). Medians of these parts do not add to the whole-call median.
What should you do?
These are our recommendations, not tested fixes.
- Set a timeout for each route from its p95 and its slowest call. A 5 s timeout leaves Jev a wide margin (slowest 297.3 ms). It sits just above Sonnet's p95 (4,298 ms), so it may cut off some Sonnet calls. It would cut off over half of Haiku's calls (median 12,543 ms). The slowest hard-task call took 92.21 s.
- Retry only inside a deadline. Consider one retry after a timeout or a retryable error, if time budget remains. Passing p95 alone does not show that a call failed. We did not measure whether retries help.
- Run independent calls in parallel, each with a deadline. In a queue, delays add; in parallel, the task waits for the slowest call. For 100 independent calls with the same continuous latency distribution, the median wait is that distribution’s 99.3rd percentile. Calculation: 0.5^(1/100). Shared congestion and limits can break these assumptions.
- Keep latency-critical steps out of a CLI. Sonnet routing spent p50 973 ms outside the model, including CLI and harness time.
A separate small study repaired one scheduler with GPT-6.1 Sol at medium effort (n = 3 per route). API median: 17.3 s, range 16.28 to 18.61 s. Codex CLI median: 61.2 s, range 59.90 to 69.51 s. These are observed ranges, not intervals. Use an API where a person waits.
- Route with rules or a small decision model. A router sits on the path of every call, so its tail repeats within a task.
How we measured
- Routing runs: 82 recorded calls per router, one at a time through Claude Code. We recomputed p50 and p95 from the per-call log with nearest rank: sorted call ceil(p × n). For n = 82, p50 uses call 41; the usual median averages calls 41 and 42. We retain the dataset’s p50 convention here. Haiku’s usual wall-time median is 12,673.5 ms; its nearest-rank p50 is 12,543 ms. Jev uses linearly interpolated quantiles; hard tasks and repeated prompts use the usual median. CLI time is wall time minus model API time.
- Rules policy: 20,000 timed decisions in process after 5,000 warm-up calls, on one Apple M3 Ultra Mac.
- Jev: a separate live run with an API key on 2026-10-06. It made 246 calls (82 decisions, 3 reps) from one Mac in 35 seconds. We use client wall time.
- Hard tasks: 24 calls per Claude Code configuration, 16 per Codex CLI configuration, with a 300 s limit. Effort ladder: 16 per cell, same limit. Repeated prompts: 10 per cell. In those three sets, each call ran one at a time, in an empty folder, with tools off.
- Ranges are the fastest and slowest call, not intervals. Ratios and chances are calculations.
- Source data: routing receipts, Jev calls, hard-task receipts, effort receipts, and repeat receipts.
Caveats
- A p95 from 82 calls is near the slowest 4 calls. Nearest rank selects call 78, with 4 calls above it (calculation: ceil(0.95 × 82) = 78). Jev's 246 calls repeat the same 82 decisions 3 times. Treat each p95 as noisy.
- The routes differ: in process, direct HTTPS and a CLI. We compare time only. The Jev time includes a home network path. We did not test a server in the API’s region.
- Batches differ. The counted Claude and Codex hard-task batches ran at different hours on 2026-10-06. Effort-ladder reference cells came from those earlier batches. An earlier Codex batch had 30 blocked startup receipts; no inference ran. Those receipts are excluded from the latency table.
- Not measured: causes of any tail, retries, and Claude through its direct API.
What to read next
- The routing hub
- Routing overhead study
- What does a router cost you?
- Same prompt, ten answers
- Hard model head-to-head
See your own tails
Disclosure: I build Agent, the product behind these benchmarks.
Agent records each routing decision with its reason, its cost and its time. Try Agent to find your slow decisions.