Explainer · p95 latency
p95 latency, explained: median, tail and range for LLM calls
Definition
p95 latency marks the time at or below which about 95% of calls fall. About 1 call in 20 is slower. Small samples and tied times can change that share. It shows slow calls that the median (p50) hides. In 82 recorded routing calls, Claude Sonnet 5.5 through Claude Code had a median of 2.60 s and a p95 of 4.30 s (range 1.993 to 5.583 s).
Agent team · · 5 min read · Every number is from the public studies
Motion reduced: press Replay zoom to animate
1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.
about 96 thousand×Jev 1.13 (TypeSafe) takes about 96 thousand times as long as Deterministic routing policy (Agent, in process). Calculation: ratio of the two medians (137 ms ÷ 1.42 µs). Log scale: each gridline is 10 times the one before.
| Item | Decision time | Median to p95 | n |
|---|---|---|---|
| Deterministic routing policy (Agent, in process) | 1.42 µs | 1.42 µs–2.33 µs | 20000 |
| Jev 1.13 (TypeSafe) | 137 ms | 137 ms–196 ms | 246 |
| Claude Sonnet 5.5 (effort low, via Claude Code) | 2.6 s | 2.6 s–4.3 s | 82 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 12.54 s | 12.54 s–34.48 s | 82 |
4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n 82–20000 per row
Median; whiskers = median to 95th percentile
The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
Why the median hides the slow calls
The median marks the middle of the sample. It does not show how much slower the tail is.
In the routing overhead study, Claude Haiku 4.5 took a median 12.54 s and a p95 of 34.48 s per routing decision (n = 82; range 5.857 to 51.278 s). It used Claude Code with default thinking. Its p95 is 2.7 times its median (calculation: 34,481 ms ÷ 12,543 ms). For Sonnet 5.5 at low effort, the ratio is 1.7 (calculation: 4,298 ms ÷ 2,597 ms).
Tails also repeat. Across 48 recorded agent runs, the median was 49.5 model calls per run (range 13 to 73). For an assumed 50-call task, the chance of at least one slow call is about 92% (calculation: 1 − 0.95^50). This assumes independent calls, each with a 5% chance of exceeding the population p95. We did not test that assumption or measure this task-level rate.
How to read p50, p95, p99, IQR, CV and range
Sort call times to find a percentile.
| Measure | Tells you | Weak spot |
|---|---|---|
| p50 (median) | The typical call | Ignores slow calls |
| p95 | The time about 1 call in 20 exceeds | Rests on a few slow calls |
| p99 | The time about 1 call in 100 exceeds | Needs many calls |
| IQR (interquartile range) | The spread of the middle half | Ignores both tails |
| CV (coefficient of variation) | Standard deviation ÷ mean (calculation) | Sensitive to extreme calls |
| Min-max range | The shortest and longest time | Cannot narrow when you add calls to the same sample |
Sonnet 5.5 on a JSON-object prompt (10 calls) had a median 2.89 s, IQR 0.54 s, CV 0.28 (calculation) and range 2.68 to 5.30 s. Its rounded CV ties for the highest of 9 cells. The slowest call is 1.8 times the median (calculation).
Haiku 4.5 on the same prompt has a CV calculation that also rounds to 0.28 but an IQR of 2.21 s (n = 10; range 5.28 to 12.27 s). CVs that look equal can hide different shapes.
How many calls each measure needs
A tail estimate needs enough calls above it. Nearest rank picks one measured call; interpolation can change the value. For the same 82 Haiku calls (range 5.857–51.278 s), nearest-rank p50/p95 are 12.54/34.48 s. The routing study interpolates: 12.67/34.41 s (calculation checked against the per-call log).
- n = 82: p95 is the 78th sorted call, with 4 calls above it in this sample. The p99 is the slowest call.
- n = 10: p95 is the slowest call. We show the range.
- n = 20,000: 1,000 rank positions follow p95; 200 follow p99. Ties can reduce the number of slower calls.
Our rule of thumb is a judgment, not a result: keep at least 10 calls above the percentile you report. That means 200 calls for a p95 and 1,000 for a p99. With fewer, report the median and the range.
Our numbers, measured
The axis is logarithmic. The whisker runs from the p50 to the p95, not a confidence interval. The policy timer excludes database reads and the decision record write. Its large maximum shows that p95 still hides rare slow calls.
| Router | n | p50 | p95 | Observed min-max range | p95 ÷ p50 (calculation) |
|---|---|---|---|---|---|
| Deterministic policy, in process | 20,000 | 1.42 µs | 2.33 µs (p99 3.04 µs) | Min not reported; max 2,538.21 µs | 1.6 |
| Sonnet 5.5 (low effort), Claude Code | 82 | 2.60 s | 4.30 s | 1.993 to 5.583 s | 1.7 |
| Haiku 4.5 (default thinking), Claude Code | 82 | 12.54 s | 34.48 s | 5.857 to 51.278 s | 2.7 |
The observed ranges do not overlap. In this run the Sonnet setup was faster. The settings differ, so this does not rank the two models. See the policy against Sonnet and what a router costs.
These are 10 calls per cell from the caching and consistency study. Whiskers show the shortest and longest call, not a confidence interval. Sonnet 5.5 on the exact-number prompt had a median 6.89 s, range 5.81 to 7.81 s, IQR 0.95 s and CV 0.09 (calculation). Its slowest call took 1.1 times the median (calculation).
How to measure and report latency
- Define where the timer starts and stops. Sonnet routing includes a median 973 ms of CLI and harness time (n = 82; range 826 to 3,673 ms).
- Warm up first. The policy run discarded 5,000 calls, then timed 20,000.
- Count wrong calls and timeouts. Report the timeout limit; it is not an exact completion time. Haiku 4.5 passed 0 of 10 exact-number calls (95% Wilson interval 0% to 28%). The chart still counts all 10 measured times.
- With enough calls, report the p50 and p95. With too few, report the median and range.
- State n, the host or route and the date. Ours are single runs (2026-10-05 and 2026-10-06), and the policy ran on one Apple M3 Ultra Mac. We did not measure day-to-day drift.
Call one side ahead only when the ranges do not overlap. See what an LLM router is and the routing hub.
Frequently asked questions
What is p95 latency?
It marks the time at or below which about 95% of calls fall. About 1 call in 20 is slower; ties and small samples can change that share. The routing table above shows measured examples.
Why is p95 more useful than average latency?
An average blends fast and slow calls. A p95 shows a slow-tail threshold, but calls beyond it can take longer. Use both; these charts show medians, not means.
How many requests do I need to measure p99?
We did not test p99 stability. With nearest rank, our 82-call p99 is the slowest call (calculation). Our rule of thumb is at least 10 calls above the percentile: 1,000 calls for a p99.
Is a min-max range a confidence interval?
No. It shows observed limits; future calls can fall outside them. A p50-to-p95 whisker is not an interval either. We use 95% Wilson intervals for pass rates and report no interval for a median.
The public routing overhead extract records the timing summaries and per-run call counts. The routing extract and repeated-prompt receipts supply the other figures.
Watch the data
Routing overhead: a 1.42 µs policy vs LLM routers
A deterministic routing policy decides in 1.42 µs (median, n = 20,000); Sonnet 5.5 as a router takes 2.60 s through the CLI. Per-task costs and delays are calculations.
Transcript
- Routing overhead · policy vs LLM routers. What a routing decision costs before the work starts. Time and money per decision, then per task. Jev is timed over its API, a different route from the CLI routers.
- The latency ladder: the in-process policy decides in 1.42 µs. Jev, a direct API call, takes 137 ms. Sonnet 5.5 as a router takes 2.60 s through the CLI, Haiku 4.5 12.5 s. Chart: Time per decision · log scale · median to p95 (n = 82–20000 each). Caveat: Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
- The LLM router takes about 1.8 million times longer per decision. Jev costs $0.0337 per 1,000 decisions, provider-reported. Policy decision, median (20,000 timed): 1.42 µs (n = 20000, p95 2.33 µs). Sonnet 5.5 router ÷ policy, medians: 1.8 million× (n = 82). Jev 1.13 per 1,000 decisions (provider-reported): $0.0337 (n = 82, 137 ms median per call over its API). Caveat: OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
- Of Sonnet 5.5’s 2.60 s per decision, a median 973 ms is CLI and harness time, not the model. Chart: Where an LLM router’s time goes · median per call (n = 82 each). Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
- Route all 49.5 model calls per task: $1.67 per 1,000 tasks with Jev, $247.30 with Sonnet 5.5. Route 7 decisions: $34.97. Chart: Added routing cost per 1,000 tasks. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
- If each decision waits in line, Sonnet 5.5 adds up to 129 s per task, or 18.2 s for 7 decisions. The policy adds 70.3 µs. Chart: Added routing delay per task · upper bound. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
- Start-up tax for a one-word answer: Claude Code 2.53 s and 6,761 input tokens; Codex CLI 6.00 s and 17,051 input tokens, 13,184 of them read from the cache. Chart: CLI start-up tax · one-word answer · median of 5 runs (n = 5 each). Caveat: CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
- Route in process where a rule is enough. Every timing, cost and gap online.