Explainer · p95 latency

p95 latency, explained: median, tail and range for LLM calls

Definition

p95 latency marks the time at or below which about 95% of calls fall. About 1 call in 20 is slower. Small samples and tied times can change that share. It shows slow calls that the median (p50) hides. In 82 recorded routing calls, Claude Sonnet 5.5 through Claude Code had a median of 2.60 s and a p95 of 4.30 s (range 1.993 to 5.583 s).

Agent team · · 5 min read · Every number is from the public studies

Log scale · axis 1 µs to

Motion reduced: press Replay zoom to animate

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

about 96 thousand×Jev 1.13 (TypeSafe) takes about 96 thousand times as long as Deterministic routing policy (Agent, in process). Calculation: ratio of the two medians (137 ms ÷ 1.42 µs). Log scale: each gridline is 10 times the one before.

4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–20000 per row

Median; whiskers = median to 95th percentile

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Why the median hides the slow calls

The median marks the middle of the sample. It does not show how much slower the tail is.

In the routing overhead study, Claude Haiku 4.5 took a median 12.54 s and a p95 of 34.48 s per routing decision (n = 82; range 5.857 to 51.278 s). It used Claude Code with default thinking. Its p95 is 2.7 times its median (calculation: 34,481 ms ÷ 12,543 ms). For Sonnet 5.5 at low effort, the ratio is 1.7 (calculation: 4,298 ms ÷ 2,597 ms).

Tails also repeat. Across 48 recorded agent runs, the median was 49.5 model calls per run (range 13 to 73). For an assumed 50-call task, the chance of at least one slow call is about 92% (calculation: 1 − 0.95^50). This assumes independent calls, each with a 5% chance of exceeding the population p95. We did not test that assumption or measure this task-level rate.

How to read p50, p95, p99, IQR, CV and range

Sort call times to find a percentile.

MeasureTells youWeak spot
p50 (median)The typical callIgnores slow calls
p95The time about 1 call in 20 exceedsRests on a few slow calls
p99The time about 1 call in 100 exceedsNeeds many calls
IQR (interquartile range)The spread of the middle halfIgnores both tails
CV (coefficient of variation)Standard deviation ÷ mean (calculation)Sensitive to extreme calls
Min-max rangeThe shortest and longest timeCannot narrow when you add calls to the same sample

Sonnet 5.5 on a JSON-object prompt (10 calls) had a median 2.89 s, IQR 0.54 s, CV 0.28 (calculation) and range 2.68 to 5.30 s. Its rounded CV ties for the highest of 9 cells. The slowest call is 1.8 times the median (calculation).

Haiku 4.5 on the same prompt has a CV calculation that also rounds to 0.28 but an IQR of 2.21 s (n = 10; range 5.28 to 12.27 s). CVs that look equal can hide different shapes.

How many calls each measure needs

A tail estimate needs enough calls above it. Nearest rank picks one measured call; interpolation can change the value. For the same 82 Haiku calls (range 5.857–51.278 s), nearest-rank p50/p95 are 12.54/34.48 s. The routing study interpolates: 12.67/34.41 s (calculation checked against the per-call log).

  • n = 82: p95 is the 78th sorted call, with 4 calls above it in this sample. The p99 is the slowest call.
  • n = 10: p95 is the slowest call. We show the range.
  • n = 20,000: 1,000 rank positions follow p95; 200 follow p99. Ties can reduce the number of slower calls.

Our rule of thumb is a judgment, not a result: keep at least 10 calls above the percentile you report. That means 200 calls for a p95 and 1,000 for a p99. With fewer, report the median and the range.

Our numbers, measured

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

Time per decision · log scale: each gridline is 10 times the one before

4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–20000 per row

Median; whiskers = median to 95th percentile

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

The axis is logarithmic. The whisker runs from the p50 to the p95, not a confidence interval. The policy timer excludes database reads and the decision record write. Its large maximum shows that p95 still hides rare slow calls.

Routernp50p95Observed min-max rangep95 ÷ p50 (calculation)
Deterministic policy, in process20,0001.42 µs2.33 µs (p99 3.04 µs)Min not reported; max 2,538.21 µs1.6
Sonnet 5.5 (low effort), Claude Code822.60 s4.30 s1.993 to 5.583 s1.7
Haiku 4.5 (default thinking), Claude Code8212.54 s34.48 s5.857 to 51.278 s2.7

The observed ranges do not overlap. In this run the Sonnet setup was faster. The settings differ, so this does not rank the two models. See the policy against Sonnet and what a router costs.

  • Exact number
  • JSON object
  • Code fix
Entrance: medians race at 9.6× real timeMotion reduced: press Replay to animateThe slowest median is 13.4 s. The clock runs at the recorded speed.
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: slowest GPT-6.1 Sol (medium) · Codex CLI 13.4 s (range 12.3 s–18 s, n 10). Fastest Claude Haiku 4.5 · Claude Code 5.1 s (range 4.4 s–6.2 s, n 10). Not all run ranges overlap. JSON object: slowest Claude Haiku 4.5 · Claude Code 7 s (range 5.3 s–12.3 s, n 10). Fastest Claude Sonnet 5.5 · Claude Code 2.9 s (range 2.7 s–5.3 s, n 10). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 10 per row

Median; whiskers = fastest and slowest of 10 calls

Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

These are 10 calls per cell from the caching and consistency study. Whiskers show the shortest and longest call, not a confidence interval. Sonnet 5.5 on the exact-number prompt had a median 6.89 s, range 5.81 to 7.81 s, IQR 0.95 s and CV 0.09 (calculation). Its slowest call took 1.1 times the median (calculation).

How to measure and report latency

  1. Define where the timer starts and stops. Sonnet routing includes a median 973 ms of CLI and harness time (n = 82; range 826 to 3,673 ms).
  2. Warm up first. The policy run discarded 5,000 calls, then timed 20,000.
  3. Count wrong calls and timeouts. Report the timeout limit; it is not an exact completion time. Haiku 4.5 passed 0 of 10 exact-number calls (95% Wilson interval 0% to 28%). The chart still counts all 10 measured times.
  4. With enough calls, report the p50 and p95. With too few, report the median and range.
  5. State n, the host or route and the date. Ours are single runs (2026-10-05 and 2026-10-06), and the policy ran on one Apple M3 Ultra Mac. We did not measure day-to-day drift.

Call one side ahead only when the ranges do not overlap. See what an LLM router is and the routing hub.

Frequently asked questions

What is p95 latency?

It marks the time at or below which about 95% of calls fall. About 1 call in 20 is slower; ties and small samples can change that share. The routing table above shows measured examples.

Why is p95 more useful than average latency?

An average blends fast and slow calls. A p95 shows a slow-tail threshold, but calls beyond it can take longer. Use both; these charts show medians, not means.

How many requests do I need to measure p99?

We did not test p99 stability. With nearest rank, our 82-call p99 is the slowest call (calculation). Our rule of thumb is at least 10 calls above the percentile: 1,000 calls for a p99.

Is a min-max range a confidence interval?

No. It shows observed limits; future calls can fall outside them. A p50-to-p95 whisker is not an interval either. We use 95% Wilson intervals for pass rates and report no interval for a median.

The public routing overhead extract records the timing summaries and per-run call counts. The routing extract and repeated-prompt receipts supply the other figures.

Watch the data

Live story · 51 sRouting overhead: a 1.42 µs policy vs LLM routers

Routing overhead: a 1.42 µs policy vs LLM routers

A deterministic routing policy decides in 1.42 µs (median, n = 20,000); Sonnet 5.5 as a router takes 2.60 s through the CLI. Per-task costs and delays are calculations.

Transcript
  1. Routing overhead · policy vs LLM routers. What a routing decision costs before the work starts. Time and money per decision, then per task. Jev is timed over its API, a different route from the CLI routers.
  2. The latency ladder: the in-process policy decides in 1.42 µs. Jev, a direct API call, takes 137 ms. Sonnet 5.5 as a router takes 2.60 s through the CLI, Haiku 4.5 12.5 s. Chart: Time per decision · log scale · median to p95 (n = 82–20000 each). Caveat: Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
  3. The LLM router takes about 1.8 million times longer per decision. Jev costs $0.0337 per 1,000 decisions, provider-reported. Policy decision, median (20,000 timed): 1.42 µs (n = 20000, p95 2.33 µs). Sonnet 5.5 router ÷ policy, medians: 1.8 million× (n = 82). Jev 1.13 per 1,000 decisions (provider-reported): $0.0337 (n = 82, 137 ms median per call over its API). Caveat: OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
  4. Of Sonnet 5.5’s 2.60 s per decision, a median 973 ms is CLI and harness time, not the model. Chart: Where an LLM router’s time goes · median per call (n = 82 each). Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
  5. Route all 49.5 model calls per task: $1.67 per 1,000 tasks with Jev, $247.30 with Sonnet 5.5. Route 7 decisions: $34.97. Chart: Added routing cost per 1,000 tasks. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
  6. If each decision waits in line, Sonnet 5.5 adds up to 129 s per task, or 18.2 s for 7 decisions. The policy adds 70.3 µs. Chart: Added routing delay per task · upper bound. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
  7. Start-up tax for a one-word answer: Claude Code 2.53 s and 6,761 input tokens; Codex CLI 6.00 s and 17,051 input tokens, 13,184 of them read from the cache. Chart: CLI start-up tax · one-word answer · median of 5 runs (n = 5 each). Caveat: CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
  8. Route in process where a rule is enough. Every timing, cost and gap online.

The data behind this explainer

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.