Routing overhead: deterministic policy vs LLM routers vs Jev
What delay and what cost does each kind of router add before the real work of a call starts?
Published · 9 charts · Download the data or a carousel
1.42µs
The answer
The deterministic routing policy decided in a median 1.42 µs (p95 2.33 µs, 20,000 decisions, $0). The fastest LLM router, Claude Sonnet 5.5 through the Claude Code CLI, took a median 2.60 s per decision (p95 4.30 s, n = 82), about 1.8 million times longer; 973 ms of that median was CLI time, not model time. Haiku 4.5 with its default thinking took 12.54 s (p95 34.48 s). Jev 1.13, called directly over HTTPS from the same Mac, took a median 137 ms per decision (p95 196 ms, n = 246 calls, client wall time with the network inside it) and cost $0.0337 per 1,000 decisions (a calculation from its reported input tokens). That is a different route from the CLI routers, so it is not a model-against-model comparison. As a calculation over 48 recorded tasks (median 49.5 model calls each), routing every call would add $1.67 per 1,000 tasks and up to 7 s of waiting per task with Jev, $247.30 with Sonnet (8.2% of the work cost) and up to 129 s of waiting per task with Sonnet; routing only the 7 System One decisions cuts Sonnet to $34.97 and 18 s. CLI start-up alone, for a one-word answer: Claude Code (Haiku 4.5) took 2.53 s and sent 6,761 input tokens; Codex CLI took 6.00 s and sent 17,051 input tokens, 13,184 of them read from the cache (5 runs each, different models).
Live story
Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.
Routing overhead: a 1.42 µs policy vs LLM routers
A deterministic routing policy decides in 1.42 µs (median, n = 20,000); Sonnet 5.5 as a router takes 2.60 s through the CLI. Per-task costs and delays are calculations.
Transcript
- Routing overhead · policy vs LLM routers. What a routing decision costs before the work starts. Time and money per decision, then per task. Jev is timed over its API, a different route from the CLI routers.
- The latency ladder: the in-process policy decides in 1.42 µs. Jev, a direct API call, takes 137 ms. Sonnet 5.5 as a router takes 2.60 s through the CLI, Haiku 4.5 12.5 s. Chart: Time per decision · log scale · median to p95 (n = 82–20000 each). Caveat: Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
- The LLM router takes about 1.8 million times longer per decision. Jev costs $0.0337 per 1,000 decisions, provider-reported. Policy decision, median (20,000 timed): 1.42 µs (n = 20000, p95 2.33 µs). Sonnet 5.5 router ÷ policy, medians: 1.8 million× (n = 82). Jev 1.13 per 1,000 decisions (provider-reported): $0.0337 (n = 82, 137 ms median per call over its API). Caveat: OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
- Of Sonnet 5.5’s 2.60 s per decision, a median 973 ms is CLI and harness time, not the model. Chart: Where an LLM router’s time goes · median per call (n = 82 each). Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
- Route all 49.5 model calls per task: $1.67 per 1,000 tasks with Jev, $247.30 with Sonnet 5.5. Route 7 decisions: $34.97. Chart: Added routing cost per 1,000 tasks. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
- If each decision waits in line, Sonnet 5.5 adds up to 129 s per task, or 18.2 s for 7 decisions. The policy adds 70.3 µs. Chart: Added routing delay per task · upper bound. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
- Start-up tax for a one-word answer: Claude Code 2.53 s and 6,761 input tokens; Codex CLI 6.00 s and 17,051 input tokens, 13,184 of them read from the cache. Chart: CLI start-up tax · one-word answer · median of 5 runs (n = 5 each). Caveat: CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
- Route in process where a rule is enough. Every timing, cost and gap online.
Key numbers
about 1.8
Median LLM router call ÷ median policy decision
million times · n = 82
1ms
Recorded rule-based System One decision, record write included
median, 2 ms p95 (millisecond resolution) · n = 419
49.5
Model calls per task (each one a routing decision)
median (13 to 73) · n = 48
137ms
Jev 1.13 (TypeSafe): median decision time over the API
p95 196 ms · n = 246
about 19
Median Claude Sonnet 5.5 call through the CLI ÷ median Jev call over the API
times · n = 82
about 92
Median Claude Haiku 4.5 call through the CLI ÷ median Jev call over the API
times · n = 82
1,690ms
Claude Code time outside the model on a one-word answer
median (1,533 to 1,811) · n = 5
Routing hub: Jev, LLM routers, the deterministic policy and gateways on one page
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.
Time per decision · log scale: each gridline is 10 times the one before
| Item | Decision time | Median to p95 | n |
|---|---|---|---|
| Deterministic routing policy (Agent, in process) | 1.42 µs | 1.42 µs–2.33 µs | 20000 |
| Jev 1.13 (TypeSafe) | 137 ms | 137 ms–196 ms | 246 |
| Claude Sonnet 5.5 (effort low, via Claude Code) | 2.6 s | 2.6 s–4.3 s | 82 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 12.54 s | 12.54 s–34.48 s | 82 |
4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n 82–20000 per row
Median; whiskers = median to 95th percentile
The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
- Model API time
- CLI and harness time
| Item | Model API time | CLI and harness time | Median to p95 | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 (effort low, via Claude Code) | 1.6 s | 973 ms | Model API time: 1.6 s–2.58 s; CLI and harness time: 973 ms–1.28 s | 82 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 10.51 s | 1.7 s | Model API time: 10.51 s–32.13 s; CLI and harness time: 1.7 s–2.68 s | 82 |
2 rows, 2 series: Model API time, CLI and harness time. Model API time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 10.51 s (median to p95 10.51 s–32.13 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 1.6 s (median to p95 1.6 s–2.58 s, n 82). Not all run ranges overlap. CLI and harness time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 1.7 s (median to p95 1.7 s–2.68 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 973 ms (median to p95 973 ms–1.28 s, n 82). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n = 82 per row
Median per call; whiskers = median to 95th percentile
Model API time is the API duration the CLI reports; CLI and harness time is wall time minus that, per call. Medians of the parts do not add up to the median of the whole. Jev is not split: its API reports no server time. The whisker is the median to the 95th percentile, not a confidence interval.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing
Every interval overlaps every other: this chart does not order these rows.
| Item | Completed | 95% interval | n |
|---|---|---|---|
| Deterministic routing policy (Agent, in process) | 100% | 100%–100% | 20000 |
| Jev 1.13 (TypeSafe) | 100% | 98%–100% | 246 |
| Claude Sonnet 5.5 (effort low, via Claude Code) | 100% | 96%–100% | 82 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 100% | 96%–100% | 82 |
4 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln 82–20000 per row4 of 4 at 100%: this task set cannot separate them.
Completed calls ÷ calls; whiskers = 95% Wilson interval
A completed call returned a decision, right or wrong (accuracy is in the routing study). Whiskers are 95% Wilson intervals.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
| Item | Cost per 1,000 decisions | n |
|---|---|---|
| Deterministic routing policy (Agent, in process) | $0 | 20000 |
| Jev 1.13 (TypeSafe) | $0.034 | 82 |
2 rows. Highest Jev 1.13 (TypeSafe) $0.034 (n 82). Lowest Deterministic routing policy (Agent, in process) $0 (n 20000).
Notesn 82–20000 per row
USD per 1,000 decisions
The policy makes no model call, so it costs nothing per decision. Jev’s figure is the cost its provider reported for the recorded production run (82 decisions). The input tokens of the live run give the same figure at the published price. The Claude routers are in the next chart: their cost is derived from list prices.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev 1.13 list price
Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the lowest value): a ratio of list-price calculations, not a measurement.
| Item | Cost per 1,000 decisions (list price) | n |
|---|---|---|
| Jev 1.13 (TypeSafe) | $0.034 | 246 |
| Claude Sonnet 5.5 (effort low, via Claude Code) | $5.00 | 82 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | $8.92 | 82 |
List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 (thinking on, via Claude Code) $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).
Notesn 82–246 per row
List price × the tokens each route reported, USD per 1,000 decisions
A calculation, not a bill: the Claude calls ran on a subscription. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free). Same tokens and prices as the routing study.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions
- Every model call routed (49.5 per task)
- Only System One decisions (7 per task) (square)
Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).
| Item | Every model call routed (49.5 per task) | Only System One decisions (7 per task) |
|---|---|---|
| Deterministic routing policy (Agent, in process) | $0 | $0 |
| Jev 1.13 (TypeSafe) | $1.67 | $0.24 |
| Claude Sonnet 5.5 (effort low, via Claude Code) | $247 | $34.97 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | $442 | $62.47 |
List-price calculation, not a run. 4 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $442. Lowest Deterministic routing policy (Agent, in process) $0. Only System One decisions (7 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $62.47. Lowest Deterministic routing policy (Agent, in process) $0.
Notes
Decisions per task from recorded runs × cost per decision
A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions
- Every model call routed (49.5 per task)
- Only System One decisions (7 per task) (square)
Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).
| Item | Every model call routed (49.5 per task) | Only System One decisions (7 per task) |
|---|---|---|
| Deterministic routing policy (Agent, in process) | 70.3 µs | 9.9 µs |
| Jev 1.13 (TypeSafe) | 6.8 s | 1 s |
| Claude Sonnet 5.5 (effort low, via Claude Code) | 129 s | 18.2 s |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 621 s | 87.8 s |
List-price calculation, not a run. 4 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 621 s. Fastest Deterministic routing policy (Agent, in process) 70.3 µs. Only System One decisions (7 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 87.8 s. Fastest Deterministic routing policy (Agent, in process) 9.9 µs.
Notes
Decisions per task × median decision time, if every decision waits in line
A calculation and an upper bound: it assumes each decision waits for the one before. Median recorded task wall time: 10.3 min. Jev’s delay uses its live median over the API from one Mac (network included); the Claude routers’ includes the CLI.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions
- First output event
- First model output
- Total wall time
| Item | First output event | First model output | Total wall time | Range (lowest–highest run) | n |
|---|---|---|---|---|---|
| Claude Code · Claude Haiku 4.5 | 563 ms | 1.46 s | 2.53 s | First output event: 519 ms–726 ms; First model output: 1.21 s–2.31 s; Total wall time: 2.27 s–3.38 s | 5 |
| Codex CLI (default model) | 489 ms | 5.06 s | 6 s | First output event: 354 ms–1.3 s; First model output: 4.39 s–5.48 s; Total wall time: 5.37 s–6.51 s | 5 |
2 rows, 3 series: First output event, First model output, Total wall time. First output event: slowest Claude Code · Claude Haiku 4.5 563 ms (range 519 ms–726 ms, n 5). Fastest Codex CLI (default model) 489 ms (range 354 ms–1.3 s, n 5). All run ranges overlap. First model output: slowest Codex CLI (default model) 5.06 s (range 4.39 s–5.48 s, n 5). Fastest Claude Code · Claude Haiku 4.5 1.46 s (range 1.21 s–2.31 s, n 5). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 5 per row
Median of 5 runs; whiskers = fastest and slowest run
Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.
Source: Routing overhead runs: policy microbenchmark and CLI start-up
| Item | Input tokens per call | n |
|---|---|---|
| Claude Code · Claude Haiku 4.5 | 6,761 | 5 |
| Codex CLI (default model) | 17,051 | 5 |
2 rows. Highest Codex CLI (default model) 17,051 (n 5). Lowest Claude Code · Claude Haiku 4.5 6,761 (n 5).
Notesn = 5 per row
Per call, mostly the CLI’s own system prompt and tool definitions
Claude Code sums its disjoint input, cache-read and cache-write fields. Codex CLI reports 17,051 input tokens, 13,184 of them read from the cache. The prompt itself is a few tokens.
Source: Routing overhead runs: policy microbenchmark and CLI start-up
Tables
What was measured, recorded, calculated or not measured
| Router | Kind | Latency | Cost |
|---|---|---|---|
| Deterministic routing policy (Agent, in process) | in-process rules | measured 2026-10-06: 1.42 µs median | $0 (no model call) |
| Jev 1.13 (TypeSafe) | hosted decision model | measured 2026-10-06: 137 ms median, p95 196 ms (246 calls over the API from one Mac; client wall time, no server time) | $0.0337 per 1,000 (list-price calculation from input tokens; the recorded run’s own cost figure agrees) |
| Claude Sonnet 5.5 (effort low, via Claude Code) | LLM router via agent CLI | recorded: 2.60 s median (82 calls) | $4.996 per 1,000 (list-price calculation) |
| Claude Haiku 4.5 (thinking on, via Claude Code) | LLM router via agent CLI | recorded: 12.54 s median (82 calls) | $8.924 per 1,000 (list-price calculation) |
| OpenRouter Auto Router / cheaper hosted inference | hosted LLM router / gateway | Not measured: no key in the environment. The harness is ready and runs when a key is set. | Not measured: no key in the environment. The harness is ready and runs when a key is set. |
| Clef / Clef-Flash (local) | local router model | Not measured: no local server running. | Not measured: no local server running. |
Jev live run: time per call and cost
| Measure | Value |
|---|---|
| Counted calls | 246 (3 repeats of the same 82 decisions, one at a time), 0 failed |
| Median time per call | 136.5 ms |
| 90th percentile | 171.0 ms |
| 95th percentile | 195.7 ms |
| Fastest and slowest call | 100.9 ms and 297.3 ms |
| Cold first call (new process, fresh connection) | 224.7 ms |
| Median per repeat | 130.3 ms, 141.7 ms, 137.1 ms (the first repeat without the cold call) |
| Input and output tokens per decision (mean) | 803 and 148 (output tokens are free at the published price) |
| Cost per 1,000 decisions (calculation) | $0.0337 = 803 input tokens × 1,000 × $0.042 per million |
| Server-side time | not available: the API sends no timing header or field |
Method
- Protocol declared before any measurement.
- Deterministic policy: the platform’s production routing decision (plan, model ladder and effort) on its default policy, timed in process on one Apple M3 Ultra Mac with Node 25: 5,000 warm-up calls, then 20,000 timed decisions over 64 synthetic routing contexts, plus a 200,000-call batch for throughput. Database reads and the decision record write are out of scope; the recorded rule-based System One decisions (millisecond resolution, record write included) are shown as a stat.
- LLM routers: the recorded routing runs of the routing study (Claude Sonnet 5.5 (effort low, via Claude Code), 82 calls; Claude Haiku 4.5 (thinking on, via Claude Code), 82 calls), one call at a time. Wall time per call, the API time the CLI reported, and the difference (CLI and harness time). p50 and p95 recomputed from the per-call log.
- Jev: a live run on 2026-10-06. The same 82 typed decisions as the routing study, 3 repeats, 246 counted calls sent one at a time over HTTPS to the TypeSafe API from the same Mac, on a home network. Time per call is client wall time from before the request to after the body was read, so the network is inside it; the API reports no server time. Cost per decision is a calculation: the mean input tokens its API reported × the published price. The recorded production run (82 decisions) has the same cost as a provider-reported figure.
- CLI start-up: Claude Code · Claude Haiku 4.5 5 runs, Codex CLI (default model) 5 runs, a one-word prompt, one call at a time. Times to the first output event, the first model output and exit.
- Per 1,000 tasks (a calculation): decisions per task from 48 recorded bench runs read only (2 empty runs excluded) × cost and median time per decision. Decisions are assumed to wait in line, so the delay is an upper bound.
Caveats
- The policy is timed in process and the LLM routers through a CLI: this compares the two ways of routing as deployed, not two models on equal footing. A direct API call would skip the CLI time (shown separately).
- The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
- Haiku 4.5 ran with the CLI’s default thinking, which makes it slower than Sonnet 5.5 at effort low here.
- OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
- CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
- Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
- Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
- Jev’s timing is one 35-second window on 2026-10-06, 246 calls from one machine; server load at that time is unknown. The 246 calls are 3 repeats of 82 requests, so they are not independent draws.
- Corrected 2026-10-07: an earlier version of this page counted the cached tokens twice (30,235). The CLI reports 17,051 input tokens including 13,184 read from the cache.
Sources
Routing overhead runs: policy microbenchmark and CLI start-up
In-process timing of the deterministic routing policy (20,000 timed decisions), CLI start-up with a one-word prompt (5 runs per CLI), and decision counts read from recorded bench runs. LLM router timings are reused from the routing runs.
Routing overhead per 1,000 tasks (calculation)
Decisions per task from recorded bench runs multiplied by the cost and the median time per decision. Decisions are assumed to wait in line, so the delay is an upper bound. A calculation, not a run.
Routing runs: Jev router vs LLM routing
Routing decisions recorded per case and arm.
Anthropic list prices (Claude models)
Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.
Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free.
Jev live run: 246 timed calls on the 82 routing decisions
Jev 1.13 called over HTTPS, 3 repeats of the same 82 typed decisions, one call at a time, from one Mac over a home network: client wall time, with the network inside it. The API reports no server time. Cost per 1,000 decisions is a calculation from the reported input tokens and the published price. The case sets were revised against Jev answers, so Jev has a home advantage.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “Routing overhead: deterministic policy vs LLM routers vs Jev”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/routing-overhead.
Explainers that cite this study
Read the methods and terms in the context of these recorded results.
More write-ups that cite this study (2)
Models and comparisons in this study
Write-ups on this study
A latency budget for voice agents: which LLM steps fit in one turn?
136.5 ms for Jev, 0.82 s for a small-model API, 2.79 s to 3.79 s for Codex CLI: which steps fit a voice agent latency budget? A thought experiment.
A voice agent latency budget, with measured times: what fits in one turn?
Rules and Jev 1.13 fit every budget we assumed; a Claude router through a CLI fits none. 14 measured steps vs 300, 800 and 1,500 ms. A thought experiment.
AI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
Claude Haiku 4.5 vs Sonnet 5.5: all 80 comparison rows, and where the small model loses
Haiku 4.5 vs Sonnet 5.5 on 80 rows: Sonnet ahead on 14, Haiku on none, 31 ties. Hard tasks 11/24 vs 24/24, plus speed, memory and price.
Does LLM routing save money? The saving, the router and the net
Routing would save 2.8% ($3.01) on 2,362 recorded calls (a calculation). A Sonnet router on every call costs about $11.80, so the net is a loss.
How fast is Jev? 136.5 ms per routing decision, measured live
Jev router latency over direct HTTPS: median 136.5 ms, p95 195.7 ms, range 100.9–297.3 ms, n = 246. Claude routers used a different CLI route.
How long does an AI coding agent take per task? Minutes, calls and where the time goes
Agent took a median 9.6 minutes per SWE-bench attempt (n = 33, range 1.6 to 54.1). Compare single calls, repairs and full tasks with limits.
How to turn off extended thinking in Claude Code, and what we measured for Haiku 4.5
Set MAX_THINKING_TOKENS=0, then check both counters. Our Haiku 4.5 test covers routing and hard tasks, with timings, costs and uncertainty.
Most AI model comparisons are ties: 44 of 1,196 rows show a clear gap
44 of 1,196 AI model comparison rows show a gap under our overlap rules. Most gaps are timing rows. No effort row separates quality.
Plan for p95, not the median: LLM tail latency in our runs
4.30 s p95 against a 2.60 s median for Claude Sonnet 5.5; 34.5 s against 12.5 s for Haiku 4.5. Measured LLM tail latency and what to do about it.
The cheapest LLM for classification: 100,000 decisions a day, and why Haiku cost more than Sonnet
Cost calculations for 100,000 routing decisions a day: Jev $3.37, Sonnet $732.40, Haiku $892.40. Tested on 82 cases; not general classification.
The cheapest way to run an AI coding agent: 7 levers from measured runs
7 levers that may cut an AI coding agent's bill, sized from our data: prompt cache 3.9x, Fable/Sonnet cost per pass 6.5x, and 5 more. List-price calculations.
What does routing a million AI requests a day cost? Rules vs Jev vs Claude
What 10,000 to 10 million routing decisions a day cost with rules, Jev and Claude routers, how many run at once and how long they make requests wait.
Which Claude model is fastest? It depends on the task, and on thinking
1.94 s was the lowest median on short calls (Fable 5.1). 7.75 s on hard calls (Sonnet 5.5). Haiku 4.5 took 4.43 s and 39.01 s with default thinking.
Which Claude model should you use? A task-by-task guide from our measurements
Which Claude model fits your job? Sonnet passed 24/24 hard calls (95% interval 86%–100%) at $0.01435 per pass (calculation). The top models hit a ceiling.
Why is Claude Code slow? Where the seconds go in a coding CLI call, and what to change
A one-word Claude Code call took a median 2.5 s, with 1.7 s outside the model (n = 5). Where the rest goes: thinking, effort, route. Measured splits and fixes.
AI coding benchmarks roundup, October 2026: sixteen studies, every number in one place
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
Claude vs Codex on hard tasks: GPT-6.1 Sol joins the hard set
GPT-6.1 Sol in the Codex CLI passed 16/16 hard tasks at medium and high effort. Sonnet, Opus and Fable passed 24/24. What separates them: time and cost.
Jev vs Claude Haiku vs Claude Sonnet as a router: an honest comparison
Every row of our Jev, Haiku and Sonnet router comparisons: accuracy ties, Jev is far cheaper and quicker per call over its API, on a different route.
Jev vs Clef and five open decision models: 1,085 checkable decisions, tested
Seven decision models, 1,085 checkable decisions. The hosted model led; size bought less than you think.
What does a router cost you? Rules vs Jev vs an LLM router
A rule-based router decides in 1.42 µs for $0, Jev in 136.5 ms, Sonnet in 2.60 s. What each adds per 1,000 tasks, in delay and in dollars.
More studies
All benchmarksWhat does routing a million AI requests a day cost? A calculation from measured runs
A calculation from measured runs: what 10,000 to 10 million routing decisions a day cost with rules, Jev and Claude, with median/p95 time scenarios.
Voice agent latency budget: component calculations, not a measured turn
Calculation: component times against assumed 300 ms, 800 ms and 1,500 ms budgets. No voice turn or audio was measured.
Does thinking pay for Claude Haiku 4.5? Thinking on vs off
Claude Haiku 4.5 with extended thinking on and off: 82 routing decisions and 8 hard tasks. Accuracy with 95% intervals, time and cost.