What does a router cost you? Rules vs Jev vs an LLM router
A rule-based router decides in 1.42 µs for $0, Jev in 136.5 ms, Sonnet in 2.60 s. What each adds per 1,000 tasks, in delay and in dollars.
TL;DR
- A router makes a decision before the real work starts: which model, which effort, how much context. That decision has a delay and a cost.
- Deterministic rules: a median 1.42 µs per decision (95th percentile 2.33 µs, n = 20,000) and $0.
- Claude Sonnet 5.5 as a router, through the Claude Code CLI: a median 2.60 s per decision (p95 4.30 s, n = 82). That is about 1.8 million times longer. 973 ms of the median is CLI time, not model time.
- Claude Haiku 4.5 with its default thinking: a median 12.54 s (p95 34.48 s, n = 82).
- Jev 1.13, called directly over HTTPS: a median 136.5 ms per call (p95 195.7 ms, n = 246 calls, network included) and $0.0337 per 1,000 decisions (a calculation from its reported input tokens). That is a different route from the CLI routers, so it is not a model-against-model comparison. The median is about 19 times shorter than Sonnet's and 92 times shorter than Haiku's (a calculation across the two routes).
- Per 1,000 tasks, routing every model call (49.5 per task) adds $0 with rules, $1.67 with Jev, $247.30 with Sonnet and $441.74 with Haiku. That is a calculation. If every decision waits in line, Jev adds up to 6.76 s per task and Sonnet up to 128.6 s. Routing only the 7 System One decisions per task cuts Sonnet to $34.97 and 18.2 s, and Jev to $0.24 and 0.96 s.
The full study: /benchmarks/routing-overhead.
Routing overhead: a 1.42 µs policy vs LLM routers
A deterministic routing policy decides in 1.42 µs (median, n = 20,000); Sonnet 5.5 as a router takes 2.60 s through the CLI. Per-task costs and delays are calculations.
Transcript
- Routing overhead · policy vs LLM routers. What a routing decision costs before the work starts. Time and money per decision, then per task. Jev is timed over its API, a different route from the CLI routers.
- The latency ladder: the in-process policy decides in 1.42 µs. Jev, a direct API call, takes 137 ms. Sonnet 5.5 as a router takes 2.60 s through the CLI, Haiku 4.5 12.5 s. Chart: Time per decision · log scale · median to p95 (n = 82–20000 each). Caveat: Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
- The LLM router takes about 1.8 million times longer per decision. Jev costs $0.0337 per 1,000 decisions, provider-reported. Policy decision, median (20,000 timed): 1.42 µs (n = 20000, p95 2.33 µs). Sonnet 5.5 router ÷ policy, medians: 1.8 million× (n = 82). Jev 1.13 per 1,000 decisions (provider-reported): $0.0337 (n = 82, 137 ms median per call over its API). Caveat: OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
- Of Sonnet 5.5’s 2.60 s per decision, a median 973 ms is CLI and harness time, not the model. Chart: Where an LLM router’s time goes · median per call (n = 82 each). Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
- Route all 49.5 model calls per task: $1.67 per 1,000 tasks with Jev, $247.30 with Sonnet 5.5. Route 7 decisions: $34.97. Chart: Added routing cost per 1,000 tasks. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
- If each decision waits in line, Sonnet 5.5 adds up to 129 s per task, or 18.2 s for 7 decisions. The policy adds 70.3 µs. Chart: Added routing delay per task · upper bound. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
- Start-up tax for a one-word answer: Claude Code 2.53 s and 6,761 input tokens; Codex CLI 6.00 s and 17,051 input tokens, 13,184 of them read from the cache. Chart: CLI start-up tax · one-word answer · median of 5 runs (n = 5 each). Caveat: CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
- Route in process where a rule is enough. Every timing, cost and gap online.
Three ways to route
- Rules in the process. Code reads the task and picks a model, an effort and a plan. No model call. Our production routing policy is this kind.
- A dedicated decision model. A small hosted model fills a typed schema. Jev 1.13 is this kind.
- A general LLM as the router. You ask Claude, through its CLI or API, to choose. Simple to build. Is it affordable?
Accuracy is a separate question, answered in Jev vs Claude Haiku and Sonnet as a router: on 82 typed decisions, the three model routers tie within their intervals. This post is about the price of asking.
Delay per decision
- Deterministic routing policy: median 1.42 µs, p95 2.33 µs, p99 3.04 µs, over 20,000 timed decisions after 5,000 warm-up calls. In a 200,000-call batch it made 469,409 decisions per second. The chart uses a log scale, so this point shows next to routers that take seconds.
- Claude Sonnet 5.5 (effort low, via Claude Code): median 2,597 ms, p95 4,298 ms.
- Claude Haiku 4.5 (thinking on, via Claude Code): median 12,543 ms, p95 34,481 ms.
- Jev 1.13 (a direct HTTPS call): median 136.5 ms, p95 195.7 ms, fastest 100.9 ms, slowest 297.3 ms, over 246 calls. The first call of the run, on a fresh connection, took 224.7 ms.
The whisker is the median to the 95th percentile, not a confidence interval. On the comparison pages the policy leads both Claude routers on decision time: each Claude median sits above the policy's 95th percentile. See rules vs Sonnet and rules vs Haiku.
Jev's row is a different route. We called its API from the same Mac over a home network: 3 repeats of the 82 routing decisions, one at a time, in a 35-second window. The network is inside the 136.5 ms, and the API reports no server time. The Claude routers went through the CLI. So the chart shows what a caller waits per decision, not model compute time. Both Claude medians still sit above Jev's 95th percentile, and Sonnet's model time alone (1,596 ms) is above Jev's whole call.
A fair note: the policy is timed in process, without its database reads and the decision record. In production, a recorded rule-based System One decision took a median 1 ms, p95 2 ms, with the record write included (n = 419, millisecond resolution). Still about three orders of magnitude below a Claude call.
Where an LLM router's time goes
We split each Claude call into the API time the CLI reported and everything else:
- Sonnet: model time median 1,596 ms, CLI and harness time median 973 ms.
- Haiku: model time median 10,508 ms, CLI and harness time median 1,698 ms.
Medians of the parts do not add up to the median of the whole. The point stands: about a second of each Sonnet decision is the CLI, not the model. A direct API call would skip that part. We did not time a direct API router in this study.
The CLI tax shows up even on a one-word answer. Claude Code with Haiku 4.5 took 2.53 s and sent 6,761 input tokens; the Codex CLI took 6.00 s and sent 17,051, with a different model (5 runs each). More in Claude vs Codex on hard tasks.
Cost per 1,000 decisions
- Deterministic policy: $0. It makes no model call.
- Jev 1.13: $0.0337 per 1,000. A calculation: 803 input tokens per decision, as its API reported, × $0.042 per million input tokens; output tokens are free. Its provider's own cost figure for the recorded production run is the same.
- Claude Sonnet 5.5: $5.00 per 1,000, a list-price calculation.
- Claude Haiku 4.5: $8.92 per 1,000, a list-price calculation. Its default thinking wrote far more output tokens than Sonnet at effort low.
Per 1,000 tasks: where it adds up
One decision is cheap. An agent makes many. In 48 recorded bench runs, a task made a median 49.5 model calls (13 to 73). If a router decides before every call, it decides 49.5 times per task. If it decides only at the System One decision points the platform records, it decides 7 times per task.
Added cost per 1,000 tasks (a calculation; the median recorded work cost was $3.03 per task):
| Router | Every call routed (49.5) | System One only (7) |
|---|---|---|
| Deterministic policy | $0 | $0 |
| Jev 1.13 | $1.67 | $0.24 |
| Claude Sonnet 5.5 | $247.30 | $34.97 |
| Claude Haiku 4.5 | $441.74 | $62.47 |
Sonnet on every call adds 8.2% to the work cost. Haiku adds about 14.6% ($0.44 on $3.03 per task, our arithmetic). Jev adds about 0.06%.
Added delay per task (a calculation and an upper bound, because it assumes each decision waits for the one before):
- Sonnet, every call: up to 128.6 s. The median recorded task took 10.3 minutes, so that is about a fifth more wall time (our arithmetic).
- Haiku, every call: up to 620.9 s, about as long as the median task itself.
- Sonnet, System One only: up to 18.2 s. Haiku: up to 87.8 s.
- Rules: 70 µs per task when every call is routed.
- Jev: up to 6.76 s with every call routed and 0.96 s with the 7 System One decisions, at its live median of 136.5 ms. About 19 times less than Sonnet's waiting (a calculation across two routes).
So what should you route with?
- Use rules for everything rules can decide. They are free, they take microseconds and they never fail to return. In our run, 20,000 of 20,000 policy decisions completed.
- Use a model only where judgment pays. Our routing study found that simple classifications saturate fast; multi-field context judgments are where model routers differ.
- If you use a model, keep it off the hot path. Seven decisions per task instead of 49.5 cuts the Sonnet bill by about seven times and the delay from minutes to seconds.
- Do not put a CLI in front of a router. About a second of every Sonnet decision was CLI overhead.
- A dedicated decision model is cheap enough to run on every call. Jev on every call costs less than Sonnet on 7 decisions per task. Its measured delay is small enough to consider: 136.5 ms per call from one Mac, up to 6.76 s per task if all 49.5 decisions wait in line. Time it from where your code runs.
Pair pages: Jev vs rules, Jev vs Sonnet and Jev vs Haiku.
How we measured
- Protocol declared before any measurement.
- Policy: the platform's production routing decision (plan, model ladder and effort) on its default policy, timed in process on one Apple M3 Ultra Mac with Node 25. 5,000 warm-up calls, then 20,000 timed decisions over 64 synthetic routing contexts, plus a 200,000-call batch for throughput.
- Claude routers: the recorded runs of the routing study, 82 calls each, one call at a time. Wall time, the API time the CLI reported, and the difference. p50 and p95 recomputed from the per-call log.
- Jev: a live run on 2026-10-06. The same 82 decisions as the routing study, 3 repeats, 246 counted calls, one at a time over HTTPS from the same Mac. Time is client wall time, network included. Cost is a calculation from the input tokens its API reported × the published price. The protocol was written before the first counted call.
- Per 1,000 tasks: decisions per task from 48 recorded bench runs (2 empty runs excluded) × cost and median time per decision. A calculation.
Caveats
- Not equal footing. The policy runs in process; the Claude routers ran through a CLI. This compares the two ways of routing as deployed, not two models.
- Upper bound on delay. Decisions that run in parallel with other work would add less.
- Calculations on runs with routing off. The per-task figures price a router onto recorded work. They are not runs with the router on.
- Jev's timing is one window. 246 calls in 35 seconds, from one machine on a home network, on 2026-10-06. Server load then is unknown. A caller near the API would likely see less. The 246 calls are 3 repeats of 82 requests, not independent draws.
- Not measured: OpenRouter's Auto Router and cheaper hosted inference (no key in the environment), and a local router server (none was running).
- Small medians shift. The routing study reports its medians from its own summary; this study recomputes them from the per-call log, so they can differ by a few tens of milliseconds.
What to read next
- Jev vs Claude Haiku vs Claude Sonnet as a router: an honest comparison
- What if every call ran on Opus?
- OpenRouter vs going direct: what the gateway costs
See every routing decision
Agent records each routing decision with its reason, its cost and its time. Try Agent and see which choices it made for your work.