Explainer · LLM router
What is an LLM router?
Definition
An LLM router is the component that decides, before a request runs, which language model should answer it, at which reasoning effort and with how much context. The router can be a set of rules in code, a small model trained to make the decision, or a general LLM that you ask to choose. A good router sends easy work to cheap, fast models and hard work to strong ones, and adds less delay and cost than it saves.
Agent team · · 5 min read · Every number is from the public studies
Interactive
What a router does with one task
The router reads the task and picks the model that runs it. Pick a router to see its measured decision time and accuracy.
| Router | Median decision time | p95 | n (time) | Exact decisions | 95% Wilson interval | n (cases) |
|---|---|---|---|---|---|---|
| Deterministic routing policy | 1.42 µs | 2.33 µs | 20000 | not measured | — | — |
| Jev 1.13 | 137 ms | 196 ms | 246 | 90% | 82%–95% | 82 |
| Claude Sonnet 5.5 | 2.6 s | 4.3 s | 82 | 94% | 87%–97% | 82 |
| Claude Haiku 4.5 | 12.54 s | 34.48 s | 82 | 89% | 80%–94% | 82 |
4 routers. Deterministic routing policy: 1.42 µs median · p95 2.33 µs (n = 20000); exact rate not measured. Jev 1.13: 137 ms median · p95 196 ms (n = 246); 90% exact (95% interval 82% to 95%, n = 82). Claude Sonnet 5.5: 2.6 s median · p95 4.3 s (n = 82); 94% exact (95% interval 87% to 97%, n = 82). Claude Haiku 4.5: 12.54 s median · p95 34.48 s (n = 82); 89% exact (95% interval 80% to 94%, n = 82).
Time at the router is drawn on a log scale (microseconds to seconds), not in real time. The model the route lights is an example decision; the measured numbers are the router's time and its exact-decision rate on 82 typed cases.
Source: Routing overhead study
Why route at all
Models differ in list price by several times, and in speed by seconds per call. Most agent workloads mix easy steps (classify a message, pick a file, summarize a log) with hard steps (write the patch, review the change). If every step runs on the strongest model, you pay its price for work a smaller model could do. If every step runs on the smallest model, the hard steps fail.
A router tries to get the best of both. It reads the request, and sometimes the state of the task, and returns a decision: a model, an effort level, and often a plan for how much context to send.
Three kinds of router
- Rules in the process. Code reads typed fields of the task and picks a model. There is no model call, so the decision is fast and free. The limit is that rules only know what you wrote down.
- A dedicated decision model. A small hosted model fills a typed schema: which model, which effort, which next step. Jev 1.13 is this kind.
- A general LLM as the router. You ask Claude or GPT, through an API or a CLI, to choose. It is easy to build, but each decision is a full model call.
A fourth thing is often called a router but is a different layer: an inference gateway such as OpenRouter forwards a request to one of many providers. It routes between providers of the same model, not between models for the task. See inference gateways and OpenRouter.
How accurate is a router?
Accuracy means the router picked an acceptable answer for every scored question of a decision. In our routing study, three routers answered the same 82 typed decisions:
- Jev 1.13 (TypeSafe): 221 of 246 live calls exactly right (90%; 74, 73 and 74 of 82 per repeat; 95% interval 82% to 95%).
- Claude Haiku 4.5: 73 of 82 (89%, 80% to 94%).
- Claude Sonnet 5.5: 77 of 82 (94%, 87% to 97%).
The three intervals overlap, so this sample does not separate them on accuracy. The case sets were tuned against Jev answers, which gives Jev a home advantage. The full study is at /benchmarks/routing-jev-vs-llm.
What a decision costs in time
A router runs before the real work, so its delay adds to every call it decides. The difference between kinds is large:
- The deterministic policy decided in a median 1.42 µs (95th percentile 2.33 µs, over 20,000 decisions).
- Claude Sonnet 5.5 through the Claude Code CLI took a median 2.60 s per decision (p95 4.30 s, n = 82). About 973 ms of that median was CLI time, not model time.
- Claude Haiku 4.5 with its default thinking took a median 12.54 s (p95 34.48 s).
- Jev 1.13, a hosted decision model, took a median 136.5 ms per call (p95 195.7 ms, n = 246 calls). We called its API directly from one Mac over a home network, so the network is inside that time.
The whisker on this chart is the median to the 95th percentile, not a confidence interval. Jev's row is a different route from the Claude rows: a direct API call, not a CLI. The chart shows what a caller waits per decision, not model compute time.
What a decision costs in money
Per 1,000 decisions, Jev cost $0.0337 (a calculation from the input tokens its API reported). The Claude routers cost $8.92 (Haiku 4.5) and $5.00 (Sonnet 5.5) at list price, a calculation from their reported tokens. The rule-based policy makes no model call, so it costs $0.
One decision is cheap. An agent makes many: in 48 recorded bench runs, a task made a median 49.5 model calls. If a router decides before every call, its cost and delay multiply by that number. The routing overhead study works out the per-task effect.
How to choose a router
- Start with rules for every decision that typed fields can answer. They are fast, free and easy to test.
- Use a small decision model for the decisions rules cannot express, when its accuracy is close to a large model's on your own cases.
- Use a general LLM as the router only where its extra accuracy is measured and worth seconds per call.
- Route at decision points, not at every call. Routing 7 decisions per task instead of every model call cuts the overhead by the same ratio.
- Measure on your own cases. Accuracy on someone else's decision set does not transfer, and a set tuned on one router favors it.
Frequently asked questions
Is an LLM router the same as OpenRouter?
No. OpenRouter is an inference gateway: it forwards a request for one model to one of several providers that serve it. An LLM router chooses which model, and which effort, should handle the task. You can use both: a router picks the model, and a gateway picks the provider.
Does a router make answers better or only cheaper?
Mostly cheaper and faster. A router saves money when it moves easy work to small models without a loss in quality. On our 82 typed decisions, the three model routers tie within their 95% intervals, so the choice between them comes down to cost and delay.
How much delay does an LLM router add?
It depends on the kind. Rules in the process decided in a median 1.42 µs. Claude Sonnet 5.5 through the Claude Code CLI took a median 2.60 s per decision, and Claude Haiku 4.5 with its default thinking took a median 12.54 s.
Should a router run before every model call?
Usually not. A task in our runs made a median 49.5 model calls, so a router on every call multiplies its delay and cost by that number. Routing only at the few decision points that change the outcome keeps the overhead small.
Watch the data
Routing overhead: a 1.42 µs policy vs LLM routers
A deterministic routing policy decides in 1.42 µs (median, n = 20,000); Sonnet 5.5 as a router takes 2.60 s through the CLI. Per-task costs and delays are calculations.
Transcript
- Routing overhead · policy vs LLM routers. What a routing decision costs before the work starts. Time and money per decision, then per task. Jev is timed over its API, a different route from the CLI routers.
- The latency ladder: the in-process policy decides in 1.42 µs. Jev, a direct API call, takes 137 ms. Sonnet 5.5 as a router takes 2.60 s through the CLI, Haiku 4.5 12.5 s. Chart: Time per decision · log scale · median to p95 (n = 82–20000 each). Caveat: Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
- The LLM router takes about 1.8 million times longer per decision. Jev costs $0.0337 per 1,000 decisions, provider-reported. Policy decision, median (20,000 timed): 1.42 µs (n = 20000, p95 2.33 µs). Sonnet 5.5 router ÷ policy, medians: 1.8 million× (n = 82). Jev 1.13 per 1,000 decisions (provider-reported): $0.0337 (n = 82, 137 ms median per call over its API). Caveat: OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
- Of Sonnet 5.5’s 2.60 s per decision, a median 973 ms is CLI and harness time, not the model. Chart: Where an LLM router’s time goes · median per call (n = 82 each). Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
- Route all 49.5 model calls per task: $1.67 per 1,000 tasks with Jev, $247.30 with Sonnet 5.5. Route 7 decisions: $34.97. Chart: Added routing cost per 1,000 tasks. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
- If each decision waits in line, Sonnet 5.5 adds up to 129 s per task, or 18.2 s for 7 decisions. The policy adds 70.3 µs. Chart: Added routing delay per task · upper bound. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
- Start-up tax for a one-word answer: Claude Code 2.53 s and 6,761 input tokens; Codex CLI 6.00 s and 17,051 input tokens, 13,184 of them read from the cache. Chart: CLI start-up tax · one-word answer · median of 5 runs (n = 5 each). Caveat: CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
- Route in process where a rule is enough. Every timing, cost and gap online.