Explainer · LLM router

What is an LLM router?

Definition

An LLM router is the component that decides, before a request runs, which language model should answer it, at which reasoning effort and with how much context. The router can be a set of rules in code, a small model trained to make the decision, or a general LLM that you ask to choose. A good router sends easy work to cheap, fast models and hard work to strong ones, and adds less delay and cost than it saves.

Agent team · · 5 min read · Every number is from the public studies

Interactive

What a router does with one task

The router reads the task and picks the model that runs it. Pick a router to see its measured decision time and accuracy.

MeasuredRoute: illustration
Router

4 routers. Deterministic routing policy: 1.42 µs median · p95 2.33 µs (n = 20000); exact rate not measured. Jev 1.13: 137 ms median · p95 196 ms (n = 246); 90% exact (95% interval 82% to 95%, n = 82). Claude Sonnet 5.5: 2.6 s median · p95 4.3 s (n = 82); 94% exact (95% interval 87% to 97%, n = 82). Claude Haiku 4.5: 12.54 s median · p95 34.48 s (n = 82); 89% exact (95% interval 80% to 94%, n = 82).

Time at the router is drawn on a log scale (microseconds to seconds), not in real time. The model the route lights is an example decision; the measured numbers are the router's time and its exact-decision rate on 82 typed cases.

Source: Routing overhead study

Why route at all

Models differ in list price by several times, and in speed by seconds per call. Most agent workloads mix easy steps (classify a message, pick a file, summarize a log) with hard steps (write the patch, review the change). If every step runs on the strongest model, you pay its price for work a smaller model could do. If every step runs on the smallest model, the hard steps fail.

A router tries to get the best of both. It reads the request, and sometimes the state of the task, and returns a decision: a model, an effort level, and often a plan for how much context to send.

Three kinds of router

  1. Rules in the process. Code reads typed fields of the task and picks a model. There is no model call, so the decision is fast and free. The limit is that rules only know what you wrote down.
  2. A dedicated decision model. A small hosted model fills a typed schema: which model, which effort, which next step. Jev 1.13 is this kind.
  3. A general LLM as the router. You ask Claude or GPT, through an API or a CLI, to choose. It is easy to build, but each decision is a full model call.

A fourth thing is often called a router but is a different layer: an inference gateway such as OpenRouter forwards a request to one of many providers. It routes between providers of the same model, not between models for the task. See inference gateways and OpenRouter.

How accurate is a router?

Accuracy means the router picked an acceptable answer for every scored question of a decision. In our routing study, three routers answered the same 82 typed decisions:

Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 82 per row

Share of asked cases where every scored question was acceptable

Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

  • Jev 1.13 (TypeSafe): 221 of 246 live calls exactly right (90%; 74, 73 and 74 of 82 per repeat; 95% interval 82% to 95%).
  • Claude Haiku 4.5: 73 of 82 (89%, 80% to 94%).
  • Claude Sonnet 5.5: 77 of 82 (94%, 87% to 97%).

The three intervals overlap, so this sample does not separate them on accuracy. The case sets were tuned against Jev answers, which gives Jev a home advantage. The full study is at /benchmarks/routing-jev-vs-llm.

What a decision costs in time

A router runs before the real work, so its delay adds to every call it decides. The difference between kinds is large:

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

Time per decision · log scale: each gridline is 10 times the one before

4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–20000 per row

Median; whiskers = median to 95th percentile

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

  • The deterministic policy decided in a median 1.42 µs (95th percentile 2.33 µs, over 20,000 decisions).
  • Claude Sonnet 5.5 through the Claude Code CLI took a median 2.60 s per decision (p95 4.30 s, n = 82). About 973 ms of that median was CLI time, not model time.
  • Claude Haiku 4.5 with its default thinking took a median 12.54 s (p95 34.48 s).
  • Jev 1.13, a hosted decision model, took a median 136.5 ms per call (p95 195.7 ms, n = 246 calls). We called its API directly from one Mac over a home network, so the network is inside that time.

The whisker on this chart is the median to the 95th percentile, not a confidence interval. Jev's row is a different route from the Claude rows: a direct API call, not a CLI. The chart shows what a caller waits per decision, not model compute time.

What a decision costs in money

Calculation
Largest value is 260x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).

Notesn 82–246 per row

List price × reported tokens per decision

List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions

Per 1,000 decisions, Jev cost $0.0337 (a calculation from the input tokens its API reported). The Claude routers cost $8.92 (Haiku 4.5) and $5.00 (Sonnet 5.5) at list price, a calculation from their reported tokens. The rule-based policy makes no model call, so it costs $0.

One decision is cheap. An agent makes many: in 48 recorded bench runs, a task made a median 49.5 model calls. If a router decides before every call, its cost and delay multiply by that number. The routing overhead study works out the per-task effect.

How to choose a router

  • Start with rules for every decision that typed fields can answer. They are fast, free and easy to test.
  • Use a small decision model for the decisions rules cannot express, when its accuracy is close to a large model's on your own cases.
  • Use a general LLM as the router only where its extra accuracy is measured and worth seconds per call.
  • Route at decision points, not at every call. Routing 7 decisions per task instead of every model call cuts the overhead by the same ratio.
  • Measure on your own cases. Accuracy on someone else's decision set does not transfer, and a set tuned on one router favors it.

Frequently asked questions

Is an LLM router the same as OpenRouter?

No. OpenRouter is an inference gateway: it forwards a request for one model to one of several providers that serve it. An LLM router chooses which model, and which effort, should handle the task. You can use both: a router picks the model, and a gateway picks the provider.

Does a router make answers better or only cheaper?

Mostly cheaper and faster. A router saves money when it moves easy work to small models without a loss in quality. On our 82 typed decisions, the three model routers tie within their 95% intervals, so the choice between them comes down to cost and delay.

How much delay does an LLM router add?

It depends on the kind. Rules in the process decided in a median 1.42 µs. Claude Sonnet 5.5 through the Claude Code CLI took a median 2.60 s per decision, and Claude Haiku 4.5 with its default thinking took a median 12.54 s.

Should a router run before every model call?

Usually not. A task in our runs made a median 49.5 model calls, so a router on every call multiplies its delay and cost by that number. Routing only at the few decision points that change the outcome keeps the overhead small.

Watch the data

Live story · 51 sRouting overhead: a 1.42 µs policy vs LLM routers

Routing overhead: a 1.42 µs policy vs LLM routers

A deterministic routing policy decides in 1.42 µs (median, n = 20,000); Sonnet 5.5 as a router takes 2.60 s through the CLI. Per-task costs and delays are calculations.

Transcript
  1. Routing overhead · policy vs LLM routers. What a routing decision costs before the work starts. Time and money per decision, then per task. Jev is timed over its API, a different route from the CLI routers.
  2. The latency ladder: the in-process policy decides in 1.42 µs. Jev, a direct API call, takes 137 ms. Sonnet 5.5 as a router takes 2.60 s through the CLI, Haiku 4.5 12.5 s. Chart: Time per decision · log scale · median to p95 (n = 82–20000 each). Caveat: Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
  3. The LLM router takes about 1.8 million times longer per decision. Jev costs $0.0337 per 1,000 decisions, provider-reported. Policy decision, median (20,000 timed): 1.42 µs (n = 20000, p95 2.33 µs). Sonnet 5.5 router ÷ policy, medians: 1.8 million× (n = 82). Jev 1.13 per 1,000 decisions (provider-reported): $0.0337 (n = 82, 137 ms median per call over its API). Caveat: OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
  4. Of Sonnet 5.5’s 2.60 s per decision, a median 973 ms is CLI and harness time, not the model. Chart: Where an LLM router’s time goes · median per call (n = 82 each). Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
  5. Route all 49.5 model calls per task: $1.67 per 1,000 tasks with Jev, $247.30 with Sonnet 5.5. Route 7 decisions: $34.97. Chart: Added routing cost per 1,000 tasks. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
  6. If each decision waits in line, Sonnet 5.5 adds up to 129 s per task, or 18.2 s for 7 decisions. The policy adds 70.3 µs. Chart: Added routing delay per task · upper bound. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
  7. Start-up tax for a one-word answer: Claude Code 2.53 s and 6,761 input tokens; Codex CLI 6.00 s and 17,051 input tokens, 13,184 of them read from the cache. Chart: CLI start-up tax · one-word answer · median of 5 runs (n = 5 each). Caveat: CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
  8. Route in process where a rule is enough. Every timing, cost and gap online.

The data behind this explainer

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.