• Routing
  • Model Routing
  • Latency
  • LLM pricing

What does a router cost you? Rules vs Jev vs an LLM router

A rule-based router decides in 1.42 µs for $0, Jev in 136.5 ms, Sonnet in 2.60 s. What each adds per 1,000 tasks, in delay and in dollars.

TL;DR

  • A router makes a decision before the real work starts: which model, which effort, how much context. That decision has a delay and a cost.
  • Deterministic rules: a median 1.42 µs per decision (95th percentile 2.33 µs, n = 20,000) and $0.
  • Claude Sonnet 5.5 as a router, through the Claude Code CLI: a median 2.60 s per decision (p95 4.30 s, n = 82). That is about 1.8 million times longer. 973 ms of the median is CLI time, not model time.
  • Claude Haiku 4.5 with its default thinking: a median 12.54 s (p95 34.48 s, n = 82).
  • Jev 1.13, called directly over HTTPS: a median 136.5 ms per call (p95 195.7 ms, n = 246 calls, network included) and $0.0337 per 1,000 decisions (a calculation from its reported input tokens). That is a different route from the CLI routers, so it is not a model-against-model comparison. The median is about 19 times shorter than Sonnet's and 92 times shorter than Haiku's (a calculation across the two routes).
  • Per 1,000 tasks, routing every model call (49.5 per task) adds $0 with rules, $1.67 with Jev, $247.30 with Sonnet and $441.74 with Haiku. That is a calculation. If every decision waits in line, Jev adds up to 6.76 s per task and Sonnet up to 128.6 s. Routing only the 7 System One decisions per task cuts Sonnet to $34.97 and 18.2 s, and Jev to $0.24 and 0.96 s.

The full study: /benchmarks/routing-overhead.

Live story · 51 sRouting overhead: a 1.42 µs policy vs LLM routers

Routing overhead: a 1.42 µs policy vs LLM routers

A deterministic routing policy decides in 1.42 µs (median, n = 20,000); Sonnet 5.5 as a router takes 2.60 s through the CLI. Per-task costs and delays are calculations.

Transcript
  1. Routing overhead · policy vs LLM routers. What a routing decision costs before the work starts. Time and money per decision, then per task. Jev is timed over its API, a different route from the CLI routers.
  2. The latency ladder: the in-process policy decides in 1.42 µs. Jev, a direct API call, takes 137 ms. Sonnet 5.5 as a router takes 2.60 s through the CLI, Haiku 4.5 12.5 s. Chart: Time per decision · log scale · median to p95 (n = 82–20000 each). Caveat: Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
  3. The LLM router takes about 1.8 million times longer per decision. Jev costs $0.0337 per 1,000 decisions, provider-reported. Policy decision, median (20,000 timed): 1.42 µs (n = 20000, p95 2.33 µs). Sonnet 5.5 router ÷ policy, medians: 1.8 million× (n = 82). Jev 1.13 per 1,000 decisions (provider-reported): $0.0337 (n = 82, 137 ms median per call over its API). Caveat: OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
  4. Of Sonnet 5.5’s 2.60 s per decision, a median 973 ms is CLI and harness time, not the model. Chart: Where an LLM router’s time goes · median per call (n = 82 each). Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
  5. Route all 49.5 model calls per task: $1.67 per 1,000 tasks with Jev, $247.30 with Sonnet 5.5. Route 7 decisions: $34.97. Chart: Added routing cost per 1,000 tasks. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
  6. If each decision waits in line, Sonnet 5.5 adds up to 129 s per task, or 18.2 s for 7 decisions. The policy adds 70.3 µs. Chart: Added routing delay per task · upper bound. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
  7. Start-up tax for a one-word answer: Claude Code 2.53 s and 6,761 input tokens; Codex CLI 6.00 s and 17,051 input tokens, 13,184 of them read from the cache. Chart: CLI start-up tax · one-word answer · median of 5 runs (n = 5 each). Caveat: CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
  8. Route in process where a rule is enough. Every timing, cost and gap online.

Three ways to route

  1. Rules in the process. Code reads the task and picks a model, an effort and a plan. No model call. Our production routing policy is this kind.
  2. A dedicated decision model. A small hosted model fills a typed schema. Jev 1.13 is this kind.
  3. A general LLM as the router. You ask Claude, through its CLI or API, to choose. Simple to build. Is it affordable?

Accuracy is a separate question, answered in Jev vs Claude Haiku and Sonnet as a router: on 82 typed decisions, the three model routers tie within their intervals. This post is about the price of asking.

Delay per decision

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

Time per decision · log scale: each gridline is 10 times the one before

4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–20000 per row

Median; whiskers = median to 95th percentile

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

  • Deterministic routing policy: median 1.42 µs, p95 2.33 µs, p99 3.04 µs, over 20,000 timed decisions after 5,000 warm-up calls. In a 200,000-call batch it made 469,409 decisions per second. The chart uses a log scale, so this point shows next to routers that take seconds.
  • Claude Sonnet 5.5 (effort low, via Claude Code): median 2,597 ms, p95 4,298 ms.
  • Claude Haiku 4.5 (thinking on, via Claude Code): median 12,543 ms, p95 34,481 ms.
  • Jev 1.13 (a direct HTTPS call): median 136.5 ms, p95 195.7 ms, fastest 100.9 ms, slowest 297.3 ms, over 246 calls. The first call of the run, on a fresh connection, took 224.7 ms.

The whisker is the median to the 95th percentile, not a confidence interval. On the comparison pages the policy leads both Claude routers on decision time: each Claude median sits above the policy's 95th percentile. See rules vs Sonnet and rules vs Haiku.

Jev's row is a different route. We called its API from the same Mac over a home network: 3 repeats of the 82 routing decisions, one at a time, in a 35-second window. The network is inside the 136.5 ms, and the API reports no server time. The Claude routers went through the CLI. So the chart shows what a caller waits per decision, not model compute time. Both Claude medians still sit above Jev's 95th percentile, and Sonnet's model time alone (1,596 ms) is above Jev's whole call.

A fair note: the policy is timed in process, without its database reads and the decision record. In production, a recorded rule-based System One decision took a median 1 ms, p95 2 ms, with the record write included (n = 419, millisecond resolution). Still about three orders of magnitude below a Claude call.

Where an LLM router's time goes

  • Model API time
  • CLI and harness time
Entrance: medians race at 7.5× real timeMotion reduced: press Replay to animateThe slowest median is 10.51 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

2 rows, 2 series: Model API time, CLI and harness time. Model API time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 10.51 s (median to p95 10.51 s–32.13 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 1.6 s (median to p95 1.6 s–2.58 s, n 82). Not all run ranges overlap. CLI and harness time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 1.7 s (median to p95 1.7 s–2.68 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 973 ms (median to p95 973 ms–1.28 s, n 82). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n = 82 per row

Median per call; whiskers = median to 95th percentile

Model API time is the API duration the CLI reports; CLI and harness time is wall time minus that, per call. Medians of the parts do not add up to the median of the whole. Jev is not split: its API reports no server time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing

We split each Claude call into the API time the CLI reported and everything else:

  • Sonnet: model time median 1,596 ms, CLI and harness time median 973 ms.
  • Haiku: model time median 10,508 ms, CLI and harness time median 1,698 ms.

Medians of the parts do not add up to the median of the whole. The point stands: about a second of each Sonnet decision is the CLI, not the model. A direct API call would skip that part. We did not time a direct API router in this study.

The CLI tax shows up even on a one-word answer. Claude Code with Haiku 4.5 took 2.53 s and sent 6,761 input tokens; the Codex CLI took 6.00 s and sent 17,051, with a different model (5 runs each). More in Claude vs Codex on hard tasks.

Cost per 1,000 decisions

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)

2 rows. Highest Jev 1.13 (TypeSafe) $0.034 (n 82). Lowest Deterministic routing policy (Agent, in process) $0 (n 20000).

Notesn 82–20000 per row

USD per 1,000 decisions

The policy makes no model call, so it costs nothing per decision. Jev’s figure is the cost its provider reported for the recorded production run (82 decisions). The input tokens of the live run give the same figure at the published price. The Claude routers are in the next chart: their cost is derived from list prices.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev 1.13 list price

Calculation
Largest value is 260x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 (thinking on, via Claude Code) $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).

Notesn 82–246 per row

List price × the tokens each route reported, USD per 1,000 decisions

A calculation, not a bill: the Claude calls ran on a subscription. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free). Same tokens and prices as the routing study.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions

  • Deterministic policy: $0. It makes no model call.
  • Jev 1.13: $0.0337 per 1,000. A calculation: 803 input tokens per decision, as its API reported, × $0.042 per million input tokens; output tokens are free. Its provider's own cost figure for the recorded production run is the same.
  • Claude Sonnet 5.5: $5.00 per 1,000, a list-price calculation.
  • Claude Haiku 4.5: $8.92 per 1,000, a list-price calculation. Its default thinking wrote far more output tokens than Sonnet at effort low.

Per 1,000 tasks: where it adds up

One decision is cheap. An agent makes many. In 48 recorded bench runs, a task made a median 49.5 model calls (13 to 73). If a router decides before every call, it decides 49.5 times per task. If it decides only at the System One decision points the platform records, it decides 7 times per task.

Calculation
  • Every model call routed (49.5 per task)
  • Only System One decisions (7 per task) (square)
In chart order.
Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).

List-price calculation, not a run. 4 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $442. Lowest Deterministic routing policy (Agent, in process) $0. Only System One decisions (7 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $62.47. Lowest Deterministic routing policy (Agent, in process) $0.

Notes

Decisions per task from recorded runs × cost per decision

A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions

Added cost per 1,000 tasks (a calculation; the median recorded work cost was $3.03 per task):

RouterEvery call routed (49.5)System One only (7)
Deterministic policy$0$0
Jev 1.13$1.67$0.24
Claude Sonnet 5.5$247.30$34.97
Claude Haiku 4.5$441.74$62.47

Sonnet on every call adds 8.2% to the work cost. Haiku adds about 14.6% ($0.44 on $3.03 per task, our arithmetic). Jev adds about 0.06%.

Calculation
  • Every model call routed (49.5 per task)
  • Only System One decisions (7 per task) (square)
In chart order.
Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).

List-price calculation, not a run. 4 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 621 s. Fastest Deterministic routing policy (Agent, in process) 70.3 µs. Only System One decisions (7 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 87.8 s. Fastest Deterministic routing policy (Agent, in process) 9.9 µs.

Notes

Decisions per task × median decision time, if every decision waits in line

A calculation and an upper bound: it assumes each decision waits for the one before. Median recorded task wall time: 10.3 min. Jev’s delay uses its live median over the API from one Mac (network included); the Claude routers’ includes the CLI.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions

Added delay per task (a calculation and an upper bound, because it assumes each decision waits for the one before):

  • Sonnet, every call: up to 128.6 s. The median recorded task took 10.3 minutes, so that is about a fifth more wall time (our arithmetic).
  • Haiku, every call: up to 620.9 s, about as long as the median task itself.
  • Sonnet, System One only: up to 18.2 s. Haiku: up to 87.8 s.
  • Rules: 70 µs per task when every call is routed.
  • Jev: up to 6.76 s with every call routed and 0.96 s with the 7 System One decisions, at its live median of 136.5 ms. About 19 times less than Sonnet's waiting (a calculation across two routes).

So what should you route with?

  • Use rules for everything rules can decide. They are free, they take microseconds and they never fail to return. In our run, 20,000 of 20,000 policy decisions completed.
  • Use a model only where judgment pays. Our routing study found that simple classifications saturate fast; multi-field context judgments are where model routers differ.
  • If you use a model, keep it off the hot path. Seven decisions per task instead of 49.5 cuts the Sonnet bill by about seven times and the delay from minutes to seconds.
  • Do not put a CLI in front of a router. About a second of every Sonnet decision was CLI overhead.
  • A dedicated decision model is cheap enough to run on every call. Jev on every call costs less than Sonnet on 7 decisions per task. Its measured delay is small enough to consider: 136.5 ms per call from one Mac, up to 6.76 s per task if all 49.5 decisions wait in line. Time it from where your code runs.

Pair pages: Jev vs rules, Jev vs Sonnet and Jev vs Haiku.

How we measured

  • Protocol declared before any measurement.
  • Policy: the platform's production routing decision (plan, model ladder and effort) on its default policy, timed in process on one Apple M3 Ultra Mac with Node 25. 5,000 warm-up calls, then 20,000 timed decisions over 64 synthetic routing contexts, plus a 200,000-call batch for throughput.
  • Claude routers: the recorded runs of the routing study, 82 calls each, one call at a time. Wall time, the API time the CLI reported, and the difference. p50 and p95 recomputed from the per-call log.
  • Jev: a live run on 2026-10-06. The same 82 decisions as the routing study, 3 repeats, 246 counted calls, one at a time over HTTPS from the same Mac. Time is client wall time, network included. Cost is a calculation from the input tokens its API reported × the published price. The protocol was written before the first counted call.
  • Per 1,000 tasks: decisions per task from 48 recorded bench runs (2 empty runs excluded) × cost and median time per decision. A calculation.

Caveats

  • Not equal footing. The policy runs in process; the Claude routers ran through a CLI. This compares the two ways of routing as deployed, not two models.
  • Upper bound on delay. Decisions that run in parallel with other work would add less.
  • Calculations on runs with routing off. The per-task figures price a router onto recorded work. They are not runs with the router on.
  • Jev's timing is one window. 246 calls in 35 seconds, from one machine on a home network, on 2026-10-06. Server load then is unknown. A caller near the API would likely see less. The 246 calls are 3 repeats of 82 requests, not independent draws.
  • Not measured: OpenRouter's Auto Router and cheaper hosted inference (no key in the environment), and a local router server (none was running).
  • Small medians shift. The routing study reports its medians from its own summary; this study recomputes them from the per-call log, so they can differ by a few tens of milliseconds.

See every routing decision

Agent records each routing decision with its reason, its cost and its time. Try Agent and see which choices it made for your work.

The data behind this post

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.