• Routing
  • Latency
  • Overhead
  • Jev
  • LLM Router
  • CLI
  • Calculation

Routing overhead: deterministic policy vs LLM routers vs Jev

What delay and what cost does each kind of router add before the real work of a call starts?

Published · 9 charts · Download the data or a carousel

1.42µs

n = 20000

Deterministic routing policy: median decision time · p95 2.33 µs, p99 3.04 µs

Timer resolution 0.041 µs; 469,409 decisions per second in a 200,000-call batch.

The answer

The deterministic routing policy decided in a median 1.42 µs (p95 2.33 µs, 20,000 decisions, $0). The fastest LLM router, Claude Sonnet 5.5 through the Claude Code CLI, took a median 2.60 s per decision (p95 4.30 s, n = 82), about 1.8 million times longer; 973 ms of that median was CLI time, not model time. Haiku 4.5 with its default thinking took 12.54 s (p95 34.48 s). Jev 1.13, called directly over HTTPS from the same Mac, took a median 137 ms per decision (p95 196 ms, n = 246 calls, client wall time with the network inside it) and cost $0.0337 per 1,000 decisions (a calculation from its reported input tokens). That is a different route from the CLI routers, so it is not a model-against-model comparison. As a calculation over 48 recorded tasks (median 49.5 model calls each), routing every call would add $1.67 per 1,000 tasks and up to 7 s of waiting per task with Jev, $247.30 with Sonnet (8.2% of the work cost) and up to 129 s of waiting per task with Sonnet; routing only the 7 System One decisions cuts Sonnet to $34.97 and 18 s. CLI start-up alone, for a one-word answer: Claude Code (Haiku 4.5) took 2.53 s and sent 6,761 input tokens; Codex CLI took 6.00 s and sent 17,051 input tokens, 13,184 of them read from the cache (5 runs each, different models).

Live story

Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.

Live story · 51 sRouting overhead: a 1.42 µs policy vs LLM routers

Routing overhead: a 1.42 µs policy vs LLM routers

A deterministic routing policy decides in 1.42 µs (median, n = 20,000); Sonnet 5.5 as a router takes 2.60 s through the CLI. Per-task costs and delays are calculations.

Transcript
  1. Routing overhead · policy vs LLM routers. What a routing decision costs before the work starts. Time and money per decision, then per task. Jev is timed over its API, a different route from the CLI routers.
  2. The latency ladder: the in-process policy decides in 1.42 µs. Jev, a direct API call, takes 137 ms. Sonnet 5.5 as a router takes 2.60 s through the CLI, Haiku 4.5 12.5 s. Chart: Time per decision · log scale · median to p95 (n = 82–20000 each). Caveat: Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
  3. The LLM router takes about 1.8 million times longer per decision. Jev costs $0.0337 per 1,000 decisions, provider-reported. Policy decision, median (20,000 timed): 1.42 µs (n = 20000, p95 2.33 µs). Sonnet 5.5 router ÷ policy, medians: 1.8 million× (n = 82). Jev 1.13 per 1,000 decisions (provider-reported): $0.0337 (n = 82, 137 ms median per call over its API). Caveat: OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
  4. Of Sonnet 5.5’s 2.60 s per decision, a median 973 ms is CLI and harness time, not the model. Chart: Where an LLM router’s time goes · median per call (n = 82 each). Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
  5. Route all 49.5 model calls per task: $1.67 per 1,000 tasks with Jev, $247.30 with Sonnet 5.5. Route 7 decisions: $34.97. Chart: Added routing cost per 1,000 tasks. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
  6. If each decision waits in line, Sonnet 5.5 adds up to 129 s per task, or 18.2 s for 7 decisions. The policy adds 70.3 µs. Chart: Added routing delay per task · upper bound. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
  7. Start-up tax for a one-word answer: Claude Code 2.53 s and 6,761 input tokens; Codex CLI 6.00 s and 17,051 input tokens, 13,184 of them read from the cache. Chart: CLI start-up tax · one-word answer · median of 5 runs (n = 5 each). Caveat: CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
  8. Route in process where a rule is enough. Every timing, cost and gap online.

Key numbers

about 1.8

Median LLM router call ÷ median policy decision

million times · n = 82

1ms

Recorded rule-based System One decision, record write included

median, 2 ms p95 (millisecond resolution) · n = 419

49.5

Model calls per task (each one a routing decision)

median (13 to 73) · n = 48

137ms

Jev 1.13 (TypeSafe): median decision time over the API

p95 196 ms · n = 246

about 19

Median Claude Sonnet 5.5 call through the CLI ÷ median Jev call over the API

times · n = 82

about 92

Median Claude Haiku 4.5 call through the CLI ÷ median Jev call over the API

times · n = 82

1,690ms

Claude Code time outside the model on a one-word answer

median (1,533 to 1,811) · n = 5

Routing hub: Jev, LLM routers, the deterministic policy and gateways on one page

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

Time per decision · log scale: each gridline is 10 times the one before

4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–20000 per row

Median; whiskers = median to 95th percentile

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
  • Model API time
  • CLI and harness time
Entrance: medians race at 7.5× real timeMotion reduced: press Replay to animateThe slowest median is 10.51 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

2 rows, 2 series: Model API time, CLI and harness time. Model API time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 10.51 s (median to p95 10.51 s–32.13 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 1.6 s (median to p95 1.6 s–2.58 s, n 82). Not all run ranges overlap. CLI and harness time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 1.7 s (median to p95 1.7 s–2.68 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 973 ms (median to p95 973 ms–1.28 s, n 82). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n = 82 per row

Median per call; whiskers = median to 95th percentile

Model API time is the API duration the CLI reports; CLI and harness time is wall time minus that, per call. Medians of the parts do not add up to the median of the whole. Jev is not split: its API reports no server time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing

Share card (PNG)
Every rate is 95% or more
Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

Every interval overlaps every other: this chart does not order these rows.

4 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln 82–20000 per row4 of 4 at 100%: this task set cannot separate them.

Completed calls ÷ calls; whiskers = 95% Wilson interval

A completed call returned a decision, right or wrong (accuracy is in the routing study). Whiskers are 95% Wilson intervals.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)

2 rows. Highest Jev 1.13 (TypeSafe) $0.034 (n 82). Lowest Deterministic routing policy (Agent, in process) $0 (n 20000).

Notesn 82–20000 per row

USD per 1,000 decisions

The policy makes no model call, so it costs nothing per decision. Jev’s figure is the cost its provider reported for the recorded production run (82 decisions). The input tokens of the live run give the same figure at the published price. The Claude routers are in the next chart: their cost is derived from list prices.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev 1.13 list price

Share card (PNG)
Calculation
Largest value is 260x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 (thinking on, via Claude Code) $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).

Notesn 82–246 per row

List price × the tokens each route reported, USD per 1,000 decisions

A calculation, not a bill: the Claude calls ran on a subscription. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free). Same tokens and prices as the routing study.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
Calculation
  • Every model call routed (49.5 per task)
  • Only System One decisions (7 per task) (square)
In chart order.
Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).

List-price calculation, not a run. 4 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $442. Lowest Deterministic routing policy (Agent, in process) $0. Only System One decisions (7 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $62.47. Lowest Deterministic routing policy (Agent, in process) $0.

Notes

Decisions per task from recorded runs × cost per decision

A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
Calculation
  • Every model call routed (49.5 per task)
  • Only System One decisions (7 per task) (square)
In chart order.
Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).

List-price calculation, not a run. 4 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 621 s. Fastest Deterministic routing policy (Agent, in process) 70.3 µs. Only System One decisions (7 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 87.8 s. Fastest Deterministic routing policy (Agent, in process) 9.9 µs.

Notes

Decisions per task × median decision time, if every decision waits in line

A calculation and an upper bound: it assumes each decision waits for the one before. Median recorded task wall time: 10.3 min. Jev’s delay uses its live median over the API from one Mac (network included); the Claude routers’ includes the CLI.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
  • First output event
  • First model output
  • Total wall time
Entrance: medians race at 4.3× real timeMotion reduced: press Replay to animateThe slowest median is 6 s. The clock runs at the recorded speed.
Claude Code · Claude Haiku 4.5
Codex CLI (default model)

2 rows, 3 series: First output event, First model output, Total wall time. First output event: slowest Claude Code · Claude Haiku 4.5 563 ms (range 519 ms–726 ms, n 5). Fastest Codex CLI (default model) 489 ms (range 354 ms–1.3 s, n 5). All run ranges overlap. First model output: slowest Codex CLI (default model) 5.06 s (range 4.39 s–5.48 s, n 5). Fastest Claude Code · Claude Haiku 4.5 1.46 s (range 1.21 s–2.31 s, n 5). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 5 per row

Median of 5 runs; whiskers = fastest and slowest run

Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.

Source: Routing overhead runs: policy microbenchmark and CLI start-up

Share card (PNG)
Claude Code · Claude Haiku 4.5
Codex CLI (default model)

2 rows. Highest Codex CLI (default model) 17,051 (n 5). Lowest Claude Code · Claude Haiku 4.5 6,761 (n 5).

Notesn = 5 per row

Per call, mostly the CLI’s own system prompt and tool definitions

Claude Code sums its disjoint input, cache-read and cache-write fields. Codex CLI reports 17,051 input tokens, 13,184 of them read from the cache. The prompt itself is a few tokens.

Source: Routing overhead runs: policy microbenchmark and CLI start-up

Share card (PNG)

Tables

What was measured, recorded, calculated or not measured

RouterKindLatencyCost
Deterministic routing policy (Agent, in process)in-process rulesmeasured 2026-10-06: 1.42 µs median$0 (no model call)
Jev 1.13 (TypeSafe)hosted decision modelmeasured 2026-10-06: 137 ms median, p95 196 ms (246 calls over the API from one Mac; client wall time, no server time)$0.0337 per 1,000 (list-price calculation from input tokens; the recorded run’s own cost figure agrees)
Claude Sonnet 5.5 (effort low, via Claude Code)LLM router via agent CLIrecorded: 2.60 s median (82 calls)$4.996 per 1,000 (list-price calculation)
Claude Haiku 4.5 (thinking on, via Claude Code)LLM router via agent CLIrecorded: 12.54 s median (82 calls)$8.924 per 1,000 (list-price calculation)
OpenRouter Auto Router / cheaper hosted inferencehosted LLM router / gatewayNot measured: no key in the environment. The harness is ready and runs when a key is set.Not measured: no key in the environment. The harness is ready and runs when a key is set.
Clef / Clef-Flash (local)local router modelNot measured: no local server running.Not measured: no local server running.

Jev live run: time per call and cost

MeasureValue
Counted calls246 (3 repeats of the same 82 decisions, one at a time), 0 failed
Median time per call136.5 ms
90th percentile171.0 ms
95th percentile195.7 ms
Fastest and slowest call100.9 ms and 297.3 ms
Cold first call (new process, fresh connection)224.7 ms
Median per repeat130.3 ms, 141.7 ms, 137.1 ms (the first repeat without the cold call)
Input and output tokens per decision (mean)803 and 148 (output tokens are free at the published price)
Cost per 1,000 decisions (calculation)$0.0337 = 803 input tokens × 1,000 × $0.042 per million
Server-side timenot available: the API sends no timing header or field

Method

  1. Protocol declared before any measurement.
  2. Deterministic policy: the platform’s production routing decision (plan, model ladder and effort) on its default policy, timed in process on one Apple M3 Ultra Mac with Node 25: 5,000 warm-up calls, then 20,000 timed decisions over 64 synthetic routing contexts, plus a 200,000-call batch for throughput. Database reads and the decision record write are out of scope; the recorded rule-based System One decisions (millisecond resolution, record write included) are shown as a stat.
  3. LLM routers: the recorded routing runs of the routing study (Claude Sonnet 5.5 (effort low, via Claude Code), 82 calls; Claude Haiku 4.5 (thinking on, via Claude Code), 82 calls), one call at a time. Wall time per call, the API time the CLI reported, and the difference (CLI and harness time). p50 and p95 recomputed from the per-call log.
  4. Jev: a live run on 2026-10-06. The same 82 typed decisions as the routing study, 3 repeats, 246 counted calls sent one at a time over HTTPS to the TypeSafe API from the same Mac, on a home network. Time per call is client wall time from before the request to after the body was read, so the network is inside it; the API reports no server time. Cost per decision is a calculation: the mean input tokens its API reported × the published price. The recorded production run (82 decisions) has the same cost as a provider-reported figure.
  5. CLI start-up: Claude Code · Claude Haiku 4.5 5 runs, Codex CLI (default model) 5 runs, a one-word prompt, one call at a time. Times to the first output event, the first model output and exit.
  6. Per 1,000 tasks (a calculation): decisions per task from 48 recorded bench runs read only (2 empty runs excluded) × cost and median time per decision. Decisions are assumed to wait in line, so the delay is an upper bound.

Caveats

  • The policy is timed in process and the LLM routers through a CLI: this compares the two ways of routing as deployed, not two models on equal footing. A direct API call would skip the CLI time (shown separately).
  • The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
  • Haiku 4.5 ran with the CLI’s default thinking, which makes it slower than Sonnet 5.5 at effort low here.
  • OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
  • CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
  • Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
  • Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
  • Jev’s timing is one 35-second window on 2026-10-06, 246 calls from one machine; server load at that time is unknown. The 246 calls are 3 repeats of 82 requests, so they are not independent draws.
  • Corrected 2026-10-07: an earlier version of this page counted the cached tokens twice (30,235). The CLI reports 17,051 input tokens including 13,184 read from the cache.

Sources

  • Routing overhead runs: policy microbenchmark and CLI start-up

    Our recorded runs ·

    In-process timing of the deterministic routing policy (20,000 timed decisions), CLI start-up with a one-word prompt (5 runs per CLI), and decision counts read from recorded bench runs. LLM router timings are reused from the routing runs.

    Raw data: routing-overhead/results.json

  • Routing overhead per 1,000 tasks (calculation)

    Calculation ·

    Decisions per task from recorded bench runs multiplied by the cost and the median time per decision. Decisions are assumed to wait in line, so the delay is an upper bound. A calculation, not a run.

    Raw data: routing-overhead/results.json

  • Routing runs: Jev router vs LLM routing

    Our recorded runs ·

    Routing decisions recorded per case and arm.

    Raw data: routing/receipts.json

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • Jev 1.13 list price

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free.

  • Jev live run: 246 timed calls on the 82 routing decisions

    Our recorded runs ·

    Jev 1.13 called over HTTPS, 3 repeats of the same 82 typed decisions, one call at a time, from one Mac over a home network: client wall time, with the network inside it. The API reports no server time. Cost per 1,000 decisions is a calculation from the reported input tokens and the published price. The case sets were revised against Jev answers, so Jev has a home advantage.

    Raw data: jev-live/summary.json, jev-live/calls.json

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Routing overhead: deterministic policy vs LLM routers vs Jev”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/routing-overhead.

Explainers that cite this study

Read the methods and terms in the context of these recorded results.

More write-ups that cite this study (2)

Models and comparisons in this study

More studies

All benchmarks
  • Claude Haiku
  • Extended Thinking

Does thinking pay for Claude Haiku 4.5? Thinking on vs off

Claude Haiku 4.5 with extended thinking on and off: 82 routing decisions and 8 hard tasks. Accuracy with 95% intervals, time and cost.

87% (71/82)Claude Haiku 4.5 (thinking off): exact routing decisions · n = 82

6 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.