{"i":11,"study":{"slug":"routing-overhead","title":"Routing overhead: deterministic policy vs LLM routers vs Jev","seoTitle":"Routing overhead: rules vs LLM routers vs Jev","description":"How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.","question":"What delay and what cost does each kind of router add before the real work of a call starts?","answer":"The deterministic routing policy decided in a median 1.42 µs (p95 2.33 µs, 20,000 decisions, $0). The fastest LLM router, Claude Sonnet 5.5 through the Claude Code CLI, took a median 2.60 s per decision (p95 4.30 s, n = 82), about 1.8 million times longer; 973 ms of that median was CLI time, not model time. Haiku 4.5 with its default thinking took 12.54 s (p95 34.48 s). Jev 1.13, called directly over HTTPS from the same Mac, took a median 137 ms per decision (p95 196 ms, n = 246 calls, client wall time with the network inside it) and cost $0.0337 per 1,000 decisions (a calculation from its reported input tokens). That is a different route from the CLI routers, so it is not a model-against-model comparison. As a calculation over 48 recorded tasks (median 49.5 model calls each), routing every call would add $1.67 per 1,000 tasks and up to 7 s of waiting per task with Jev, $247.30 with Sonnet (8.2% of the work cost) and up to 129 s of waiting per task with Sonnet; routing only the 7 System One decisions cuts Sonnet to $34.97 and 18 s. CLI start-up alone, for a one-word answer: Claude Code (Haiku 4.5) took 2.53 s and sent 6,761 input tokens; Codex CLI took 6.00 s and sent 17,051 input tokens, 13,184 of them read from the cache (5 runs each, different models).","date":"2026-10-06","updated":"2026-10-06","tags":["routing","latency","overhead","jev","llm-router","cli","calculation"],"caveats":["The policy is timed in process and the LLM routers through a CLI: this compares the two ways of routing as deployed, not two models on equal footing. A direct API call would skip the CLI time (shown separately).","The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.","Haiku 4.5 ran with the CLI’s default thinking, which makes it slower than Sonnet 5.5 at effort low here.","OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.","CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.","Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.","Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.","Jev’s timing is one 35-second window on 2026-10-06, 246 calls from one machine; server load at that time is unknown. The 246 calls are 3 repeats of 82 requests, so they are not independent draws.","Corrected 2026-10-07: an earlier version of this page counted the cached tokens twice (30,235). The CLI reports 17,051 input tokens including 13,184 read from the cache."],"sourceIds":["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"],"stats":{"$k":["id","label","value","unit","display","n","note"],"$r":[["router-overhead-policy-p50","Deterministic routing policy: median decision time",0.00142,"ms","1.42 µs (p95 2.33 µs, p99 3.04 µs)",20000,"Timer resolution 0.041 µs; 469,409 decisions per second in a 200,000-call batch."],["router-overhead-speedup","Median LLM router call ÷ median policy decision",1828873,"ratio","about 1.8 million times",82,"\u0001"],["router-overhead-system-one-rule-arm","Recorded rule-based System One decision, record write included",1,"ms","1 ms median, 2 ms p95 (millisecond resolution)",419,"\u0001"],["router-overhead-decisions-per-task","Model calls per task (each one a routing decision)",49.5,"calls","49.5 median (13 to 73)",48,"\u0001"],["router-overhead-jev-p50","Jev 1.13 (TypeSafe): median decision time over the API",136.5,"ms","137 ms (p95 196 ms)",246,"Client wall time from one Mac over a home network, 246 calls in a 35-second window. The API reports no server time."],["router-overhead-sonnet-vs-jev","Median Claude Sonnet 5.5 call through the CLI ÷ median Jev call over the API",19,"ratio","about 19 times",82,"A calculation across two routes (CLI vs a direct API call from one Mac), not a model-against-model comparison."],["router-overhead-haiku-vs-jev","Median Claude Haiku 4.5 call through the CLI ÷ median Jev call over the API",91.9,"ratio","about 92 times",82,"A calculation across two routes (CLI vs a direct API call from one Mac), not a model-against-model comparison."],["cli-startup-claude-harness-ms","Claude Code time outside the model on a one-word answer",1690,"ms","1,690 ms median (1,533 to 1,811)",5,"\u0001"]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","whisker","series","note","sourceIds"],"$r":[["router-overhead-decision-latency","Time to make one routing decision","Median; whiskers = median to 95th percentile","dot-range","ms","Time per decision","p50-p95",[{"name":"Decision time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Deterministic routing policy (Agent, in process)",0.00142,0.00142,0.00233,20000,true],["Jev 1.13 (TypeSafe)",136.5,136.5,195.7,246,"\u0001"],["Claude Sonnet 5.5 (effort low, via Claude Code)",2597,2597,4298,82,"\u0001"],["Claude Haiku 4.5 (thinking on, via Claude Code)",12543,12543,34481,82,"\u0001"]]}}],"The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.",["agent-routing-overhead","agent-routing","agent-jev-live"]],["router-overhead-cli-vs-model-time","Where an LLM router’s time goes: model vs CLI","Median per call; whiskers = median to 95th percentile","dot-range","ms","Time per call","p50-p95",[{"name":"Model API time","points":[{"label":"Claude Sonnet 5.5 (effort low, via Claude Code)","value":1596,"lo":1596,"hi":2583,"n":82},{"label":"Claude Haiku 4.5 (thinking on, via Claude Code)","value":10508,"lo":10508,"hi":32132,"n":82}]},{"name":"CLI and harness time","points":[{"label":"Claude Sonnet 5.5 (effort low, via Claude Code)","value":973,"lo":973,"hi":1277,"n":82},{"label":"Claude Haiku 4.5 (thinking on, via Claude Code)","value":1698,"lo":1698,"hi":2677,"n":82}]}],"Model API time is the API duration the CLI reports; CLI and harness time is wall time minus that, per call. Medians of the parts do not add up to the median of the whole. Jev is not split: its API reports no server time. The whisker is the median to the 95th percentile, not a confidence interval.",["agent-routing-overhead","agent-routing"]],["router-overhead-completed","Routing calls that returned a decision","Completed calls ÷ calls; whiskers = 95% Wilson interval","dot-range","rate","Completed","ci95",[{"name":"Completed","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Deterministic routing policy (Agent, in process)",1,0.9998,1,20000,true],["Jev 1.13 (TypeSafe)",1,0.9846,1,246,"\u0001"],["Claude Sonnet 5.5 (effort low, via Claude Code)",1,0.9552,1,82,"\u0001"],["Claude Haiku 4.5 (thinking on, via Claude Code)",1,0.9552,1,82,"\u0001"]]}}],"A completed call returned a decision, right or wrong (accuracy is in the routing study). Whiskers are 95% Wilson intervals.",["agent-routing-overhead","agent-routing","agent-jev-live"]],["router-overhead-cost-reported","Cost per 1,000 routing decisions: no model call vs provider-reported","USD per 1,000 decisions","bar","usd","USD per 1,000 decisions","\u0001",[{"name":"Cost per 1,000 decisions","points":[{"label":"Deterministic routing policy (Agent, in process)","value":0,"n":20000,"highlight":true},{"label":"Jev 1.13 (TypeSafe)","value":0.0337,"n":82,"highlight":false}]}],"The policy makes no model call, so it costs nothing per decision. Jev’s figure is the cost its provider reported for the recorded production run (82 decisions). The input tokens of the live run give the same figure at the published price. The Claude routers are in the next chart: their cost is derived from list prices.",["agent-routing-overhead","agent-routing","price-jev"]],["router-overhead-cost-list-price","Cost per 1,000 routing decisions for the model routers (calculation)","List price × the tokens each route reported, USD per 1,000 decisions","bar","usd","USD per 1,000 decisions","\u0001",[{"name":"Cost per 1,000 decisions (list price)","points":{"$k":["label","value","n"],"$r":[["Jev 1.13 (TypeSafe)",0.0337,246],["Claude Sonnet 5.5 (effort low, via Claude Code)",4.996,82],["Claude Haiku 4.5 (thinking on, via Claude Code)",8.924,82]]}}],"A calculation, not a bill: the Claude calls ran on a subscription. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free). Same tokens and prices as the routing study.",["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]],["router-overhead-cost-per-1000-tasks","Added routing cost per 1,000 tasks (calculation)","Decisions per task from recorded runs × cost per decision","grouped-bar","usd","USD per 1,000 tasks","\u0001",[{"name":"Every model call routed (49.5 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0],["Jev 1.13 (TypeSafe)",1.67],["Claude Sonnet 5.5 (effort low, via Claude Code)",247.3],["Claude Haiku 4.5 (thinking on, via Claude Code)",441.74]]}},{"name":"Only System One decisions (7 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0],["Jev 1.13 (TypeSafe)",0.24],["Claude Sonnet 5.5 (effort low, via Claude Code)",34.97],["Claude Haiku 4.5 (thinking on, via Claude Code)",62.47]]}}],"A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).",["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]],["router-overhead-delay-per-task","Added routing delay per task (calculation)","Decisions per task × median decision time, if every decision waits in line","grouped-bar","seconds","Seconds per task","\u0001",[{"name":"Every model call routed (49.5 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0.0000703],["Jev 1.13 (TypeSafe)",6.7568],["Claude Sonnet 5.5 (effort low, via Claude Code)",128.5515],["Claude Haiku 4.5 (thinking on, via Claude Code)",620.8785]]}},{"name":"Only System One decisions (7 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0.0000099],["Jev 1.13 (TypeSafe)",0.9555],["Claude Sonnet 5.5 (effort low, via Claude Code)",18.179],["Claude Haiku 4.5 (thinking on, via Claude Code)",87.801]]}}],"A calculation and an upper bound: it assumes each decision waits for the one before. Median recorded task wall time: 10.3 min. Jev’s delay uses its live median over the API from one Mac (network included); the Claude routers’ includes the CLI.",["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]],["cli-startup-tax","CLI start-up tax on a one-word answer","Median of 5 runs; whiskers = fastest and slowest run","dot-range","ms","Time to output","minmax",{"$k":["name","points"],"$r":[["First output event",[{"label":"Claude Code · Claude Haiku 4.5","value":563,"lo":519,"hi":726,"n":5},{"label":"Codex CLI (default model)","value":489,"lo":354,"hi":1304,"n":5}]],["First model output",[{"label":"Claude Code · Claude Haiku 4.5","value":1461,"lo":1206,"hi":2308,"n":5},{"label":"Codex CLI (default model)","value":5059,"lo":4391,"hi":5478,"n":5}]],["Total wall time",[{"label":"Claude Code · Claude Haiku 4.5","value":2529,"lo":2273,"hi":3382,"n":5},{"label":"Codex CLI (default model)","value":5999,"lo":5367,"hi":6506,"n":5}]]]},"Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.",["agent-routing-overhead"]],["cli-startup-input-tokens","Input tokens a CLI sends for a one-word answer","Per call, mostly the CLI’s own system prompt and tool definitions","bar","tokens","Input tokens per call","\u0001",[{"name":"Input tokens per call","points":[{"label":"Claude Code · Claude Haiku 4.5","value":6761,"n":5},{"label":"Codex CLI (default model)","value":17051,"n":5}]}],"Claude Code sums its disjoint input, cache-read and cache-write fields. Codex CLI reports 17,051 input tokens, 13,184 of them read from the cache. The prompt itself is a few tokens.",["agent-routing-overhead"]]]},"related":["routing-jev-vs-llm","cli-model-latency-tokens","inference-provider-index"]}}