Routing overhead: deterministic policy vs LLM routers vs Jev
How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.
4 studies · 8 comparisons · updated October 6, 2026
A router decides which model and effort each call gets. It can be a rule in the same process, a hosted decision model such as Jev, a large model asked for its opinion, or a gateway such as OpenRouter that sends one API to many providers. Here is what each one adds in time, cost and accuracy, with the sample size behind every number.
Motion reduced: press Replay zoom to animate
1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.
about 96 thousand×Jev 1.13 (TypeSafe) takes about 96 thousand times as long as Deterministic routing policy (Agent, in process). Calculation: ratio of the two medians (137 ms ÷ 1.42 µs). Log scale: each gridline is 10 times the one before.
| Item | Decision time | Median to p95 | n |
|---|---|---|---|
| Deterministic routing policy (Agent, in process) | 1.42 µs | 1.42 µs–2.33 µs | 20000 |
| Jev 1.13 (TypeSafe) | 137 ms | 137 ms–196 ms | 246 |
| Claude Sonnet 5.5 (effort low, via Claude Code) | 2.6 s | 2.6 s–4.3 s | 82 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 12.54 s | 12.54 s–34.48 s | 82 |
4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.
Median; whiskers = median to 95th percentile
The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
Rules in process
Microseconds and no tokens. They only know what the rules know.
Decision models and LLM routers
They read the task. Each decision costs a model call: seconds through a CLI.
Gateways and cheaper inference
One API, many providers. Closed models cost the same per token everywhere; open-weight prices vary up to 12.6x between providers.
Drawn live in the page from the same data. Play one, or download it as a video from the player.
A deterministic routing policy decides in 1.42 µs (median, n = 20,000); Sonnet 5.5 as a router takes 2.60 s through the CLI. Per-task costs and delays are calculations.
Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.
Exact decisions and key accuracy on the same 82 routing cases, with 95% Wilson intervals. Overlapping intervals mean the sample cannot separate the routers.
Every interval overlaps every other: this chart does not order these rows.
| Item | Exact rate | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 90% | 82%–95% | 82 |
| Claude Haiku 4.5 | 89% | 80%–94% | 82 |
| Claude Sonnet 5.5 | 94% | 87%–97% | 82 |
3 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.
Share of asked cases where every scored question was acceptable
Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.
Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
Every interval overlaps every other: this chart does not order these rows.
| Item | Key accuracy | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 95% | 91%–97% | 194 |
| Claude Haiku 4.5 | 94% | 90%–97% | 194 |
| Claude Sonnet 5.5 | 97% | 94%–99% | 194 |
3 rows. Highest Claude Sonnet 5.5 97% (95% interval 94%–99%, n 194). Lowest Claude Haiku 4.5 94% (95% interval 90%–97%, n 194). All intervals overlap.
Each open question the router was asked; an unanswered question counts as wrong
Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.
Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
Each case was routed by both. Cases both got right, or both got wrong, say nothing about which is better; the exact McNemar test reads only the cases where they disagree.
Every case once: both right, only one right, both wrong
Only the 7 discordant decisions count: 3 vs 4. Exact McNemar p = 1: no evidence of a difference.
One square per paired decision (n = 82). Squares start in one grid and sort into the four outcomes; their order inside a quadrant carries no meaning.
| Pair | Both right | Only first right | Only second right | Both wrong | Exact McNemar p |
|---|---|---|---|---|---|
| Claude Haiku 4.5 vs Jev 1.13 (TypeSafe) | 70 | 3 | 4 | 5 | 1 |
| Claude Sonnet 5.5 vs Jev 1.13 (TypeSafe) | 73 | 4 | 1 | 4 | 0.375 |
Counts of cases from the table; the exact McNemar p reads only the discordant casesn = 82 cases per pair
Claude Haiku 4.5 vs Jev 1.13 (TypeSafe): 3 vs 4 discordant cases of 82, exact McNemar p = 1. Claude Sonnet 5.5 vs Jev 1.13 (TypeSafe): 4 vs 1 discordant cases of 82, exact McNemar p = 0.375.
| Item | Cost per 1,000 decisions | n |
|---|---|---|
| Deterministic routing policy (Agent, in process) | $0 | 20000 |
| Jev 1.13 (TypeSafe) | $0.034 | 82 |
2 rows. Highest Jev 1.13 (TypeSafe) $0.034 (n 82). Lowest Deterministic routing policy (Agent, in process) $0 (n 20000).
USD per 1,000 decisions
The policy makes no model call, so it costs nothing per decision. Jev’s figure is the cost its provider reported for the recorded production run (82 decisions). The input tokens of the live run give the same figure at the published price. The Claude routers are in the next chart: their cost is derived from list prices.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev 1.13 list price
Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the highlighted row): a ratio of list-price calculations, not a measurement.
| Item | Cost | n |
|---|---|---|
| Jev 1.13 (TypeSafe) | $0.034 | 246 |
| Claude Haiku 4.5 | $8.92 | 82 |
| Claude Sonnet 5.5 | $5.00 | 82 |
List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).
List price × reported tokens per decision
List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.
Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions
Where a routed call’s time goes, then calculations from recorded decision counts per task (not runs). The delay is an upper bound: it assumes each decision waits for the one before.
| Item | Model API time | CLI and harness time | Median to p95 | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 (effort low, via Claude Code) | 1.6 s | 973 ms | Model API time: 1.6 s–2.58 s; CLI and harness time: 973 ms–1.28 s | 82 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 10.51 s | 1.7 s | Model API time: 10.51 s–32.13 s; CLI and harness time: 1.7 s–2.68 s | 82 |
2 rows, 2 series: Model API time, CLI and harness time. Model API time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 10.51 s (median to p95 10.51 s–32.13 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 1.6 s (median to p95 1.6 s–2.58 s, n 82). Not all run ranges overlap. CLI and harness time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 1.7 s (median to p95 1.7 s–2.68 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 973 ms (median to p95 973 ms–1.28 s, n 82). Not all run ranges overlap.
Median per call; whiskers = median to 95th percentile
Model API time is the API duration the CLI reports; CLI and harness time is wall time minus that, per call. Medians of the parts do not add up to the median of the whole. Jev is not split: its API reports no server time. The whisker is the median to the 95th percentile, not a confidence interval.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing
Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).
| Item | Every model call routed (49.5 per task) | Only System One decisions (7 per task) |
|---|---|---|
| Deterministic routing policy (Agent, in process) | $0 | $0 |
| Jev 1.13 (TypeSafe) | $1.67 | $0.24 |
| Claude Sonnet 5.5 (effort low, via Claude Code) | $247 | $34.97 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | $442 | $62.47 |
List-price calculation, not a run. 4 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $442. Lowest Deterministic routing policy (Agent, in process) $0. Only System One decisions (7 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $62.47. Lowest Deterministic routing policy (Agent, in process) $0.
Decisions per task from recorded runs × cost per decision
A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions
Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).
| Item | Every model call routed (49.5 per task) | Only System One decisions (7 per task) |
|---|---|---|
| Deterministic routing policy (Agent, in process) | 70.3 µs | 9.9 µs |
| Jev 1.13 (TypeSafe) | 6.8 s | 1 s |
| Claude Sonnet 5.5 (effort low, via Claude Code) | 129 s | 18.2 s |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 621 s | 87.8 s |
List-price calculation, not a run. 4 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 621 s. Fastest Deterministic routing policy (Agent, in process) 70.3 µs. Only System One decisions (7 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 87.8 s. Fastest Deterministic routing policy (Agent, in process) 9.9 µs.
Decisions per task × median decision time, if every decision waits in line
A calculation and an upper bound: it assumes each decision waits for the one before. Median recorded task wall time: 10.3 min. Jev’s delay uses its live median over the API from one Mac (network included); the Claude routers’ includes the CLI.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions
| Item | First output event | First model output | Total wall time | Range (lowest–highest run) | n |
|---|---|---|---|---|---|
| Claude Code · Claude Haiku 4.5 | 563 ms | 1.46 s | 2.53 s | First output event: 519 ms–726 ms; First model output: 1.21 s–2.31 s; Total wall time: 2.27 s–3.38 s | 5 |
| Codex CLI (default model) | 489 ms | 5.06 s | 6 s | First output event: 354 ms–1.3 s; First model output: 4.39 s–5.48 s; Total wall time: 5.37 s–6.51 s | 5 |
2 rows, 3 series: First output event, First model output, Total wall time. First output event: slowest Claude Code · Claude Haiku 4.5 563 ms (range 519 ms–726 ms, n 5). Fastest Codex CLI (default model) 489 ms (range 354 ms–1.3 s, n 5). All run ranges overlap. First model output: slowest Codex CLI (default model) 5.06 s (range 4.39 s–5.48 s, n 5). Fastest Claude Code · Claude Haiku 4.5 1.46 s (range 1.21 s–2.31 s, n 5). Not all run ranges overlap.
Median of 5 runs; whiskers = fastest and slowest run
Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.
Source: Routing overhead runs: policy microbenchmark and CLI start-up
Scale routing cost and waiting to your own decision volume
Hover or focus a bar for its ratio to all Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.
| Item | Repriced cost |
|---|---|
| all Fable 5.1 | $370 |
| all Opus 5.5 | $171 |
| policy (Opus strong, Haiku ancillary) | $162 |
| all Sonnet 5.5 | $109 |
| split (Sonnet main line, Haiku ancillary) | $106 |
| all Haiku 4.5 | $54.27 |
List-price calculation, not a run. 6 rows. Highest all Fable 5.1 $370. Lowest all Haiku 4.5 $54.27.
50 benchmark runs, 2,362 model calls, repriced
Calculation, not a run: every recorded call ran on Sonnet 5.5 with routing off. Same tokens on every model; a different model or mix would take a different path.
Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Anthropic list prices (Claude models)
Haiku 4.5
Sonnet 5.5
Opus 5.5
Fable 5.1
One panel per series, all on the same axis.
| Item | Haiku 4.5 | Sonnet 5.5 | Opus 5.5 | Fable 5.1 |
|---|---|---|---|---|
| act (strong) | $28.66 | $57.31 | $78.86 | $152 |
| research (strong) | $14.31 | $28.62 | $47.24 | $106 |
| verify (strong) | $4.74 | $9.48 | $18.90 | $47.18 |
| review (strong) | $3.36 | $6.71 | $13.21 | $32.78 |
| memory and onboarding (economy) | $3.11 | $6.22 | $12.32 | $30.65 |
| other (standard) | $0.1 | $0.19 | $0.39 | $0.96 |
List-price calculation, not a run. 6 rows, 4 series: Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1. Haiku 4.5: highest act (strong) $28.66. Lowest other (standard) $0.1. Sonnet 5.5: highest act (strong) $57.31. Lowest other (standard) $0.19.
Calculation, not a run. The tier in brackets is the routing policy tier for that stage.
Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Anthropic list prices (Claude models)
Unknown stays unknown. A router without a measured number says why.
Exact decisions with the 95% interval, tokens, cost and decision time
—no accuracy run
74/82exact decisions
90% · 95% CI 82%–95% · n = 82
live API run 2026-10-06: 3 repeats of the same 82 decisions, 246 counted calls one at a time from one Mac. All calls: 221 of 246 exact and 552 of 582 questions; the counts shown are on the 82-decision scale of the interval. The recorded production run of 2026-10-05 scored 74 of 82. Cost is a calculation from the reported input tokens
77/82exact decisions
94% · 95% CI 87%–97% · n = 82
this benchmark, Claude Code CLI on a subscription account, one call per decision
73/82exact decisions
89% · 95% CI 80%–94% · n = 82
this benchmark, Claude Code CLI on a subscription account, one call per decision
—no accuracy run
—exact decisions: not measured
not measured: No local Clef server was running and installing a 6-20 GB model was out of scope for this run.
| Router | Kind | Decision time | Cost |
|---|---|---|---|
| Deterministic routing policy (Agent, in process) | in-process rules | Measuredmeasured 2026-10-06: 1.42 µs median | Measured$0 (no model call) |
| Jev 1.13 (TypeSafe) | hosted decision model | Measuredmeasured 2026-10-06: 137 ms median, p95 196 ms (246 calls over the API from one Mac; client wall time, no server time) | Calculation$0.0337 per 1,000 (list-price calculation from input tokens; the recorded run’s own cost figure agrees) |
| Claude Sonnet 5.5 (effort low, via Claude Code) | LLM router via agent CLI | Measuredrecorded: 2.60 s median (82 calls) | Calculation$4.996 per 1,000 (list-price calculation) |
| Claude Haiku 4.5 (thinking on, via Claude Code) | LLM router via agent CLI | Measuredrecorded: 12.54 s median (82 calls) | Calculation$8.924 per 1,000 (list-price calculation) |
| OpenRouter Auto Router / cheaper hosted inference | hosted LLM router / gateway | Not measuredNot measured: no key in the environment. The harness is ready and runs when a key is set. | Not measuredNot measured: no key in the environment. The harness is ready and runs when a key is set. |
| Clef / Clef-Flash (local) | local router model | Not measuredNot measured: no local server running. | Not measuredNot measured: no local server running. |
Meter: 95% Wilson interval on exact decisions (82 cases)Each value is measured, reported, a calculation or not measured, as marked
6 routers. Decision time and cost are marked measured, reported, a calculation or not measured; exact decisions carry their 95% interval.
Each tile opens its study. The small text is the sample size and the 95% interval where one exists.
Routing overhead: deterministic policy vs LLM routers vs Jev
1.42µs
Deterministic routing policy: median decision time
p95 2.33 µs, p99 3.04 µs · n = 20000
Routing overhead: deterministic policy vs LLM routers vs Jev
about 1.8
Median LLM router call ÷ median policy decision
million times · n = 82
Jev vs Claude as a router: accuracy and cost
90% (74/82)
Jev 1.13 (TypeSafe): exact decisions
95% CI 82%–95% · n = 82
Jev vs Claude as a router: accuracy and cost
94% (77/82)
Claude Sonnet 5.5: exact decisions
95% CI 87%–97% · n = 82
Jev vs Claude as a router: accuracy and cost
$0.0337
Jev cost per 1,000 decisions
n = 246
Routing overhead: deterministic policy vs LLM routers vs Jev
49.5
Model calls per task (each one a routing decision)
median (13 to 73) · n = 48
Inference provider index: 27 models, 52 providers
5.5%
Credit-purchase fee on OpenRouter’s Standard plan (third-party-reported)
($0.80 minimum by card)
Inference provider index: 27 models, 52 providers
12.6x
Largest standard-tier price spread (DeepSeek V4 Flash 0423)
(most expensive: Cloudflare; cheapest: StreamLake (fp8)) · n = 15
Reported by OpenRouter’s public API. The gateway’s own delay is not measured: no key in the environment.
0%per-token markup on 10 of 10 models
5.5% with the card credit feeCalculation
| Item | Per-token markup (input) | Per-token markup (output) | With the 5.5% card credit fee (input) |
|---|---|---|---|
| Claude Haiku 4.5 | 0% | 0% | 5.5% |
| Claude Sonnet 5 | 0% | 0% | 5.5% |
| Claude Sonnet 5.5 | 0% | 0% | 5.5% |
| Claude Opus 4.8 | 0% | 0% | 5.5% |
| Claude Opus 5 | 0% | 0% | 5.5% |
| Claude Opus 5.5 | 0% | 0% | 5.5% |
| Claude Fable 5.1 | 0% | 0% | 5.5% |
| GPT-6 Luna | 0% | 0% | 5.5% |
| Gemini 3.8 Flash | 0% | 0% | 5.5% |
| Gemini 3.5 Flash | 0% | 0% | 5.5% |
List-price calculation, not a run. 10 rows, 3 series: Per-token markup (input), Per-token markup (output), With the 5.5% card credit fee (input). Per-token markup (input): all at 0%. Per-token markup (output): all at 0%.
Per-token markup, and the markup after the Standard credit-purchase fee (calculation)
A calculation: OpenRouter list price ÷ first-party list price − 1, then × (1 + 5.5%) for credits bought by card on the Standard plan (purchases large enough that the minimum fee does not apply). Fees as published on 2026-10-06; first-party prices from the vendors’ list prices.
Sources: OpenRouter public API: models and provider endpoints (snapshot), OpenRouter pricing and fees, Anthropic list prices (Claude models), OpenAI list prices, Google Gemini list prices
| Item | Price spread | n |
|---|---|---|
| DeepSeek V4 Flash 0423 | 12.6x | 15 |
| DeepSeek V4 Pro 0423 | 11.2x | 15 |
| gpt-oss-120b | 6.9x | 20 |
| Llama 3.3 70B Instruct | 6.7x | 10 |
| GLM 5.3 | 4.7x | 32 |
| Kimi K3 | 1.8x | 19 |
| Llama 4 Maverick | 1.7x | 3 |
| Claude Haiku 4.5 | 1x | 4 |
| Claude Sonnet 5 | 1x | 5 |
| Claude Sonnet 5.5 | 1x | 5 |
| Claude Opus 4.8 | 1x | 5 |
| Claude Opus 5 | 1x | 5 |
| Claude Opus 5.5 | 1x | 5 |
| Claude Fable 5.1 | 1x | 4 |
| GPT-6 Sol | 1x | 2 |
| GPT-6 Luna | 1x | 2 |
| GPT-6 Astra | 1x | 2 |
| GPT-5.5 | 1x | 2 |
| Gemini 3.8 Flash | 1x | 2 |
| Gemini 3.5 Flash | 1x | 2 |
| Gemini 3.5 Flash Lite | 1x | 2 |
| Gemini 3.1 Pro Preview | 1x | 2 |
List-price calculation, not a run. 22 rows. Highest DeepSeek V4 Flash 0423 12.6x (n 15). Lowest Gemini 3.1 Pro Preview 1x (n 2).
Standard tier, blended price (3 input : 1 output), most expensive provider ÷ cheapest provider
A calculation on prices reported by OpenRouter’s public API, snapshot 2026-10-06. n = providers with a standard-tier endpoint. 1x means every provider charges the same. Highlighted: a spread of 4x or more.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Every provider, priced by model Price a workload via the cheapest provider
Router · Agent
Agent’s rule-based routing: an in-process policy picks the model and effort for each call from the task stage and signals. No model call, so no token cost.
Router · TypeSafe
A small routing model from TypeSafe that answers typed routing decisions over an API. Priced on input tokens only. Measured here as a router (accuracy, live time per call and cost per 1,000 decisions) and in the System One arena.
Gateway · OpenRouter
A gateway that routes one API to many inference providers. Its per-token price is compared with first-party list prices; its fee is charged when credits are bought.
A side wins a row only when the 95% intervals or run ranges do not overlap. Price rows never name a winner. The bar shows wins, ties and unclear rows.
Jev 1.13 ahead on 2 · 15 ties · 6 unclear
Jev 1.13 ahead on 2 · 15 ties · 6 unclear
Deterministic routing policy ahead on 1 · 1 tie · 5 unclear
Deterministic routing policy ahead on 1 · 1 tie · 4 unclear
Deterministic routing policy ahead on 1 · 1 tie · 4 unclear
14 ties
4 ties
2 ties
How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.
Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.
Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.
194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.
136.5 ms for Jev, 0.82 s for a small-model API, 2.79 s to 3.79 s for Codex CLI: which steps fit a voice agent latency budget? A thought experiment.
Rules and Jev 1.13 fit every budget we assumed; a Claude router through a CLI fits none. 14 measured steps vs 300, 800 and 1,500 ms. A thought experiment.
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.