Does LLM routing save money? The saving, the router and the net
Routing would save 2.8% ($3.01) on 2,362 recorded calls (a calculation). A Sonnet router on every call costs about $11.80, so the net is a loss.
TL;DR
- Short answer: on our recorded mix, routing would save 2.8% before the router's own cost. Only a near-free router keeps that saving. Dollar figures in this list are calculations, not runs. The saving uses recorded tokens. Each router's cost uses the tokens that router reported.
- The saving (calculation): on 2,362 recorded calls, Sonnet on the main line and Haiku on side jobs cost $105.53 against $108.54 for all Sonnet. The difference is $3.01, a floor (see Caveats).
- The router's cost and the net (calculation): assume one decision per call. A Sonnet router costs about $11.80, so the net is a loss of about $8.79. Haiku costs $21.08. Jev costs $0.08 and keeps $2.93. Rules cost $0.
- The break-even (calculation): a router on every call must cost less than about $1.27 per 1,000 decisions.
- The real money is on the main line (calculation): act and research hold $85.93 of $108.54 (79%). On hard tasks, Haiku passed 11/24 and Sonnet 24/24.
Studies: Jev vs Claude as a router, routing overhead, routing hub.
Routing overhead: a 1.42 µs policy vs LLM routers
A deterministic routing policy decides in 1.42 µs (median, n = 20,000); Sonnet 5.5 as a router takes 2.60 s through the CLI. Per-task costs and delays are calculations.
Transcript
- Routing overhead · policy vs LLM routers. What a routing decision costs before the work starts. Time and money per decision, then per task. Jev is timed over its API, a different route from the CLI routers.
- The latency ladder: the in-process policy decides in 1.42 µs. Jev, a direct API call, takes 137 ms. Sonnet 5.5 as a router takes 2.60 s through the CLI, Haiku 4.5 12.5 s. Chart: Time per decision · log scale · median to p95 (n = 82–20000 each). Caveat: Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
- The LLM router takes about 1.8 million times longer per decision. Jev costs $0.0337 per 1,000 decisions, provider-reported. Policy decision, median (20,000 timed): 1.42 µs (n = 20000, p95 2.33 µs). Sonnet 5.5 router ÷ policy, medians: 1.8 million× (n = 82). Jev 1.13 per 1,000 decisions (provider-reported): $0.0337 (n = 82, 137 ms median per call over its API). Caveat: OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
- Of Sonnet 5.5’s 2.60 s per decision, a median 973 ms is CLI and harness time, not the model. Chart: Where an LLM router’s time goes · median per call (n = 82 each). Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
- Route all 49.5 model calls per task: $1.67 per 1,000 tasks with Jev, $247.30 with Sonnet 5.5. Route 7 decisions: $34.97. Chart: Added routing cost per 1,000 tasks. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
- If each decision waits in line, Sonnet 5.5 adds up to 129 s per task, or 18.2 s for 7 decisions. The policy adds 70.3 µs. Chart: Added routing delay per task · upper bound. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
- Start-up tax for a one-word answer: Claude Code 2.53 s and 6,761 input tokens; Codex CLI 6.00 s and 17,051 input tokens, 13,184 of them read from the cache. Chart: CLI start-up tax · one-word answer · median of 5 runs (n = 5 each). Caveat: CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
- Route in process where a rule is enough. Every timing, cost and gap online.
The question: saving minus the router
A router saves money when it moves work to a cheaper model, and costs money on each decision. The net is what counts:
net saving = (cost on one model − cost on the mix) − (decisions × cost per decision)
What if every call ran on Opus? prices the mixes.
Step 1: the gross saving is $3.01
The data: 2,362 recorded model calls from 50 benchmark runs, all on Sonnet 5.5 through Claude Code with routing off. The split moves the learning and memory calls (the side jobs) to Haiku. Everything else stays on Sonnet.
All Sonnet 5.5 costs $108.54. The split costs $105.53. Saving (calculation): $3.01, or 2.8%. It is a floor, not a point estimate (see Caveats).
Step 2: the router costs between $0 and $21.08
Rules cost $0: no model call. Jev's cost is the 803 input tokens its API reported per decision × its published price (246 live calls). Its recorded run's provider-reported cost is the same. The Claude figures are list price × reported tokens (82 decisions each). With one decision per recorded call, router cost = cost per 1,000 × 2.362 (calculation):
| Router | Cost per 1,000 decisions | Router cost, 2,362 decisions | Net (calculation): $3.01 − router cost |
|---|---|---|---|
| Rules (deterministic policy) | $0 | $0 | +$3.01 |
| Jev 1.13 | $0.0337 | $0.08 | +$2.93 |
| Claude Sonnet 5.5 | $4.996 | $11.80 | −$8.79 |
| Claude Haiku 4.5 | $8.924 | $21.08 | −$18.07 |
With a Sonnet router, the split costs $117.33, or 1.08x the all-Sonnet bill, not 0.97x (calculation).
Accuracy does not separate the routers. On 82 labelled cases, exact decisions were Jev 90% (74 of 82; 95% interval 82% to 95%), Haiku 89% (80% to 94%) and Sonnet 94% (87% to 97%). The intervals overlap.
Step 3: the break-even price per decision
The break-even is the saving divided by the decisions (calculation): $3.01 ÷ 2,362 × 1,000 = $1.27 per 1,000 decisions. Sonnet is about 3.9x above it, Haiku about 7x and Jev about 38x below (calculation).
In both repriced mixes, a call's tier comes only from its purpose label. A rule reads that label for $0. A model router may earn its cost where the choice needs more than the label. This calculation does not test that.
The same cost per 1,000 tasks
The overhead study reads the same runs, counted per task. 48 of the 50 runs made model calls. It counts the 105 compaction calls as decisions, so its units differ from the 2,362 calls above: a median 49.5 model calls per task. Its $3.03 median cost per task is the platform's own recorded estimate. We show it apart and do not net it against the first set.
Routing every call adds, per 1,000 tasks (calculation): Jev $1.67, Sonnet $247.30, Haiku $441.74. Routing only the 7 System One decisions per task cuts Sonnet to $34.97.
If each decision waits in line, a router adds up to 128.6 s per task with Sonnet through the CLI and 6.8 s with Jev. These are upper bounds (calculation). Jev's median was 136.5 ms per call (p95 195.7 ms, 246 calls, over HTTPS), a different route from the CLI.
Where a real saving must come from
At Sonnet prices, act cost $57.31 and research $28.62: $85.93 of $108.54, about 79% (calculation). The split moves only the learning and memory calls of the memory and onboarding stage ($6.22) to Haiku. The chart's Haiku bar prices the whole stage ($3.11 gap). The 13 onboarding calls stay on Sonnet (about $0.10 of the gap).
Act and research on Haiku would cut $42.96, about 40% (calculation: $85.93 − $28.66 − $14.31). That is a ceiling, not a plan. The hard-task data warns against it:
On eight hard tasks with strict validators, Haiku 4.5 passed 11 of 24 calls (46%, 95% Wilson interval 28% to 65%). Sonnet 5.5 passed 24 of 24 (86% to 100%). The intervals do not overlap, so Sonnet is ahead on this set. These are not the pipeline's calls: read this as a warning, not a measure. Retries are not priced.
Routing up costs more, with no measured gain
The policy mix puts the strong stages on Opus 5.5. It costs $161.62, 1.49x all-Sonnet (calculation). Sonnet and Opus both passed 24 of 24 hard-task calls, a ceiling, so the data shows no quality gain. When is Opus worth the price?
What we recommend
- Price your calls by stage before you add a router. Route by rule when the stage decides.
- Compute the break-even: expected saving ÷ number of decisions.
- Decide less often. Fewer decisions means fewer router costs.
- If a model must decide on every call, price a dedicated decision model first. Jev's median was 136.5 ms over HTTPS. We did not measure a general model through a direct API.
- Route down only where the cheaper model passes your checks.
How we measured
- Economics data: we repriced the recorded calls at Anthropic list prices from 2026-09-21. Each purpose label has one fixed tier from the platform routing policy. The live policy also adapts. Nets and break-evens are our arithmetic.
- Router cost: Jev 1.13: 246 live calls on 2026-10-06. Sonnet 5.5 (effort low) and Haiku 4.5 (thinking on) through the Claude Code CLI: 82 decisions each.
Caveats
- Repricing is not a run, and one decision per call is an assumption. A different model takes a different path and different tokens. The router costs come from 82 routing cases, not these calls. The $3.01 has no interval: it is arithmetic on one recorded set.
- $3.01 is a floor on two counts. Cache writes are not recorded. The platform's own recorded cost for the same 2,362 calls is $143.09, about 32% above $108.54 (calculation). The gap matches pricing all uncached input as one-hour cache writes at twice the input price (calculation). On that basis the split saves about $4.44, or 3.1% (calculation). The 105 compaction calls (recorded cost $10.02) have no recorded tokens, so they are left out, although the split would move them to Haiku. A Sonnet router at $11.80 costs more than either saving.
- Jev. Its speed is one 35-second window from one Mac on a home network: 3 repeats of 82 requests, not independent draws. The Claude routers ran through a CLI, which adds tokens and start-up time. We did not measure a direct API call to Claude.
- Home advantage. We revised the 82-case routing set against Jev answers, so intervals are wide.
- Limits. 2 calls (one act, one research) had more input than Haiku 4.5 takes. Sonnet, Opus and Fable all passed 24 of 24, so the hard set cannot rank them.
What to read next
- What does a router cost you?
- What is an LLM router? and Jev vs Haiku vs Sonnet as a router
- Compare: Jev vs Sonnet, rules vs Sonnet
- Price your own mix: AI cost calculator
See what your routing saves
Agent records each routing decision with its reason, cost and time. Try Agent and see your own net.