What does routing a million AI requests a day cost? Rules vs Jev vs Claude
What 10,000 to 10 million routing decisions a day cost with rules, Jev and Claude routers, how many run at once and how long they make requests wait.
Jev vs Claude as a router: accuracy and cost
Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.
Transcript
- Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.
- Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
- Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
- Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
- Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
- Open benchmarks: intervals, sources and every failure kept.
TL;DR
- This is a calculation, not a run. We multiplied numbers we measured earlier. We made no new model call and ran nothing at scale. Rate limits, load and queueing are not measured.
- Cost. At 1 million routing decisions a day, the rule-based policy adds $0 in model-call cost and Jev costs $33.70. A Sonnet 5.5 router costs $7,324 and a Haiku 4.5 router costs $8,924. These model-call calculations fix the token mix; no statistical cost bounds are available. At 10 million a day: Jev $337, Sonnet $73,240 and Haiku $89,240. Each model cost per decision rests on 82 recorded decisions (n = 82).
- In flight (calculation, median/p95 scenarios). 1 million a day is 11.57 decisions per second. In the median-time scenario, a Sonnet router has about 30 decisions in flight at once (50 at its 95th percentile time). A Haiku router has about 147 (398) and Jev about 1.6 (2.3). Time n: Claude 82 each, Jev 246 calls, policy 20,000. The policy scenario gives 0.000016.
- Waiting (calculation). Assume every request waits and every decision takes the median time. At 1 million a day, Jev adds 37.9 hours of waiting a day, a Sonnet router 722 and a Haiku router 3,521. The policy adds 1.42 s in total.
- Route less. The ratio of task medians was 7 ÷ 49.5 = 14.1% (calculation, n = 48). The pooled share was 17.0% (calculation). If the router uses that ratio, the Sonnet router costs $1,035.72 a day, not $7,324.
- The routes differ. Jev ran over direct HTTPS. The Claude routers ran through the Claude Code CLI. The policy ran in process. Read this as three deployments, not three models on equal footing.
The study with every input and formula: agent.sasid.ai/benchmarks/routing-at-scale.
The question
An agent makes a routing decision before the real work of a call starts. It picks a model, an effort or a plan. At 10 calls a day, nobody asks what that decision costs. At a million, the decision is a line in your bill, a queue in your system and a delay in front of every request.
We ask three questions at 10,000, 100,000, 1 million and 10 million decisions a day:
- What does each router cost per day?
- How many decisions are in flight at once?
- How much waiting does it add?
We compare four routers. New to the idea? Read what an LLM router is.
- The rule-based policy is Agent's production routing decision (plan, model ladder and effort), timed in process.
- Jev 1.13 is a hosted decision model. We timed it live over direct HTTPS.
- Claude Sonnet 5.5 (effort low) and Claude Haiku 4.5 (default thinking) ran as routers through the Claude Code CLI.
The measured inputs
Everything below multiplies these numbers. We did not change them.
| Router | Cost per 1,000 decisions | Median time per decision (p95; recorded range) | n (cost, time) | Route |
|---|---|---|---|---|
| Rule-based policy | $0 in model calls | 1.42 µs (2.33 µs; max 2,538.21 µs, min unavailable) | 20,000 | In process |
| Jev 1.13 | $0.0337 (calculation from reported run cost) | 136.5 ms (195.7 ms; 100.9–297.3 ms) | 82, 246 | Direct HTTPS, one call at a time |
| Claude Sonnet 5.5 router | $7.32 (list-price calculation) | 2.60 s (4.30 s; 1.993–5.583 s) | 82, 82 | Claude Code CLI |
| Claude Haiku 4.5 router | $8.92 (list-price calculation) | 12.67 s (34.41 s; 5.857–51.278 s) | 82, 82 | Claude Code CLI |
These costs use the routing-at-scale input table, including one-hour cache writes.
Per decision, Sonnet costs about 217 times as much as Jev, and Haiku about 265 times as much (calculation from the input costs). The recorded ranges are not confidence intervals. The p95 is a point on the time distribution. It is not a confidence interval.
Cost per day
The formula is simple: daily cost = decisions a day × cost per decision.
| Decisions a day | Policy | Jev | Sonnet router | Haiku router |
|---|---|---|---|---|
| 10,000 | $0 | $0.337 | $73.24 | $89.24 |
| 100,000 | $0 | $3.37 | $732.40 | $892.40 |
| 1 million | $0 | $33.70 | $7,324 | $8,924 |
| 10 million | $0 | $337 | $73,240 | $89,240 |
The chart has no policy bar, because a log axis cannot show $0. At 1 million a day, a year costs $12,301 with Jev, $2.67 million with Sonnet and $3.26 million with Haiku (calculation: cost per day × 365).
Two checks on these costs:
- The Claude cost is a list-price calculation on subscription calls. Sonnet used one-hour cache writes. The old estimate used five-minute write rates and gave $5.00 per 1,000. The corrected $7.324 per 1,000 agrees with the CLI estimate at this precision. It is not an invoice.
- Scale it against the work. Agent's recorded bench tasks give $0.0611 per model call as a ratio of medians (calculation) (a notional list-price cost, $3.026 per task ÷ 49.5 calls, n = 48 tasks). If a router decided every call, it would add 12.0% with Sonnet, 14.6% with Haiku and 0.06% with Jev (calculation).
How many decisions are in flight
Little's law uses the mean time: mean decisions in flight = decisions per second × mean time per decision. At 1 million a day the rate is 1,000,000 ÷ 86,400 = 11.57 decisions per second, at a steady rate over 24 hours.
The table assumes every decision takes the median time. Brackets use the p95 time. Neither scenario measures average concurrency or guarantees capacity.
| Decisions a day | Per second | Jev | Sonnet router | Haiku router |
|---|---|---|---|---|
| 10,000 | 0.116 | 0.016 (0.023) | 0.30 (0.50) | 1.5 (4.0) |
| 100,000 | 1.16 | 0.16 (0.23) | 3.0 (5.0) | 14.7 (39.8) |
| 1 million | 11.57 | 1.6 (2.3) | 30.1 (49.7) | 146.7 (398.3) |
| 10 million | 115.7 | 15.8 (22.7) | 300.7 (497.5) | 1,467 (3,983) |
Read the Sonnet column as work for your system. At 10 million a day, the median-time scenario gives about 301 routing calls in flight. In our runs, each routing call was one CLI process, so that is about 301 processes under that assumption. Actual pool size needs the mean time, traffic peaks and a load test.
The rule-based policy runs in one process in this microbenchmark. One process decided 469,409 times per second in one 200,000-iteration throughput batch of the routing-overhead study. 10 million a day needs 115.7 per second, which is 0.025% of that (calculation). The batch left out database reads and the decision-record write.
How much waiting does it add
If each request waits and every decision takes the assumed median or p95 time, waiting hours a day = decisions a day × time per decision ÷ 3,600.
At 1 million decisions a day, the Jev scenario gives 37.9 hours (54.4 at the 95th percentile time). A Sonnet scenario gives 721.7 (1,194). The Haiku scenario gives 3,521 (9,559). Per request, the wait is the median time: 136.5 ms, 2.60 s and 12.67 s. The policy adds 1.42 s in total.
Parallel work can hide some delay. These scenarios are not an upper bound: neither the median nor p95 bounds the mean time. Actual waiting totals need the mean.
Route fewer decisions
Not every model call needs a router. In our 48 recorded bench runs, a task made a median 49.5 model calls. The median System One count was 7: typed choices an agent asks a router to make, such as what kind of failure just happened. The scenario uses 7 ÷ 49.5 = 14.1%, a calculation from two medians, not the observed pooled share.
| 1 million model calls a day | Every call routed | Only System One decisions routed |
|---|---|---|
| Policy | $0 | $0 |
| Jev | $33.70 | $4.77 |
| Sonnet router | $7,324 | $1,035.72 |
| Haiku router | $8,924 | $1,262 |
For the Sonnet router, the median-time scenario for routing only those decisions cuts the in-flight count from 30.1 to 4.25. That scenario cuts waiting from 722 to 102 hours a day (calculation). Pooled over the runs, System One decisions are 17.0% of calls (419 of 2,467; calculation). That share raises the scenario costs by about 20%. Your traffic will differ.
What a cheap router does not tell you
This calculation prices calls. It does not price mistakes. In the routing study, the three model routers answered these typed decisions exactly right as follows. Jev: 74 of 82 (90%, 95% interval 82% to 95%). Claude Sonnet 5.5: 77 of 82 (94%, 87% to 97%). Claude Haiku 4.5: 73 of 82 (89%, 80% to 94%). The intervals overlap, so accuracy does not separate them. Exact paired McNemar p-values against recorded Jev are 0.375 for Sonnet and 1 for Haiku. The live Jev run scored 221/246 across three repeats of the same 82 cases (case-level Wilson 95% interval 81.9%–95.0%, calculation with the pooled rate rounded to 74 successes at n = 82 distinct cases). Some purposes hit a ceiling; these tuned cases do not prove equal quality on new traffic.
We tuned those case sets against Jev's answers, so Jev has a home advantage. A router that picks wrongly costs the work it misroutes. We did not price that. See the routing study and the Jev vs Claude Sonnet 5.5 comparison.
What to do with this
These are suggestions from the numbers. We have not tested all of them.
- Size by calls in flight, not by calls a day. Multiply the decisions per second by the time per decision on your route. Use your mean time and peak rate. Test capacity under load; p95 alone cannot guarantee it.
- Put rules first. The policy decided in a median 1.42 µs with $0 in model-call cost. Use a model router only for choices that rules cannot make. We did not test a hybrid. The System One scenario is the closest number we have.
- Route only the decisions that need judgment. In this scenario that cuts the Sonnet router bill by 7.1 times (calculation: $7,324 ÷ $1,035.72).
- Time your own route. Wall time minus CLI-reported API time had a median of 975 ms for Sonnet (range 826–3,673 ms, n = 82). For Haiku it was 1,701 ms (range 1,356–3,872 ms, n = 82). These differences are calculations; the ranges are not intervals. A direct API call avoids the CLI process, but its time may differ. We did not time a direct call.
- Price your own tokens. Our Sonnet cost was $7.324 per 1,000 with one-hour cache writes, matching the CLI estimate at this precision. The AI cost calculator takes your token mix.
- Ask for rate limits before you ramp up. We did not measure them. They may bind before the bill does.
- Test accuracy on your own decisions. Overlapping intervals at n = 82 do not prove that two routers are equal.
For the per-task view of the same routers, read What does a router cost you?.
How we calculated
We made no new call. Inputs come from the routing receipts, overhead results and live Jev summary:
- Cost per decision: routing receipts, reported tokens and the study's price table. Jev scales the provider-reported run cost (calculation). Claude uses one-hour cache-write rates; the old Sonnet
listCostPer1000Usdfield used five-minute rates. The live Jev run, priced at the published input price, also gives $0.0337 per 1,000 (calculation). - Time per decision: the routing-overhead results for the policy (
latencyUs) and the Claude run ranges. Their median and p95 come from routing receipts (latencyMs.wallP50andwallP95), checked against call rows. The overhead summary uses different quantiles, so we do not mix its median and p95 into this calculation. For Jev, the live run: 246 calls (82 typed decisions × 3 repetitions) over direct HTTPS from one Mac,latencyMs.callsmedian and p95. - System One share: routing-overhead
perTask.systemOneDecisions.median÷perTask.modelCalls.median= 7 ÷ 49.5, over 48 recorded runs. Calls ranged from 13 to 73 per task; System One decisions from 2 to 24. These are ranges, not intervals. Routing was off; all System One decisions used rules.
| Quantity | Formula |
|---|---|
| Daily cost | decisions a day × cost per 1,000 decisions ÷ 1,000 |
| Decisions per second | decisions a day ÷ 86,400 |
| Decisions in flight (scenario) | decisions per second × assumed median or p95 time in seconds |
| Waiting hours a day (scenario) | decisions a day × assumed median or p95 time in seconds ÷ 3,600 |
| Cost per year | cost per day × 365 |
Every value, for every scenario, volume and router, is in the study's grid table and downloads as JSON and CSV.
Caveats
- Nothing ran at scale. We did not measure rate limits, queueing or slowdown under load. The calculation does not show that a provider accepts this many decisions a day, or this many at once.
- The routes differ. The policy ran in process. Jev ran over direct HTTPS from one Mac on a home network, so its time includes the network. The Claude routers ran through a CLI that adds time and its own prompt. We did not time a direct Claude API call.
- Little's law needs the mean time. The runs record the median and the 95th percentile for the CLI routers. Neither quantile bounds the mean. These are scenarios, not mean concurrency, actual waiting totals or confidence bounds. For Jev, the mean (142.1 ms) is 4.1% above the median (calculation).
- A steady rate is an assumption. Real traffic has peaks.
- The samples are limited. Each Claude router has 82 decisions, Jev's live time has 246 calls and the policy has 20,000. Cost per decision is one estimate with no interval. The projections fix the token mix and have no statistical cost bounds. Jev's 246 calls repeat 82 cases.
- The policy time is a microbenchmark. It leaves out database reads and the decision-record write, and it needs a server. Individual timing samples are unavailable, so we cannot rebuild the policy quantiles or full range.
- We did not isolate the Mac. Other jobs may have run during the timed calls, so a wall time may include contention.
- The System One share comes from one agent. 48 bench runs of Agent. Your share may differ.
- We did not price servers, retries or engineering time.
What to read next
- What does a router cost you?: delay and cost per 1,000 tasks.
- Jev vs Claude as routers on routing decisions: accuracy and cost per decision.
- The routing hub: every router, measured, recorded, calculated or not measured.
- What is an LLM router?
Count what your router costs
Agent records the tokens, time and cost of every model call, so you can see what routing adds on your own tasks. Try Agent.
Disclosure: I build Agent, the product behind these benchmarks. Its platform uses the rule-based policy and Jev in production, so both are home-team routers here.