• Routing
  • Model Routing
  • LLM Router
  • Jev

What does routing a million AI requests a day cost? Rules vs Jev vs Claude

What 10,000 to 10 million routing decisions a day cost with rules, Jev and Claude routers, how many run at once and how long they make requests wait.

Live story · 34 sJev vs Claude as a router: accuracy and cost

Jev vs Claude as a router: accuracy and cost

Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.

Transcript
  1. Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.
  2. Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
  3. Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
  4. Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
  5. Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
  6. Open benchmarks: intervals, sources and every failure kept.

TL;DR

  • This is a calculation, not a run. We multiplied numbers we measured earlier. We made no new model call and ran nothing at scale. Rate limits, load and queueing are not measured.
  • Cost. At 1 million routing decisions a day, the rule-based policy adds $0 in model-call cost and Jev costs $33.70. A Sonnet 5.5 router costs $7,324 and a Haiku 4.5 router costs $8,924. These model-call calculations fix the token mix; no statistical cost bounds are available. At 10 million a day: Jev $337, Sonnet $73,240 and Haiku $89,240. Each model cost per decision rests on 82 recorded decisions (n = 82).
  • In flight (calculation, median/p95 scenarios). 1 million a day is 11.57 decisions per second. In the median-time scenario, a Sonnet router has about 30 decisions in flight at once (50 at its 95th percentile time). A Haiku router has about 147 (398) and Jev about 1.6 (2.3). Time n: Claude 82 each, Jev 246 calls, policy 20,000. The policy scenario gives 0.000016.
  • Waiting (calculation). Assume every request waits and every decision takes the median time. At 1 million a day, Jev adds 37.9 hours of waiting a day, a Sonnet router 722 and a Haiku router 3,521. The policy adds 1.42 s in total.
  • Route less. The ratio of task medians was 7 ÷ 49.5 = 14.1% (calculation, n = 48). The pooled share was 17.0% (calculation). If the router uses that ratio, the Sonnet router costs $1,035.72 a day, not $7,324.
  • The routes differ. Jev ran over direct HTTPS. The Claude routers ran through the Claude Code CLI. The policy ran in process. Read this as three deployments, not three models on equal footing.

The study with every input and formula: agent.sasid.ai/benchmarks/routing-at-scale.

The question

An agent makes a routing decision before the real work of a call starts. It picks a model, an effort or a plan. At 10 calls a day, nobody asks what that decision costs. At a million, the decision is a line in your bill, a queue in your system and a delay in front of every request.

We ask three questions at 10,000, 100,000, 1 million and 10 million decisions a day:

  1. What does each router cost per day?
  2. How many decisions are in flight at once?
  3. How much waiting does it add?

We compare four routers. New to the idea? Read what an LLM router is.

  • The rule-based policy is Agent's production routing decision (plan, model ladder and effort), timed in process.
  • Jev 1.13 is a hosted decision model. We timed it live over direct HTTPS.
  • Claude Sonnet 5.5 (effort low) and Claude Haiku 4.5 (default thinking) ran as routers through the Claude Code CLI.

The measured inputs

Everything below multiplies these numbers. We did not change them.

RouterCost per 1,000 decisionsMedian time per decision (p95; recorded range)n (cost, time)Route
Rule-based policy$0 in model calls1.42 µs (2.33 µs; max 2,538.21 µs, min unavailable)20,000In process
Jev 1.13$0.0337 (calculation from reported run cost)136.5 ms (195.7 ms; 100.9–297.3 ms)82, 246Direct HTTPS, one call at a time
Claude Sonnet 5.5 router$7.32 (list-price calculation)2.60 s (4.30 s; 1.993–5.583 s)82, 82Claude Code CLI
Claude Haiku 4.5 router$8.92 (list-price calculation)12.67 s (34.41 s; 5.857–51.278 s)82, 82Claude Code CLI

These costs use the routing-at-scale input table, including one-hour cache writes.

Per decision, Sonnet costs about 217 times as much as Jev, and Haiku about 265 times as much (calculation from the input costs). The recorded ranges are not confidence intervals. The p95 is a point on the time distribution. It is not a confidence interval.

Cost per day

The formula is simple: daily cost = decisions a day × cost per decision.

Calculation

Jev 1.13 (TypeSafe)

10,000 a day
100,000 a day
1 million a day
10 million a day

Claude Sonnet 5.5 (low) · Claude Code

10,000 a day
100,000 a day
1 million a day
10 million a day

Claude Haiku 4.5 · Claude Code

10,000 a day
100,000 a day
1 million a day
10 million a day

One panel per series, all on the same axis.

List-price calculation, not a run. 4 rows, 3 series: Jev 1.13 (TypeSafe), Claude Sonnet 5.5 (low) · Claude Code, Claude Haiku 4.5 · Claude Code. Jev 1.13 (TypeSafe): highest 10 million a day $337 (n 82). Lowest 10,000 a day $0.34 (n 82). Claude Sonnet 5.5 (low) · Claude Code: highest 10 million a day $73,240 (n 82). Lowest 10,000 a day $73.24 (n 82).

Notesn = 82 per row

Decisions a day × cost per decision, USD per day, every decision routed. The rule-based policy costs $0 and is not on the log axis

This is a calculation, not a run. Daily cost = decisions a day × cost per decision. Cost per 1,000 decisions: Jev $0.0337 (calculation from provider-reported run cost, n = 82), Sonnet 5.5 $7.32, Haiku 4.5 $8.92. The Claude figures use one-hour cache-write rates and list price × the tokens the CLI reported (Sonnet n = 82, Haiku n = 82). The rule-based policy makes no model call, so it costs $0 at every volume and has no bar: a log axis cannot show zero. The chart prices model calls only, not servers, retries or the work itself.

Sources: Routing at scale calculation, Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price

Decisions a dayPolicyJevSonnet routerHaiku router
10,000$0$0.337$73.24$89.24
100,000$0$3.37$732.40$892.40
1 million$0$33.70$7,324$8,924
10 million$0$337$73,240$89,240

The chart has no policy bar, because a log axis cannot show $0. At 1 million a day, a year costs $12,301 with Jev, $2.67 million with Sonnet and $3.26 million with Haiku (calculation: cost per day × 365).

Two checks on these costs:

  • The Claude cost is a list-price calculation on subscription calls. Sonnet used one-hour cache writes. The old estimate used five-minute write rates and gave $5.00 per 1,000. The corrected $7.324 per 1,000 agrees with the CLI estimate at this precision. It is not an invoice.
  • Scale it against the work. Agent's recorded bench tasks give $0.0611 per model call as a ratio of medians (calculation) (a notional list-price cost, $3.026 per task ÷ 49.5 calls, n = 48 tasks). If a router decided every call, it would add 12.0% with Sonnet, 14.6% with Haiku and 0.06% with Jev (calculation).

How many decisions are in flight

Little's law uses the mean time: mean decisions in flight = decisions per second × mean time per decision. At 1 million a day the rate is 1,000,000 ÷ 86,400 = 11.57 decisions per second, at a steady rate over 24 hours.

The table assumes every decision takes the median time. Brackets use the p95 time. Neither scenario measures average concurrency or guarantees capacity.

Decisions a dayPer secondJevSonnet routerHaiku router
10,0000.1160.016 (0.023)0.30 (0.50)1.5 (4.0)
100,0001.160.16 (0.23)3.0 (5.0)14.7 (39.8)
1 million11.571.6 (2.3)30.1 (49.7)146.7 (398.3)
10 million115.715.8 (22.7)300.7 (497.5)1,467 (3,983)
Calculation

Jev 1.13 (TypeSafe)

10,000 a day
100,000 a day
1 million a day
10 million a day

Claude Sonnet 5.5 (low) · Claude Code

10,000 a day
100,000 a day
1 million a day
10 million a day

Claude Haiku 4.5 · Claude Code

10,000 a day
100,000 a day
1 million a day
10 million a day

One panel per series, all on the same axis; whiskers are the median to p95 (not an interval).

List-price calculation, not a run. 4 rows, 3 series: Jev 1.13 (TypeSafe), Claude Sonnet 5.5 (low) · Claude Code, Claude Haiku 4.5 · Claude Code. Jev 1.13 (TypeSafe): highest 10 million a day 15.8 (median to p95 15.8–22.7, n 246). Lowest 10,000 a day 0 (median to p95 0–0, n 246). Not all run ranges overlap. Claude Sonnet 5.5 (low) · Claude Code: highest 10 million a day 300.7 (median to p95 300.7–497.5, n 82). Lowest 10,000 a day 0.3 (median to p95 0.3–0.5, n 82). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–246 per row

Decisions per second × assumed time per decision; bar = median time, whisker = the same calculation at the 95th percentile. The in-process policy is in the table

This is a calculation, not a run. In flight = decisions a day ÷ 86,400 × time per decision, at a steady rate over 24 hours. The bar uses the median time. The whisker uses the 95th percentile time and is not a confidence interval. These are assumed-time scenarios. Actual average concurrency requires the mean time. The inputs table lists the times. Jev ran over direct HTTPS and the Claude routers ran through a CLI, so the routes differ. Traffic has peaks, so multiply by the peak rate ÷ the average rate to size a peak. The rule-based policy has no row. It runs in process, so at 1 million decisions a day the median-time scenario gives 0.000016 decisions in flight, a small share of one decision and not a connection. Labels round small values. The grid table lists every value.

Sources: Routing at scale calculation, Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Read the Sonnet column as work for your system. At 10 million a day, the median-time scenario gives about 301 routing calls in flight. In our runs, each routing call was one CLI process, so that is about 301 processes under that assumption. Actual pool size needs the mean time, traffic peaks and a load test.

The rule-based policy runs in one process in this microbenchmark. One process decided 469,409 times per second in one 200,000-iteration throughput batch of the routing-overhead study. 10 million a day needs 115.7 per second, which is 0.025% of that (calculation). The batch left out database reads and the decision-record write.

How much waiting does it add

If each request waits and every decision takes the assumed median or p95 time, waiting hours a day = decisions a day × time per decision ÷ 3,600.

Calculation
Largest value is 93x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 · Claude Code 3,520.6 (median to p95 3,520.6–9,559.2, n 82). Lowest Jev 1.13 (TypeSafe) 37.9 (median to p95 37.9–54.4, n 246). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–246 per row

Decisions × time per decision, in hours of waiting added across all requests; whisker = the same calculation at the 95th percentile

This is a calculation, not a run. Hours of waiting a day = decisions a day × time per decision ÷ 3,600, if each request waits for its decision. The bar uses the median time. The whisker uses the 95th percentile time and is not a confidence interval. Total waiting equals decisions × the mean time. The public extracts do not contain the mean for the CLI routers, so the true total is not known. The rule-based policy has no bar: at this volume the median-time scenario gives 1.42 s in total. A decision outside the request’s critical path may add no waiting for that request.

Sources: Routing at scale calculation, Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

At 1 million decisions a day, the Jev scenario gives 37.9 hours (54.4 at the 95th percentile time). A Sonnet scenario gives 721.7 (1,194). The Haiku scenario gives 3,521 (9,559). Per request, the wait is the median time: 136.5 ms, 2.60 s and 12.67 s. The policy adds 1.42 s in total.

Parallel work can hide some delay. These scenarios are not an upper bound: neither the median nor p95 bounds the mean time. Actual waiting totals need the mean.

Route fewer decisions

Not every model call needs a router. In our 48 recorded bench runs, a task made a median 49.5 model calls. The median System One count was 7: typed choices an agent asks a router to make, such as what kind of failure just happened. The scenario uses 7 ÷ 49.5 = 14.1%, a calculation from two medians, not the observed pooled share.

Calculation
  • Every model call routed
  • Only System One decisions routed (14.1% of calls) (square)
In chart order.
Deterministic routing policy
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (low) · Claude Code
Claude Haiku 4.5 · Claude Code

Gap labels, Only System One decisions routed (14.1% of calls) vs Every model call routed: Only System One decisions routed (14.1% of calls) is x% higher (+) or lower (−) than Every model call routed, calculated from the two values shown (the change counted from Every model call routed’s value).

List-price calculation, not a run. 4 rows, 2 series: Every model call routed, Only System One decisions routed (14.1% of calls). Every model call routed: highest Claude Haiku 4.5 · Claude Code $8,924 (n 82). Lowest Deterministic routing policy $0 (n 20000). Only System One decisions routed (14.1% of calls): highest Claude Haiku 4.5 · Claude Code $1,262 (n 82). Lowest Deterministic routing policy $0 (n 20000).

Notesn 82–20000 per row

Every call is one decision; System One decisions are 7 of 49.5 model calls per task in the recorded runs

This is a calculation, not a run. It assumes 1,000,000 model calls a day. Route every call: 1,000,000 decisions. Route only System One decisions: 141,414 decisions, which is 7 ÷ 49.5 = 14.1% of calls. System One decisions are the typed choices an agent asks a router to make. The 7 and the 49.5 are medians over 48 recorded bench runs. Routing was off in those runs, so each model call counts as one decision a router could make. Pooled over the runs, System One decisions are 17.0% of calls (419 of 2,467). The policy is the third scenario: it decides every call in process at no model cost.

Sources: Routing at scale calculation, Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price

1 million model calls a dayEvery call routedOnly System One decisions routed
Policy$0$0
Jev$33.70$4.77
Sonnet router$7,324$1,035.72
Haiku router$8,924$1,262

For the Sonnet router, the median-time scenario for routing only those decisions cuts the in-flight count from 30.1 to 4.25. That scenario cuts waiting from 722 to 102 hours a day (calculation). Pooled over the runs, System One decisions are 17.0% of calls (419 of 2,467; calculation). That share raises the scenario costs by about 20%. Your traffic will differ.

What a cheap router does not tell you

This calculation prices calls. It does not price mistakes. In the routing study, the three model routers answered these typed decisions exactly right as follows. Jev: 74 of 82 (90%, 95% interval 82% to 95%). Claude Sonnet 5.5: 77 of 82 (94%, 87% to 97%). Claude Haiku 4.5: 73 of 82 (89%, 80% to 94%). The intervals overlap, so accuracy does not separate them. Exact paired McNemar p-values against recorded Jev are 0.375 for Sonnet and 1 for Haiku. The live Jev run scored 221/246 across three repeats of the same 82 cases (case-level Wilson 95% interval 81.9%–95.0%, calculation with the pooled rate rounded to 74 successes at n = 82 distinct cases). Some purposes hit a ceiling; these tuned cases do not prove equal quality on new traffic.

We tuned those case sets against Jev's answers, so Jev has a home advantage. A router that picks wrongly costs the work it misroutes. We did not price that. See the routing study and the Jev vs Claude Sonnet 5.5 comparison.

What to do with this

These are suggestions from the numbers. We have not tested all of them.

  1. Size by calls in flight, not by calls a day. Multiply the decisions per second by the time per decision on your route. Use your mean time and peak rate. Test capacity under load; p95 alone cannot guarantee it.
  2. Put rules first. The policy decided in a median 1.42 µs with $0 in model-call cost. Use a model router only for choices that rules cannot make. We did not test a hybrid. The System One scenario is the closest number we have.
  3. Route only the decisions that need judgment. In this scenario that cuts the Sonnet router bill by 7.1 times (calculation: $7,324 ÷ $1,035.72).
  4. Time your own route. Wall time minus CLI-reported API time had a median of 975 ms for Sonnet (range 826–3,673 ms, n = 82). For Haiku it was 1,701 ms (range 1,356–3,872 ms, n = 82). These differences are calculations; the ranges are not intervals. A direct API call avoids the CLI process, but its time may differ. We did not time a direct call.
  5. Price your own tokens. Our Sonnet cost was $7.324 per 1,000 with one-hour cache writes, matching the CLI estimate at this precision. The AI cost calculator takes your token mix.
  6. Ask for rate limits before you ramp up. We did not measure them. They may bind before the bill does.
  7. Test accuracy on your own decisions. Overlapping intervals at n = 82 do not prove that two routers are equal.

For the per-task view of the same routers, read What does a router cost you?.

How we calculated

We made no new call. Inputs come from the routing receipts, overhead results and live Jev summary:

  • Cost per decision: routing receipts, reported tokens and the study's price table. Jev scales the provider-reported run cost (calculation). Claude uses one-hour cache-write rates; the old Sonnet listCostPer1000Usd field used five-minute rates. The live Jev run, priced at the published input price, also gives $0.0337 per 1,000 (calculation).
  • Time per decision: the routing-overhead results for the policy (latencyUs) and the Claude run ranges. Their median and p95 come from routing receipts (latencyMs.wallP50 and wallP95), checked against call rows. The overhead summary uses different quantiles, so we do not mix its median and p95 into this calculation. For Jev, the live run: 246 calls (82 typed decisions × 3 repetitions) over direct HTTPS from one Mac, latencyMs.calls median and p95.
  • System One share: routing-overhead perTask.systemOneDecisions.median ÷ perTask.modelCalls.median = 7 ÷ 49.5, over 48 recorded runs. Calls ranged from 13 to 73 per task; System One decisions from 2 to 24. These are ranges, not intervals. Routing was off; all System One decisions used rules.
QuantityFormula
Daily costdecisions a day × cost per 1,000 decisions ÷ 1,000
Decisions per seconddecisions a day ÷ 86,400
Decisions in flight (scenario)decisions per second × assumed median or p95 time in seconds
Waiting hours a day (scenario)decisions a day × assumed median or p95 time in seconds ÷ 3,600
Cost per yearcost per day × 365

Every value, for every scenario, volume and router, is in the study's grid table and downloads as JSON and CSV.

Caveats

  • Nothing ran at scale. We did not measure rate limits, queueing or slowdown under load. The calculation does not show that a provider accepts this many decisions a day, or this many at once.
  • The routes differ. The policy ran in process. Jev ran over direct HTTPS from one Mac on a home network, so its time includes the network. The Claude routers ran through a CLI that adds time and its own prompt. We did not time a direct Claude API call.
  • Little's law needs the mean time. The runs record the median and the 95th percentile for the CLI routers. Neither quantile bounds the mean. These are scenarios, not mean concurrency, actual waiting totals or confidence bounds. For Jev, the mean (142.1 ms) is 4.1% above the median (calculation).
  • A steady rate is an assumption. Real traffic has peaks.
  • The samples are limited. Each Claude router has 82 decisions, Jev's live time has 246 calls and the policy has 20,000. Cost per decision is one estimate with no interval. The projections fix the token mix and have no statistical cost bounds. Jev's 246 calls repeat 82 cases.
  • The policy time is a microbenchmark. It leaves out database reads and the decision-record write, and it needs a server. Individual timing samples are unavailable, so we cannot rebuild the policy quantiles or full range.
  • We did not isolate the Mac. Other jobs may have run during the timed calls, so a wall time may include contention.
  • The System One share comes from one agent. 48 bench runs of Agent. Your share may differ.
  • We did not price servers, retries or engineering time.

Count what your router costs

Agent records the tokens, time and cost of every model call, so you can see what routing adds on your own tasks. Try Agent.

Disclosure: I build Agent, the product behind these benchmarks. Its platform uses the rule-based policy and Jev in production, so both are home-team routers here.

The data behind this post

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.