{"i":27,"study":{"slug":"routing-at-scale","title":"What does routing a million AI requests a day cost? A calculation from measured runs","seoTitle":"LLM router cost at scale: 1 million decisions a day","description":"A calculation from measured runs: what 10,000 to 10 million routing decisions a day cost with rules, Jev and Claude, with median/p95 time scenarios.","question":"At 10,000 to 10 million routing decisions a day, what does each router cost per day, what in-flight and waiting scenarios do median and p95 times give?","answer":"This is a calculation, not a run. We multiplied measured numbers and made no model call. At 1 million routing decisions a day, the rule-based policy costs $0 and Jev costs $33.70. A Sonnet 5.5 router costs $7,324 and a Haiku 4.5 router costs $8,924. Cost samples: Jev n = 82, Sonnet n = 82, Haiku n = 82. Jev scales reported run cost. Claude uses list prices and reported tokens. These calculations fix the token mix and have no statistical cost bounds. At 10 million decisions a day, Jev costs $337.00, Sonnet $73,240 and Haiku $89,240. At 1 million a day, that is 11.57 decisions per second. At its median time, the Sonnet scenario gives about 30.1 decisions in flight at once (49.7 at its 95th percentile time). At its median time, the Haiku scenario gives about 146.7 (398.3). At its median time, the Jev scenario gives about 1.58 (2.27), timed live over direct HTTPS. If every decision took its median time, Sonnet would add 721.7 hours of waiting a day and Jev would add 37.9 hours. The policy median-time scenario adds 1.42 s. Say the router decides only the System One decisions (14.1% of model calls). The Sonnet cost at 1 million calls a day then falls from $7,324 to $1,036. Time samples: policy n = 20000, Sonnet n = 82, Haiku n = 82, Jev n = 246. These are median/p95 scenarios, not measured totals or confidence intervals. We did not measure rate limits, load or accuracy.","date":"2026-10-06","updated":"2026-10-06","tags":["thought-experiment","calculation","routing","llm-router","jev","latency","cost","scale","littles-law"],"caveats":["The routing and live Jev protocols predate the first counted calls by file birth time. The overhead protocol predates its microbenchmark output. Claude ran 82 calls per router, below its 120-call cap. Jev ran 246 counted calls, below its 300-call cap. One Haiku pilot and one Jev probe are excluded. All counted calls completed without call errors.\n\nThe policy microbenchmark used 5,000 warm-up iterations and 64 synthetic contexts; the minimum and individual timing samples are unavailable, so we cannot rebuild its quantiles or full range. No separate validator-control receipt or Sonnet pilot is retained.","The source accuracy sets hit a ceiling in some purposes. These tuned cases do not establish equal quality on new traffic. The same 82 Jev cases repeat three times; 246 latency calls are not 246 distinct tasks.","Nothing ran at scale. We did not measure rate limits, queueing or slowdown under load. The Claude routers ran one call at a time through a CLI, and the live Jev run also ran one call at a time. The calculation does not show that a provider accepts this many decisions a day, or this many at once.","The routes differ. The policy runs in process. We timed Jev over direct HTTPS from one Mac on a home network, so its time includes the network round trip. The Claude routers ran through the Claude Code CLI. The CLI adds time (a median 975 ms for Sonnet and 1,701 ms for Haiku) and its own prompt to every call. \n\nWe did not time a direct API call, so the Claude rows show the CLI route and not the best a Claude router can do.","The Claude costs are calculations on subscription calls, not invoices. Sonnet's receipts report one-hour cache writes. The original routing estimate used five-minute write rates and gave $5.00 per 1,000. This study uses the one-hour rate and gives $7.32.\n\nThe CLI estimate is $7.32 per 1,000. Mean output tokens: Haiku 1,419, Sonnet 107. Haiku used default thinking; this comparison does not isolate its effect.","Little’s law uses the mean time per decision. The public CLI extracts contain the median and p95, but no mean. Neither value bounds the mean. These are assumed-time scenarios, not average concurrency, actual waiting totals or capacity guarantees. For Jev, the mean (142.1 ms) is 4.1% above the median (calculation). \n\nThe waiting total has the same limit. Each whisker runs from the median calculation to the 95th percentile calculation. It is not a confidence interval.","A steady rate over 24 hours is an assumption. Real traffic has peaks. To size a peak, multiply the in-flight figures by the peak rate ÷ the average rate.","We did not score accuracy here. The policy picks a model and effort from typed signals. Jev and the Claude routers answer typed System One decisions. The routing study scores the three model routers on 82 decisions. We tuned the case sets against Jev’s answers, so Jev has a home advantage. A cheap router that is wrong costs more in the work it misroutes, and this calculation does not price that.","The System One share comes from 48 bench runs of one agent. The ratio of medians is 14.1%. Pooled over the runs, it is 17.0%, which raises the System One scenario cost by 20% (calculation). Your traffic may differ. The work cost per call is Agent’s notional list-price cost on Sonnet 5.5, not a bill.","The policy’s 1.42 µs median leaves out database reads and the decision-record write, and it needs a server to run on. The calculation prices model calls only. It does not price servers, retries or engineering time.","The samples are small. Each Claude router has 82 decisions (one sample per case). The Jev cost has 82, the live Jev time has 246 calls and the policy time has 20,000. A cost per decision is one estimate with no interval. The projection fixes that estimate and the token mix; it has no statistical cost bounds.","We did not isolate the Mac that ran the timed calls. Other jobs, such as local model servers, may have run at the same time, so a wall time may include contention."],"sourceIds":["calc-routing-at-scale","agent-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"],"stats":{"$k":["id","label","value","unit","display","n","note"],"$r":[["routing-at-scale-cost-per-1000-jev","Cost per 1,000 routing decisions, Jev (calculation)",0.0337,"usd","$0.0337",82,"Calculation: provider-reported total cost ÷ counted calls × 1,000. This fixes the recorded token mix; no cost interval is available."],["routing-at-scale-cost-per-1000-sonnet","Cost per 1,000 routing decisions, Sonnet 5.5 router (calculation)",7.324,"usd","$7.32",82,"A list-price calculation on the tokens the CLI reported. The calls ran on a subscription."],["routing-at-scale-cost-per-1000-haiku","Cost per 1,000 routing decisions, Haiku 4.5 router (calculation)",8.924,"usd","$8.92",82,"A list-price calculation on the tokens the CLI reported. The calls ran on a subscription."],["routing-at-scale-system-one-share","System One decisions as a share of model calls (ratio of the two medians, calculation)",0.1414,"rate","14.1%",48,"7 ÷ 49.5. A ratio of medians, not a sample proportion, so it carries no interval."],["routing-at-scale-rate-1m","Decisions per second at 1 million a day (calculation)",11.5741,"count","11.57 per second","\u0001","A steady rate over 24 hours."],["routing-at-scale-cost-1m-policy","Daily cost at 1 million decisions, rule-based policy (calculation)",0,"usd","$0",20000,"$0 a year at a flat volume."],["routing-at-scale-cost-1m-jev","Daily cost at 1 million decisions, Jev (calculation)",33.7,"usd","$33.70",82,"$12,301 a year at a flat volume."],["routing-at-scale-cost-1m-sonnet","Daily cost at 1 million decisions, Sonnet 5.5 router (calculation)",7324,"usd","$7,324",82,"$2,673,260 a year at a flat volume."],["routing-at-scale-cost-1m-haiku","Daily cost at 1 million decisions, Haiku 4.5 router (calculation)",8924,"usd","$8,924",82,"$3,257,260 a year at a flat volume."],["routing-at-scale-flight-1m-jev","Decisions in flight at 1 million a day, Jev (calculation)",1.5799,"calls","1.58 (2.27 at the 95th percentile time)",246,"Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."],["routing-at-scale-flight-1m-sonnet","Decisions in flight at 1 million a day, Sonnet 5.5 router (calculation)",30.069,"calls","30.1 (49.7 at the 95th percentile time)",82,"Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."],["routing-at-scale-flight-1m-haiku","Decisions in flight at 1 million a day, Haiku 4.5 router (calculation)",146.69,"calls","146.7 (398.3 at the 95th percentile time)",82,"Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."],["routing-at-scale-hours-1m-policy","Waiting hours a day at 1 million decisions, rule-based policy (calculation)",0.00039444,"count","1.42 s in total",20000,"Decisions × assumed median or p95 time. Actual total waiting requires the mean; these are not confidence bounds."],["routing-at-scale-hours-1m-jev","Waiting hours a day at 1 million decisions, Jev (calculation)",37.917,"count","37.9 hours (54.4 at the 95th percentile time)",246,"Decisions × assumed median or p95 time. Actual total waiting requires the mean; these are not confidence bounds."],["routing-at-scale-hours-1m-sonnet","Waiting hours a day at 1 million decisions, Sonnet 5.5 router (calculation)",721.67,"count","721.7 hours (1,194 at the 95th percentile time)",82,"Decisions × assumed median or p95 time. Actual total waiting requires the mean; these are not confidence bounds."]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds","whisker"],"$r":[["routing-at-scale-daily-cost","Daily cost of routing at 10,000 to 10 million decisions a day (calculation)","Decisions a day × cost per decision, USD per day, every decision routed. The rule-based policy costs $0 and is not on the log axis","grouped-bar","usd","USD per day",{"$k":["name","points"],"$r":[["Jev 1.13 (TypeSafe)",{"$k":["label","value","n"],"$r":[["10,000 a day",0.337,82],["100,000 a day",3.37,82],["1 million a day",33.7,82],["10 million a day",337,82]]}],["Claude Sonnet 5.5 (low) · Claude Code",{"$k":["label","value","n"],"$r":[["10,000 a day",73.24,82],["100,000 a day",732.4,82],["1 million a day",7324,82],["10 million a day",73240,82]]}],["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","n"],"$r":[["10,000 a day",89.24,82],["100,000 a day",892.4,82],["1 million a day",8924,82],["10 million a day",89240,82]]}]]},"This is a calculation, not a run. Daily cost = decisions a day × cost per decision. Cost per 1,000 decisions: Jev $0.0337 (calculation from provider-reported run cost, n = 82), Sonnet 5.5 $7.32, Haiku 4.5 $8.92. The Claude figures use one-hour cache-write rates and list price × the tokens the CLI reported (Sonnet n = 82, Haiku n = 82). \n\nThe rule-based policy makes no model call, so it costs $0 at every volume and has no bar: a log axis cannot show zero. The chart prices model calls only, not servers, retries or the work itself.",["calc-routing-at-scale","agent-routing-overhead","agent-routing","price-anthropic","price-jev"],"\u0001"],["routing-at-scale-in-flight","In-flight scenarios at 10,000 to 10 million a day (calculation)","Decisions per second × assumed time per decision; bar = median time, whisker = the same calculation at the 95th percentile. The in-process policy is in the table","grouped-bar","calls","In-flight scenarios at median / p95 time",{"$k":["name","points"],"$r":[["Jev 1.13 (TypeSafe)",{"$k":["label","value","lo","hi","n"],"$r":[["10,000 a day",0.015799,0.015799,0.02265,246],["100,000 a day",0.15799,0.15799,0.2265,246],["1 million a day",1.5799,1.5799,2.265,246],["10 million a day",15.799,15.799,22.65,246]]}],["Claude Sonnet 5.5 (low) · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10,000 a day",0.30069,0.30069,0.49745,82],["100,000 a day",3.0069,3.0069,4.9745,82],["1 million a day",30.069,30.069,49.745,82],["10 million a day",300.69,300.69,497.45,82]]}],["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","lo","hi","n"],"$r":[["10,000 a day",1.4669,1.4669,3.983,82],["100,000 a day",14.669,14.669,39.83,82],["1 million a day",146.69,146.69,398.3,82],["10 million a day",1466.9,1466.9,3983,82]]}]]},"This is a calculation, not a run. In flight = decisions a day ÷ 86,400 × time per decision, at a steady rate over 24 hours. The bar uses the median time. The whisker uses the 95th percentile time and is not a confidence interval. These are assumed-time scenarios. Actual average concurrency requires the mean time. \n\nThe inputs table lists the times. Jev ran over direct HTTPS and the Claude routers ran through a CLI, so the routes differ. Traffic has peaks, so multiply by the peak rate ÷ the average rate to size a peak. \n\nThe rule-based policy has no row. It runs in process, so at 1 million decisions a day the median-time scenario gives 0.000016 decisions in flight, a small share of one decision and not a connection. Labels round small values. The grid table lists every value.",["calc-routing-at-scale","agent-routing-overhead","agent-routing","agent-jev-live"],"p50-p95"],["routing-at-scale-waiting-hours","Waiting scenarios per day at 1 million decisions a day (calculation)","Decisions × time per decision, in hours of waiting added across all requests; whisker = the same calculation at the 95th percentile","bar","count","Waiting hours at assumed median / p95 time",[{"name":"Hours of waiting per day","points":{"$k":["label","value","lo","hi","n"],"$r":[["Jev 1.13 (TypeSafe)",37.917,37.917,54.361,246],["Claude Sonnet 5.5 (low) · Claude Code",721.67,721.67,1193.9,82],["Claude Haiku 4.5 · Claude Code",3520.6,3520.6,9559.2,82]]}}],"This is a calculation, not a run. Hours of waiting a day = decisions a day × time per decision ÷ 3,600, if each request waits for its decision. The bar uses the median time. The whisker uses the 95th percentile time and is not a confidence interval. Total waiting equals decisions × the mean time. The public extracts do not contain the mean for the CLI routers, so the true total is not known. \n\nThe rule-based policy has no bar: at this volume the median-time scenario gives 1.42 s in total. A decision outside the request’s critical path may add no waiting for that request.",["calc-routing-at-scale","agent-routing-overhead","agent-routing","agent-jev-live"],"p50-p95"],["routing-at-scale-scenarios","Daily cost at 1 million model calls a day: route every call or only System One decisions (calculation)","Every call is one decision; System One decisions are 7 of 49.5 model calls per task in the recorded runs","grouped-bar","usd","USD per day",[{"name":"Every model call routed","points":{"$k":["label","value","n"],"$r":[["Deterministic routing policy",0,20000],["Jev 1.13 (TypeSafe)",33.7,82],["Claude Sonnet 5.5 (low) · Claude Code",7324,82],["Claude Haiku 4.5 · Claude Code",8924,82]]}},{"name":"Only System One decisions routed (14.1% of calls)","points":{"$k":["label","value","n"],"$r":[["Deterministic routing policy",0,20000],["Jev 1.13 (TypeSafe)",4.7657,82],["Claude Sonnet 5.5 (low) · Claude Code",1035.7,82],["Claude Haiku 4.5 · Claude Code",1262,82]]}}],"This is a calculation, not a run. It assumes 1,000,000 model calls a day. Route every call: 1,000,000 decisions. Route only System One decisions: 141,414 decisions, which is 7 ÷ 49.5 = 14.1% of calls. \n\nSystem One decisions are the typed choices an agent asks a router to make. The 7 and the 49.5 are medians over 48 recorded bench runs. \n\nRouting was off in those runs, so each model call counts as one decision a router could make. Pooled over the runs, System One decisions are 17.0% of calls (419 of 2,467). The policy is the third scenario: it decides every call in process at no model cost.",["calc-routing-at-scale","agent-routing-overhead","agent-routing","price-anthropic","price-jev"],"\u0001"]]},"related":["routing-overhead","routing-jev-vs-llm"]}}