{"i":10,"study":{"slug":"routing-jev-vs-llm","title":"Jev vs Claude as a router: accuracy and cost","seoTitle":"Jev router vs Claude Haiku and Sonnet: routing and cost","description":"Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.","question":"Should a small dedicated router or a general LLM make the platform’s typed routing decisions?","answer":"Exact decisions: Jev 1.13 (TypeSafe) 221 of 246 live calls (90%; 74, 73 and 74 of 82 per repeat; case-level interval 82% to 95%); Claude Haiku 4.5 73 of 82 (89%, 80% to 94%); Claude Sonnet 5.5 77 of 82 (94%, 87% to 97%). The intervals overlap, so accuracy does not separate the routers here. Cost does: Jev costs $0.0337 per 1,000 decisions (a calculation from its reported input tokens) against Claude Haiku 4.5 $8.92, Claude Sonnet 5.5 $5.00. Case by case against Jev’s recorded production run (one pass), the exact McNemar test finds no difference (Claude Haiku 4.5 p = 1, Claude Sonnet 5.5 p = 0.375). Per decision, Jev is about 265x cheaper than Claude Haiku 4.5 and about 148x cheaper than Claude Sonnet 5.5. Median model time per decision through the Claude Code CLI: Claude Haiku 4.5 10.7 s, Claude Sonnet 5.5 1.6 s. Jev, called directly over HTTPS from one Mac, took a median 137 ms per call (p95 196 ms, 246 calls, wall time with the network inside it). That is a different route from the CLI, so it is not a model-against-model compute comparison. The case sets were tuned against Jev answers, which gives Jev a home advantage. Thought experiment (a calculation on 2,362 recorded calls, not a run): the same tokens cost $108.54 all on Sonnet 5.5, $161.62 under the platform policy mix (1.49x) and $105.53 with Haiku on side jobs only (2.8% less), because the main coding stages hold most of the spend.","date":"2026-10-05","updated":"2026-10-06","tags":["routing","jev","claude-haiku","claude-sonnet","model-routing","thought-experiment"],"caveats":["The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.","Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.","Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.","One sample per decision; production asks a second sample when confidence is low. Confidence-gated coverage is therefore not compared.","82 cases in four small hand-labelled sets: intervals are wide.","Clef / Clef-Flash (local): not measured (No local Clef server was running and installing a 6-20 GB model was out of scope for this run).","Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.","Jev ran as a direct HTTPS call; the Claude routers ran through the Claude Code CLI. These are different routes, so speed and cost compare what a caller pays per decision, not one model against the other.","Jev’s latency is one 35-second window from one Mac over a home network. The API reports no server time. A caller near the API would see less."],"sourceIds":["agent-routing","calc-repricing","price-jev","price-anthropic","agent-jev-live"],"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["exact-jev","Jev 1.13 (TypeSafe): exact decisions",0.8984,"rate","90% (74/82)",82,[0.8191,0.9497],"Live run: 3 repeats of the same 82 decisions. 221 of 246 calls were exact (89.8%; 74, 73 and 74 of 82 per repeat). The count is shown on the 82-decision scale (89.8% of 82 is 74), the scale of the interval: repeats of one decision are not independent, so the interval is taken at n = 82, not 246."],["exact-claude-haiku","Claude Haiku 4.5: exact decisions",0.8902,"rate","89% (73/82)",82,[0.8044,0.9412],"\u0001"],["exact-claude-sonnet","Claude Sonnet 5.5: exact decisions",0.939,"rate","94% (77/82)",82,[0.8651,0.9737],"\u0001"],["jev-cost-per-1000","Jev cost per 1,000 decisions",0.0337,"usd","$0.0337",246,"\u0001","Calculation: mean reported input tokens per decision × the published input price."],["economics-policy-vs-all-sonnet","Thought experiment: policy (Opus strong, Haiku ancillary) vs all Sonnet 5.5",1.489,"ratio","$161.62 (1.49x)",2362,"\u0001","Calculation on recorded tokens, not a run."],["economics-split-vs-all-sonnet","Thought experiment: split (Sonnet main line, Haiku ancillary) vs all Sonnet 5.5",0.9723,"ratio","$105.53 (0.97x)",2362,"\u0001","Calculation on recorded tokens, not a run."],["economics-cache-share","Cache-read share of recorded input (economics data)",0.9304,"rate","93.0%",2362,"\u0001","\u0001"]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","whisker","series","note","sourceIds"],"$r":[["routing-exact-decisions","Typed routing decisions answered exactly right","Share of asked cases where every scored question was acceptable","dot-range","rate","Exact","ci95",[{"name":"Exact rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8984,0.8191,0.9497,82,true],["Claude Haiku 4.5",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5",0.939,0.8651,0.9737,82,false]]}}],"Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.",["agent-routing","agent-jev-live"]],["routing-key-accuracy","Per-question accuracy","Each open question the router was asked; an unanswered question counts as wrong","dot-range","rate","Correct answers","ci95",[{"name":"Key accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.9485,0.9077,0.9718,194,true],["Claude Haiku 4.5",0.9433,0.9013,0.968,194,false],["Claude Sonnet 5.5",0.9742,0.9411,0.9889,194,false]]}}],"Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.",["agent-routing","agent-jev-live"]],["routing-exact-by-decision","Exact rate by decision type","\u0001","grouped-bar","rate","Exact","\u0001",{"$k":["name","points"],"$r":[["Jev 1.13 (TypeSafe)",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",1,0.8241,1,18,true],["Message intent",1,0.8389,1,20,true],["Is it a rule?",1,0.7575,1,12,true],["Context shape",0.7396,0.5789,0.8675,32,true]]}],["Claude Haiku 4.5",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",0.9444,0.7424,0.9901,18,false],["Message intent",1,0.8389,1,20,false],["Is it a rule?",1,0.7575,1,12,false],["Context shape",0.75,0.5789,0.8675,32,false]]}],["Claude Sonnet 5.5",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",1,0.8241,1,18,false],["Message intent",1,0.8389,1,20,false],["Is it a rule?",1,0.7575,1,12,false],["Context shape",0.8438,0.6825,0.9314,32,false]]}]]},"A missing bar means that router was not run on that decision type. Whiskers are 95% Wilson intervals. Jev: the pooled rate of the live repeats, with the interval taken at the number of decisions of that type.",["agent-routing","agent-jev-live"]],["routing-cost-per-1000","Cost per 1,000 routing decisions","List price × reported tokens per decision","bar","usd","USD per 1,000 decisions","\u0001",[{"name":"Cost","points":{"$k":["label","value","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.0337,246,true],["Claude Haiku 4.5",8.924,82,false],["Claude Sonnet 5.5",4.996,82,false]]}}],"List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.",["agent-routing","calc-repricing","price-jev","price-anthropic","agent-jev-live"]],["routing-decision-latency","Time per routing decision","Median wall time, whisker to the 95th percentile","dot-range","ms","Time per decision","p50-p95",{"$k":["name","points"],"$r":[["Wall time (CLI)",[{"label":"Claude Haiku 4.5","value":12674,"lo":12674,"hi":34413,"n":82},{"label":"Claude Sonnet 5.5","value":2598,"lo":2598,"hi":4298,"n":82}]],["Wall time (direct API call)",[{"label":"Jev 1.13 (TypeSafe)","value":136.5,"lo":136.5,"hi":195.7,"n":246}]],["Model time (API)",[{"label":"Claude Haiku 4.5","value":10734,"lo":10734,"hi":32072,"n":82},{"label":"Claude Sonnet 5.5","value":1599,"lo":1599,"hi":2574,"n":82}]]]},"Whiskers run from p50 to p95. The Claude routers ran through the Claude Code CLI, so their wall time includes CLI start-up and the tool schema; one pass of 82 decisions each. Jev was called directly over HTTPS from one Mac on a home network: 246 calls in a 35-second window, client wall time with the network inside it. Its API reports no server time, so Jev has no model-time point. These are different routes: the chart shows what a caller waits per decision, not model compute time.",["agent-routing","agent-jev-live"]],["routing-economics-scenarios","Thought experiment: recorded agent work under different model mixes","50 benchmark runs, 2,362 model calls, repriced","bar","usd","USD for all runs","\u0001",[{"name":"Repriced cost","points":{"$k":["label","value","highlight"],"$r":[["all Fable 5.1",369.58,false],["all Opus 5.5",170.92,false],["policy (Opus strong, Haiku ancillary)",161.62,false],["all Sonnet 5.5",108.54,true],["split (Sonnet main line, Haiku ancillary)",105.53,false],["all Haiku 4.5",54.27,false]]}}],"Calculation, not a run: every recorded call ran on Sonnet 5.5 with routing off. Same tokens on every model; a different model or mix would take a different path.",["agent-routing","calc-repricing","price-anthropic"]],["routing-economics-by-stage","Thought experiment: repriced cost by pipeline stage","\u0001","grouped-bar","usd","USD","\u0001",{"$k":["name","points"],"$r":[["Haiku 4.5",{"$k":["label","value"],"$r":[["act (strong)",28.66],["research (strong)",14.31],["verify (strong)",4.74],["review (strong)",3.36],["memory and onboarding (economy)",3.11],["other (standard)",0.1]]}],["Sonnet 5.5",{"$k":["label","value"],"$r":[["act (strong)",57.31],["research (strong)",28.62],["verify (strong)",9.48],["review (strong)",6.71],["memory and onboarding (economy)",6.22],["other (standard)",0.19]]}],["Opus 5.5",{"$k":["label","value"],"$r":[["act (strong)",78.86],["research (strong)",47.24],["verify (strong)",18.9],["review (strong)",13.21],["memory and onboarding (economy)",12.32],["other (standard)",0.39]]}],["Fable 5.1",{"$k":["label","value"],"$r":[["act (strong)",152.43],["research (strong)",105.58],["verify (strong)",47.18],["review (strong)",32.78],["memory and onboarding (economy)",30.65],["other (standard)",0.96]]}]]},"Calculation, not a run. The tier in brackets is the routing policy tier for that stage.",["agent-routing","calc-repricing","price-anthropic"]]]}}}