{"$k":["slug","chart"],"$r":[["routing-jev-vs-llm",{"id":"routing-exact-decisions","title":"Typed routing decisions answered exactly right","subtitle":"Share of asked cases where every scored question was acceptable","kind":"dot-range","unit":"rate","yLabel":"Exact","whisker":"ci95","series":[{"name":"Exact rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8984,0.8191,0.9497,82,true],["Claude Haiku 4.5",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5",0.939,0.8651,0.9737,82,false]]}}],"note":"Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.","sourceIds":["agent-routing","agent-jev-live"]}],["routing-jev-vs-llm",{"id":"routing-key-accuracy","title":"Per-question accuracy","subtitle":"Each open question the router was asked; an unanswered question counts as wrong","kind":"dot-range","unit":"rate","yLabel":"Correct answers","whisker":"ci95","series":[{"name":"Key accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.9485,0.9077,0.9718,194,true],["Claude Haiku 4.5",0.9433,0.9013,0.968,194,false],["Claude Sonnet 5.5",0.9742,0.9411,0.9889,194,false]]}}],"note":"Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.","sourceIds":["agent-routing","agent-jev-live"]}],["routing-jev-vs-llm",{"id":"routing-exact-by-decision","title":"Exact rate by decision type","kind":"grouped-bar","unit":"rate","yLabel":"Exact","series":{"$k":["name","points"],"$r":[["Jev 1.13 (TypeSafe)",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",1,0.8241,1,18,true],["Message intent",1,0.8389,1,20,true],["Is it a rule?",1,0.7575,1,12,true],["Context shape",0.7396,0.5789,0.8675,32,true]]}],["Claude Haiku 4.5",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",0.9444,0.7424,0.9901,18,false],["Message intent",1,0.8389,1,20,false],["Is it a rule?",1,0.7575,1,12,false],["Context shape",0.75,0.5789,0.8675,32,false]]}],["Claude Sonnet 5.5",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",1,0.8241,1,18,false],["Message intent",1,0.8389,1,20,false],["Is it a rule?",1,0.7575,1,12,false],["Context shape",0.8438,0.6825,0.9314,32,false]]}]]},"note":"A missing bar means that router was not run on that decision type. Whiskers are 95% Wilson intervals. Jev: the pooled rate of the live repeats, with the interval taken at the number of decisions of that type.","sourceIds":["agent-routing","agent-jev-live"]}],["routing-jev-vs-llm",{"id":"routing-cost-per-1000","title":"Cost per 1,000 routing decisions","subtitle":"List price × reported tokens per decision","kind":"bar","unit":"usd","yLabel":"USD per 1,000 decisions","series":[{"name":"Cost","points":{"$k":["label","value","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.0337,246,true],["Claude Haiku 4.5",8.924,82,false],["Claude Sonnet 5.5",4.996,82,false]]}}],"note":"List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.","sourceIds":["agent-routing","calc-repricing","price-jev","price-anthropic","agent-jev-live"]}],["routing-overhead",{"id":"router-overhead-decision-latency","title":"Time to make one routing decision","subtitle":"Median; whiskers = median to 95th percentile","kind":"dot-range","unit":"ms","yLabel":"Time per decision","whisker":"p50-p95","series":[{"name":"Decision time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Deterministic routing policy (Agent, in process)",0.00142,0.00142,0.00233,20000,true],["Jev 1.13 (TypeSafe)",136.5,136.5,195.7,246,"\u0001"],["Claude Sonnet 5.5 (effort low, via Claude Code)",2597,2597,4298,82,"\u0001"],["Claude Haiku 4.5 (thinking on, via Claude Code)",12543,12543,34481,82,"\u0001"]]}}],"note":"The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.","sourceIds":["agent-routing-overhead","agent-routing","agent-jev-live"]}],["routing-overhead",{"id":"router-overhead-completed","title":"Routing calls that returned a decision","subtitle":"Completed calls ÷ calls; whiskers = 95% Wilson interval","kind":"dot-range","unit":"rate","yLabel":"Completed","whisker":"ci95","series":[{"name":"Completed","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Deterministic routing policy (Agent, in process)",1,0.9998,1,20000,true],["Jev 1.13 (TypeSafe)",1,0.9846,1,246,"\u0001"],["Claude Sonnet 5.5 (effort low, via Claude Code)",1,0.9552,1,82,"\u0001"],["Claude Haiku 4.5 (thinking on, via Claude Code)",1,0.9552,1,82,"\u0001"]]}}],"note":"A completed call returned a decision, right or wrong (accuracy is in the routing study). Whiskers are 95% Wilson intervals.","sourceIds":["agent-routing-overhead","agent-routing","agent-jev-live"]}],["routing-overhead",{"id":"router-overhead-cost-per-1000-tasks","title":"Added routing cost per 1,000 tasks (calculation)","subtitle":"Decisions per task from recorded runs × cost per decision","kind":"grouped-bar","unit":"usd","yLabel":"USD per 1,000 tasks","series":[{"name":"Every model call routed (49.5 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0],["Jev 1.13 (TypeSafe)",1.67],["Claude Sonnet 5.5 (effort low, via Claude Code)",247.3],["Claude Haiku 4.5 (thinking on, via Claude Code)",441.74]]}},{"name":"Only System One decisions (7 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0],["Jev 1.13 (TypeSafe)",0.24],["Claude Sonnet 5.5 (effort low, via Claude Code)",34.97],["Claude Haiku 4.5 (thinking on, via Claude Code)",62.47]]}}],"note":"A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).","sourceIds":["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]}],["routing-overhead",{"id":"router-overhead-delay-per-task","title":"Added routing delay per task (calculation)","subtitle":"Decisions per task × median decision time, if every decision waits in line","kind":"grouped-bar","unit":"seconds","yLabel":"Seconds per task","series":[{"name":"Every model call routed (49.5 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0.0000703],["Jev 1.13 (TypeSafe)",6.7568],["Claude Sonnet 5.5 (effort low, via Claude Code)",128.5515],["Claude Haiku 4.5 (thinking on, via Claude Code)",620.8785]]}},{"name":"Only System One decisions (7 per task)","points":{"$k":["label","value"],"$r":[["Deterministic routing policy (Agent, in process)",0.0000099],["Jev 1.13 (TypeSafe)",0.9555],["Claude Sonnet 5.5 (effort low, via Claude Code)",18.179],["Claude Haiku 4.5 (thinking on, via Claude Code)",87.801]]}}],"note":"A calculation and an upper bound: it assumes each decision waits for the one before. Median recorded task wall time: 10.3 min. Jev’s delay uses its live median over the API from one Mac (network included); the Claude routers’ includes the CLI.","sourceIds":["agent-routing-overhead","calc-routing-overhead","agent-routing","price-anthropic","price-jev","agent-jev-live"]}],["routing-holdout",{"id":"routing-holdout-exact","title":"Unseen routing decisions answered exactly right","subtitle":"Share of the 56 holdout cases where every scored question was acceptable","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Exact","series":[{"name":"Exact rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8214,0.7016,0.9,56,true],["Claude Haiku 4.5 · Claude Code",0.7857,0.6618,0.8729,56,false],["Claude Sonnet 5.5 (low) · Claude Code",0.875,0.7637,0.9381,56,false]]}}],"note":"Whiskers are 95% Wilson intervals. Cases written blind to every router answer, then frozen. Jev: its first of three repetitions.","whisker":"ci95","sourceIds":["agent-routing-holdout"]}],["routing-holdout",{"id":"routing-holdout-key-accuracy","title":"Per-question accuracy on unseen decisions","subtitle":"Each open question a router was asked; an unanswered question counts as wrong","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Correct answers","series":[{"name":"Key accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.904,0.8397,0.9442,125,true],["Claude Haiku 4.5 · Claude Code",0.816,0.739,0.8741,125,false],["Claude Sonnet 5.5 (low) · Claude Code",0.92,0.859,0.956,125,false]]}}],"note":"Whiskers are nominal 95% Wilson intervals. Context-shape cases ask up to 8 questions each, the other decision types one. Questions in one case are not independent; these intervals do not adjust for that grouping.","whisker":"ci95","sourceIds":["agent-routing-holdout"]}],["routing-holdout",{"id":"routing-holdout-by-purpose","title":"Exact rate on unseen decisions, by decision type","subtitle":"14 cases per decision type","kind":"grouped-bar","unit":"rate","polarity":"higher","yLabel":"Exact","series":{"$k":["name","points"],"$r":[["Jev 1.13 (TypeSafe)",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",0.9286,0.6853,0.9873,14,true],["Message intent",0.8571,0.6006,0.9599,14,true],["Is it a rule?",0.9286,0.6853,0.9873,14,true],["Context shape",0.5714,0.3259,0.7862,14,true]]}],["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",0.9286,0.6853,0.9873,14,false],["Message intent",0.9286,0.6853,0.9873,14,false],["Is it a rule?",0.9286,0.6853,0.9873,14,false],["Context shape",0.3571,0.1634,0.6124,14,false]]}],["Claude Sonnet 5.5 (low) · Claude Code",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",1,0.7847,1,14,false],["Message intent",1,0.7847,1,14,false],["Is it a rule?",0.9286,0.6853,0.9873,14,false],["Context shape",0.5714,0.3259,0.7862,14,false]]}]]},"note":"Whiskers are 95% Wilson intervals. With 14 cases a perfect score has an interval of 78% to 100%, so a decision type where every router scores 14 of 14 is at its ceiling and cannot rank them.","whisker":"ci95","sourceIds":["agent-routing-holdout"]}],["routing-holdout",{"id":"routing-holdout-tuned-vs-unseen","title":"Tuned case set vs unseen holdout: exact rate per router","subtitle":"Tuned set: the routing study’s 82 cases, revised against Jev answers. Holdout: 56 new cases, frozen before any router call","kind":"grouped-bar","unit":"rate","polarity":"higher","yLabel":"Exact","series":[{"name":"Tuned set (routing-jev-vs-llm)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.9024,0.8191,0.9497,82,true],["Claude Haiku 4.5 · Claude Code",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5 (low) · Claude Code",0.939,0.8651,0.9737,82,false]]}},{"name":"Unseen holdout","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8214,0.7016,0.9,56,true],["Claude Haiku 4.5 · Claude Code",0.7857,0.6618,0.8729,56,false],["Claude Sonnet 5.5 (low) · Claude Code",0.875,0.7637,0.9381,56,false]]}}],"note":"Whiskers are 95% Wilson intervals. The two case sets differ in mix and size, so a gap mixes a change of case set with any change in the router, and the data cannot separate them. A gap counts only when the two intervals do not overlap.","whisker":"ci95","sourceIds":["agent-routing-holdout","agent-routing"]}],["routing-holdout",{"id":"routing-holdout-latency","title":"Time per routing decision, by route","subtitle":"Median, whisker to the 95th percentile","kind":"dot-range","unit":"seconds","yLabel":"Seconds","series":[{"name":"Wall time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.139,0.139,0.192,168,true],["Claude Haiku 4.5 · Claude Code",9.444,9.444,25.413,56,false],["Claude Sonnet 5.5 (low) · Claude Code",2.359,2.359,3.657,56,false]]}},{"name":"Model time (API, CLI-reported)","points":[{"label":"Claude Haiku 4.5 · Claude Code","value":7.522,"lo":7.522,"hi":23.913,"n":56},{"label":"Claude Sonnet 5.5 (low) · Claude Code","value":1.485,"lo":1.485,"hi":2.377,"n":56}]}],"note":"The routes differ. Jev: one HTTPS call to the vendor API from one Mac, all three repetitions. Claude: the Claude Code CLI from process start to exit, which adds start-up time a direct API call would not. The whisker runs from the median to the 95th percentile; it is not a confidence interval. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.","whisker":"p50-p95","sourceIds":["agent-routing-holdout"]}],["routing-holdout",{"id":"routing-holdout-cost-per-1000","title":"Cost per 1,000 unseen routing decisions","subtitle":"Reported tokens per decision × list price","kind":"bar","unit":"usd","yLabel":"USD per 1,000 decisions","series":[{"name":"Cost","points":{"$k":["label","value","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.03065,168,true],["Claude Haiku 4.5 · Claude Code",7.129,56,false],["Claude Sonnet 5.5 (low) · Claude Code",7.244,56,false]]}}],"note":"Calculation, not a bill. Jev: reported input tokens × $0.042 per million, output free. Claude: CLI-reported tokens × list price; the CLI wrote its prompt cache as 1-hour writes, priced at 2× input, and adds its own system prompt and tool-schema tokens. The Claude routers ran on a subscription. At the tuned-set study’s convention (every cache write at 1.25× input), Sonnet 5.5 (low) would be $4.877 here. That study shows $4.996 for it on its own cases, where the CLI reported $7.324, so its cost row and this one differ by convention and by case mix.","sourceIds":["agent-routing-holdout","calc-repricing","price-jev","price-anthropic"]}]]}