{"i":21,"study":{"slug":"routing-holdout","title":"Jev vs Claude routers on unseen decisions: a blind holdout","seoTitle":"Jev vs Claude router accuracy on unseen decisions","description":"Jev 1.13, Claude Haiku 4.5 and Sonnet 5.5 on 56 routing decisions written blind and frozen first: exact rate with 95% intervals, tuned vs unseen.","question":"Does a router keep its accuracy on typed routing decisions that nobody tuned against its answers?","answer":"The holdout does not establish a router ranking. The author wrote 56 new decisions blind to router answers. Labels froze before the first router call. Jev 1.13 (TypeSafe): 46 of 56 (82%, 95% interval 70% to 90%). Claude Haiku 4.5 · Claude Code: 44 of 56 (79%, 95% interval 66% to 87%). Claude Sonnet 5.5 (low) · Claude Code: 49 of 56 (88%, 95% interval 76% to 94%). The intervals overlap, so the holdout does not rank the routers. The exact McNemar tests find no significant difference (p from 0.18 to 0.688). These paired tests have no multiple-test correction. Tuned set → holdout exact rate (calculation; counts and intervals in the chart): Jev 1.13 90% → 82%, Haiku 4.5 89% → 79%, Sonnet 5.5 (low) 94% → 88%. All 3 observed rates were lower. Each router’s tuned and holdout intervals overlap. The data does not show a clear drop for any router. Decision groups (calculation; counts and intervals in the table). On the 3 one-question types the observed exact rate changed by Jev 1.13 −9.5, Haiku 4.5 −5.1, Sonnet 5.5 (low) −2.4 points. These are calculations, not a ranking of drops. On context shape the observed exact rate changed by Jev 1.13 −17.9, Haiku 4.5 −39.3, Sonnet 5.5 (low) −27.2 points. These are calculations, not a ranking of drops. Every tuned and holdout pair of intervals overlaps. The data neither shows nor rules out a home advantage for Jev. The 3 drops range from 6.4 to 10.5 points (calculation). The sets differ in mix: context shape is 32 of 82 tuned cases and 14 of 56 holdout cases. Secondary scoring keeps only questions where our set accepts the second label. Our set accepts 115 of 125 (92%, 95% interval 86% to 96%). Our first label agrees on 100 of 125 (80%, 95% interval 72% to 86%). Jev 1.13: 46 of 56 (82%, 95% interval 70% to 90%). Haiku 4.5: 46 of 56 (82%, 95% interval 70% to 90%). Sonnet 5.5 (low): 52 of 56 (93%, 95% interval 83% to 97%). The weakest label is “artifacts”. Our set accepts the second label on 1 of 9 (11%, 95% interval 2% to 43%). Set-aware kappa is 0 (calculation). This label broke our protocol. Matches to our label: Jev 1.13: 9 of 9 (100%, 95% interval 70% to 100%). Haiku 4.5: 4 of 9 (44%, 95% interval 19% to 73%). Sonnet 5.5 (low): 7 of 9 (78%, 95% interval 45% to 94%). See the caveats. Exact by decision type, of 14 each (95% intervals in the chart): Failure class 13 to 14; Message intent 12 to 14; Is it a rule? 13; Context shape 5 to 8. Failure class, Message intent and Is it a rule? are near the ceiling for every router (85% or more), so most differences come from context shape. Cost per 1,000 decisions (calculation): Jev 1.13 $0.0307, Haiku 4.5 $7.13, Sonnet 5.5 (low) $7.24. Median time per decision: Jev 1.13: 139 ms (p50–p95 139 ms to 192 ms, n = 168). Haiku 4.5: 9.44 s (p50–p95 9.44 s to 25.41 s, n = 56). Sonnet 5.5 (low): 2.36 s (p50–p95 2.36 s to 3.66 s, n = 56). These ranges are not confidence intervals. Jev used direct HTTPS; Claude used its CLI. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.","date":"2026-10-06","updated":"2026-10-06","tags":["routing","jev","claude-haiku","claude-sonnet","model-routing","holdout","benchmark-method"],"caveats":["The case author and two of the three routers are Claude models, so a same-family label bias is possible. A second labeller (GPT-6.1 Sol) labelled every case blind; the secondary scoring keeps only the keys where its label is inside our acceptable set.","56 cases (14 per decision type) from one author: per-type intervals are very wide.","The routes differ: Jev ran over direct HTTPS, the Claude routers through the Claude Code CLI, which adds start-up time and tokens.","Jev figures are its first repetition; repetitions 2 and 3 and the stability across all three are reported beside it.","The artifacts label violates the protocol rule to leave judgement calls unscored. The author accepted only no on all 9 labelled cases. The second labeller chose no on 1 of 9. Frozen labels stay unchanged; sensitivity calculations are not new runs.","The protocol file was created at 21:35:45.935 UTC on 2026-10-06. It followed the case drafts and an uncounted labeller probe. It preceded the first counted labelling call at 21:35:49.3 UTC, the label freeze at 21:37:02.582 UTC and every router call. It was not written before every call.","The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.","The three one-question case sets are near a ceiling: every router scored 12 to 14 of 14 on each. The 95% Wilson interval for 14 of 14 is 78.5% to 100%. These sets provide little room to separate routers.","Per-question Wilson intervals treat questions as independent. Several questions share each context-shape case, so those intervals can understate uncertainty. Label-agreement intervals have the same limit.","Set-aware kappa uses the second label as the author label when the acceptable set contains it. This adaptive calculation can inflate agreement. Only kappaStrict compares two fixed labels.","The tuned set was revised against Jev answers. Its size and decision mix differ from the holdout. Overlapping intervals neither show nor rule out a home advantage.","The three paired McNemar tests have no multiple-test correction. Their p values do not prove equal accuracy.","Failure class, Message intent and Is it a rule?: every router answered at least 85% of the 14 cases exactly. These case sets are near the ceiling and barely separate the routers; context shape carries most of the differences.","Agreement with the second labeller counts a label inside our acceptable set: 115 of 125 questions (92%). 24 of the 125 questions accept two or three labels, so agreement with our first label alone is lower: 100 of 125 (80%).","The “artifacts” label broke our own rules. The second labeller’s answer was inside our set on 1 of 9 cases (kappa 0). We labelled “no” on every one of them. The protocol leaves out a question whose answer is a judgement call, and the production case suites accept both answers for a fresh work item. Cases where the router’s answer equalled our label: Jev 1.13 9 of 9, Haiku 4.5 4 of 9, Sonnet 5.5 (low) 7 of 9. The scoring against our labels therefore favours Jev on this question. The labels stayed frozen.\n\nWith both answers accepted for “artifacts” and the same answers scored again (a calculation), exact answers are Jev 1.13 46 of 56 (82%, 95% interval 70% to 90%); Haiku 4.5 45 of 56 (80%, 95% interval 68% to 89%); Sonnet 5.5 (low) 50 of 56 (89%, 95% interval 79% to 95%). Context shape: Jev 1.13 8 of 14 (57%, 95% interval 33% to 79%); Haiku 4.5 6 of 14 (43%, 95% interval 21% to 67%); Sonnet 5.5 (low) 9 of 14 (64%, 95% interval 39% to 84%); the overall intervals still overlap.","Label spread is narrow on some questions. One label holds at least 90% of the cases for artifacts (“no”, 9 of 9) and memories (“yes”, 9 of 10). One option has a single case for knowledge (“none”), scope (“large”) and transcript (“summary”), so those options are barely tested.","The tuned and holdout sets differ in size and mix, so a change between them cannot be assigned to the router or to the case set. The data neither shows nor rules out a home advantage."],"sourceIds":["agent-routing-holdout","agent-routing","calc-repricing","price-jev","price-anthropic"],"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["holdout-exact-jev","Jev 1.13 (TypeSafe): exact on unseen decisions",0.8214,"rate","82% (46/56)",56,[0.7016,0.9],"\u0001"],["holdout-exact-claude-haiku","Claude Haiku 4.5 · Claude Code: exact on unseen decisions",0.7857,"rate","79% (44/56)",56,[0.6618,0.8729],"\u0001"],["holdout-exact-claude-sonnet","Claude Sonnet 5.5 (low) · Claude Code: exact on unseen decisions",0.875,"rate","88% (49/56)",56,[0.7637,0.9381],"\u0001"],["holdout-jev-stability","Jev 1.13 (TypeSafe): same answers on every key in 3 repetitions",0.9464,"rate","95% (53/56)",56,[0.8539,0.9816],"\u0001"],["holdout-jev-reps-range","Jev 1.13 (TypeSafe): exact in each of 3 repetitions",0.8214,"rate","46, 46, 47 of 56",56,"\u0001","Range 46 to 47 of 56; not a confidence interval. The value is the first repetition."],["holdout-label-agreement","Second labeller (GPT-6.1 Sol): inside our acceptable label sets",0.92,"rate","92% (115/125)",125,[0.859,0.956],"Per labelled question, across all four decision types. 24 of the 125 questions accept two or three labels, so this counts a second label that is inside our set, not only our first label. First label only: 100 of 125 (80%)."],["holdout-label-agreement-first","Second labeller (GPT-6.1 Sol): equal to our first label",0.8,"rate","80% (100/125)",125,[0.7214,0.8607],"Per labelled question. Our first label is the least inclusive one in our set. The inside-our-set figure is holdout-label-agreement."],["holdout-gap-jev","Jev 1.13 (TypeSafe): holdout minus tuned-set exact rate",-0.081,"rate","−8.1 points",56,"\u0001","Calculation. Tuned 90% (82% to 95%, n = 82); holdout 82% (70% to 90%, n = 56). The intervals overlap."],["holdout-gap-claude-haiku","Claude Haiku 4.5 · Claude Code: holdout minus tuned-set exact rate",-0.1045,"rate","−10.5 points",56,"\u0001","Calculation. Tuned 89% (80% to 94%, n = 82); holdout 79% (66% to 87%, n = 56). The intervals overlap."],["holdout-gap-claude-sonnet","Claude Sonnet 5.5 (low) · Claude Code: holdout minus tuned-set exact rate",-0.064,"rate","−6.4 points",56,"\u0001","Calculation. Tuned 94% (87% to 97%, n = 82); holdout 88% (76% to 94%, n = 56). The intervals overlap."],["holdout-jev-cost-per-1000","Jev 1.13 (TypeSafe): cost per 1,000 unseen decisions",0.03065,"usd","$0.0307",168,"\u0001","Calculation: reported input tokens × $0.042 per million."]]},"charts":{"$k":["id","title","subtitle","kind","unit","polarity","yLabel","series","note","whisker","sourceIds"],"$r":[["routing-holdout-exact","Unseen routing decisions answered exactly right","Share of the 56 holdout cases where every scored question was acceptable","dot-range","rate","higher","Exact",[{"name":"Exact rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8214,0.7016,0.9,56,true],["Claude Haiku 4.5 · Claude Code",0.7857,0.6618,0.8729,56,false],["Claude Sonnet 5.5 (low) · Claude Code",0.875,0.7637,0.9381,56,false]]}}],"Whiskers are 95% Wilson intervals. Cases written blind to every router answer, then frozen. Jev: its first of three repetitions.","ci95",["agent-routing-holdout"]],["routing-holdout-key-accuracy","Per-question accuracy on unseen decisions","Each open question a router was asked; an unanswered question counts as wrong","dot-range","rate","higher","Correct answers",[{"name":"Key accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.904,0.8397,0.9442,125,true],["Claude Haiku 4.5 · Claude Code",0.816,0.739,0.8741,125,false],["Claude Sonnet 5.5 (low) · Claude Code",0.92,0.859,0.956,125,false]]}}],"Whiskers are nominal 95% Wilson intervals. Context-shape cases ask up to 8 questions each, the other decision types one. Questions in one case are not independent; these intervals do not adjust for that grouping.","ci95",["agent-routing-holdout"]],["routing-holdout-by-purpose","Exact rate on unseen decisions, by decision type","14 cases per decision type","grouped-bar","rate","higher","Exact",{"$k":["name","points"],"$r":[["Jev 1.13 (TypeSafe)",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",0.9286,0.6853,0.9873,14,true],["Message intent",0.8571,0.6006,0.9599,14,true],["Is it a rule?",0.9286,0.6853,0.9873,14,true],["Context shape",0.5714,0.3259,0.7862,14,true]]}],["Claude Haiku 4.5 · Claude Code",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",0.9286,0.6853,0.9873,14,false],["Message intent",0.9286,0.6853,0.9873,14,false],["Is it a rule?",0.9286,0.6853,0.9873,14,false],["Context shape",0.3571,0.1634,0.6124,14,false]]}],["Claude Sonnet 5.5 (low) · Claude Code",{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Failure class",1,0.7847,1,14,false],["Message intent",1,0.7847,1,14,false],["Is it a rule?",0.9286,0.6853,0.9873,14,false],["Context shape",0.5714,0.3259,0.7862,14,false]]}]]},"Whiskers are 95% Wilson intervals. With 14 cases a perfect score has an interval of 78% to 100%, so a decision type where every router scores 14 of 14 is at its ceiling and cannot rank them.","ci95",["agent-routing-holdout"]],["routing-holdout-tuned-vs-unseen","Tuned case set vs unseen holdout: exact rate per router","Tuned set: the routing study’s 82 cases, revised against Jev answers. Holdout: 56 new cases, frozen before any router call","grouped-bar","rate","higher","Exact",[{"name":"Tuned set (routing-jev-vs-llm)","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.9024,0.8191,0.9497,82,true],["Claude Haiku 4.5 · Claude Code",0.8902,0.8044,0.9412,82,false],["Claude Sonnet 5.5 (low) · Claude Code",0.939,0.8651,0.9737,82,false]]}},{"name":"Unseen holdout","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.8214,0.7016,0.9,56,true],["Claude Haiku 4.5 · Claude Code",0.7857,0.6618,0.8729,56,false],["Claude Sonnet 5.5 (low) · Claude Code",0.875,0.7637,0.9381,56,false]]}}],"Whiskers are 95% Wilson intervals. The two case sets differ in mix and size, so a gap mixes a change of case set with any change in the router, and the data cannot separate them. A gap counts only when the two intervals do not overlap.","ci95",["agent-routing-holdout","agent-routing"]],["routing-holdout-latency","Time per routing decision, by route","Median, whisker to the 95th percentile","dot-range","seconds","\u0001","Seconds",[{"name":"Wall time","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.139,0.139,0.192,168,true],["Claude Haiku 4.5 · Claude Code",9.444,9.444,25.413,56,false],["Claude Sonnet 5.5 (low) · Claude Code",2.359,2.359,3.657,56,false]]}},{"name":"Model time (API, CLI-reported)","points":[{"label":"Claude Haiku 4.5 · Claude Code","value":7.522,"lo":7.522,"hi":23.913,"n":56},{"label":"Claude Sonnet 5.5 (low) · Claude Code","value":1.485,"lo":1.485,"hi":2.377,"n":56}]}],"The routes differ. Jev: one HTTPS call to the vendor API from one Mac, all three repetitions. Claude: the Claude Code CLI from process start to exit, which adds start-up time a direct API call would not. The whisker runs from the median to the 95th percentile; it is not a confidence interval. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.","p50-p95",["agent-routing-holdout"]],["routing-holdout-cost-per-1000","Cost per 1,000 unseen routing decisions","Reported tokens per decision × list price","bar","usd","\u0001","USD per 1,000 decisions",[{"name":"Cost","points":{"$k":["label","value","n","highlight"],"$r":[["Jev 1.13 (TypeSafe)",0.03065,168,true],["Claude Haiku 4.5 · Claude Code",7.129,56,false],["Claude Sonnet 5.5 (low) · Claude Code",7.244,56,false]]}}],"Calculation, not a bill. Jev: reported input tokens × $0.042 per million, output free. Claude: CLI-reported tokens × list price; the CLI wrote its prompt cache as 1-hour writes, priced at 2× input, and adds its own system prompt and tool-schema tokens. The Claude routers ran on a subscription. At the tuned-set study’s convention (every cache write at 1.25× input), Sonnet 5.5 (low) would be $4.877 here. That study shows $4.996 for it on its own cases, where the CLI reported $7.324, so its cost row and this one differ by convention and by case mix.","\u0001",["agent-routing-holdout","calc-repricing","price-jev","price-anthropic"]]]},"related":["routing-jev-vs-llm","routing-overhead"]}}