{"i":7,"comparison":{"slug":"jev-1-13-vs-claude-haiku-4-5","a":"jev-1-13","b":"claude-haiku-4-5","title":"Jev 1.13 vs Claude Haiku 4.5","seoTitle":"Jev 1.13 vs Claude Haiku 4.5: measured benchmarks","description":"Jev 1.13 vs Claude Haiku 4.5: 17 measured metrics from 3 studies (Typed routing decisions answered exactly right; more), with sample sizes and intervals.","verdict":"Jev 1.13 and Claude Haiku 4.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 12,543 ms; Time per routing decision, by route (Wall time), 0.14 s vs 9.44 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Typed routing decisions answered exactly right",0.8984,0.8902,"rate","90%","89% (73/82)","tie","The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Haiku 4.5 80% to 94%), so this sample cannot separate them.","routing-jev-vs-llm",82,"routing-exact-decisions",82,82,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.8191,0.9497],[0.8044,0.9412],"\u0001"],["Per-question accuracy",0.9485,0.9433,"rate","95% (184/194)","94% (183/194)","tie","The 95% intervals overlap (Jev 1.13 91% to 97%; Claude Haiku 4.5 90% to 97%), so this sample cannot separate them.","routing-jev-vs-llm",194,"routing-key-accuracy",194,194,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.9077,0.9718],[0.9013,0.968],"\u0001"],["Exact rate by decision type: Failure class",1,0.9444,"rate","100% (18/18)","94% (17/18)","tie","The 95% intervals overlap (Jev 1.13 82% to 100%; Claude Haiku 4.5 74% to 99%), so this sample cannot separate them.","routing-jev-vs-llm",18,"routing-exact-by-decision",18,18,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.8241,1],[0.7424,0.9901],"\u0001"],["Exact rate by decision type: Message intent",1,1,"rate","100% (20/20)","100% (20/20)","tie","The 95% intervals overlap (Jev 1.13 84% to 100%; Claude Haiku 4.5 84% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",20,"routing-exact-by-decision",20,20,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.8389,1],[0.8389,1],"\u0001"],["Exact rate by decision type: Is it a rule?",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Jev 1.13 76% to 100%; Claude Haiku 4.5 76% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",12,"routing-exact-by-decision",12,12,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Exact rate by decision type: Context shape",0.7396,0.75,"rate","74%","75% (24/32)","tie","The 95% intervals overlap (Jev 1.13 58% to 87%; Claude Haiku 4.5 58% to 87%), so this sample cannot separate them.","routing-jev-vs-llm",32,"routing-exact-by-decision",32,32,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.5789,0.8675],[0.5789,0.8675],"\u0001"],["Cost per 1,000 routing decisions",0.0337,8.924,"usd","$0.034","$8.92","unclear","No interval or range was recorded for either side, so the gap ($0.034 vs $8.92, 265x) is not tested against run-to-run variation.","routing-jev-vs-llm","\u0001","routing-cost-per-1000",246,82,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Time to make one routing decision",136.5,12543,"ms","137 ms","12,543 ms","a","Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Haiku 4.5 12,543 ms to 34,481 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone.","routing-overhead","\u0001","router-overhead-decision-latency",246,82,"routing overhead per decision · TypeSafe API","thinking on · via Claude Code · routing overhead per decision","range","p50-p95",[136.5,195.7],[12543,34481],"\u0001"],["Routing calls that returned a decision",1,1,"rate","100% (246/246)","100% (82/82)","tie","The 95% intervals overlap (Jev 1.13 98% to 100%; Claude Haiku 4.5 96% to 100%), so this sample cannot separate them.","routing-overhead","\u0001","router-overhead-completed",246,82,"routing overhead per decision · TypeSafe API","thinking on · via Claude Code · routing overhead per decision","ci95","ci95",[0.9846,1],[0.9552,1],"\u0001"],["Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))",1.67,441.74,"usd","$1.67","$441.74","unclear","No interval or range was recorded for either side, so the gap ($1.67 vs $441.74, 265x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-cost-per-1000-tasks","\u0001","\u0001","calculation per 1,000 tasks from recorded decision counts · TypeSafe API","thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))",0.24,62.47,"usd","$0.24","$62.47","unclear","No interval or range was recorded for either side, so the gap ($0.24 vs $62.47, 260x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-cost-per-1000-tasks","\u0001","\u0001","calculation per 1,000 tasks from recorded decision counts · TypeSafe API","thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Every model call routed (49.5 per task))",6.7568,620.8785,"seconds","6.76 s","620.9 s","unclear","No interval or range was recorded for either side, so the gap (6.76 s vs 620.9 s, 92x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-delay-per-task","\u0001","\u0001","calculation per task from recorded decision counts, decisions in line · TypeSafe API","thinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Only System One decisions (7 per task))",0.9555,87.801,"seconds","0.96 s","87.8 s","unclear","No interval or range was recorded for either side, so the gap (0.96 s vs 87.8 s, 92x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-delay-per-task","\u0001","\u0001","calculation per task from recorded decision counts, decisions in line · TypeSafe API","thinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Unseen routing decisions answered exactly right",0.8214,0.7857,"rate","82% (46/56)","79% (44/56)","tie","The 95% intervals overlap (Jev 1.13 70% to 90%; Claude Haiku 4.5 66% to 87%), so this sample cannot separate them.","routing-holdout",56,"routing-holdout-exact",56,56,"","Claude Code","ci95","ci95",[0.7016,0.9],[0.6618,0.8729],"\u0001"],["Per-question accuracy on unseen decisions",0.904,0.816,"rate","90% (113/125)","82% (102/125)","tie","The 95% intervals overlap (Jev 1.13 84% to 94%; Claude Haiku 4.5 74% to 87%), so this sample cannot separate them.","routing-holdout",125,"routing-holdout-key-accuracy",125,125,"","Claude Code","ci95","ci95",[0.8397,0.9442],[0.739,0.8741],"\u0001"],["Exact rate on unseen decisions, by decision type: Failure class",0.9286,0.9286,"rate","93% (13/14)","93% (13/14)","tie","The 95% intervals overlap (Jev 1.13 69% to 99%; Claude Haiku 4.5 69% to 99%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code","ci95","ci95",[0.6853,0.9873],[0.6853,0.9873],"\u0001"],["Exact rate on unseen decisions, by decision type: Message intent",0.8571,0.9286,"rate","86% (12/14)","93% (13/14)","tie","The 95% intervals overlap (Jev 1.13 60% to 96%; Claude Haiku 4.5 69% to 99%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code","ci95","ci95",[0.6006,0.9599],[0.6853,0.9873],"\u0001"],["Exact rate on unseen decisions, by decision type: Is it a rule?",0.9286,0.9286,"rate","93% (13/14)","93% (13/14)","tie","The 95% intervals overlap (Jev 1.13 69% to 99%; Claude Haiku 4.5 69% to 99%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code","ci95","ci95",[0.6853,0.9873],[0.6853,0.9873],"\u0001"],["Exact rate on unseen decisions, by decision type: Context shape",0.5714,0.3571,"rate","57% (8/14)","36% (5/14)","tie","The 95% intervals overlap (Jev 1.13 33% to 79%; Claude Haiku 4.5 16% to 61%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code","ci95","ci95",[0.3259,0.7862],[0.1634,0.6124],"\u0001"],["Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm))",0.9024,0.8902,"rate","90% (74/82)","89% (73/82)","tie","The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Haiku 4.5 80% to 94%), so this sample cannot separate them.","routing-holdout",82,"routing-holdout-tuned-vs-unseen",82,82,"","Claude Code","ci95","ci95",[0.8191,0.9497],[0.8044,0.9412],"\u0001"],["Tuned case set vs unseen holdout: exact rate per router (Unseen holdout)",0.8214,0.7857,"rate","82% (46/56)","79% (44/56)","tie","The 95% intervals overlap (Jev 1.13 70% to 90%; Claude Haiku 4.5 66% to 87%), so this sample cannot separate them.","routing-holdout",56,"routing-holdout-tuned-vs-unseen",56,56,"","Claude Code","ci95","ci95",[0.7016,0.9],[0.6618,0.8729],"\u0001"],["Time per routing decision, by route (Wall time)",0.139,9.444,"seconds","0.14 s","9.44 s","a","Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Haiku 4.5 9.44 s to 25.4 s); not a confidence interval.","routing-holdout","\u0001","routing-holdout-latency",168,56,"","Claude Code","range","p50-p95",[0.139,0.192],[9.444,25.413],"\u0001"],["Cost per 1,000 unseen routing decisions",0.03065,7.129,"usd","$0.031","$7.13","unclear","No interval or range was recorded for either side, so the gap ($0.031 vs $7.13, 233x) is not tested against run-to-run variation.","routing-holdout","\u0001","routing-holdout-cost-per-1000",168,56,"","Claude Code","\u0001","\u0001","\u0001","\u0001",true]]}}}