{"i":8,"comparison":{"slug":"jev-1-13-vs-claude-sonnet-5-5","a":"jev-1-13","b":"claude-sonnet-5-5","title":"Jev 1.13 vs Claude Sonnet 5.5","seoTitle":"Jev 1.13 vs Claude Sonnet 5.5: measured benchmarks","description":"Jev 1.13 vs Claude Sonnet 5.5: 17 measured metrics from 3 studies (Typed routing decisions answered exactly right; more), with sample sizes and intervals.","verdict":"Jev 1.13 and Claude Sonnet 5.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 2,597 ms; Time per routing decision, by route (Wall time), 0.14 s vs 2.36 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange","calculation"],"$r":[["Typed routing decisions answered exactly right",0.8984,0.939,"rate","90%","94% (77/82)","tie","The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.","routing-jev-vs-llm",82,"routing-exact-decisions",82,82,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.8191,0.9497],[0.8651,0.9737],"\u0001"],["Per-question accuracy",0.9485,0.9742,"rate","95% (184/194)","97% (189/194)","tie","The 95% intervals overlap (Jev 1.13 91% to 97%; Claude Sonnet 5.5 94% to 99%), so this sample cannot separate them.","routing-jev-vs-llm",194,"routing-key-accuracy",194,194,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.9077,0.9718],[0.9411,0.9889],"\u0001"],["Exact rate by decision type: Failure class",1,1,"rate","100% (18/18)","100% (18/18)","tie","The 95% intervals overlap (Jev 1.13 82% to 100%; Claude Sonnet 5.5 82% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",18,"routing-exact-by-decision",18,18,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.8241,1],[0.8241,1],"\u0001"],["Exact rate by decision type: Message intent",1,1,"rate","100% (20/20)","100% (20/20)","tie","The 95% intervals overlap (Jev 1.13 84% to 100%; Claude Sonnet 5.5 84% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",20,"routing-exact-by-decision",20,20,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.8389,1],[0.8389,1],"\u0001"],["Exact rate by decision type: Is it a rule?",1,1,"rate","100% (12/12)","100% (12/12)","tie","The 95% intervals overlap (Jev 1.13 76% to 100%; Claude Sonnet 5.5 76% to 100%), so this sample cannot separate them.","routing-jev-vs-llm",12,"routing-exact-by-decision",12,12,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.7575,1],[0.7575,1],"\u0001"],["Exact rate by decision type: Context shape",0.7396,0.8438,"rate","74%","84% (27/32)","tie","The 95% intervals overlap (Jev 1.13 58% to 87%; Claude Sonnet 5.5 68% to 93%), so this sample cannot separate them.","routing-jev-vs-llm",32,"routing-exact-by-decision",32,32,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","ci95","ci95",[0.5789,0.8675],[0.6825,0.9314],"\u0001"],["Cost per 1,000 routing decisions",0.0337,4.996,"usd","$0.034","$5.00","unclear","No interval or range was recorded for either side, so the gap ($0.034 vs $5.00, 148x) is not tested against run-to-run variation.","routing-jev-vs-llm","\u0001","routing-cost-per-1000",246,82,"typed routing decisions · TypeSafe API","typed routing decisions · via Claude Code","\u0001","\u0001","\u0001","\u0001",true],["Time to make one routing decision",136.5,2597,"ms","137 ms","2,597 ms","a","Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Sonnet 5.5 2,597 ms to 4,298 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone.","routing-overhead","\u0001","router-overhead-decision-latency",246,82,"routing overhead per decision · TypeSafe API","effort low · via Claude Code · routing overhead per decision","range","p50-p95",[136.5,195.7],[2597,4298],"\u0001"],["Routing calls that returned a decision",1,1,"rate","100% (246/246)","100% (82/82)","tie","The 95% intervals overlap (Jev 1.13 98% to 100%; Claude Sonnet 5.5 96% to 100%), so this sample cannot separate them.","routing-overhead","\u0001","router-overhead-completed",246,82,"routing overhead per decision · TypeSafe API","effort low · via Claude Code · routing overhead per decision","ci95","ci95",[0.9846,1],[0.9552,1],"\u0001"],["Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))",1.67,247.3,"usd","$1.67","$247.30","unclear","No interval or range was recorded for either side, so the gap ($1.67 vs $247.30, 148x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-cost-per-1000-tasks","\u0001","\u0001","calculation per 1,000 tasks from recorded decision counts · TypeSafe API","effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))",0.24,34.97,"usd","$0.24","$34.97","unclear","No interval or range was recorded for either side, so the gap ($0.24 vs $34.97, 146x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-cost-per-1000-tasks","\u0001","\u0001","calculation per 1,000 tasks from recorded decision counts · TypeSafe API","effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Every model call routed (49.5 per task))",6.7568,128.5515,"seconds","6.76 s","128.6 s","unclear","No interval or range was recorded for either side, so the gap (6.76 s vs 128.6 s, 19x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-delay-per-task","\u0001","\u0001","calculation per task from recorded decision counts, decisions in line · TypeSafe API","effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Added routing delay per task (calculation) (Only System One decisions (7 per task))",0.9555,18.179,"seconds","0.96 s","18.2 s","unclear","No interval or range was recorded for either side, so the gap (0.96 s vs 18.2 s, 19x) is not tested against run-to-run variation.","routing-overhead","\u0001","router-overhead-delay-per-task","\u0001","\u0001","calculation per task from recorded decision counts, decisions in line · TypeSafe API","effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line","\u0001","\u0001","\u0001","\u0001",true],["Unseen routing decisions answered exactly right",0.8214,0.875,"rate","82% (46/56)","88% (49/56)","tie","The 95% intervals overlap (Jev 1.13 70% to 90%; Claude Sonnet 5.5 76% to 94%), so this sample cannot separate them.","routing-holdout",56,"routing-holdout-exact",56,56,"","Claude Code · effort low","ci95","ci95",[0.7016,0.9],[0.7637,0.9381],"\u0001"],["Per-question accuracy on unseen decisions",0.904,0.92,"rate","90% (113/125)","92% (115/125)","tie","The 95% intervals overlap (Jev 1.13 84% to 94%; Claude Sonnet 5.5 86% to 96%), so this sample cannot separate them.","routing-holdout",125,"routing-holdout-key-accuracy",125,125,"","Claude Code · effort low","ci95","ci95",[0.8397,0.9442],[0.859,0.956],"\u0001"],["Exact rate on unseen decisions, by decision type: Failure class",0.9286,1,"rate","93% (13/14)","100% (14/14)","tie","The 95% intervals overlap (Jev 1.13 69% to 99%; Claude Sonnet 5.5 78% to 100%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code · effort low","ci95","ci95",[0.6853,0.9873],[0.7847,1],"\u0001"],["Exact rate on unseen decisions, by decision type: Message intent",0.8571,1,"rate","86% (12/14)","100% (14/14)","tie","The 95% intervals overlap (Jev 1.13 60% to 96%; Claude Sonnet 5.5 78% to 100%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code · effort low","ci95","ci95",[0.6006,0.9599],[0.7847,1],"\u0001"],["Exact rate on unseen decisions, by decision type: Is it a rule?",0.9286,0.9286,"rate","93% (13/14)","93% (13/14)","tie","The 95% intervals overlap (Jev 1.13 69% to 99%; Claude Sonnet 5.5 69% to 99%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code · effort low","ci95","ci95",[0.6853,0.9873],[0.6853,0.9873],"\u0001"],["Exact rate on unseen decisions, by decision type: Context shape",0.5714,0.5714,"rate","57% (8/14)","57% (8/14)","tie","The 95% intervals overlap (Jev 1.13 33% to 79%; Claude Sonnet 5.5 33% to 79%), so this sample cannot separate them.","routing-holdout",14,"routing-holdout-by-purpose",14,14,"","Claude Code · effort low","ci95","ci95",[0.3259,0.7862],[0.3259,0.7862],"\u0001"],["Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm))",0.9024,0.939,"rate","90% (74/82)","94% (77/82)","tie","The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.","routing-holdout",82,"routing-holdout-tuned-vs-unseen",82,82,"","Claude Code · effort low","ci95","ci95",[0.8191,0.9497],[0.8651,0.9737],"\u0001"],["Tuned case set vs unseen holdout: exact rate per router (Unseen holdout)",0.8214,0.875,"rate","82% (46/56)","88% (49/56)","tie","The 95% intervals overlap (Jev 1.13 70% to 90%; Claude Sonnet 5.5 76% to 94%), so this sample cannot separate them.","routing-holdout",56,"routing-holdout-tuned-vs-unseen",56,56,"","Claude Code · effort low","ci95","ci95",[0.7016,0.9],[0.7637,0.9381],"\u0001"],["Time per routing decision, by route (Wall time)",0.139,2.359,"seconds","0.14 s","2.36 s","a","Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Sonnet 5.5 2.36 s to 3.66 s); not a confidence interval.","routing-holdout","\u0001","routing-holdout-latency",168,56,"","Claude Code · effort low","range","p50-p95",[0.139,0.192],[2.359,3.657],"\u0001"],["Cost per 1,000 unseen routing decisions",0.03065,7.244,"usd","$0.031","$7.24","unclear","No interval or range was recorded for either side, so the gap ($0.031 vs $7.24, 236x) is not tested against run-to-run variation.","routing-holdout","\u0001","routing-holdout-cost-per-1000",168,56,"","Claude Code · effort low","\u0001","\u0001","\u0001","\u0001",true]]}}}