[{"i":11,"story":{"id":"jev-vs-llm-router","title":"Jev vs Claude as a router: accuracy and cost","description":"Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.","studySlug":"routing-jev-vs-llm","chartIds":["routing-exact-decisions","routing-cost-per-1000","routing-economics-scenarios"],"durationSeconds":34.2,"transcript":["Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.","Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.","Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.","Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.","Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.","Open benchmarks: intervals, sources and every failure kept."]}},{"i":18,"story":{"id":"jev-vs-sonnet-router","title":"Jev 1.13 vs Claude Sonnet 5.5: what the measurements say","description":"23 comparison rows from 3 studies: 2 rows favour Jev 1.13, 0 favour Sonnet 5.5, 21 are ties or unclear. Cost rows are calculations.","studySlug":"routing-jev-vs-llm","placement":"compare","chartIds":[],"durationSeconds":35.4,"transcript":["Comparison · 23 rows · 3 studies. Jev 1.13 vs Sonnet 5.5. A winner only where the 95% intervals or run ranges do not overlap.","23 comparison rows from 3 studies: Jev 1.13 ahead on 2, Sonnet 5.5 ahead on 0. The rest do not separate them. Rows where Jev 1.13 is ahead: 2 (of 23). Rows where Sonnet 5.5 is ahead: 0 (of 23). Ties or unclear: 21 (15 ties · 6 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.","Jev vs LLM routers: pass rate 90% vs 94%, tie: 95% intervals overlap. None of the 7 rows separates them. Table: Jev vs LLM routers · 5 of 7 rows · typed routing decisions · n = 12–194 per side. Source study: Jev vs Claude as a router: accuracy and cost. Rows shown: Typed routing decisions answered exactly right; Per-question accuracy; Exact rate by decision type: Failure class; Exact rate by decision type: Message intent; Exact rate by decision type: Is it a rule? Recorded settings: typed routing decisions · TypeSafe API; typed routing decisions · via Claude Code. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.","Routing overhead: success rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 6 rows separates them. Table: Routing overhead · 5 of 6 rows · routing overhead per decision vs effort low · n = 246 vs 82. Source study: Routing overhead: deterministic policy vs LLM routers vs Jev. Rows shown: Time to make one routing decision; Routing calls that returned a decision; Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)); Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)); Added routing delay per task (calculation) (Every model call routed (49.5 per task)). Recorded settings: routing overhead per decision · TypeSafe API; effort low · via Claude Code · routing overhead per decision; calculation per 1,000 tasks from recorded decision counts · TypeSafe API; effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts; calculation per task from recorded decision counts, decisions in line · TypeSafe API; effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line. Includes a calculation, not a bill or a new run. Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.","Unseen routing decisions: pass rate 82% vs 88%, tie: 95% intervals overlap. 1 of 10 rows separates them. Table: Unseen routing decisions · 5 of 10 rows · Claude Code · n = 14–168 vs 14–125. Source study: Jev vs Claude routers on unseen decisions: a blind holdout. Rows shown: Unseen routing decisions answered exactly right; Per-question accuracy on unseen decisions; Exact rate on unseen decisions, by decision type: Failure class; Exact rate on unseen decisions, by decision type: Message intent; Time per routing decision, by route (Wall time). Recorded settings: Claude Code · effort low. Caveat: 56 cases (14 per decision type) from one author: per-type intervals are very wide.","No winner where the data shows none. Showing 15 of 23 rows; every row and its reason online."]}}]