17 measured metrics · 6 calculated · 3 studies
Jev 1.13vsClaude Sonnet 5.5
Jev 1.13 ahead on 2; 15 ties, 6 unclear. A side is ahead only where the intervals or ranges do not overlap.
The verdict
Jev 1.13 and Claude Sonnet 5.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 2,597 ms; Time per routing decision, by route (Wall time), 0.14 s vs 2.36 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | Jev 1.13 | Claude Sonnet 5.5 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Typed routing decisions answered exactly right | 90%typed routing decisions · TypeSafe API | 94% (77/82)typed routing decisions · via Claude Code | 82 | 95% CI: 82%–95% vs 87%–97% | Tie | The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Per-question accuracy | 95% (184/194)typed routing decisions · TypeSafe API | 97% (189/194)typed routing decisions · via Claude Code | 194 | 95% CI: 91%–97% vs 94%–99% | Tie | The 95% intervals overlap (Jev 1.13 91% to 97%; Claude Sonnet 5.5 94% to 99%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Failure class | 100% (18/18)typed routing decisions · TypeSafe API | 100% (18/18)typed routing decisions · via Claude Code | 18 | 95% CI: 82%–100% vs 82%–100% | Tie | The 95% intervals overlap (Jev 1.13 82% to 100%; Claude Sonnet 5.5 82% to 100%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Message intent | 100% (20/20)typed routing decisions · TypeSafe API | 100% (20/20)typed routing decisions · via Claude Code | 20 | 95% CI: 84%–100% vs 84%–100% | Tie | The 95% intervals overlap (Jev 1.13 84% to 100%; Claude Sonnet 5.5 84% to 100%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Is it a rule? | 100% (12/12)typed routing decisions · TypeSafe API | 100% (12/12)typed routing decisions · via Claude Code | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Jev 1.13 76% to 100%; Claude Sonnet 5.5 76% to 100%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Context shape | 74%typed routing decisions · TypeSafe API | 84% (27/32)typed routing decisions · via Claude Code | 32 | 95% CI: 58%–87% vs 68%–93% | Tie | The 95% intervals overlap (Jev 1.13 58% to 87%; Claude Sonnet 5.5 68% to 93%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
Marks: 95% intervals (Wilson for rates)n is shown per side on every row
6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Exact rate by decision type: Context shape, 1.1x (Claude Sonnet 5.5 larger).
Watch it build
A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.
Jev 1.13 vs Claude Sonnet 5.5: what the measurements say
23 comparison rows from 3 studies: 2 rows favour Jev 1.13, 0 favour Sonnet 5.5, 21 are ties or unclear. Cost rows are calculations.
Transcript
- Comparison · 23 rows · 3 studies. Jev 1.13 vs Sonnet 5.5. A winner only where the 95% intervals or run ranges do not overlap.
- 23 comparison rows from 3 studies: Jev 1.13 ahead on 2, Sonnet 5.5 ahead on 0. The rest do not separate them. Rows where Jev 1.13 is ahead: 2 (of 23). Rows where Sonnet 5.5 is ahead: 0 (of 23). Ties or unclear: 21 (15 ties · 6 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
- Jev vs LLM routers: pass rate 90% vs 94%, tie: 95% intervals overlap. None of the 7 rows separates them. Table: Jev vs LLM routers · 5 of 7 rows · typed routing decisions · n = 12–194 per side. Source study: Jev vs Claude as a router: accuracy and cost. Rows shown: Typed routing decisions answered exactly right; Per-question accuracy; Exact rate by decision type: Failure class; Exact rate by decision type: Message intent; Exact rate by decision type: Is it a rule? Recorded settings: typed routing decisions · TypeSafe API; typed routing decisions · via Claude Code. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
- Routing overhead: success rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 6 rows separates them. Table: Routing overhead · 5 of 6 rows · routing overhead per decision vs effort low · n = 246 vs 82. Source study: Routing overhead: deterministic policy vs LLM routers vs Jev. Rows shown: Time to make one routing decision; Routing calls that returned a decision; Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)); Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)); Added routing delay per task (calculation) (Every model call routed (49.5 per task)). Recorded settings: routing overhead per decision · TypeSafe API; effort low · via Claude Code · routing overhead per decision; calculation per 1,000 tasks from recorded decision counts · TypeSafe API; effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts; calculation per task from recorded decision counts, decisions in line · TypeSafe API; effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line. Includes a calculation, not a bill or a new run. Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
- Unseen routing decisions: pass rate 82% vs 88%, tie: 95% intervals overlap. 1 of 10 rows separates them. Table: Unseen routing decisions · 5 of 10 rows · Claude Code · n = 14–168 vs 14–125. Source study: Jev vs Claude routers on unseen decisions: a blind holdout. Rows shown: Unseen routing decisions answered exactly right; Per-question accuracy on unseen decisions; Exact rate on unseen decisions, by decision type: Failure class; Exact rate on unseen decisions, by decision type: Message intent; Time per routing decision, by route (Wall time). Recorded settings: Claude Code · effort low. Caveat: 56 cases (14 per decision type) from one author: per-type intervals are very wide.
- No winner where the data shows none. Showing 15 of 23 rows; every row and its reason online.
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- Jev 1.13
- Claude Sonnet 5.5
- 95% interval
- median to p95 (not an interval)
- where the two overlap
- hollow: list-price calculation
These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.
Jev vs Claude as a router: accuracy and cost
- Typed routing decisions answered exactly right90%n 8294% (77/82)n 82TieTyped routing decisions answered exactly right: Jev 1.13 90% (n 82, 95% interval 82%–95%); Claude Sonnet 5.5 94% (77/82) (n 82, 95% interval 87%–97%). Tie.
- Per-question accuracy95% (184/194)n 19497% (189/194)n 194TiePer-question accuracy: Jev 1.13 95% (184/194) (n 194, 95% interval 91%–97%); Claude Sonnet 5.5 97% (189/194) (n 194, 95% interval 94%–99%). Tie.
Calculation: at these rates, about 826 runs per side would separate them.
- Exact rate by decision type: Failure class100% (18/18)n 18100% (18/18)n 18TieExact rate by decision type: Failure class: Jev 1.13 100% (18/18) (n 18, 95% interval 82%–100%); Claude Sonnet 5.5 100% (18/18) (n 18, 95% interval 82%–100%). Tie.
- Exact rate by decision type: Message intent100% (20/20)n 20100% (20/20)n 20TieExact rate by decision type: Message intent: Jev 1.13 100% (20/20) (n 20, 95% interval 84%–100%); Claude Sonnet 5.5 100% (20/20) (n 20, 95% interval 84%–100%). Tie.
- Exact rate by decision type: Is it a rule?100% (12/12)n 12100% (12/12)n 12TieExact rate by decision type: Is it a rule?: Jev 1.13 100% (12/12) (n 12, 95% interval 76%–100%); Claude Sonnet 5.5 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
- Exact rate by decision type: Context shape74%n 3284% (27/32)n 32TieExact rate by decision type: Context shape: Jev 1.13 74% (n 32, 95% interval 58%–87%); Claude Sonnet 5.5 84% (27/32) (n 32, 95% interval 68%–93%). Tie.
- Cost per 1,000 routing decisionsCalculation$0.034n 246$5.00n 82UnclearCost per 1,000 routing decisions, calculation: Jev 1.13 $0.034 (n 246); Claude Sonnet 5.5 $5.00 (n 82). Unclear.
Routing overhead: deterministic policy vs LLM routers vs Jev
- Time to make one routing decision137 msn 2462,597 msn 82Jev 1.13 aheadTime to make one routing decision: Jev 1.13 137 ms (n 246, median to p95 137 ms–196 ms); Claude Sonnet 5.5 2,597 ms (n 82, median to p95 2.6 s–4.3 s). Jev 1.13 ahead.
- Routing calls that returned a decision100% (246/246)n 246100% (82/82)n 82TieRouting calls that returned a decision: Jev 1.13 100% (246/246) (n 246, 95% interval 98%–100%); Claude Sonnet 5.5 100% (82/82) (n 82, 95% interval 96%–100%). Tie.
- Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))Calculation$1.67$247.30UnclearAdded routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)), calculation: Jev 1.13 $1.67; Claude Sonnet 5.5 $247.30. Unclear.
- Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))Calculation$0.24$34.97UnclearAdded routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)), calculation: Jev 1.13 $0.24; Claude Sonnet 5.5 $34.97. Unclear.
- Added routing delay per task (calculation) (Every model call routed (49.5 per task))Calculation6.76 s128.6 sUnclearAdded routing delay per task (calculation) (Every model call routed (49.5 per task)), calculation: Jev 1.13 6.76 s; Claude Sonnet 5.5 128.6 s. Unclear.
- Added routing delay per task (calculation) (Only System One decisions (7 per task))Calculation0.96 s18.2 sUnclearAdded routing delay per task (calculation) (Only System One decisions (7 per task)), calculation: Jev 1.13 0.96 s; Claude Sonnet 5.5 18.2 s. Unclear.
Jev vs Claude routers on unseen decisions: a blind holdout
- Unseen routing decisions answered exactly right82% (46/56)n 5688% (49/56)n 56TieUnseen routing decisions answered exactly right: Jev 1.13 82% (46/56) (n 56, 95% interval 70%–90%); Claude Sonnet 5.5 88% (49/56) (n 56, 95% interval 76%–94%). Tie.
Calculation: at these rates, about 675 runs per side would separate them.
- Per-question accuracy on unseen decisions90% (113/125)n 12592% (115/125)n 125TiePer-question accuracy on unseen decisions: Jev 1.13 90% (113/125) (n 125, 95% interval 84%–94%); Claude Sonnet 5.5 92% (115/125) (n 125, 95% interval 86%–96%). Tie.
Calculation: at these rates, about 4,756 runs per side would separate them.
- Exact rate on unseen decisions, by decision type: Failure class93% (13/14)n 14100% (14/14)n 14TieExact rate on unseen decisions, by decision type: Failure class: Jev 1.13 93% (13/14) (n 14, 95% interval 69%–99%); Claude Sonnet 5.5 100% (14/14) (n 14, 95% interval 78%–100%). Tie.
Calculation: at these rates, about 106 runs per side would separate them.
- Exact rate on unseen decisions, by decision type: Message intent86% (12/14)n 14100% (14/14)n 14TieExact rate on unseen decisions, by decision type: Message intent: Jev 1.13 86% (12/14) (n 14, 95% interval 60%–96%); Claude Sonnet 5.5 100% (14/14) (n 14, 95% interval 78%–100%). Tie.
Calculation: at these rates, about 53 runs per side would separate them.
- Exact rate on unseen decisions, by decision type: Is it a rule?93% (13/14)n 1493% (13/14)n 14TieExact rate on unseen decisions, by decision type: Is it a rule?: Jev 1.13 93% (13/14) (n 14, 95% interval 69%–99%); Claude Sonnet 5.5 93% (13/14) (n 14, 95% interval 69%–99%). Tie.
- Exact rate on unseen decisions, by decision type: Context shape57% (8/14)n 1457% (8/14)n 14TieExact rate on unseen decisions, by decision type: Context shape: Jev 1.13 57% (8/14) (n 14, 95% interval 33%–79%); Claude Sonnet 5.5 57% (8/14) (n 14, 95% interval 33%–79%). Tie.
- Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm))90% (74/82)n 8294% (77/82)n 82TieTuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm)): Jev 1.13 90% (74/82) (n 82, 95% interval 82%–95%); Claude Sonnet 5.5 94% (77/82) (n 82, 95% interval 87%–97%). Tie.
Calculation: at these rates, about 795 runs per side would separate them.
- Tuned case set vs unseen holdout: exact rate per router (Unseen holdout)82% (46/56)n 5688% (49/56)n 56TieTuned case set vs unseen holdout: exact rate per router (Unseen holdout): Jev 1.13 82% (46/56) (n 56, 95% interval 70%–90%); Claude Sonnet 5.5 88% (49/56) (n 56, 95% interval 76%–94%). Tie.
Calculation: at these rates, about 675 runs per side would separate them.
- Time per routing decision, by route (Wall time)0.14 sn 1682.36 sn 56Jev 1.13 aheadTime per routing decision, by route (Wall time): Jev 1.13 0.14 s (n 168, median to p95 0.1 s–0.2 s); Claude Sonnet 5.5 2.36 s (n 56, median to p95 2.4 s–3.7 s). Jev 1.13 ahead.
- Cost per 1,000 unseen routing decisionsCalculation$0.031n 168$7.24n 56UnclearCost per 1,000 unseen routing decisions, calculation: Jev 1.13 $0.031 (n 168); Claude Sonnet 5.5 $7.24 (n 56). Unclear.
| Metric | Jev 1.13 | Claude Sonnet 5.5 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Typed routing decisions answered exactly right | 90%typed routing decisions · TypeSafe API | 94% (77/82)typed routing decisions · via Claude Code | 82 | 95% CI: 82%–95% vs 87%–97% | Tie | The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Per-question accuracy | 95% (184/194)typed routing decisions · TypeSafe API | 97% (189/194)typed routing decisions · via Claude Code | 194 | 95% CI: 91%–97% vs 94%–99% | Tie | The 95% intervals overlap (Jev 1.13 91% to 97%; Claude Sonnet 5.5 94% to 99%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Failure class | 100% (18/18)typed routing decisions · TypeSafe API | 100% (18/18)typed routing decisions · via Claude Code | 18 | 95% CI: 82%–100% vs 82%–100% | Tie | The 95% intervals overlap (Jev 1.13 82% to 100%; Claude Sonnet 5.5 82% to 100%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Message intent | 100% (20/20)typed routing decisions · TypeSafe API | 100% (20/20)typed routing decisions · via Claude Code | 20 | 95% CI: 84%–100% vs 84%–100% | Tie | The 95% intervals overlap (Jev 1.13 84% to 100%; Claude Sonnet 5.5 84% to 100%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Is it a rule? | 100% (12/12)typed routing decisions · TypeSafe API | 100% (12/12)typed routing decisions · via Claude Code | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Jev 1.13 76% to 100%; Claude Sonnet 5.5 76% to 100%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Context shape | 74%typed routing decisions · TypeSafe API | 84% (27/32)typed routing decisions · via Claude Code | 32 | 95% CI: 58%–87% vs 68%–93% | Tie | The 95% intervals overlap (Jev 1.13 58% to 87%; Claude Sonnet 5.5 68% to 93%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Cost per 1,000 routing decisionsCalculation | $0.034typed routing decisions · TypeSafe API | $5.00typed routing decisions · via Claude Code | 246 / 82 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.034 vs $5.00, 148x) is not tested against run-to-run variation. | Jev vs Claude as a router: accuracy and cost |
| Time to make one routing decision | 137 msrouting overhead per decision · TypeSafe API | 2,597 mseffort low · via Claude Code · routing overhead per decision | 246 / 82 | p50–p95: 137 ms–196 ms vs 2.6 s–4.3 s | Jev 1.13 ahead | Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Sonnet 5.5 2,597 ms to 4,298 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Routing calls that returned a decision | 100% (246/246)routing overhead per decision · TypeSafe API | 100% (82/82)effort low · via Claude Code · routing overhead per decision | 246 / 82 | 95% CI: 98%–100% vs 96%–100% | Tie | The 95% intervals overlap (Jev 1.13 98% to 100%; Claude Sonnet 5.5 96% to 100%), so this sample cannot separate them. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))Calculation | $1.67calculation per 1,000 tasks from recorded decision counts · TypeSafe API | $247.30effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($1.67 vs $247.30, 148x) is not tested against run-to-run variation. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))Calculation | $0.24calculation per 1,000 tasks from recorded decision counts · TypeSafe API | $34.97effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.24 vs $34.97, 146x) is not tested against run-to-run variation. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Added routing delay per task (calculation) (Every model call routed (49.5 per task))Calculation | 6.76 scalculation per task from recorded decision counts, decisions in line · TypeSafe API | 128.6 seffort low · via Claude Code · calculation per task from recorded decision counts, decisions in line | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap (6.76 s vs 128.6 s, 19x) is not tested against run-to-run variation. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Added routing delay per task (calculation) (Only System One decisions (7 per task))Calculation | 0.96 scalculation per task from recorded decision counts, decisions in line · TypeSafe API | 18.2 seffort low · via Claude Code · calculation per task from recorded decision counts, decisions in line | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap (0.96 s vs 18.2 s, 19x) is not tested against run-to-run variation. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Unseen routing decisions answered exactly right | 82% (46/56) | 88% (49/56)Claude Code · effort low | 56 | 95% CI: 70%–90% vs 76%–94% | Tie | The 95% intervals overlap (Jev 1.13 70% to 90%; Claude Sonnet 5.5 76% to 94%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Per-question accuracy on unseen decisions | 90% (113/125) | 92% (115/125)Claude Code · effort low | 125 | 95% CI: 84%–94% vs 86%–96% | Tie | The 95% intervals overlap (Jev 1.13 84% to 94%; Claude Sonnet 5.5 86% to 96%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Exact rate on unseen decisions, by decision type: Failure class | 93% (13/14) | 100% (14/14)Claude Code · effort low | 14 | 95% CI: 69%–99% vs 78%–100% | Tie | The 95% intervals overlap (Jev 1.13 69% to 99%; Claude Sonnet 5.5 78% to 100%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Exact rate on unseen decisions, by decision type: Message intent | 86% (12/14) | 100% (14/14)Claude Code · effort low | 14 | 95% CI: 60%–96% vs 78%–100% | Tie | The 95% intervals overlap (Jev 1.13 60% to 96%; Claude Sonnet 5.5 78% to 100%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Exact rate on unseen decisions, by decision type: Is it a rule? | 93% (13/14) | 93% (13/14)Claude Code · effort low | 14 | 95% CI: 69%–99% vs 69%–99% | Tie | The 95% intervals overlap (Jev 1.13 69% to 99%; Claude Sonnet 5.5 69% to 99%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Exact rate on unseen decisions, by decision type: Context shape | 57% (8/14) | 57% (8/14)Claude Code · effort low | 14 | 95% CI: 33%–79% vs 33%–79% | Tie | The 95% intervals overlap (Jev 1.13 33% to 79%; Claude Sonnet 5.5 33% to 79%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm)) | 90% (74/82) | 94% (77/82)Claude Code · effort low | 82 | 95% CI: 82%–95% vs 87%–97% | Tie | The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Tuned case set vs unseen holdout: exact rate per router (Unseen holdout) | 82% (46/56) | 88% (49/56)Claude Code · effort low | 56 | 95% CI: 70%–90% vs 76%–94% | Tie | The 95% intervals overlap (Jev 1.13 70% to 90%; Claude Sonnet 5.5 76% to 94%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Time per routing decision, by route (Wall time) | 0.14 s | 2.36 sClaude Code · effort low | 168 / 56 | p50–p95: 0.1 s–0.2 s vs 2.4 s–3.7 s | Jev 1.13 ahead | Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Sonnet 5.5 2.36 s to 3.66 s); not a confidence interval. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Cost per 1,000 unseen routing decisionsCalculation | $0.031 | $7.24Claude Code · effort low | 168 / 56 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.031 vs $7.24, 236x) is not tested against run-to-run variation. | Jev vs Claude routers on unseen decisions: a blind holdout |
Marks: 95% intervals (Wilson for rates); median to p95 bands (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
23 rows from 3 studies. Jev 1.13 ahead on 2; 15 ties, 6 unclear. A side is ahead only where the intervals or ranges do not overlap.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick Jev 1.13
- Cost per 1,000 routing decisions: $0.034 vs $5.00. A list-price calculation, not a measured difference. Calculation
- Time to make one routing decision: 137 ms vs 2,597 ms. Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Sonnet 5.5 2,597 ms to 4,298 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone.
- Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)): $1.67 vs $247.30. A list-price calculation, not a measured difference. Calculation
- Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)): $0.24 vs $34.97. A list-price calculation, not a measured difference. Calculation
- Added routing delay per task (calculation) (Every model call routed (49.5 per task)): 6.76 s vs 128.6 s. A list-price calculation, not a measured difference. Calculation
- Added routing delay per task (calculation) (Only System One decisions (7 per task)): 0.96 s vs 18.2 s. A list-price calculation, not a measured difference. Calculation
- Time per routing decision, by route (Wall time): 0.14 s vs 2.36 s. Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Sonnet 5.5 2.36 s to 3.66 s); not a confidence interval.
- Cost per 1,000 unseen routing decisions: $0.031 vs $7.24. A list-price calculation, not a measured difference. Calculation
When to pick Claude Sonnet 5.5
No row in this data puts Claude Sonnet 5.5 ahead of Jev 1.13. Pick on other grounds (price, access, the tasks you run), or measure your own workload.
Side by side
The study charts, showing only these two. Open a study for every configuration.
| Item | Exact rate | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 90% | 82%–95% | 82 |
| Claude Sonnet 5.5 | 94% | 87%–97% | 82 |
2 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Jev 1.13 (TypeSafe) 90% (95% interval 82%–95%, n 82). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 82 per row
Share of asked cases where every scored question was acceptable
Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.
Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
| Item | Key accuracy | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 95% | 91%–97% | 194 |
| Claude Sonnet 5.5 | 97% | 94%–99% | 194 |
2 rows. Highest Claude Sonnet 5.5 97% (95% interval 94%–99%, n 194). Lowest Jev 1.13 (TypeSafe) 95% (95% interval 91%–97%, n 194). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 194 per row
Each open question the router was asked; an unanswered question counts as wrong
Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.
Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | Jev 1.13 (TypeSafe) | Claude Haiku 4.5 | Claude Sonnet 5.5 | 95% interval | n |
|---|---|---|---|---|---|
| Failure class | 100% | 94% | 100% | Jev 1.13 (TypeSafe): 82%–100%; Claude Haiku 4.5: 74%–99%; Claude Sonnet 5.5: 82%–100% | 18 |
| Context shape | 74% | 75% | 84% | Jev 1.13 (TypeSafe): 58%–87%; Claude Haiku 4.5: 58%–87%; Claude Sonnet 5.5: 68%–93% | 32 |
2 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5, Claude Sonnet 5.5. Jev 1.13 (TypeSafe): highest Failure class 100% (95% interval 82%–100%, n 18). Lowest Context shape 74% (95% interval 58%–87%, n 32). All intervals overlap. Claude Haiku 4.5: highest Failure class 94% (95% interval 74%–99%, n 18). Lowest Context shape 75% (95% interval 58%–87%, n 32). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 18–32 per row
A missing bar means that router was not run on that decision type. Whiskers are 95% Wilson intervals. Jev: the pooled rate of the live repeats, with the interval taken at the number of decisions of that type.
Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
| Item | Cost | n |
|---|---|---|
| Jev 1.13 (TypeSafe) | $0.034 | 246 |
| Claude Sonnet 5.5 | $5.00 | 82 |
List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 $5.00 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).
Notesn 82–246 per row
List price × reported tokens per decision
List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.
Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions
1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.
| Item | Decision time | Median to p95 | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 137 ms | 137 ms–196 ms | 246 |
| Claude Sonnet 5.5 (effort low, via Claude Code) | 2.6 s | 2.6 s–4.3 s | 82 |
2 rows. Slowest Claude Sonnet 5.5 (effort low, via Claude Code) 2.6 s (median to p95 2.6 s–4.3 s, n 82). Fastest Jev 1.13 (TypeSafe) 137 ms (median to p95 137 ms–196 ms, n 246). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n 82–246 per row
Median; whiskers = median to 95th percentile
The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
| Item | Completed | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 100% | 98%–100% | 246 |
| Claude Sonnet 5.5 (effort low, via Claude Code) | 100% | 96%–100% | 82 |
2 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln 82–246 per row
Completed calls ÷ calls; whiskers = 95% Wilson interval
A completed call returned a decision, right or wrong (accuracy is in the routing study). Whiskers are 95% Wilson intervals.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
- Every model call routed (49.5 per task)
- Only System One decisions (7 per task) (square)
Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).
| Item | Every model call routed (49.5 per task) | Only System One decisions (7 per task) |
|---|---|---|
| Jev 1.13 (TypeSafe) | $1.67 | $0.24 |
| Claude Sonnet 5.5 (effort low, via Claude Code) | $247 | $34.97 |
List-price calculation, not a run. 2 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): highest Claude Sonnet 5.5 (effort low, via Claude Code) $247. Lowest Jev 1.13 (TypeSafe) $1.67. Only System One decisions (7 per task): highest Claude Sonnet 5.5 (effort low, via Claude Code) $34.97. Lowest Jev 1.13 (TypeSafe) $0.24.
Notes
Decisions per task from recorded runs × cost per decision
A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions
- Every model call routed (49.5 per task)
- Only System One decisions (7 per task) (square)
Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).
| Item | Every model call routed (49.5 per task) | Only System One decisions (7 per task) |
|---|---|---|
| Jev 1.13 (TypeSafe) | 6.8 s | 1 s |
| Claude Sonnet 5.5 (effort low, via Claude Code) | 129 s | 18.2 s |
List-price calculation, not a run. 2 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): slowest Claude Sonnet 5.5 (effort low, via Claude Code) 129 s. Fastest Jev 1.13 (TypeSafe) 6.8 s. Only System One decisions (7 per task): slowest Claude Sonnet 5.5 (effort low, via Claude Code) 18.2 s. Fastest Jev 1.13 (TypeSafe) 1 s.
Notes
Decisions per task × median decision time, if every decision waits in line
A calculation and an upper bound: it assumes each decision waits for the one before. Median recorded task wall time: 10.3 min. Jev’s delay uses its live median over the API from one Mac (network included); the Claude routers’ includes the CLI.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions
| Item | Exact rate | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 82% | 70%–90% | 56 |
| Claude Sonnet 5.5 (low) · Claude Code | 88% | 76%–94% | 56 |
2 rows. Highest Claude Sonnet 5.5 (low) · Claude Code 88% (95% interval 76%–94%, n 56). Lowest Jev 1.13 (TypeSafe) 82% (95% interval 70%–90%, n 56). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 56 per row
Share of the 56 holdout cases where every scored question was acceptable
Whiskers are 95% Wilson intervals. Cases written blind to every router answer, then frozen. Jev: its first of three repetitions.
Source: Routing on unseen holdout decisions
| Item | Key accuracy | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 90% | 84%–94% | 125 |
| Claude Sonnet 5.5 (low) · Claude Code | 92% | 86%–96% | 125 |
2 rows. Highest Claude Sonnet 5.5 (low) · Claude Code 92% (95% interval 86%–96%, n 125). Lowest Jev 1.13 (TypeSafe) 90% (95% interval 84%–94%, n 125). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 125 per row
Each open question a router was asked; an unanswered question counts as wrong
Whiskers are nominal 95% Wilson intervals. Context-shape cases ask up to 8 questions each, the other decision types one. Questions in one case are not independent; these intervals do not adjust for that grouping.
Source: Routing on unseen holdout decisions
Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | Jev 1.13 (TypeSafe) | Claude Haiku 4.5 · Claude Code | Claude Sonnet 5.5 (low) · Claude Code | 95% interval | n |
|---|---|---|---|---|---|
| Failure class | 93% | 93% | 100% | Jev 1.13 (TypeSafe): 69%–99%; Claude Haiku 4.5 · Claude Code: 69%–99%; Claude Sonnet 5.5 (low) · Claude Code: 78%–100% | 14 |
| Message intent | 86% | 93% | 100% | Jev 1.13 (TypeSafe): 60%–96%; Claude Haiku 4.5 · Claude Code: 69%–99%; Claude Sonnet 5.5 (low) · Claude Code: 78%–100% | 14 |
| Is it a rule? | 93% | 93% | 93% | Jev 1.13 (TypeSafe): 69%–99%; Claude Haiku 4.5 · Claude Code: 69%–99%; Claude Sonnet 5.5 (low) · Claude Code: 69%–99% | 14 |
| Context shape | 57% | 36% | 57% | Jev 1.13 (TypeSafe): 33%–79%; Claude Haiku 4.5 · Claude Code: 16%–61%; Claude Sonnet 5.5 (low) · Claude Code: 33%–79% | 14 |
4 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 (low) · Claude Code. Jev 1.13 (TypeSafe): highest Failure class 93% (95% interval 69%–99%, n 14). Lowest Context shape 57% (95% interval 33%–79%, n 14). All intervals overlap. Claude Haiku 4.5 · Claude Code: highest Failure class 93% (95% interval 69%–99%, n 14). Lowest Context shape 36% (95% interval 16%–61%, n 14). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 14 per row
14 cases per decision type
Whiskers are 95% Wilson intervals. With 14 cases a perfect score has an interval of 78% to 100%, so a decision type where every router scores 14 of 14 is at its ceiling and cannot rank them.
Source: Routing on unseen holdout decisions
- Tuned set (routing-jev-vs-llm)
- Unseen holdout (square)
Gap labels, Unseen holdout vs Tuned set (routing-jev-vs-llm): Unseen holdout is x percentage points higher (+) or lower (−) than Tuned set (routing-jev-vs-llm), calculated from the two values shown; lines are the 95% Wilson interval.
| Item | Tuned set (routing-jev-vs-llm) | Unseen holdout | 95% interval | n |
|---|---|---|---|---|
| Jev 1.13 (TypeSafe) | 90% | 82% | Tuned set (routing-jev-vs-llm): 82%–95%; Unseen holdout: 70%–90% | 82 |
| Claude Sonnet 5.5 (low) · Claude Code | 94% | 88% | Tuned set (routing-jev-vs-llm): 87%–97%; Unseen holdout: 76%–94% | 82 |
2 rows, 2 series: Tuned set (routing-jev-vs-llm), Unseen holdout. Tuned set (routing-jev-vs-llm): highest Claude Sonnet 5.5 (low) · Claude Code 94% (95% interval 87%–97%, n 82). Lowest Jev 1.13 (TypeSafe) 90% (95% interval 82%–95%, n 82). All intervals overlap. Unseen holdout: highest Claude Sonnet 5.5 (low) · Claude Code 88% (95% interval 76%–94%, n 56). Lowest Jev 1.13 (TypeSafe) 82% (95% interval 70%–90%, n 56). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 56–82 per row
Tuned set: the routing study’s 82 cases, revised against Jev answers. Holdout: 56 new cases, frozen before any router call
Whiskers are 95% Wilson intervals. The two case sets differ in mix and size, so a gap mixes a change of case set with any change in the router, and the data cannot separate them. A gap counts only when the two intervals do not overlap.
Sources: Routing on unseen holdout decisions, Routing runs: Jev router vs LLM routing
- Wall time
- Model time (API, CLI-reported)
| Item | Wall time | Model time (API, CLI-reported) | Median to p95 | n |
|---|---|---|---|---|
| Jev 1.13 (TypeSafe) | 0.1 s | — | Wall time: 0.1 s–0.2 s | 168 |
| Claude Sonnet 5.5 (low) · Claude Code | 2.4 s | 1.5 s | Wall time: 2.4 s–3.7 s; Model time (API, CLI-reported): 1.5 s–2.4 s | 56 |
2 rows, 2 series: Wall time, Model time (API, CLI-reported). Wall time: slowest Claude Sonnet 5.5 (low) · Claude Code 2.4 s (median to p95 2.4 s–3.7 s, n 56). Fastest Jev 1.13 (TypeSafe) 0.1 s (median to p95 0.1 s–0.2 s, n 168). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n 56–168 per row
Median, whisker to the 95th percentile
The routes differ. Jev: one HTTPS call to the vendor API from one Mac, all three repetitions. Claude: the Claude Code CLI from process start to exit, which adds start-up time a direct API call would not. The whisker runs from the median to the 95th percentile; it is not a confidence interval. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.
Source: Routing on unseen holdout decisions
| Item | Cost | n |
|---|---|---|
| Jev 1.13 (TypeSafe) | $0.031 | 168 |
| Claude Sonnet 5.5 (low) · Claude Code | $7.24 | 56 |
List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 (low) · Claude Code $7.24 (n 56). Lowest Jev 1.13 (TypeSafe) $0.031 (n 168).
Notesn 56–168 per row
Reported tokens per decision × list price
Calculation, not a bill. Jev: reported input tokens × $0.042 per million, output free. Claude: CLI-reported tokens × list price; the CLI wrote its prompt cache as 1-hour writes, priced at 2× input, and adds its own system prompt and tool-schema tokens. The Claude routers ran on a subscription. At the tuned-set study’s convention (every cache write at 1.25× input), Sonnet 5.5 (low) would be $4.877 here. That study shows $4.996 for it on its own cases, where the CLI reported $7.324, so its cost row and this one differ by convention and by case mix.
Sources: Routing on unseen holdout decisions, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models)
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Which is better, Jev 1.13 or Claude Sonnet 5.5?
- Jev 1.13 and Claude Sonnet 5.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 2,597 ms; Time per routing decision, by route (Wall time), 0.14 s vs 2.36 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).
- How were Jev 1.13 and Claude Sonnet 5.5 measured?
- They share 17 measured metrics and 6 list-price calculations from 3 public studies: Jev vs Claude as a router: accuracy and cost; Routing overhead: deterministic policy vs LLM routers vs Jev; Jev vs Claude routers on unseen decisions: a blind holdout. Every row names its configuration, its sample size and its interval or range.
- How do Jev 1.13 and Claude Sonnet 5.5 compare on time to make one routing decision?
- Jev 1.13: 137 ms (routing overhead per decision · TypeSafe API; n = 246; p50 to p95 137 ms to 196 ms). Claude Sonnet 5.5: 2,597 ms (effort low · via Claude Code · routing overhead per decision; n = 82; p50 to p95 2.6 s to 4.3 s). Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Sonnet 5.5 2,597 ms to 4,298 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone.
- How do Jev 1.13 and Claude Sonnet 5.5 compare on time per routing decision, by route (Wall time)?
- Jev 1.13: 0.14 s (n = 168; p50 to p95 0.1 s to 0.2 s). Claude Sonnet 5.5: 2.36 s (Claude Code · effort low; n = 56; p50 to p95 2.4 s to 3.7 s). Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Sonnet 5.5 2.36 s to 3.66 s); not a confidence interval.
- How do Jev 1.13 and Claude Sonnet 5.5 compare on typed routing decisions answered exactly right?
- Jev 1.13: 90% (typed routing decisions · TypeSafe API; n = 82; 95% interval 82% to 95%). Claude Sonnet 5.5: 94% (77/82) (typed routing decisions · via Claude Code; n = 82; 95% interval 87% to 97%). The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.
- How do Jev 1.13 and Claude Sonnet 5.5 compare on per-question accuracy?
- Jev 1.13: 95% (184/194) (typed routing decisions · TypeSafe API; n = 194; 95% interval 91% to 97%). Claude Sonnet 5.5: 97% (189/194) (typed routing decisions · via Claude Code; n = 194; 95% interval 94% to 99%). The 95% intervals overlap (Jev 1.13 91% to 97%; Claude Sonnet 5.5 94% to 99%), so this sample cannot separate them.
- How do Jev 1.13 and Claude Sonnet 5.5 compare on exact rate by decision type: Failure class?
- Jev 1.13: 100% (18/18) (typed routing decisions · TypeSafe API; n = 18; 95% interval 82% to 100%). Claude Sonnet 5.5: 100% (18/18) (typed routing decisions · via Claude Code; n = 18; 95% interval 82% to 100%). The 95% intervals overlap (Jev 1.13 82% to 100%; Claude Sonnet 5.5 82% to 100%), so this sample cannot separate them.
The studies behind this page
Jev vs Claude as a router: accuracy and cost
Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.
Routing overhead: deterministic policy vs LLM routers vs Jev
How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.
Jev vs Claude routers on unseen decisions: a blind holdout
Jev 1.13, Claude Haiku 4.5 and Sonnet 5.5 on 56 routing decisions written blind and frozen first: exact rate with 95% intervals, tuned vs unseen.