17 measured metrics · 6 calculated · 3 studies
Jev 1.13vsClaude Haiku 4.5
Jev 1.13 ahead on 2; 15 ties, 6 unclear. A side is ahead only where the intervals or ranges do not overlap.
The verdict
Jev 1.13 and Claude Haiku 4.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 12,543 ms; Time per routing decision, by route (Wall time), 0.14 s vs 9.44 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | Jev 1.13 | Claude Haiku 4.5 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Typed routing decisions answered exactly right | 90%typed routing decisions · TypeSafe API | 89% (73/82)typed routing decisions · via Claude Code | 82 | 95% CI: 82%–95% vs 80%–94% | Tie | The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Haiku 4.5 80% to 94%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Per-question accuracy | 95% (184/194)typed routing decisions · TypeSafe API | 94% (183/194)typed routing decisions · via Claude Code | 194 | 95% CI: 91%–97% vs 90%–97% | Tie | The 95% intervals overlap (Jev 1.13 91% to 97%; Claude Haiku 4.5 90% to 97%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Failure class | 100% (18/18)typed routing decisions · TypeSafe API | 94% (17/18)typed routing decisions · via Claude Code | 18 | 95% CI: 82%–100% vs 74%–99% | Tie | The 95% intervals overlap (Jev 1.13 82% to 100%; Claude Haiku 4.5 74% to 99%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Message intent | 100% (20/20)typed routing decisions · TypeSafe API | 100% (20/20)typed routing decisions · via Claude Code | 20 | 95% CI: 84%–100% vs 84%–100% | Tie | The 95% intervals overlap (Jev 1.13 84% to 100%; Claude Haiku 4.5 84% to 100%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Is it a rule? | 100% (12/12)typed routing decisions · TypeSafe API | 100% (12/12)typed routing decisions · via Claude Code | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Jev 1.13 76% to 100%; Claude Haiku 4.5 76% to 100%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Context shape | 74%typed routing decisions · TypeSafe API | 75% (24/32)typed routing decisions · via Claude Code | 32 | 95% CI: 58%–87% vs 58%–87% | Tie | The 95% intervals overlap (Jev 1.13 58% to 87%; Claude Haiku 4.5 58% to 87%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
Marks: 95% intervals (Wilson for rates)n is shown per side on every row
6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Exact rate by decision type: Failure class, 1.1x (Jev 1.13 larger).
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- Jev 1.13
- Claude Haiku 4.5
- 95% interval
- median to p95 (not an interval)
- where the two overlap
- hollow: list-price calculation
These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.
Jev vs Claude as a router: accuracy and cost
- Typed routing decisions answered exactly right90%n 8289% (73/82)n 82TieTyped routing decisions answered exactly right: Jev 1.13 90% (n 82, 95% interval 82%–95%); Claude Haiku 4.5 89% (73/82) (n 82, 95% interval 80%–94%). Tie.
- Per-question accuracy95% (184/194)n 19494% (183/194)n 194TiePer-question accuracy: Jev 1.13 95% (184/194) (n 194, 95% interval 91%–97%); Claude Haiku 4.5 94% (183/194) (n 194, 95% interval 90%–97%). Tie.
Calculation: at these rates, more than 5,000 runs per side would be needed.
- Exact rate by decision type: Failure class100% (18/18)n 1894% (17/18)n 18TieExact rate by decision type: Failure class: Jev 1.13 100% (18/18) (n 18, 95% interval 82%–100%); Claude Haiku 4.5 94% (17/18) (n 18, 95% interval 74%–99%). Tie.
Calculation: at these rates, about 135 runs per side would separate them.
- Exact rate by decision type: Message intent100% (20/20)n 20100% (20/20)n 20TieExact rate by decision type: Message intent: Jev 1.13 100% (20/20) (n 20, 95% interval 84%–100%); Claude Haiku 4.5 100% (20/20) (n 20, 95% interval 84%–100%). Tie.
- Exact rate by decision type: Is it a rule?100% (12/12)n 12100% (12/12)n 12TieExact rate by decision type: Is it a rule?: Jev 1.13 100% (12/12) (n 12, 95% interval 76%–100%); Claude Haiku 4.5 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
- Exact rate by decision type: Context shape74%n 3275% (24/32)n 32TieExact rate by decision type: Context shape: Jev 1.13 74% (n 32, 95% interval 58%–87%); Claude Haiku 4.5 75% (24/32) (n 32, 95% interval 58%–87%). Tie.
- Cost per 1,000 routing decisionsCalculation$0.034n 246$8.92n 82UnclearCost per 1,000 routing decisions, calculation: Jev 1.13 $0.034 (n 246); Claude Haiku 4.5 $8.92 (n 82). Unclear.
Routing overhead: deterministic policy vs LLM routers vs Jev
- Time to make one routing decision137 msn 24612,543 msn 82Jev 1.13 aheadTime to make one routing decision: Jev 1.13 137 ms (n 246, median to p95 137 ms–196 ms); Claude Haiku 4.5 12,543 ms (n 82, median to p95 12.54 s–34.48 s). Jev 1.13 ahead.
- Routing calls that returned a decision100% (246/246)n 246100% (82/82)n 82TieRouting calls that returned a decision: Jev 1.13 100% (246/246) (n 246, 95% interval 98%–100%); Claude Haiku 4.5 100% (82/82) (n 82, 95% interval 96%–100%). Tie.
- Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))Calculation$1.67$441.74UnclearAdded routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)), calculation: Jev 1.13 $1.67; Claude Haiku 4.5 $441.74. Unclear.
- Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))Calculation$0.24$62.47UnclearAdded routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)), calculation: Jev 1.13 $0.24; Claude Haiku 4.5 $62.47. Unclear.
- Added routing delay per task (calculation) (Every model call routed (49.5 per task))Calculation6.76 s620.9 sUnclearAdded routing delay per task (calculation) (Every model call routed (49.5 per task)), calculation: Jev 1.13 6.76 s; Claude Haiku 4.5 620.9 s. Unclear.
- Added routing delay per task (calculation) (Only System One decisions (7 per task))Calculation0.96 s87.8 sUnclearAdded routing delay per task (calculation) (Only System One decisions (7 per task)), calculation: Jev 1.13 0.96 s; Claude Haiku 4.5 87.8 s. Unclear.
Jev vs Claude routers on unseen decisions: a blind holdout
- Unseen routing decisions answered exactly right82% (46/56)n 5679% (44/56)n 56TieUnseen routing decisions answered exactly right: Jev 1.13 82% (46/56) (n 56, 95% interval 70%–90%); Claude Haiku 4.5 79% (44/56) (n 56, 95% interval 66%–87%). Tie.
Calculation: at these rates, about 1,883 runs per side would separate them.
- Per-question accuracy on unseen decisions90% (113/125)n 12582% (102/125)n 125TiePer-question accuracy on unseen decisions: Jev 1.13 90% (113/125) (n 125, 95% interval 84%–94%); Claude Haiku 4.5 82% (102/125) (n 125, 95% interval 74%–87%). Tie.
Calculation: at these rates, about 231 runs per side would separate them.
- Exact rate on unseen decisions, by decision type: Failure class93% (13/14)n 1493% (13/14)n 14TieExact rate on unseen decisions, by decision type: Failure class: Jev 1.13 93% (13/14) (n 14, 95% interval 69%–99%); Claude Haiku 4.5 93% (13/14) (n 14, 95% interval 69%–99%). Tie.
- Exact rate on unseen decisions, by decision type: Message intent86% (12/14)n 1493% (13/14)n 14TieExact rate on unseen decisions, by decision type: Message intent: Jev 1.13 86% (12/14) (n 14, 95% interval 60%–96%); Claude Haiku 4.5 93% (13/14) (n 14, 95% interval 69%–99%). Tie.
Calculation: at these rates, about 270 runs per side would separate them.
- Exact rate on unseen decisions, by decision type: Is it a rule?93% (13/14)n 1493% (13/14)n 14TieExact rate on unseen decisions, by decision type: Is it a rule?: Jev 1.13 93% (13/14) (n 14, 95% interval 69%–99%); Claude Haiku 4.5 93% (13/14) (n 14, 95% interval 69%–99%). Tie.
- Exact rate on unseen decisions, by decision type: Context shape57% (8/14)n 1436% (5/14)n 14TieExact rate on unseen decisions, by decision type: Context shape: Jev 1.13 57% (8/14) (n 14, 95% interval 33%–79%); Claude Haiku 4.5 36% (5/14) (n 14, 95% interval 16%–61%). Tie.
Calculation: at these rates, about 82 runs per side would separate them.
- Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm))90% (74/82)n 8289% (73/82)n 82TieTuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm)): Jev 1.13 90% (74/82) (n 82, 95% interval 82%–95%); Claude Haiku 4.5 89% (73/82) (n 82, 95% interval 80%–94%). Tie.
Calculation: at these rates, more than 5,000 runs per side would be needed.
- Tuned case set vs unseen holdout: exact rate per router (Unseen holdout)82% (46/56)n 5679% (44/56)n 56TieTuned case set vs unseen holdout: exact rate per router (Unseen holdout): Jev 1.13 82% (46/56) (n 56, 95% interval 70%–90%); Claude Haiku 4.5 79% (44/56) (n 56, 95% interval 66%–87%). Tie.
Calculation: at these rates, about 1,883 runs per side would separate them.
- Time per routing decision, by route (Wall time)0.14 sn 1689.44 sn 56Jev 1.13 aheadTime per routing decision, by route (Wall time): Jev 1.13 0.14 s (n 168, median to p95 0.1 s–0.2 s); Claude Haiku 4.5 9.44 s (n 56, median to p95 9.4 s–25.4 s). Jev 1.13 ahead.
- Cost per 1,000 unseen routing decisionsCalculation$0.031n 168$7.13n 56UnclearCost per 1,000 unseen routing decisions, calculation: Jev 1.13 $0.031 (n 168); Claude Haiku 4.5 $7.13 (n 56). Unclear.
| Metric | Jev 1.13 | Claude Haiku 4.5 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Typed routing decisions answered exactly right | 90%typed routing decisions · TypeSafe API | 89% (73/82)typed routing decisions · via Claude Code | 82 | 95% CI: 82%–95% vs 80%–94% | Tie | The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Haiku 4.5 80% to 94%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Per-question accuracy | 95% (184/194)typed routing decisions · TypeSafe API | 94% (183/194)typed routing decisions · via Claude Code | 194 | 95% CI: 91%–97% vs 90%–97% | Tie | The 95% intervals overlap (Jev 1.13 91% to 97%; Claude Haiku 4.5 90% to 97%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Failure class | 100% (18/18)typed routing decisions · TypeSafe API | 94% (17/18)typed routing decisions · via Claude Code | 18 | 95% CI: 82%–100% vs 74%–99% | Tie | The 95% intervals overlap (Jev 1.13 82% to 100%; Claude Haiku 4.5 74% to 99%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Message intent | 100% (20/20)typed routing decisions · TypeSafe API | 100% (20/20)typed routing decisions · via Claude Code | 20 | 95% CI: 84%–100% vs 84%–100% | Tie | The 95% intervals overlap (Jev 1.13 84% to 100%; Claude Haiku 4.5 84% to 100%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Is it a rule? | 100% (12/12)typed routing decisions · TypeSafe API | 100% (12/12)typed routing decisions · via Claude Code | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Jev 1.13 76% to 100%; Claude Haiku 4.5 76% to 100%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Context shape | 74%typed routing decisions · TypeSafe API | 75% (24/32)typed routing decisions · via Claude Code | 32 | 95% CI: 58%–87% vs 58%–87% | Tie | The 95% intervals overlap (Jev 1.13 58% to 87%; Claude Haiku 4.5 58% to 87%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Cost per 1,000 routing decisionsCalculation | $0.034typed routing decisions · TypeSafe API | $8.92typed routing decisions · via Claude Code | 246 / 82 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.034 vs $8.92, 265x) is not tested against run-to-run variation. | Jev vs Claude as a router: accuracy and cost |
| Time to make one routing decision | 137 msrouting overhead per decision · TypeSafe API | 12,543 msthinking on · via Claude Code · routing overhead per decision | 246 / 82 | p50–p95: 137 ms–196 ms vs 12.54 s–34.48 s | Jev 1.13 ahead | Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Haiku 4.5 12,543 ms to 34,481 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Routing calls that returned a decision | 100% (246/246)routing overhead per decision · TypeSafe API | 100% (82/82)thinking on · via Claude Code · routing overhead per decision | 246 / 82 | 95% CI: 98%–100% vs 96%–100% | Tie | The 95% intervals overlap (Jev 1.13 98% to 100%; Claude Haiku 4.5 96% to 100%), so this sample cannot separate them. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))Calculation | $1.67calculation per 1,000 tasks from recorded decision counts · TypeSafe API | $441.74thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($1.67 vs $441.74, 265x) is not tested against run-to-run variation. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))Calculation | $0.24calculation per 1,000 tasks from recorded decision counts · TypeSafe API | $62.47thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.24 vs $62.47, 260x) is not tested against run-to-run variation. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Added routing delay per task (calculation) (Every model call routed (49.5 per task))Calculation | 6.76 scalculation per task from recorded decision counts, decisions in line · TypeSafe API | 620.9 sthinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap (6.76 s vs 620.9 s, 92x) is not tested against run-to-run variation. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Added routing delay per task (calculation) (Only System One decisions (7 per task))Calculation | 0.96 scalculation per task from recorded decision counts, decisions in line · TypeSafe API | 87.8 sthinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap (0.96 s vs 87.8 s, 92x) is not tested against run-to-run variation. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Unseen routing decisions answered exactly right | 82% (46/56) | 79% (44/56)Claude Code | 56 | 95% CI: 70%–90% vs 66%–87% | Tie | The 95% intervals overlap (Jev 1.13 70% to 90%; Claude Haiku 4.5 66% to 87%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Per-question accuracy on unseen decisions | 90% (113/125) | 82% (102/125)Claude Code | 125 | 95% CI: 84%–94% vs 74%–87% | Tie | The 95% intervals overlap (Jev 1.13 84% to 94%; Claude Haiku 4.5 74% to 87%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Exact rate on unseen decisions, by decision type: Failure class | 93% (13/14) | 93% (13/14)Claude Code | 14 | 95% CI: 69%–99% vs 69%–99% | Tie | The 95% intervals overlap (Jev 1.13 69% to 99%; Claude Haiku 4.5 69% to 99%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Exact rate on unseen decisions, by decision type: Message intent | 86% (12/14) | 93% (13/14)Claude Code | 14 | 95% CI: 60%–96% vs 69%–99% | Tie | The 95% intervals overlap (Jev 1.13 60% to 96%; Claude Haiku 4.5 69% to 99%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Exact rate on unseen decisions, by decision type: Is it a rule? | 93% (13/14) | 93% (13/14)Claude Code | 14 | 95% CI: 69%–99% vs 69%–99% | Tie | The 95% intervals overlap (Jev 1.13 69% to 99%; Claude Haiku 4.5 69% to 99%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Exact rate on unseen decisions, by decision type: Context shape | 57% (8/14) | 36% (5/14)Claude Code | 14 | 95% CI: 33%–79% vs 16%–61% | Tie | The 95% intervals overlap (Jev 1.13 33% to 79%; Claude Haiku 4.5 16% to 61%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm)) | 90% (74/82) | 89% (73/82)Claude Code | 82 | 95% CI: 82%–95% vs 80%–94% | Tie | The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Haiku 4.5 80% to 94%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Tuned case set vs unseen holdout: exact rate per router (Unseen holdout) | 82% (46/56) | 79% (44/56)Claude Code | 56 | 95% CI: 70%–90% vs 66%–87% | Tie | The 95% intervals overlap (Jev 1.13 70% to 90%; Claude Haiku 4.5 66% to 87%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Time per routing decision, by route (Wall time) | 0.14 s | 9.44 sClaude Code | 168 / 56 | p50–p95: 0.1 s–0.2 s vs 9.4 s–25.4 s | Jev 1.13 ahead | Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Haiku 4.5 9.44 s to 25.4 s); not a confidence interval. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Cost per 1,000 unseen routing decisionsCalculation | $0.031 | $7.13Claude Code | 168 / 56 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.031 vs $7.13, 233x) is not tested against run-to-run variation. | Jev vs Claude routers on unseen decisions: a blind holdout |
Marks: 95% intervals (Wilson for rates); median to p95 bands (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
23 rows from 3 studies. Jev 1.13 ahead on 2; 15 ties, 6 unclear. A side is ahead only where the intervals or ranges do not overlap.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick Jev 1.13
- Cost per 1,000 routing decisions: $0.034 vs $8.92. A list-price calculation, not a measured difference. Calculation
- Time to make one routing decision: 137 ms vs 12,543 ms. Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Haiku 4.5 12,543 ms to 34,481 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone.
- Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)): $1.67 vs $441.74. A list-price calculation, not a measured difference. Calculation
- Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)): $0.24 vs $62.47. A list-price calculation, not a measured difference. Calculation
- Added routing delay per task (calculation) (Every model call routed (49.5 per task)): 6.76 s vs 620.9 s. A list-price calculation, not a measured difference. Calculation
- Added routing delay per task (calculation) (Only System One decisions (7 per task)): 0.96 s vs 87.8 s. A list-price calculation, not a measured difference. Calculation
- Time per routing decision, by route (Wall time): 0.14 s vs 9.44 s. Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Haiku 4.5 9.44 s to 25.4 s); not a confidence interval.
- Cost per 1,000 unseen routing decisions: $0.031 vs $7.13. A list-price calculation, not a measured difference. Calculation
When to pick Claude Haiku 4.5
No row in this data puts Claude Haiku 4.5 ahead of Jev 1.13. Pick on other grounds (price, access, the tasks you run), or measure your own workload.
Side by side
The study charts, showing only these two. Open a study for every configuration.
| Item | Exact rate | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 90% | 82%–95% | 82 |
| Claude Haiku 4.5 | 89% | 80%–94% | 82 |
2 rows. Highest Jev 1.13 (TypeSafe) 90% (95% interval 82%–95%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 82 per row
Share of asked cases where every scored question was acceptable
Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.
Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
| Item | Key accuracy | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 95% | 91%–97% | 194 |
| Claude Haiku 4.5 | 94% | 90%–97% | 194 |
2 rows. Highest Jev 1.13 (TypeSafe) 95% (95% interval 91%–97%, n 194). Lowest Claude Haiku 4.5 94% (95% interval 90%–97%, n 194). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 194 per row
Each open question the router was asked; an unanswered question counts as wrong
Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.
Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | Jev 1.13 (TypeSafe) | Claude Haiku 4.5 | Claude Sonnet 5.5 | 95% interval | n |
|---|---|---|---|---|---|
| Failure class | 100% | 94% | 100% | Jev 1.13 (TypeSafe): 82%–100%; Claude Haiku 4.5: 74%–99%; Claude Sonnet 5.5: 82%–100% | 18 |
| Message intent | 100% | 100% | 100% | Jev 1.13 (TypeSafe): 84%–100%; Claude Haiku 4.5: 84%–100%; Claude Sonnet 5.5: 84%–100% | 20 |
| Context shape | 74% | 75% | 84% | Jev 1.13 (TypeSafe): 58%–87%; Claude Haiku 4.5: 58%–87%; Claude Sonnet 5.5: 68%–93% | 32 |
3 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5, Claude Sonnet 5.5. Jev 1.13 (TypeSafe): highest Failure class 100% (95% interval 82%–100%, n 18). Lowest Context shape 74% (95% interval 58%–87%, n 32). All intervals overlap. Claude Haiku 4.5: highest Message intent 100% (95% interval 84%–100%, n 20). Lowest Context shape 75% (95% interval 58%–87%, n 32). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 18–32 per row
A missing bar means that router was not run on that decision type. Whiskers are 95% Wilson intervals. Jev: the pooled rate of the live repeats, with the interval taken at the number of decisions of that type.
Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
| Item | Cost | n |
|---|---|---|
| Jev 1.13 (TypeSafe) | $0.034 | 246 |
| Claude Haiku 4.5 | $8.92 | 82 |
List-price calculation, not a run. 2 rows. Highest Claude Haiku 4.5 $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).
Notesn 82–246 per row
List price × reported tokens per decision
List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.
Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions
1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.
Time per decision · log scale: each gridline is 10 times the one before
| Item | Decision time | Median to p95 | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 137 ms | 137 ms–196 ms | 246 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 12.54 s | 12.54 s–34.48 s | 82 |
2 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Jev 1.13 (TypeSafe) 137 ms (median to p95 137 ms–196 ms, n 246). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n 82–246 per row
Median; whiskers = median to 95th percentile
The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
| Item | Completed | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 100% | 98%–100% | 246 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 100% | 96%–100% | 82 |
2 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln 82–246 per row
Completed calls ÷ calls; whiskers = 95% Wilson interval
A completed call returned a decision, right or wrong (accuracy is in the routing study). Whiskers are 95% Wilson intervals.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
- Every model call routed (49.5 per task)
- Only System One decisions (7 per task) (square)
Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).
| Item | Every model call routed (49.5 per task) | Only System One decisions (7 per task) |
|---|---|---|
| Jev 1.13 (TypeSafe) | $1.67 | $0.24 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | $442 | $62.47 |
List-price calculation, not a run. 2 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $442. Lowest Jev 1.13 (TypeSafe) $1.67. Only System One decisions (7 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $62.47. Lowest Jev 1.13 (TypeSafe) $0.24.
Notes
Decisions per task from recorded runs × cost per decision
A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions
- Every model call routed (49.5 per task)
- Only System One decisions (7 per task) (square)
Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).
| Item | Every model call routed (49.5 per task) | Only System One decisions (7 per task) |
|---|---|---|
| Jev 1.13 (TypeSafe) | 6.8 s | 1 s |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 621 s | 87.8 s |
List-price calculation, not a run. 2 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 621 s. Fastest Jev 1.13 (TypeSafe) 6.8 s. Only System One decisions (7 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 87.8 s. Fastest Jev 1.13 (TypeSafe) 1 s.
Notes
Decisions per task × median decision time, if every decision waits in line
A calculation and an upper bound: it assumes each decision waits for the one before. Median recorded task wall time: 10.3 min. Jev’s delay uses its live median over the API from one Mac (network included); the Claude routers’ includes the CLI.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions
| Item | Exact rate | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 82% | 70%–90% | 56 |
| Claude Haiku 4.5 · Claude Code | 79% | 66%–87% | 56 |
2 rows. Highest Jev 1.13 (TypeSafe) 82% (95% interval 70%–90%, n 56). Lowest Claude Haiku 4.5 · Claude Code 79% (95% interval 66%–87%, n 56). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 56 per row
Share of the 56 holdout cases where every scored question was acceptable
Whiskers are 95% Wilson intervals. Cases written blind to every router answer, then frozen. Jev: its first of three repetitions.
Source: Routing on unseen holdout decisions
| Item | Key accuracy | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 90% | 84%–94% | 125 |
| Claude Haiku 4.5 · Claude Code | 82% | 74%–87% | 125 |
2 rows. Highest Jev 1.13 (TypeSafe) 90% (95% interval 84%–94%, n 125). Lowest Claude Haiku 4.5 · Claude Code 82% (95% interval 74%–87%, n 125). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 125 per row
Each open question a router was asked; an unanswered question counts as wrong
Whiskers are nominal 95% Wilson intervals. Context-shape cases ask up to 8 questions each, the other decision types one. Questions in one case are not independent; these intervals do not adjust for that grouping.
Source: Routing on unseen holdout decisions
Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | Jev 1.13 (TypeSafe) | Claude Haiku 4.5 · Claude Code | Claude Sonnet 5.5 (low) · Claude Code | 95% interval | n |
|---|---|---|---|---|---|
| Failure class | 93% | 93% | 100% | Jev 1.13 (TypeSafe): 69%–99%; Claude Haiku 4.5 · Claude Code: 69%–99%; Claude Sonnet 5.5 (low) · Claude Code: 78%–100% | 14 |
| Message intent | 86% | 93% | 100% | Jev 1.13 (TypeSafe): 60%–96%; Claude Haiku 4.5 · Claude Code: 69%–99%; Claude Sonnet 5.5 (low) · Claude Code: 78%–100% | 14 |
| Context shape | 57% | 36% | 57% | Jev 1.13 (TypeSafe): 33%–79%; Claude Haiku 4.5 · Claude Code: 16%–61%; Claude Sonnet 5.5 (low) · Claude Code: 33%–79% | 14 |
3 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 (low) · Claude Code. Jev 1.13 (TypeSafe): highest Failure class 93% (95% interval 69%–99%, n 14). Lowest Context shape 57% (95% interval 33%–79%, n 14). All intervals overlap. Claude Haiku 4.5 · Claude Code: highest Failure class 93% (95% interval 69%–99%, n 14). Lowest Context shape 36% (95% interval 16%–61%, n 14). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 14 per row
14 cases per decision type
Whiskers are 95% Wilson intervals. With 14 cases a perfect score has an interval of 78% to 100%, so a decision type where every router scores 14 of 14 is at its ceiling and cannot rank them.
Source: Routing on unseen holdout decisions
- Tuned set (routing-jev-vs-llm)
- Unseen holdout (square)
Gap labels, Unseen holdout vs Tuned set (routing-jev-vs-llm): Unseen holdout is x percentage points higher (+) or lower (−) than Tuned set (routing-jev-vs-llm), calculated from the two values shown; lines are the 95% Wilson interval.
| Item | Tuned set (routing-jev-vs-llm) | Unseen holdout | 95% interval | n |
|---|---|---|---|---|
| Jev 1.13 (TypeSafe) | 90% | 82% | Tuned set (routing-jev-vs-llm): 82%–95%; Unseen holdout: 70%–90% | 82 |
| Claude Haiku 4.5 · Claude Code | 89% | 79% | Tuned set (routing-jev-vs-llm): 80%–94%; Unseen holdout: 66%–87% | 82 |
2 rows, 2 series: Tuned set (routing-jev-vs-llm), Unseen holdout. Tuned set (routing-jev-vs-llm): highest Jev 1.13 (TypeSafe) 90% (95% interval 82%–95%, n 82). Lowest Claude Haiku 4.5 · Claude Code 89% (95% interval 80%–94%, n 82). All intervals overlap. Unseen holdout: highest Jev 1.13 (TypeSafe) 82% (95% interval 70%–90%, n 56). Lowest Claude Haiku 4.5 · Claude Code 79% (95% interval 66%–87%, n 56). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 56–82 per row
Tuned set: the routing study’s 82 cases, revised against Jev answers. Holdout: 56 new cases, frozen before any router call
Whiskers are 95% Wilson intervals. The two case sets differ in mix and size, so a gap mixes a change of case set with any change in the router, and the data cannot separate them. A gap counts only when the two intervals do not overlap.
Sources: Routing on unseen holdout decisions, Routing runs: Jev router vs LLM routing
- Wall time
- Model time (API, CLI-reported)
Seconds · log scale: each gridline is 10 times the one before
| Item | Wall time | Model time (API, CLI-reported) | Median to p95 | n |
|---|---|---|---|---|
| Jev 1.13 (TypeSafe) | 0.1 s | — | Wall time: 0.1 s–0.2 s | 168 |
| Claude Haiku 4.5 · Claude Code | 9.4 s | 7.5 s | Wall time: 9.4 s–25.4 s; Model time (API, CLI-reported): 7.5 s–23.9 s | 56 |
2 rows, 2 series: Wall time, Model time (API, CLI-reported). Wall time: slowest Claude Haiku 4.5 · Claude Code 9.4 s (median to p95 9.4 s–25.4 s, n 56). Fastest Jev 1.13 (TypeSafe) 0.1 s (median to p95 0.1 s–0.2 s, n 168). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n 56–168 per row
Median, whisker to the 95th percentile
The routes differ. Jev: one HTTPS call to the vendor API from one Mac, all three repetitions. Claude: the Claude Code CLI from process start to exit, which adds start-up time a direct API call would not. The whisker runs from the median to the 95th percentile; it is not a confidence interval. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.
Source: Routing on unseen holdout decisions
| Item | Cost | n |
|---|---|---|
| Jev 1.13 (TypeSafe) | $0.031 | 168 |
| Claude Haiku 4.5 · Claude Code | $7.13 | 56 |
List-price calculation, not a run. 2 rows. Highest Claude Haiku 4.5 · Claude Code $7.13 (n 56). Lowest Jev 1.13 (TypeSafe) $0.031 (n 168).
Notesn 56–168 per row
Reported tokens per decision × list price
Calculation, not a bill. Jev: reported input tokens × $0.042 per million, output free. Claude: CLI-reported tokens × list price; the CLI wrote its prompt cache as 1-hour writes, priced at 2× input, and adds its own system prompt and tool-schema tokens. The Claude routers ran on a subscription. At the tuned-set study’s convention (every cache write at 1.25× input), Sonnet 5.5 (low) would be $4.877 here. That study shows $4.996 for it on its own cases, where the CLI reported $7.324, so its cost row and this one differ by convention and by case mix.
Sources: Routing on unseen holdout decisions, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models)
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Which is better, Jev 1.13 or Claude Haiku 4.5?
- Jev 1.13 and Claude Haiku 4.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 12,543 ms; Time per routing decision, by route (Wall time), 0.14 s vs 9.44 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).
- How were Jev 1.13 and Claude Haiku 4.5 measured?
- They share 17 measured metrics and 6 list-price calculations from 3 public studies: Jev vs Claude as a router: accuracy and cost; Routing overhead: deterministic policy vs LLM routers vs Jev; Jev vs Claude routers on unseen decisions: a blind holdout. Every row names its configuration, its sample size and its interval or range.
- How do Jev 1.13 and Claude Haiku 4.5 compare on time to make one routing decision?
- Jev 1.13: 137 ms (routing overhead per decision · TypeSafe API; n = 246; p50 to p95 137 ms to 196 ms). Claude Haiku 4.5: 12,543 ms (thinking on · via Claude Code · routing overhead per decision; n = 82; p50 to p95 12.54 s to 34.48 s). Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Haiku 4.5 12,543 ms to 34,481 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone.
- How do Jev 1.13 and Claude Haiku 4.5 compare on time per routing decision, by route (Wall time)?
- Jev 1.13: 0.14 s (n = 168; p50 to p95 0.1 s to 0.2 s). Claude Haiku 4.5: 9.44 s (Claude Code; n = 56; p50 to p95 9.4 s to 25.4 s). Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Haiku 4.5 9.44 s to 25.4 s); not a confidence interval.
- How do Jev 1.13 and Claude Haiku 4.5 compare on typed routing decisions answered exactly right?
- Jev 1.13: 90% (typed routing decisions · TypeSafe API; n = 82; 95% interval 82% to 95%). Claude Haiku 4.5: 89% (73/82) (typed routing decisions · via Claude Code; n = 82; 95% interval 80% to 94%). The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Haiku 4.5 80% to 94%), so this sample cannot separate them.
- How do Jev 1.13 and Claude Haiku 4.5 compare on per-question accuracy?
- Jev 1.13: 95% (184/194) (typed routing decisions · TypeSafe API; n = 194; 95% interval 91% to 97%). Claude Haiku 4.5: 94% (183/194) (typed routing decisions · via Claude Code; n = 194; 95% interval 90% to 97%). The 95% intervals overlap (Jev 1.13 91% to 97%; Claude Haiku 4.5 90% to 97%), so this sample cannot separate them.
- How do Jev 1.13 and Claude Haiku 4.5 compare on exact rate by decision type: Failure class?
- Jev 1.13: 100% (18/18) (typed routing decisions · TypeSafe API; n = 18; 95% interval 82% to 100%). Claude Haiku 4.5: 94% (17/18) (typed routing decisions · via Claude Code; n = 18; 95% interval 74% to 99%). The 95% intervals overlap (Jev 1.13 82% to 100%; Claude Haiku 4.5 74% to 99%), so this sample cannot separate them.
The studies behind this page
Jev vs Claude as a router: accuracy and cost
Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.
Routing overhead: deterministic policy vs LLM routers vs Jev
How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.
Jev vs Claude routers on unseen decisions: a blind holdout
Jev 1.13, Claude Haiku 4.5 and Sonnet 5.5 on 56 routing decisions written blind and frozen first: exact rate with 95% intervals, tuned vs unseen.