17 measured metrics · 6 calculated · 3 studies

Jev 1.13vsClaude Haiku 4.5

Jev 1.13 ahead on 2; 15 ties, 6 unclear. A side is ahead only where the intervals or ranges do not overlap.

The verdict

Jev 1.13 and Claude Haiku 4.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 12,543 ms; Time per routing decision, by route (Wall time), 0.14 s vs 9.44 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates)n is shown per side on every row

6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Exact rate by decision type: Failure class, 1.1x (Jev 1.13 larger).

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Jev 1.13
  • Claude Haiku 4.5
  • 95% interval
  • median to p95 (not an interval)
  • where the two overlap
  • hollow: list-price calculation

These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.

Jev vs Claude as a router: accuracy and cost

6 ties · 1 unclear
  • Typed routing decisions answered exactly right: Jev 1.13 90% (n 82, 95% interval 82%–95%); Claude Haiku 4.5 89% (73/82) (n 82, 95% interval 80%–94%). Tie.
  • Per-question accuracy: Jev 1.13 95% (184/194) (n 194, 95% interval 91%–97%); Claude Haiku 4.5 94% (183/194) (n 194, 95% interval 90%–97%). Tie.

    Calculation: at these rates, more than 5,000 runs per side would be needed.

  • Exact rate by decision type: Failure class: Jev 1.13 100% (18/18) (n 18, 95% interval 82%–100%); Claude Haiku 4.5 94% (17/18) (n 18, 95% interval 74%–99%). Tie.

    Calculation: at these rates, about 135 runs per side would separate them.

  • Exact rate by decision type: Message intent: Jev 1.13 100% (20/20) (n 20, 95% interval 84%–100%); Claude Haiku 4.5 100% (20/20) (n 20, 95% interval 84%–100%). Tie.
  • Exact rate by decision type: Is it a rule?: Jev 1.13 100% (12/12) (n 12, 95% interval 76%–100%); Claude Haiku 4.5 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
  • Exact rate by decision type: Context shape: Jev 1.13 74% (n 32, 95% interval 58%–87%); Claude Haiku 4.5 75% (24/32) (n 32, 95% interval 58%–87%). Tie.
  • Cost per 1,000 routing decisions, calculation: Jev 1.13 $0.034 (n 246); Claude Haiku 4.5 $8.92 (n 82). Unclear.

Routing overhead: deterministic policy vs LLM routers vs Jev

Jev 1.13 1 · 1 tie · 4 unclear
  • Time to make one routing decision: Jev 1.13 137 ms (n 246, median to p95 137 ms–196 ms); Claude Haiku 4.5 12,543 ms (n 82, median to p95 12.54 s–34.48 s). Jev 1.13 ahead.
  • Routing calls that returned a decision: Jev 1.13 100% (246/246) (n 246, 95% interval 98%–100%); Claude Haiku 4.5 100% (82/82) (n 82, 95% interval 96%–100%). Tie.
  • Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)), calculation: Jev 1.13 $1.67; Claude Haiku 4.5 $441.74. Unclear.
  • Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)), calculation: Jev 1.13 $0.24; Claude Haiku 4.5 $62.47. Unclear.
  • Added routing delay per task (calculation) (Every model call routed (49.5 per task)), calculation: Jev 1.13 6.76 s; Claude Haiku 4.5 620.9 s. Unclear.
  • Added routing delay per task (calculation) (Only System One decisions (7 per task)), calculation: Jev 1.13 0.96 s; Claude Haiku 4.5 87.8 s. Unclear.

Jev vs Claude routers on unseen decisions: a blind holdout

Jev 1.13 1 · 8 ties · 1 unclear
  • Unseen routing decisions answered exactly right: Jev 1.13 82% (46/56) (n 56, 95% interval 70%–90%); Claude Haiku 4.5 79% (44/56) (n 56, 95% interval 66%–87%). Tie.

    Calculation: at these rates, about 1,883 runs per side would separate them.

  • Per-question accuracy on unseen decisions: Jev 1.13 90% (113/125) (n 125, 95% interval 84%–94%); Claude Haiku 4.5 82% (102/125) (n 125, 95% interval 74%–87%). Tie.

    Calculation: at these rates, about 231 runs per side would separate them.

  • Exact rate on unseen decisions, by decision type: Failure class: Jev 1.13 93% (13/14) (n 14, 95% interval 69%–99%); Claude Haiku 4.5 93% (13/14) (n 14, 95% interval 69%–99%). Tie.
  • Exact rate on unseen decisions, by decision type: Message intent: Jev 1.13 86% (12/14) (n 14, 95% interval 60%–96%); Claude Haiku 4.5 93% (13/14) (n 14, 95% interval 69%–99%). Tie.

    Calculation: at these rates, about 270 runs per side would separate them.

  • Exact rate on unseen decisions, by decision type: Is it a rule?: Jev 1.13 93% (13/14) (n 14, 95% interval 69%–99%); Claude Haiku 4.5 93% (13/14) (n 14, 95% interval 69%–99%). Tie.
  • Exact rate on unseen decisions, by decision type: Context shape: Jev 1.13 57% (8/14) (n 14, 95% interval 33%–79%); Claude Haiku 4.5 36% (5/14) (n 14, 95% interval 16%–61%). Tie.

    Calculation: at these rates, about 82 runs per side would separate them.

  • Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm)): Jev 1.13 90% (74/82) (n 82, 95% interval 82%–95%); Claude Haiku 4.5 89% (73/82) (n 82, 95% interval 80%–94%). Tie.

    Calculation: at these rates, more than 5,000 runs per side would be needed.

  • Tuned case set vs unseen holdout: exact rate per router (Unseen holdout): Jev 1.13 82% (46/56) (n 56, 95% interval 70%–90%); Claude Haiku 4.5 79% (44/56) (n 56, 95% interval 66%–87%). Tie.

    Calculation: at these rates, about 1,883 runs per side would separate them.

  • Time per routing decision, by route (Wall time): Jev 1.13 0.14 s (n 168, median to p95 0.1 s–0.2 s); Claude Haiku 4.5 9.44 s (n 56, median to p95 9.4 s–25.4 s). Jev 1.13 ahead.
  • Cost per 1,000 unseen routing decisions, calculation: Jev 1.13 $0.031 (n 168); Claude Haiku 4.5 $7.13 (n 56). Unclear.

Marks: 95% intervals (Wilson for rates); median to p95 bands (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

23 rows from 3 studies. Jev 1.13 ahead on 2; 15 ties, 6 unclear. A side is ahead only where the intervals or ranges do not overlap.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Jev 1.13

  • Cost per 1,000 routing decisions: $0.034 vs $8.92. A list-price calculation, not a measured difference. Calculation
  • Time to make one routing decision: 137 ms vs 12,543 ms. Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Haiku 4.5 12,543 ms to 34,481 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone.
  • Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)): $1.67 vs $441.74. A list-price calculation, not a measured difference. Calculation
  • Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)): $0.24 vs $62.47. A list-price calculation, not a measured difference. Calculation
  • Added routing delay per task (calculation) (Every model call routed (49.5 per task)): 6.76 s vs 620.9 s. A list-price calculation, not a measured difference. Calculation
  • Added routing delay per task (calculation) (Only System One decisions (7 per task)): 0.96 s vs 87.8 s. A list-price calculation, not a measured difference. Calculation
  • Time per routing decision, by route (Wall time): 0.14 s vs 9.44 s. Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Haiku 4.5 9.44 s to 25.4 s); not a confidence interval.
  • Cost per 1,000 unseen routing decisions: $0.031 vs $7.13. A list-price calculation, not a measured difference. Calculation

When to pick Claude Haiku 4.5

No row in this data puts Claude Haiku 4.5 ahead of Jev 1.13. Pick on other grounds (price, access, the tasks you run), or measure your own workload.

Side by side

The study charts, showing only these two. Open a study for every configuration.

Jev 1.13 (TypeSafe)
Claude Haiku 4.5

2 rows. Highest Jev 1.13 (TypeSafe) 90% (95% interval 82%–95%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 82 per row

Share of asked cases where every scored question was acceptable

Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Jev 1.13 (TypeSafe)
Claude Haiku 4.5

2 rows. Highest Jev 1.13 (TypeSafe) 95% (95% interval 91%–97%, n 194). Lowest Claude Haiku 4.5 94% (95% interval 90%–97%, n 194). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 194 per row

Each open question the router was asked; an unanswered question counts as wrong

Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Jev 1.13 (TypeSafe)

Failure class
Message intent
Context shape

Claude Haiku 4.5

Failure class
Message intent
Context shape

Claude Sonnet 5.5

Failure class
Message intent
Context shape

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

3 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5, Claude Sonnet 5.5. Jev 1.13 (TypeSafe): highest Failure class 100% (95% interval 82%–100%, n 18). Lowest Context shape 74% (95% interval 58%–87%, n 32). All intervals overlap. Claude Haiku 4.5: highest Message intent 100% (95% interval 84%–100%, n 20). Lowest Context shape 75% (95% interval 58%–87%, n 32). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 18–32 per row

A missing bar means that router was not run on that decision type. Whiskers are 95% Wilson intervals. Jev: the pooled rate of the live repeats, with the interval taken at the number of decisions of that type.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Calculation
Largest value is 260x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5

List-price calculation, not a run. 2 rows. Highest Claude Haiku 4.5 $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).

Notesn 82–246 per row

List price × reported tokens per decision

List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Jev 1.13 or Claude Haiku 4.5?
Jev 1.13 and Claude Haiku 4.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 12,543 ms; Time per routing decision, by route (Wall time), 0.14 s vs 9.44 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).
How were Jev 1.13 and Claude Haiku 4.5 measured?
They share 17 measured metrics and 6 list-price calculations from 3 public studies: Jev vs Claude as a router: accuracy and cost; Routing overhead: deterministic policy vs LLM routers vs Jev; Jev vs Claude routers on unseen decisions: a blind holdout. Every row names its configuration, its sample size and its interval or range.
How do Jev 1.13 and Claude Haiku 4.5 compare on time to make one routing decision?
Jev 1.13: 137 ms (routing overhead per decision · TypeSafe API; n = 246; p50 to p95 137 ms to 196 ms). Claude Haiku 4.5: 12,543 ms (thinking on · via Claude Code · routing overhead per decision; n = 82; p50 to p95 12.54 s to 34.48 s). Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Haiku 4.5 12,543 ms to 34,481 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone.
How do Jev 1.13 and Claude Haiku 4.5 compare on time per routing decision, by route (Wall time)?
Jev 1.13: 0.14 s (n = 168; p50 to p95 0.1 s to 0.2 s). Claude Haiku 4.5: 9.44 s (Claude Code; n = 56; p50 to p95 9.4 s to 25.4 s). Claude Haiku 4.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Haiku 4.5 9.44 s to 25.4 s); not a confidence interval.
How do Jev 1.13 and Claude Haiku 4.5 compare on typed routing decisions answered exactly right?
Jev 1.13: 90% (typed routing decisions · TypeSafe API; n = 82; 95% interval 82% to 95%). Claude Haiku 4.5: 89% (73/82) (typed routing decisions · via Claude Code; n = 82; 95% interval 80% to 94%). The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Haiku 4.5 80% to 94%), so this sample cannot separate them.
How do Jev 1.13 and Claude Haiku 4.5 compare on per-question accuracy?
Jev 1.13: 95% (184/194) (typed routing decisions · TypeSafe API; n = 194; 95% interval 91% to 97%). Claude Haiku 4.5: 94% (183/194) (typed routing decisions · via Claude Code; n = 194; 95% interval 90% to 97%). The 95% intervals overlap (Jev 1.13 91% to 97%; Claude Haiku 4.5 90% to 97%), so this sample cannot separate them.
How do Jev 1.13 and Claude Haiku 4.5 compare on exact rate by decision type: Failure class?
Jev 1.13: 100% (18/18) (typed routing decisions · TypeSafe API; n = 18; 95% interval 82% to 100%). Claude Haiku 4.5: 94% (17/18) (typed routing decisions · via Claude Code; n = 18; 95% interval 74% to 99%). The 95% intervals overlap (Jev 1.13 82% to 100%; Claude Haiku 4.5 74% to 99%), so this sample cannot separate them.

The studies behind this page

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

  • Routing
  • Jev

Jev vs Claude routers on unseen decisions: a blind holdout

Jev 1.13, Claude Haiku 4.5 and Sonnet 5.5 on 56 routing decisions written blind and frozen first: exact rate with 95% intervals, tuned vs unseen.

82% (46/56)Jev 1.13 (TypeSafe): exact on unseen decisions · n = 56

6 chartsUpdated October 6, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.