17 measured metrics · 6 calculated · 3 studies

Jev 1.13vsClaude Sonnet 5.5

Jev 1.13 ahead on 2; 15 ties, 6 unclear. A side is ahead only where the intervals or ranges do not overlap.

The verdict

Jev 1.13 and Claude Sonnet 5.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 2,597 ms; Time per routing decision, by route (Wall time), 0.14 s vs 2.36 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates)n is shown per side on every row

6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Exact rate by decision type: Context shape, 1.1x (Claude Sonnet 5.5 larger).

Watch it build

A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.

Live story · 35 sJev 1.13 vs Claude Sonnet 5.5: what the measurements say

Jev 1.13 vs Claude Sonnet 5.5: what the measurements say

23 comparison rows from 3 studies: 2 rows favour Jev 1.13, 0 favour Sonnet 5.5, 21 are ties or unclear. Cost rows are calculations.

Transcript
  1. Comparison · 23 rows · 3 studies. Jev 1.13 vs Sonnet 5.5. A winner only where the 95% intervals or run ranges do not overlap.
  2. 23 comparison rows from 3 studies: Jev 1.13 ahead on 2, Sonnet 5.5 ahead on 0. The rest do not separate them. Rows where Jev 1.13 is ahead: 2 (of 23). Rows where Sonnet 5.5 is ahead: 0 (of 23). Ties or unclear: 21 (15 ties · 6 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
  3. Jev vs LLM routers: pass rate 90% vs 94%, tie: 95% intervals overlap. None of the 7 rows separates them. Table: Jev vs LLM routers · 5 of 7 rows · typed routing decisions · n = 12–194 per side. Source study: Jev vs Claude as a router: accuracy and cost. Rows shown: Typed routing decisions answered exactly right; Per-question accuracy; Exact rate by decision type: Failure class; Exact rate by decision type: Message intent; Exact rate by decision type: Is it a rule? Recorded settings: typed routing decisions · TypeSafe API; typed routing decisions · via Claude Code. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
  4. Routing overhead: success rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 6 rows separates them. Table: Routing overhead · 5 of 6 rows · routing overhead per decision vs effort low · n = 246 vs 82. Source study: Routing overhead: deterministic policy vs LLM routers vs Jev. Rows shown: Time to make one routing decision; Routing calls that returned a decision; Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)); Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)); Added routing delay per task (calculation) (Every model call routed (49.5 per task)). Recorded settings: routing overhead per decision · TypeSafe API; effort low · via Claude Code · routing overhead per decision; calculation per 1,000 tasks from recorded decision counts · TypeSafe API; effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts; calculation per task from recorded decision counts, decisions in line · TypeSafe API; effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line. Includes a calculation, not a bill or a new run. Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
  5. Unseen routing decisions: pass rate 82% vs 88%, tie: 95% intervals overlap. 1 of 10 rows separates them. Table: Unseen routing decisions · 5 of 10 rows · Claude Code · n = 14–168 vs 14–125. Source study: Jev vs Claude routers on unseen decisions: a blind holdout. Rows shown: Unseen routing decisions answered exactly right; Per-question accuracy on unseen decisions; Exact rate on unseen decisions, by decision type: Failure class; Exact rate on unseen decisions, by decision type: Message intent; Time per routing decision, by route (Wall time). Recorded settings: Claude Code · effort low. Caveat: 56 cases (14 per decision type) from one author: per-type intervals are very wide.
  6. No winner where the data shows none. Showing 15 of 23 rows; every row and its reason online.

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Jev 1.13
  • Claude Sonnet 5.5
  • 95% interval
  • median to p95 (not an interval)
  • where the two overlap
  • hollow: list-price calculation

These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.

Jev vs Claude as a router: accuracy and cost

6 ties · 1 unclear
  • Typed routing decisions answered exactly right: Jev 1.13 90% (n 82, 95% interval 82%–95%); Claude Sonnet 5.5 94% (77/82) (n 82, 95% interval 87%–97%). Tie.
  • Per-question accuracy: Jev 1.13 95% (184/194) (n 194, 95% interval 91%–97%); Claude Sonnet 5.5 97% (189/194) (n 194, 95% interval 94%–99%). Tie.

    Calculation: at these rates, about 826 runs per side would separate them.

  • Exact rate by decision type: Failure class: Jev 1.13 100% (18/18) (n 18, 95% interval 82%–100%); Claude Sonnet 5.5 100% (18/18) (n 18, 95% interval 82%–100%). Tie.
  • Exact rate by decision type: Message intent: Jev 1.13 100% (20/20) (n 20, 95% interval 84%–100%); Claude Sonnet 5.5 100% (20/20) (n 20, 95% interval 84%–100%). Tie.
  • Exact rate by decision type: Is it a rule?: Jev 1.13 100% (12/12) (n 12, 95% interval 76%–100%); Claude Sonnet 5.5 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
  • Exact rate by decision type: Context shape: Jev 1.13 74% (n 32, 95% interval 58%–87%); Claude Sonnet 5.5 84% (27/32) (n 32, 95% interval 68%–93%). Tie.
  • Cost per 1,000 routing decisions, calculation: Jev 1.13 $0.034 (n 246); Claude Sonnet 5.5 $5.00 (n 82). Unclear.

Routing overhead: deterministic policy vs LLM routers vs Jev

Jev 1.13 1 · 1 tie · 4 unclear
  • Time to make one routing decision: Jev 1.13 137 ms (n 246, median to p95 137 ms–196 ms); Claude Sonnet 5.5 2,597 ms (n 82, median to p95 2.6 s–4.3 s). Jev 1.13 ahead.
  • Routing calls that returned a decision: Jev 1.13 100% (246/246) (n 246, 95% interval 98%–100%); Claude Sonnet 5.5 100% (82/82) (n 82, 95% interval 96%–100%). Tie.
  • Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)), calculation: Jev 1.13 $1.67; Claude Sonnet 5.5 $247.30. Unclear.
  • Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)), calculation: Jev 1.13 $0.24; Claude Sonnet 5.5 $34.97. Unclear.
  • Added routing delay per task (calculation) (Every model call routed (49.5 per task)), calculation: Jev 1.13 6.76 s; Claude Sonnet 5.5 128.6 s. Unclear.
  • Added routing delay per task (calculation) (Only System One decisions (7 per task)), calculation: Jev 1.13 0.96 s; Claude Sonnet 5.5 18.2 s. Unclear.

Jev vs Claude routers on unseen decisions: a blind holdout

Jev 1.13 1 · 8 ties · 1 unclear
  • Unseen routing decisions answered exactly right: Jev 1.13 82% (46/56) (n 56, 95% interval 70%–90%); Claude Sonnet 5.5 88% (49/56) (n 56, 95% interval 76%–94%). Tie.

    Calculation: at these rates, about 675 runs per side would separate them.

  • Per-question accuracy on unseen decisions: Jev 1.13 90% (113/125) (n 125, 95% interval 84%–94%); Claude Sonnet 5.5 92% (115/125) (n 125, 95% interval 86%–96%). Tie.

    Calculation: at these rates, about 4,756 runs per side would separate them.

  • Exact rate on unseen decisions, by decision type: Failure class: Jev 1.13 93% (13/14) (n 14, 95% interval 69%–99%); Claude Sonnet 5.5 100% (14/14) (n 14, 95% interval 78%–100%). Tie.

    Calculation: at these rates, about 106 runs per side would separate them.

  • Exact rate on unseen decisions, by decision type: Message intent: Jev 1.13 86% (12/14) (n 14, 95% interval 60%–96%); Claude Sonnet 5.5 100% (14/14) (n 14, 95% interval 78%–100%). Tie.

    Calculation: at these rates, about 53 runs per side would separate them.

  • Exact rate on unseen decisions, by decision type: Is it a rule?: Jev 1.13 93% (13/14) (n 14, 95% interval 69%–99%); Claude Sonnet 5.5 93% (13/14) (n 14, 95% interval 69%–99%). Tie.
  • Exact rate on unseen decisions, by decision type: Context shape: Jev 1.13 57% (8/14) (n 14, 95% interval 33%–79%); Claude Sonnet 5.5 57% (8/14) (n 14, 95% interval 33%–79%). Tie.
  • Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm)): Jev 1.13 90% (74/82) (n 82, 95% interval 82%–95%); Claude Sonnet 5.5 94% (77/82) (n 82, 95% interval 87%–97%). Tie.

    Calculation: at these rates, about 795 runs per side would separate them.

  • Tuned case set vs unseen holdout: exact rate per router (Unseen holdout): Jev 1.13 82% (46/56) (n 56, 95% interval 70%–90%); Claude Sonnet 5.5 88% (49/56) (n 56, 95% interval 76%–94%). Tie.

    Calculation: at these rates, about 675 runs per side would separate them.

  • Time per routing decision, by route (Wall time): Jev 1.13 0.14 s (n 168, median to p95 0.1 s–0.2 s); Claude Sonnet 5.5 2.36 s (n 56, median to p95 2.4 s–3.7 s). Jev 1.13 ahead.
  • Cost per 1,000 unseen routing decisions, calculation: Jev 1.13 $0.031 (n 168); Claude Sonnet 5.5 $7.24 (n 56). Unclear.

Marks: 95% intervals (Wilson for rates); median to p95 bands (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs

23 rows from 3 studies. Jev 1.13 ahead on 2; 15 ties, 6 unclear. A side is ahead only where the intervals or ranges do not overlap.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Jev 1.13

  • Cost per 1,000 routing decisions: $0.034 vs $5.00. A list-price calculation, not a measured difference. Calculation
  • Time to make one routing decision: 137 ms vs 2,597 ms. Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Sonnet 5.5 2,597 ms to 4,298 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone.
  • Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)): $1.67 vs $247.30. A list-price calculation, not a measured difference. Calculation
  • Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)): $0.24 vs $34.97. A list-price calculation, not a measured difference. Calculation
  • Added routing delay per task (calculation) (Every model call routed (49.5 per task)): 6.76 s vs 128.6 s. A list-price calculation, not a measured difference. Calculation
  • Added routing delay per task (calculation) (Only System One decisions (7 per task)): 0.96 s vs 18.2 s. A list-price calculation, not a measured difference. Calculation
  • Time per routing decision, by route (Wall time): 0.14 s vs 2.36 s. Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Sonnet 5.5 2.36 s to 3.66 s); not a confidence interval.
  • Cost per 1,000 unseen routing decisions: $0.031 vs $7.24. A list-price calculation, not a measured difference. Calculation

When to pick Claude Sonnet 5.5

No row in this data puts Claude Sonnet 5.5 ahead of Jev 1.13. Pick on other grounds (price, access, the tasks you run), or measure your own workload.

Side by side

The study charts, showing only these two. Open a study for every configuration.

Jev 1.13 (TypeSafe)
Claude Sonnet 5.5

2 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Jev 1.13 (TypeSafe) 90% (95% interval 82%–95%, n 82). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 82 per row

Share of asked cases where every scored question was acceptable

Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Jev 1.13 (TypeSafe)
Claude Sonnet 5.5

2 rows. Highest Claude Sonnet 5.5 97% (95% interval 94%–99%, n 194). Lowest Jev 1.13 (TypeSafe) 95% (95% interval 91%–97%, n 194). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 194 per row

Each open question the router was asked; an unanswered question counts as wrong

Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Jev 1.13 (TypeSafe)

Failure class
Context shape

Claude Haiku 4.5

Failure class
Context shape

Claude Sonnet 5.5

Failure class
Context shape

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

2 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5, Claude Sonnet 5.5. Jev 1.13 (TypeSafe): highest Failure class 100% (95% interval 82%–100%, n 18). Lowest Context shape 74% (95% interval 58%–87%, n 32). All intervals overlap. Claude Haiku 4.5: highest Failure class 94% (95% interval 74%–99%, n 18). Lowest Context shape 75% (95% interval 58%–87%, n 32). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 18–32 per row

A missing bar means that router was not run on that decision type. Whiskers are 95% Wilson intervals. Jev: the pooled rate of the live repeats, with the interval taken at the number of decisions of that type.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Calculation
Largest value is 150x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5

List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 $5.00 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).

Notesn 82–246 per row

List price × reported tokens per decision

List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Jev 1.13 or Claude Sonnet 5.5?
Jev 1.13 and Claude Sonnet 5.5 share 17 measured metrics and 6 list-price calculations from 3 studies. Jev 1.13 leads on 2 rows: Time to make one routing decision, 137 ms vs 2,597 ms; Time per routing decision, by route (Wall time), 0.14 s vs 2.36 s. On those rows the p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 15 ties and 6 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. 13 rows ran the two sides through different routes (for example TypeSafe API vs Claude Code), so they compare route + model pairs, not models alone; the contexts name the route. Some rows rest on small samples (n = 12 at the smallest).
How were Jev 1.13 and Claude Sonnet 5.5 measured?
They share 17 measured metrics and 6 list-price calculations from 3 public studies: Jev vs Claude as a router: accuracy and cost; Routing overhead: deterministic policy vs LLM routers vs Jev; Jev vs Claude routers on unseen decisions: a blind holdout. Every row names its configuration, its sample size and its interval or range.
How do Jev 1.13 and Claude Sonnet 5.5 compare on time to make one routing decision?
Jev 1.13: 137 ms (routing overhead per decision · TypeSafe API; n = 246; p50 to p95 137 ms to 196 ms). Claude Sonnet 5.5: 2,597 ms (effort low · via Claude Code · routing overhead per decision; n = 82; p50 to p95 2.6 s to 4.3 s). Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 137 ms to 196 ms; Claude Sonnet 5.5 2,597 ms to 4,298 ms); not a confidence interval. The two sides ran through different routes (TypeSafe API vs Claude Code), so this row compares routes, not models alone.
How do Jev 1.13 and Claude Sonnet 5.5 compare on time per routing decision, by route (Wall time)?
Jev 1.13: 0.14 s (n = 168; p50 to p95 0.1 s to 0.2 s). Claude Sonnet 5.5: 2.36 s (Claude Code · effort low; n = 56; p50 to p95 2.4 s to 3.7 s). Claude Sonnet 5.5’s median is above Jev 1.13’s 95th percentile (p50–p95 bands: Jev 1.13 0.14 s to 0.19 s; Claude Sonnet 5.5 2.36 s to 3.66 s); not a confidence interval.
How do Jev 1.13 and Claude Sonnet 5.5 compare on typed routing decisions answered exactly right?
Jev 1.13: 90% (typed routing decisions · TypeSafe API; n = 82; 95% interval 82% to 95%). Claude Sonnet 5.5: 94% (77/82) (typed routing decisions · via Claude Code; n = 82; 95% interval 87% to 97%). The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them.
How do Jev 1.13 and Claude Sonnet 5.5 compare on per-question accuracy?
Jev 1.13: 95% (184/194) (typed routing decisions · TypeSafe API; n = 194; 95% interval 91% to 97%). Claude Sonnet 5.5: 97% (189/194) (typed routing decisions · via Claude Code; n = 194; 95% interval 94% to 99%). The 95% intervals overlap (Jev 1.13 91% to 97%; Claude Sonnet 5.5 94% to 99%), so this sample cannot separate them.
How do Jev 1.13 and Claude Sonnet 5.5 compare on exact rate by decision type: Failure class?
Jev 1.13: 100% (18/18) (typed routing decisions · TypeSafe API; n = 18; 95% interval 82% to 100%). Claude Sonnet 5.5: 100% (18/18) (typed routing decisions · via Claude Code; n = 18; 95% interval 82% to 100%). The 95% intervals overlap (Jev 1.13 82% to 100%; Claude Sonnet 5.5 82% to 100%), so this sample cannot separate them.

The studies behind this page

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

  • Routing
  • Jev

Jev vs Claude routers on unseen decisions: a blind holdout

Jev 1.13, Claude Haiku 4.5 and Sonnet 5.5 on 56 routing decisions written blind and frozen first: exact rate with 95% intervals, tuned vs unseen.

82% (46/56)Jev 1.13 (TypeSafe): exact on unseen decisions · n = 56

6 chartsUpdated October 6, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.