• Routing
  • Model Routing
  • Jev
  • Claude Haiku

Jev vs Claude routers on unseen decisions: a blind holdout test

56 new routing decisions, written blind and frozen before any router call. Jev 1.13 vs Claude Haiku 4.5 vs Sonnet 5.5: accuracy, 95% intervals, cost, speed.

TL;DR

  • Accuracy on unseen decisions: no winner. On 56 new routing decisions, Jev 1.13 answered 46 exactly right (82%, 95% interval 70% to 90%). Claude Haiku 4.5 answered 44 (79%, 66% to 87%) and Claude Sonnet 5.5 at effort low 49 (88%, 76% to 94%). The intervals overlap. The paired McNemar test finds no significant difference in any pair (p = 0.18 to 0.69).
  • Every router scored lower on the holdout, and the data cannot say why. Jev went from 90% on the tuned case set to 82% (−8.1 points), Haiku from 89% to 79% (−10.5) and Sonnet from 94% to 88% (−6.4). These changes are calculations. Each pair of intervals overlaps. The data neither shows nor rules out a home advantage for Jev: the two sets differ in mix, and split by decision type the largest drop is Jev's on the three one-question types and Haiku's on context shape.
  • A second labeller from another vendor checked our labels, and one of our labels broke our own rules. GPT-6.1 Sol labelled every case blind. Its label was accepted within our acceptable sets on 115 of 125 questions (92%, 95% interval 86% to 96%); first label only, it matched on 100 of 125 (80%, 72% to 86%). Our "artifacts" label broke our protocol and matched Jev's answers more often than the Claude routers' (Jev 9 of 9, 95% interval 70% to 100%; Haiku 4 of 9, 19% to 73%; Sonnet 7 of 9, 45% to 94%). With both answers accepted for it (calculation), the counts are Jev 46 of 56 (95% interval 70% to 90%), Haiku 45 of 56 (68% to 89%) and Sonnet 50 of 56 (79% to 95%). On the 115 agreed questions (calculation), exact cases are Jev 46 of 56 (70% to 90%), Haiku 46 of 56 (70% to 90%) and Sonnet 52 of 56 (83% to 97%). These intervals still overlap.
  • Cost and measured time differ; accuracy has no clear ranking. Jev costs $0.0307 per 1,000 decisions. Haiku costs $7.13 and Sonnet $7.24 (list-price calculations, about 230 times as much). Jev took a median 139 ms (p50–p95 139–192 ms, n = 168) over direct HTTPS. Sonnet took 2.36 s (2.36–3.66 s, n = 56) and Haiku 9.44 s (9.44–25.4 s, n = 56) through the Claude Code CLI. These p50–p95 ranges are not confidence intervals. The routes differ, and the Mac also ran other local jobs, so the times may carry contention.
  • Three of the four decision types are near the ceiling. Every router got 12 to 14 of 14 on them (95% intervals shown below). Context shape has up to 8 questions per case: Jev and Sonnet 8 of 14 (33% to 79%), Haiku 5 of 14 (16% to 61%).

Disclosure: I build Agent. Its platform uses Jev for some typed decisions in production. Jev is a third-party model from TypeSafe; we do not build or sell it.

Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 (low) · Claude Code 88% (95% interval 76%–94%, n 56). Lowest Claude Haiku 4.5 · Claude Code 79% (95% interval 66%–87%, n 56). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 56 per row

Share of the 56 holdout cases where every scored question was acceptable

Whiskers are 95% Wilson intervals. Cases written blind to every router answer, then frozen. Jev: its first of three repetitions.

Source: Routing on unseen holdout decisions

Why we built a holdout

Our first routing study, Jev vs LLM routing, had a known weakness. Its 82 cases and question texts were revised in fix waves against Jev answers. That can favour Jev. We said so in that study's caveats. Then we built a holdout to test it.

The question for this study: does a router keep its accuracy on typed routing decisions that nobody tuned against its answers?

How we built the holdout

We followed a written protocol. Its file time is 21:35:45 UTC, after the case drafts and before the first counted call. Every later change has a dated line in its log, including the correction of the time we first claimed (21:30 UTC).

  1. Protocol before counted labelling. We wrote the question, the routers, the call caps, the scoring and the stop rules into a protocol file at 21:35:45 UTC (file time). That is after the case drafts and one uncounted probe of the second labeller, and before the first counted labelling call (21:35:49 UTC), the freeze and every router call. The first version said it was declared at 21:30 UTC, before any labelling call. File times did not support that, so we corrected it.
  2. 56 new cases, 14 per decision type. The four types are failure class, message intent, "is it a rule?" and context shape. Each case has the same shape and the same questions as the production cases. Each case is one that production really asks a router about. Only the questions that the platform's rules leave open are labelled: 125 questions in all.
  3. Blind to every router answer. The case author did not open any router answer, old or new, until the labels were frozen. This is our statement. File times are consistent with it, but no file can prove it.
  4. No paraphrases. A script compared each new case with the 166 existing cases. 17 new cases were too close in topic or wording, so we rewrote them before the freeze.
  5. A second, blind labeller. GPT-6.1 Sol in the Codex CLI labelled every case without our labels. It saw what a router sees: the case state and the questions.
  6. Freeze. We recorded the case file and its checksum at 21:37 UTC. The first router call, an uncounted probe, came 48 seconds later.
  7. Run. Jev 1.13 over direct HTTPS, with the production request body: 3 repetitions of 56 cases, 168 calls, no retries. Claude Haiku 4.5 (CLI default effort) and Claude Sonnet 5.5 (effort low) through the Claude Code CLI, with the production system text and answer schema: 56 calls each. One uncounted probe per route. Every call succeeded.
  8. Score. The repository's own decision-eval runner scored every router. A case is exact when every scored question is acceptable. Key accuracy scores each question on its own.

We publish the case ids, the labels and who answered each case exactly. We do not publish the case text.

Accuracy on unseen decisions

  • Jev 1.13 (TypeSafe): 46 of 56, 82%, interval 70% to 90%.
  • Claude Haiku 4.5: 44 of 56, 79%, interval 66% to 87%.
  • Claude Sonnet 5.5 (effort low): 49 of 56, 88%, interval 76% to 94%.

Per question: Jev 113 of 125 (90%, 95% interval 84% to 94%), Haiku 102 of 125 (82%, 74% to 87%), Sonnet 115 of 125 (92%, 86% to 96%). These intervals overlap too. Questions within one case are not independent.

Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 (low) · Claude Code 92% (95% interval 86%–96%, n 125). Lowest Claude Haiku 4.5 · Claude Code 82% (95% interval 74%–87%, n 125). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 125 per row

Each open question a router was asked; an unanswered question counts as wrong

Whiskers are nominal 95% Wilson intervals. Context-shape cases ask up to 8 questions each, the other decision types one. Questions in one case are not independent; these intervals do not adjust for that grouping.

Source: Routing on unseen holdout decisions

Sonnet has the highest point estimate. Is the gap real? For two routers on the same cases, McNemar's test is the right check. It looks only at the cases where the two routers disagree.

PairBoth rightOnly first rightOnly second rightBoth wrongExact p
Jev vs Haiku 4.5424280.688
Jev vs Sonnet 5.5424730.549
Haiku 4.5 vs Sonnet 5.5422750.180

None of these p-values meets the 0.05 threshold. The three paired tests have no multiple-test correction. With 56 cases, this test cannot tell the three routers apart.

Tuned set vs unseen holdout

  • Tuned set (routing-jev-vs-llm)
  • Unseen holdout (square)
In chart order.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

Gap labels, Unseen holdout vs Tuned set (routing-jev-vs-llm): Unseen holdout is x percentage points higher (+) or lower (−) than Tuned set (routing-jev-vs-llm), calculated from the two values shown; lines are the 95% Wilson interval.

3 rows, 2 series: Tuned set (routing-jev-vs-llm), Unseen holdout. Tuned set (routing-jev-vs-llm): highest Claude Sonnet 5.5 (low) · Claude Code 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 · Claude Code 89% (95% interval 80%–94%, n 82). All intervals overlap. Unseen holdout: highest Claude Sonnet 5.5 (low) · Claude Code 88% (95% interval 76%–94%, n 56). Lowest Claude Haiku 4.5 · Claude Code 79% (95% interval 66%–87%, n 56). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 56–82 per row

Tuned set: the routing study’s 82 cases, revised against Jev answers. Holdout: 56 new cases, frozen before any router call

Whiskers are 95% Wilson intervals. The two case sets differ in mix and size, so a gap mixes a change of case set with any change in the router, and the data cannot separate them. A gap counts only when the two intervals do not overlap.

Sources: Routing on unseen holdout decisions, Routing runs: Jev router vs LLM routing

Here is the tuned-set chart from the first study. Its Jev point pools three repetitions (221 of 246 exact calls across 82 cases, 95% case-level interval 82% to 95%). The comparison table below uses Jev's first repetition, 74 of 82, as does the holdout (calculation). The chart uses case-level intervals because repetitions are not independent:

Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 82 per row

Share of asked cases where every scored question was acceptable

Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

RouterTuned set (82 cases)Unseen holdout (56 cases)Change (calculation)
Jev 1.1374, 90% (82% to 95%)46, 82% (70% to 90%)−8.1 points
Claude Haiku 4.573, 89% (80% to 94%)44, 79% (66% to 87%)−10.5 points
Claude Sonnet 5.577, 94% (87% to 97%)49, 88% (76% to 94%)−6.4 points

Each router's two intervals overlap, so the data shows no clear drop for any one router. All three routers scored lower on the holdout, by 6.4 to 10.5 points. The data neither shows nor rules out a home advantage for Jev. Two facts stop us from reading the aggregate drops either way:

  • The drops are close and the intervals are wide. Jev's drop (8.1 points) is larger than Sonnet's (6.4) and smaller than Haiku's (10.5). The differences between these observed drops are calculations, from 1.7 to 4.1 points. They are not a tested ranking of drops.
  • The two sets differ in mix. Context shape is 32 of the 82 tuned cases and 14 of the 56 holdout cases. The aggregate compares different blends of decision types, so it can hide movements in opposite directions by type.

So we split the comparison by decision group. The three one-question types are failure class, message intent and "is it a rule?". Context shape asks up to 8 questions per case. This is a calculation from the routing study's per-type records and this study:

RouterDecision groupTuned setUnseen holdoutChange (calculation)
Jev 1.13One-question types50 of 50, 100.0% (92.9% to 100.0%)38 of 42, 90.5% (77.9% to 96.2%)−9.5 points
Jev 1.13Context shape24 of 32, 75.0% (57.9% to 86.7%)8 of 14, 57.1% (32.6% to 78.6%)−17.9 points
Claude Haiku 4.5One-question types49 of 50, 98.0% (89.5% to 99.6%)39 of 42, 92.9% (81.0% to 97.5%)−5.1 points
Claude Haiku 4.5Context shape24 of 32, 75.0% (57.9% to 86.7%)5 of 14, 35.7% (16.3% to 61.2%)−39.3 points
Claude Sonnet 5.5One-question types50 of 50, 100.0% (92.9% to 100.0%)41 of 42, 97.6% (87.7% to 99.6%)−2.4 points
Claude Sonnet 5.5Context shape27 of 32, 84.4% (68.2% to 93.1%)8 of 14, 57.1% (32.6% to 78.6%)−27.2 points

On the one-question types, Jev has the largest observed drop (−9.5 points) and Sonnet the smallest (−2.4). On context shape, Haiku has the largest observed drop (−39.3) and Jev the smallest (−17.9). These calculations describe point estimates, not a ranking of drops. The split points in both directions. Every pair of intervals overlaps. Haiku's context-shape pair overlaps by only 3.3 points (the tuned interval starts at 57.9%, the holdout interval ends at 61.2%), but our rule holds: a gap counts only when the intervals do not overlap.

The author builds Agent, which uses Jev, and this study exists to answer the doubt about a home advantage. These data do not settle it. A larger holdout, with more context-shape cases, would test it better. Read any change between the two sets as a mix of case set and router. The data cannot separate them.

Where the routers miss

Jev 1.13 (TypeSafe)

Failure class
Message intent
Is it a rule?
Context shape

Claude Haiku 4.5 · Claude Code

Failure class
Message intent
Is it a rule?
Context shape

Claude Sonnet 5.5 (low) · Claude Code

Failure class
Message intent
Is it a rule?
Context shape

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

4 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 (low) · Claude Code. Jev 1.13 (TypeSafe): highest Failure class 93% (95% interval 69%–99%, n 14). Lowest Context shape 57% (95% interval 33%–79%, n 14). All intervals overlap. Claude Haiku 4.5 · Claude Code: highest Failure class 93% (95% interval 69%–99%, n 14). Lowest Context shape 36% (95% interval 16%–61%, n 14). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 14 per row

14 cases per decision type

Whiskers are 95% Wilson intervals. With 14 cases a perfect score has an interval of 78% to 100%, so a decision type where every router scores 14 of 14 is at its ceiling and cannot rank them.

Source: Routing on unseen holdout decisions

Exact answers per decision type, of 14 each (95% Wilson intervals):

Decision typeJev 1.13Haiku 4.5Sonnet 5.5
Failure class13 (69% to 99%)13 (69% to 99%)14 (78% to 100%)
Message intent12 (60% to 96%)13 (69% to 99%)14 (78% to 100%)
Is it a rule?13 (69% to 99%)13 (69% to 99%)13 (69% to 99%)
Context shape8 (33% to 79%)5 (16% to 61%)8 (33% to 79%)

Three types are near the ceiling for every router. A perfect 14 of 14 still has an interval of 78% to 100%, so these sets cannot rank the routers. Context shape asks up to 8 questions per case: which turn this is, how much transcript, which knowledge, memories, examples, size and complexity. One wrong question makes the whole case wrong. That is where most misses are.

Jev missed 10 cases in its first repetition: one failure class case (holdout-09), two message intent cases (holdout-25 and holdout-28), one "is it a rule?" case (holdout-32) and six context shape cases. On every question Jev got wrong in those 10 cases, the second labeller's label was inside our acceptable set.

Jev was also stable. It gave the same answers on every question in all three repetitions for 53 of 56 cases (95%, interval 85% to 98%). Its exact count was 46, 46 and 47 of 56 in the three repetitions (95% intervals 70% to 90%, 70% to 90%, and 72% to 91%).

Did our labels favour any router?

The case author is a Claude model, and two of the three routers are Claude models. A same-family bias in the labels is possible. So we asked a model from another vendor to label every case blind.

GPT-6.1 Sol's label was accepted within our acceptable sets on 115 of 125 questions (92%, interval 86% to 96%). Our sets are wide on some questions: 24 of the 125 questions accept two or three labels. On our first label only, it matched on 100 of 125 (80%, interval 72% to 86%). These are question-level intervals; questions within a case are not independent. First-label agreement differs on five questions: knowledge 11 of 14 (95% interval 52% to 92%), scope and complexity each 10 of 14 (45% to 88%), turn 3 of 4 (30% to 95%) and transcript 2 of 5 (12% to 77%).

Agreement was 14 of 14 (95% interval 78% to 100%) for failure class, message intent and "is it a rule?", both ways. In context shape, two questions held all 10 disagreements inside our sets:

  • Artifacts (does this turn need the work item's branches, files and links listed?): 1 of 9 (95% interval 2% to 44%). On a fresh work item we labelled "no", because a fresh item has no artifacts yet. GPT-6.1 Sol said "yes" on 8 of 9 (95% interval 57% to 98%). Cohen's kappa is 0 (calculation).
  • Examples (does this turn need approved examples of finished work?): 11 of 13 (95% interval 58% to 96%), kappa 0.68 (calculation).

The artifacts question matters for fairness, and our label broke our own rules. The protocol says to leave out a question whose answer is a judgement call. The production case suites accept both "no" and "yes" for a fresh work item. We labelled only "no" on all 9 fresh-item cases. Jev matched our "no" label on 9 of 9 (95% interval 70% to 100%). Haiku matched on 4 of 9 (19% to 73%) and Sonnet on 7 of 9 (45% to 94%). So the scoring against our labels favours Jev on this question. We kept the labels frozen, as the protocol says, and report the effect two ways. The first scores only the questions where the second label is inside our set. The second accepts both answers for artifacts and scores the same answers again (calculation).

RouterExact, our labels (95% interval)Exact, artifacts accepted both ways (calculation)Exact, questions inside our sets only (calculation)
Jev 1.1346 of 56 (70% to 90%)46 of 56 (70% to 90%)46 of 56 (70% to 90%)
Claude Haiku 4.544 of 56 (66% to 87%)45 of 56 (68% to 89%)46 of 56 (70% to 90%)
Claude Sonnet 5.549 of 56 (76% to 94%)50 of 56 (79% to 95%)52 of 56 (83% to 97%)

With artifacts accepted both ways, context shape goes from 8, 5 and 8 of 14 to 8 (95% interval 33% to 79%), 6 (21% to 67%) and 9 (39% to 84%) of 14 for Jev, Haiku and Sonnet (calculation). On either rescoring, Sonnet's point estimate is higher still (50 or 52 of 56). The intervals still overlap, so neither rescoring establishes a ranking.

Speed and cost

  • Wall time
  • Model time (API, CLI-reported)
Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

Seconds · log scale: each gridline is 10 times the one before

3 rows, 2 series: Wall time, Model time (API, CLI-reported). Wall time: slowest Claude Haiku 4.5 · Claude Code 9.4 s (median to p95 9.4 s–25.4 s, n 56). Fastest Jev 1.13 (TypeSafe) 0.1 s (median to p95 0.1 s–0.2 s, n 168). Not all run ranges overlap. Model time (API, CLI-reported): slowest Claude Haiku 4.5 · Claude Code 7.5 s (median to p95 7.5 s–23.9 s, n 56). Fastest Claude Sonnet 5.5 (low) · Claude Code 1.5 s (median to p95 1.5 s–2.4 s, n 56). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 56–168 per row

Median, whisker to the 95th percentile

The routes differ. Jev: one HTTPS call to the vendor API from one Mac, all three repetitions. Claude: the Claude Code CLI from process start to exit, which adds start-up time a direct API call would not. The whisker runs from the median to the 95th percentile; it is not a confidence interval. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.

Source: Routing on unseen holdout decisions

  • Jev 1.13: median 139 ms, 95th percentile 192 ms, over direct HTTPS from one Mac (168 calls).
  • Claude Sonnet 5.5: median 2.36 s, 95th percentile 3.66 s (56 calls). CLI-reported API time was a median 1.49 s (p50–p95 1.49–2.38 s, n = 56). The median per-call wall-minus-API time was 0.861 s (calculation); subtraction of the two medians does not give that figure.
  • Claude Haiku 4.5: median 9.44 s, 95th percentile 25.4 s (56 calls). Haiku ran with the CLI default extended thinking. This run cannot isolate the effect of thinking on time.

The p50–p95 ranges are not confidence intervals. The routes differ. A direct API call to Claude would skip the CLI start-up time. We did not measure that route here. The Mac also ran other local jobs during the run, so wall times, the CLI times most of all, may carry contention. We did not rerun the timing.

Calculation
Largest value is 240x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Sonnet 5.5 (low) · Claude Code $7.24 (n 56). Lowest Jev 1.13 (TypeSafe) $0.031 (n 168).

Notesn 56–168 per row

Reported tokens per decision × list price

Calculation, not a bill. Jev: reported input tokens × $0.042 per million, output free. Claude: CLI-reported tokens × list price; the CLI wrote its prompt cache as 1-hour writes, priced at 2× input, and adds its own system prompt and tool-schema tokens. The Claude routers ran on a subscription. At the tuned-set study’s convention (every cache write at 1.25× input), Sonnet 5.5 (low) would be $4.877 here. That study shows $4.996 for it on its own cases, where the CLI reported $7.324, so its cost row and this one differ by convention and by case mix.

Sources: Routing on unseen holdout decisions, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models)

Per 1,000 decisions, at list prices (a calculation, not a bill; n = 168 Jev calls and 56 calls per Claude router). Mean token counts below are calculations too:

  • Jev 1.13: $0.0307. About 730 input tokens per decision, and Jev output is free at its list price.
  • Claude Haiku 4.5: $7.13. About 1,775 input and 1,071 output tokens per decision, including thinking. The receipt does not split Haiku output into answer and thinking tokens.
  • Claude Sonnet 5.5: $7.24. About 1,705 input tokens, almost all written to the prompt cache, and 90.5 output tokens.

The CLI wrote Sonnet's prompt cache as 1-hour writes, which cost twice the input price. Our first routing study priced all cache writes at 1.25 times the input price. At that rate, Sonnet would cost $4.88 per 1,000 here. That study shows $4.996 for Sonnet on its own cases, where the CLI reported $7.32, so the two studies' Sonnet cost rows differ by convention and by case mix. At the primary convention, the Claude cost estimates are 232.6 to 236.3 times Jev's (calculation). At the tuned-set convention, they are 159.1 to 232.6 times Jev's (calculation).

The comparison pages line up every shared metric for each pair: Jev vs Claude Sonnet 5.5 and Jev vs Claude Haiku 4.5.

How to evaluate a router: what we recommend

  1. Keep a holdout set. Write it before you look at any router answer, freeze it, and never fix it after a run. Fix waves on one set leak the router's habits into the labels.
  2. Label twice, with two model families. Report agreement and kappa per question, and say whether agreement counts any label inside your acceptable set or your first label only. A question with low agreement, like artifacts here, is a label problem before it is a router problem. Follow your own rule for judgement-call questions: leave them out, or accept both answers.
  3. Score on the agreed questions too. If the ranking changes when you drop the disputed questions, the labels decide the result, not the routers.
  4. Pair the test. Use McNemar's test on the same cases. A few points of difference on 56 cases may reflect sampling variation. A non-significant result does not prove equal accuracy.
  5. Find the ceiling. Simple single-choice decisions saturate fast. Multi-question decisions, like context shape, are where routers differ, so put more cases there.
  6. Report speed and cost by route. A router's price and delay can differ by two orders of magnitude while the study finds no significant accuracy difference.

Limits

  • Same-family labels. The case author and two routers are Claude models. The second labeller checks agreement across model families. Secondary scoring tests label sensitivity; it does not measure the size of a same-family bias.
  • Small sets. 56 cases from one author, 14 per decision type. Per-type intervals are very wide.
  • One repetition counts. The Jev figures are its first repetition. Production makes one call per decision.
  • Routes differ. Jev ran over direct HTTPS, the Claude routers through a CLI. Speed and cost compare deployments, not bare models.
  • Shared host. The Mac also ran other local jobs during the run, so wall times may carry contention, the CLI times most of all. We did not rerun the timing.
  • One label broke our rules. The artifacts label favours Jev under the primary scoring. The effect is shown above and no conclusion changes, but the label stays as we froze it.
  • Narrow label spread. On some context-shape questions one label holds almost all the cases: artifacts is "no" on 9 of 9, and memories is "yes" on 9 of 10. Some options have one case each: knowledge "none", scope "large" and transcript "summary". Those options are barely tested.
  • Blindness is our statement. File times are consistent with it, but no file can prove that the author opened no router answer before the freeze. The protocol file was written after the case drafts, and the time we first claimed for it was wrong.
  • Effort. Haiku ran at the CLI default (extended thinking); Sonnet at effort low, as production asks. Other settings can change both speed and accuracy.

Data and next steps

Every number above comes from the study page, with its method, sources and every case: Jev vs Claude routers on unseen decisions. The tuned-set study is Jev vs LLM routing, and the routing hub collects all routing results.

Try Agent, the product behind these benchmarks.

The data behind this post

  • Routing
  • Jev

Jev vs Claude routers on unseen decisions: a blind holdout

Jev 1.13, Claude Haiku 4.5 and Sonnet 5.5 on 56 routing decisions written blind and frozen first: exact rate with 95% intervals, tuned vs unseen.

82% (46/56)Jev 1.13 (TypeSafe): exact on unseen decisions · n = 56

6 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.