Jev vs Claude routers on unseen decisions: a blind holdout test
56 new routing decisions, written blind and frozen before any router call. Jev 1.13 vs Claude Haiku 4.5 vs Sonnet 5.5: accuracy, 95% intervals, cost, speed.
TL;DR
- Accuracy on unseen decisions: no winner. On 56 new routing decisions, Jev 1.13 answered 46 exactly right (82%, 95% interval 70% to 90%). Claude Haiku 4.5 answered 44 (79%, 66% to 87%) and Claude Sonnet 5.5 at effort low 49 (88%, 76% to 94%). The intervals overlap. The paired McNemar test finds no significant difference in any pair (p = 0.18 to 0.69).
- Every router scored lower on the holdout, and the data cannot say why. Jev went from 90% on the tuned case set to 82% (−8.1 points), Haiku from 89% to 79% (−10.5) and Sonnet from 94% to 88% (−6.4). These changes are calculations. Each pair of intervals overlaps. The data neither shows nor rules out a home advantage for Jev: the two sets differ in mix, and split by decision type the largest drop is Jev's on the three one-question types and Haiku's on context shape.
- A second labeller from another vendor checked our labels, and one of our labels broke our own rules. GPT-6.1 Sol labelled every case blind. Its label was accepted within our acceptable sets on 115 of 125 questions (92%, 95% interval 86% to 96%); first label only, it matched on 100 of 125 (80%, 72% to 86%). Our "artifacts" label broke our protocol and matched Jev's answers more often than the Claude routers' (Jev 9 of 9, 95% interval 70% to 100%; Haiku 4 of 9, 19% to 73%; Sonnet 7 of 9, 45% to 94%). With both answers accepted for it (calculation), the counts are Jev 46 of 56 (95% interval 70% to 90%), Haiku 45 of 56 (68% to 89%) and Sonnet 50 of 56 (79% to 95%). On the 115 agreed questions (calculation), exact cases are Jev 46 of 56 (70% to 90%), Haiku 46 of 56 (70% to 90%) and Sonnet 52 of 56 (83% to 97%). These intervals still overlap.
- Cost and measured time differ; accuracy has no clear ranking. Jev costs $0.0307 per 1,000 decisions. Haiku costs $7.13 and Sonnet $7.24 (list-price calculations, about 230 times as much). Jev took a median 139 ms (p50–p95 139–192 ms, n = 168) over direct HTTPS. Sonnet took 2.36 s (2.36–3.66 s, n = 56) and Haiku 9.44 s (9.44–25.4 s, n = 56) through the Claude Code CLI. These p50–p95 ranges are not confidence intervals. The routes differ, and the Mac also ran other local jobs, so the times may carry contention.
- Three of the four decision types are near the ceiling. Every router got 12 to 14 of 14 on them (95% intervals shown below). Context shape has up to 8 questions per case: Jev and Sonnet 8 of 14 (33% to 79%), Haiku 5 of 14 (16% to 61%).
Disclosure: I build Agent. Its platform uses Jev for some typed decisions in production. Jev is a third-party model from TypeSafe; we do not build or sell it.
Why we built a holdout
Our first routing study, Jev vs LLM routing, had a known weakness. Its 82 cases and question texts were revised in fix waves against Jev answers. That can favour Jev. We said so in that study's caveats. Then we built a holdout to test it.
The question for this study: does a router keep its accuracy on typed routing decisions that nobody tuned against its answers?
How we built the holdout
We followed a written protocol. Its file time is 21:35:45 UTC, after the case drafts and before the first counted call. Every later change has a dated line in its log, including the correction of the time we first claimed (21:30 UTC).
- Protocol before counted labelling. We wrote the question, the routers, the call caps, the scoring and the stop rules into a protocol file at 21:35:45 UTC (file time). That is after the case drafts and one uncounted probe of the second labeller, and before the first counted labelling call (21:35:49 UTC), the freeze and every router call. The first version said it was declared at 21:30 UTC, before any labelling call. File times did not support that, so we corrected it.
- 56 new cases, 14 per decision type. The four types are failure class, message intent, "is it a rule?" and context shape. Each case has the same shape and the same questions as the production cases. Each case is one that production really asks a router about. Only the questions that the platform's rules leave open are labelled: 125 questions in all.
- Blind to every router answer. The case author did not open any router answer, old or new, until the labels were frozen. This is our statement. File times are consistent with it, but no file can prove it.
- No paraphrases. A script compared each new case with the 166 existing cases. 17 new cases were too close in topic or wording, so we rewrote them before the freeze.
- A second, blind labeller. GPT-6.1 Sol in the Codex CLI labelled every case without our labels. It saw what a router sees: the case state and the questions.
- Freeze. We recorded the case file and its checksum at 21:37 UTC. The first router call, an uncounted probe, came 48 seconds later.
- Run. Jev 1.13 over direct HTTPS, with the production request body: 3 repetitions of 56 cases, 168 calls, no retries. Claude Haiku 4.5 (CLI default effort) and Claude Sonnet 5.5 (effort low) through the Claude Code CLI, with the production system text and answer schema: 56 calls each. One uncounted probe per route. Every call succeeded.
- Score. The repository's own decision-eval runner scored every router. A case is exact when every scored question is acceptable. Key accuracy scores each question on its own.
We publish the case ids, the labels and who answered each case exactly. We do not publish the case text.
Accuracy on unseen decisions
- Jev 1.13 (TypeSafe): 46 of 56, 82%, interval 70% to 90%.
- Claude Haiku 4.5: 44 of 56, 79%, interval 66% to 87%.
- Claude Sonnet 5.5 (effort low): 49 of 56, 88%, interval 76% to 94%.
Per question: Jev 113 of 125 (90%, 95% interval 84% to 94%), Haiku 102 of 125 (82%, 74% to 87%), Sonnet 115 of 125 (92%, 86% to 96%). These intervals overlap too. Questions within one case are not independent.
Sonnet has the highest point estimate. Is the gap real? For two routers on the same cases, McNemar's test is the right check. It looks only at the cases where the two routers disagree.
| Pair | Both right | Only first right | Only second right | Both wrong | Exact p |
|---|---|---|---|---|---|
| Jev vs Haiku 4.5 | 42 | 4 | 2 | 8 | 0.688 |
| Jev vs Sonnet 5.5 | 42 | 4 | 7 | 3 | 0.549 |
| Haiku 4.5 vs Sonnet 5.5 | 42 | 2 | 7 | 5 | 0.180 |
None of these p-values meets the 0.05 threshold. The three paired tests have no multiple-test correction. With 56 cases, this test cannot tell the three routers apart.
Tuned set vs unseen holdout
Here is the tuned-set chart from the first study. Its Jev point pools three repetitions (221 of 246 exact calls across 82 cases, 95% case-level interval 82% to 95%). The comparison table below uses Jev's first repetition, 74 of 82, as does the holdout (calculation). The chart uses case-level intervals because repetitions are not independent:
| Router | Tuned set (82 cases) | Unseen holdout (56 cases) | Change (calculation) |
|---|---|---|---|
| Jev 1.13 | 74, 90% (82% to 95%) | 46, 82% (70% to 90%) | −8.1 points |
| Claude Haiku 4.5 | 73, 89% (80% to 94%) | 44, 79% (66% to 87%) | −10.5 points |
| Claude Sonnet 5.5 | 77, 94% (87% to 97%) | 49, 88% (76% to 94%) | −6.4 points |
Each router's two intervals overlap, so the data shows no clear drop for any one router. All three routers scored lower on the holdout, by 6.4 to 10.5 points. The data neither shows nor rules out a home advantage for Jev. Two facts stop us from reading the aggregate drops either way:
- The drops are close and the intervals are wide. Jev's drop (8.1 points) is larger than Sonnet's (6.4) and smaller than Haiku's (10.5). The differences between these observed drops are calculations, from 1.7 to 4.1 points. They are not a tested ranking of drops.
- The two sets differ in mix. Context shape is 32 of the 82 tuned cases and 14 of the 56 holdout cases. The aggregate compares different blends of decision types, so it can hide movements in opposite directions by type.
So we split the comparison by decision group. The three one-question types are failure class, message intent and "is it a rule?". Context shape asks up to 8 questions per case. This is a calculation from the routing study's per-type records and this study:
| Router | Decision group | Tuned set | Unseen holdout | Change (calculation) |
|---|---|---|---|---|
| Jev 1.13 | One-question types | 50 of 50, 100.0% (92.9% to 100.0%) | 38 of 42, 90.5% (77.9% to 96.2%) | −9.5 points |
| Jev 1.13 | Context shape | 24 of 32, 75.0% (57.9% to 86.7%) | 8 of 14, 57.1% (32.6% to 78.6%) | −17.9 points |
| Claude Haiku 4.5 | One-question types | 49 of 50, 98.0% (89.5% to 99.6%) | 39 of 42, 92.9% (81.0% to 97.5%) | −5.1 points |
| Claude Haiku 4.5 | Context shape | 24 of 32, 75.0% (57.9% to 86.7%) | 5 of 14, 35.7% (16.3% to 61.2%) | −39.3 points |
| Claude Sonnet 5.5 | One-question types | 50 of 50, 100.0% (92.9% to 100.0%) | 41 of 42, 97.6% (87.7% to 99.6%) | −2.4 points |
| Claude Sonnet 5.5 | Context shape | 27 of 32, 84.4% (68.2% to 93.1%) | 8 of 14, 57.1% (32.6% to 78.6%) | −27.2 points |
On the one-question types, Jev has the largest observed drop (−9.5 points) and Sonnet the smallest (−2.4). On context shape, Haiku has the largest observed drop (−39.3) and Jev the smallest (−17.9). These calculations describe point estimates, not a ranking of drops. The split points in both directions. Every pair of intervals overlaps. Haiku's context-shape pair overlaps by only 3.3 points (the tuned interval starts at 57.9%, the holdout interval ends at 61.2%), but our rule holds: a gap counts only when the intervals do not overlap.
The author builds Agent, which uses Jev, and this study exists to answer the doubt about a home advantage. These data do not settle it. A larger holdout, with more context-shape cases, would test it better. Read any change between the two sets as a mix of case set and router. The data cannot separate them.
Where the routers miss
Exact answers per decision type, of 14 each (95% Wilson intervals):
| Decision type | Jev 1.13 | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| Failure class | 13 (69% to 99%) | 13 (69% to 99%) | 14 (78% to 100%) |
| Message intent | 12 (60% to 96%) | 13 (69% to 99%) | 14 (78% to 100%) |
| Is it a rule? | 13 (69% to 99%) | 13 (69% to 99%) | 13 (69% to 99%) |
| Context shape | 8 (33% to 79%) | 5 (16% to 61%) | 8 (33% to 79%) |
Three types are near the ceiling for every router. A perfect 14 of 14 still has an interval of 78% to 100%, so these sets cannot rank the routers. Context shape asks up to 8 questions per case: which turn this is, how much transcript, which knowledge, memories, examples, size and complexity. One wrong question makes the whole case wrong. That is where most misses are.
Jev missed 10 cases in its first repetition: one failure class case (holdout-09), two message intent cases (holdout-25 and holdout-28), one "is it a rule?" case (holdout-32) and six context shape cases. On every question Jev got wrong in those 10 cases, the second labeller's label was inside our acceptable set.
Jev was also stable. It gave the same answers on every question in all three repetitions for 53 of 56 cases (95%, interval 85% to 98%). Its exact count was 46, 46 and 47 of 56 in the three repetitions (95% intervals 70% to 90%, 70% to 90%, and 72% to 91%).
Did our labels favour any router?
The case author is a Claude model, and two of the three routers are Claude models. A same-family bias in the labels is possible. So we asked a model from another vendor to label every case blind.
GPT-6.1 Sol's label was accepted within our acceptable sets on 115 of 125 questions (92%, interval 86% to 96%). Our sets are wide on some questions: 24 of the 125 questions accept two or three labels. On our first label only, it matched on 100 of 125 (80%, interval 72% to 86%). These are question-level intervals; questions within a case are not independent. First-label agreement differs on five questions: knowledge 11 of 14 (95% interval 52% to 92%), scope and complexity each 10 of 14 (45% to 88%), turn 3 of 4 (30% to 95%) and transcript 2 of 5 (12% to 77%).
Agreement was 14 of 14 (95% interval 78% to 100%) for failure class, message intent and "is it a rule?", both ways. In context shape, two questions held all 10 disagreements inside our sets:
- Artifacts (does this turn need the work item's branches, files and links listed?): 1 of 9 (95% interval 2% to 44%). On a fresh work item we labelled "no", because a fresh item has no artifacts yet. GPT-6.1 Sol said "yes" on 8 of 9 (95% interval 57% to 98%). Cohen's kappa is 0 (calculation).
- Examples (does this turn need approved examples of finished work?): 11 of 13 (95% interval 58% to 96%), kappa 0.68 (calculation).
The artifacts question matters for fairness, and our label broke our own rules. The protocol says to leave out a question whose answer is a judgement call. The production case suites accept both "no" and "yes" for a fresh work item. We labelled only "no" on all 9 fresh-item cases. Jev matched our "no" label on 9 of 9 (95% interval 70% to 100%). Haiku matched on 4 of 9 (19% to 73%) and Sonnet on 7 of 9 (45% to 94%). So the scoring against our labels favours Jev on this question. We kept the labels frozen, as the protocol says, and report the effect two ways. The first scores only the questions where the second label is inside our set. The second accepts both answers for artifacts and scores the same answers again (calculation).
| Router | Exact, our labels (95% interval) | Exact, artifacts accepted both ways (calculation) | Exact, questions inside our sets only (calculation) |
|---|---|---|---|
| Jev 1.13 | 46 of 56 (70% to 90%) | 46 of 56 (70% to 90%) | 46 of 56 (70% to 90%) |
| Claude Haiku 4.5 | 44 of 56 (66% to 87%) | 45 of 56 (68% to 89%) | 46 of 56 (70% to 90%) |
| Claude Sonnet 5.5 | 49 of 56 (76% to 94%) | 50 of 56 (79% to 95%) | 52 of 56 (83% to 97%) |
With artifacts accepted both ways, context shape goes from 8, 5 and 8 of 14 to 8 (95% interval 33% to 79%), 6 (21% to 67%) and 9 (39% to 84%) of 14 for Jev, Haiku and Sonnet (calculation). On either rescoring, Sonnet's point estimate is higher still (50 or 52 of 56). The intervals still overlap, so neither rescoring establishes a ranking.
Speed and cost
- Jev 1.13: median 139 ms, 95th percentile 192 ms, over direct HTTPS from one Mac (168 calls).
- Claude Sonnet 5.5: median 2.36 s, 95th percentile 3.66 s (56 calls). CLI-reported API time was a median 1.49 s (p50–p95 1.49–2.38 s, n = 56). The median per-call wall-minus-API time was 0.861 s (calculation); subtraction of the two medians does not give that figure.
- Claude Haiku 4.5: median 9.44 s, 95th percentile 25.4 s (56 calls). Haiku ran with the CLI default extended thinking. This run cannot isolate the effect of thinking on time.
The p50–p95 ranges are not confidence intervals. The routes differ. A direct API call to Claude would skip the CLI start-up time. We did not measure that route here. The Mac also ran other local jobs during the run, so wall times, the CLI times most of all, may carry contention. We did not rerun the timing.
Per 1,000 decisions, at list prices (a calculation, not a bill; n = 168 Jev calls and 56 calls per Claude router). Mean token counts below are calculations too:
- Jev 1.13: $0.0307. About 730 input tokens per decision, and Jev output is free at its list price.
- Claude Haiku 4.5: $7.13. About 1,775 input and 1,071 output tokens per decision, including thinking. The receipt does not split Haiku output into answer and thinking tokens.
- Claude Sonnet 5.5: $7.24. About 1,705 input tokens, almost all written to the prompt cache, and 90.5 output tokens.
The CLI wrote Sonnet's prompt cache as 1-hour writes, which cost twice the input price. Our first routing study priced all cache writes at 1.25 times the input price. At that rate, Sonnet would cost $4.88 per 1,000 here. That study shows $4.996 for Sonnet on its own cases, where the CLI reported $7.32, so the two studies' Sonnet cost rows differ by convention and by case mix. At the primary convention, the Claude cost estimates are 232.6 to 236.3 times Jev's (calculation). At the tuned-set convention, they are 159.1 to 232.6 times Jev's (calculation).
The comparison pages line up every shared metric for each pair: Jev vs Claude Sonnet 5.5 and Jev vs Claude Haiku 4.5.
How to evaluate a router: what we recommend
- Keep a holdout set. Write it before you look at any router answer, freeze it, and never fix it after a run. Fix waves on one set leak the router's habits into the labels.
- Label twice, with two model families. Report agreement and kappa per question, and say whether agreement counts any label inside your acceptable set or your first label only. A question with low agreement, like artifacts here, is a label problem before it is a router problem. Follow your own rule for judgement-call questions: leave them out, or accept both answers.
- Score on the agreed questions too. If the ranking changes when you drop the disputed questions, the labels decide the result, not the routers.
- Pair the test. Use McNemar's test on the same cases. A few points of difference on 56 cases may reflect sampling variation. A non-significant result does not prove equal accuracy.
- Find the ceiling. Simple single-choice decisions saturate fast. Multi-question decisions, like context shape, are where routers differ, so put more cases there.
- Report speed and cost by route. A router's price and delay can differ by two orders of magnitude while the study finds no significant accuracy difference.
Limits
- Same-family labels. The case author and two routers are Claude models. The second labeller checks agreement across model families. Secondary scoring tests label sensitivity; it does not measure the size of a same-family bias.
- Small sets. 56 cases from one author, 14 per decision type. Per-type intervals are very wide.
- One repetition counts. The Jev figures are its first repetition. Production makes one call per decision.
- Routes differ. Jev ran over direct HTTPS, the Claude routers through a CLI. Speed and cost compare deployments, not bare models.
- Shared host. The Mac also ran other local jobs during the run, so wall times may carry contention, the CLI times most of all. We did not rerun the timing.
- One label broke our rules. The artifacts label favours Jev under the primary scoring. The effect is shown above and no conclusion changes, but the label stays as we froze it.
- Narrow label spread. On some context-shape questions one label holds almost all the cases: artifacts is "no" on 9 of 9, and memories is "yes" on 9 of 10. Some options have one case each: knowledge "none", scope "large" and transcript "summary". Those options are barely tested.
- Blindness is our statement. File times are consistent with it, but no file can prove that the author opened no router answer before the freeze. The protocol file was written after the case drafts, and the time we first claimed for it was wrong.
- Effort. Haiku ran at the CLI default (extended thinking); Sonnet at effort low, as production asks. Other settings can change both speed and accuracy.
Data and next steps
Every number above comes from the study page, with its method, sources and every case: Jev vs Claude routers on unseen decisions. The tuned-set study is Jev vs LLM routing, and the routing hub collects all routing results.
Try Agent, the product behind these benchmarks.