Jev vs Claude routers on unseen decisions: a blind holdout
Does a router keep its accuracy on typed routing decisions that nobody tuned against its answers?
Published · 6 charts · Download the data or a carousel
82%
The answer
The holdout does not establish a router ranking. The author wrote 56 new decisions blind to router answers. Labels froze before the first router call. Jev 1.13 (TypeSafe): 46 of 56 (82%, 95% interval 70% to 90%). Claude Haiku 4.5 · Claude Code: 44 of 56 (79%, 95% interval 66% to 87%). Claude Sonnet 5.5 (low) · Claude Code: 49 of 56 (88%, 95% interval 76% to 94%). The intervals overlap, so the holdout does not rank the routers. The exact McNemar tests find no significant difference (p from 0.18 to 0.688). These paired tests have no multiple-test correction. Tuned set → holdout exact rate (calculation; counts and intervals in the chart): Jev 1.13 90% → 82%, Haiku 4.5 89% → 79%, Sonnet 5.5 (low) 94% → 88%. All 3 observed rates were lower. Each router’s tuned and holdout intervals overlap. The data does not show a clear drop for any router. Decision groups (calculation; counts and intervals in the table). On the 3 one-question types the observed exact rate changed by Jev 1.13 −9.5, Haiku 4.5 −5.1, Sonnet 5.5 (low) −2.4 points. These are calculations, not a ranking of drops. On context shape the observed exact rate changed by Jev 1.13 −17.9, Haiku 4.5 −39.3, Sonnet 5.5 (low) −27.2 points. These are calculations, not a ranking of drops. Every tuned and holdout pair of intervals overlaps. The data neither shows nor rules out a home advantage for Jev. The 3 drops range from 6.4 to 10.5 points (calculation). The sets differ in mix: context shape is 32 of 82 tuned cases and 14 of 56 holdout cases. Secondary scoring keeps only questions where our set accepts the second label. Our set accepts 115 of 125 (92%, 95% interval 86% to 96%). Our first label agrees on 100 of 125 (80%, 95% interval 72% to 86%). Jev 1.13: 46 of 56 (82%, 95% interval 70% to 90%). Haiku 4.5: 46 of 56 (82%, 95% interval 70% to 90%). Sonnet 5.5 (low): 52 of 56 (93%, 95% interval 83% to 97%). The weakest label is “artifacts”. Our set accepts the second label on 1 of 9 (11%, 95% interval 2% to 43%). Set-aware kappa is 0 (calculation). This label broke our protocol. Matches to our label: Jev 1.13: 9 of 9 (100%, 95% interval 70% to 100%). Haiku 4.5: 4 of 9 (44%, 95% interval 19% to 73%). Sonnet 5.5 (low): 7 of 9 (78%, 95% interval 45% to 94%). See the caveats. Exact by decision type, of 14 each (95% intervals in the chart): Failure class 13 to 14; Message intent 12 to 14; Is it a rule? 13; Context shape 5 to 8. Failure class, Message intent and Is it a rule? are near the ceiling for every router (85% or more), so most differences come from context shape. Cost per 1,000 decisions (calculation): Jev 1.13 $0.0307, Haiku 4.5 $7.13, Sonnet 5.5 (low) $7.24. Median time per decision: Jev 1.13: 139 ms (p50–p95 139 ms to 192 ms, n = 168). Haiku 4.5: 9.44 s (p50–p95 9.44 s to 25.41 s, n = 56). Sonnet 5.5 (low): 2.36 s (p50–p95 2.36 s to 3.66 s, n = 56). These ranges are not confidence intervals. Jev used direct HTTPS; Claude used its CLI. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.
Key numbers
79% (44/56)
Claude Haiku 4.5 · Claude Code: exact on unseen decisions
95% CI 66%–87% · n = 56
88% (49/56)
Claude Sonnet 5.5 (low) · Claude Code: exact on unseen decisions
95% CI 76%–94% · n = 56
95% (53/56)
Jev 1.13 (TypeSafe): same answers on every key in 3 repetitions
95% CI 85%–98% · n = 56
46,
Jev 1.13 (TypeSafe): exact in each of 3 repetitions
46, 47 of 56 · n = 56
92% (115/125)
Second labeller (GPT-6.1 Sol): inside our acceptable label sets
95% CI 86%–96% · n = 125
80% (100/125)
Second labeller (GPT-6.1 Sol): equal to our first label
95% CI 72%–86% · n = 125
−8.1 points
Jev 1.13 (TypeSafe): holdout minus tuned-set exact rate
n = 56
−10.5 points
Claude Haiku 4.5 · Claude Code: holdout minus tuned-set exact rate
n = 56
−6.4 points
Claude Sonnet 5.5 (low) · Claude Code: holdout minus tuned-set exact rate
n = 56
$0.0307
Jev 1.13 (TypeSafe): cost per 1,000 unseen decisions
n = 168
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
Every interval overlaps every other: this chart does not order these rows.
| Item | Exact rate | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 82% | 70%–90% | 56 |
| Claude Haiku 4.5 · Claude Code | 79% | 66%–87% | 56 |
| Claude Sonnet 5.5 (low) · Claude Code | 88% | 76%–94% | 56 |
3 rows. Highest Claude Sonnet 5.5 (low) · Claude Code 88% (95% interval 76%–94%, n 56). Lowest Claude Haiku 4.5 · Claude Code 79% (95% interval 66%–87%, n 56). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 56 per row
Share of the 56 holdout cases where every scored question was acceptable
Whiskers are 95% Wilson intervals. Cases written blind to every router answer, then frozen. Jev: its first of three repetitions.
Every interval overlaps every other: this chart does not order these rows.
| Item | Key accuracy | 95% interval | n |
|---|---|---|---|
| Jev 1.13 (TypeSafe) | 90% | 84%–94% | 125 |
| Claude Haiku 4.5 · Claude Code | 82% | 74%–87% | 125 |
| Claude Sonnet 5.5 (low) · Claude Code | 92% | 86%–96% | 125 |
3 rows. Highest Claude Sonnet 5.5 (low) · Claude Code 92% (95% interval 86%–96%, n 125). Lowest Claude Haiku 4.5 · Claude Code 82% (95% interval 74%–87%, n 125). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 125 per row
Each open question a router was asked; an unanswered question counts as wrong
Whiskers are nominal 95% Wilson intervals. Context-shape cases ask up to 8 questions each, the other decision types one. Questions in one case are not independent; these intervals do not adjust for that grouping.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | Jev 1.13 (TypeSafe) | Claude Haiku 4.5 · Claude Code | Claude Sonnet 5.5 (low) · Claude Code | 95% interval | n |
|---|---|---|---|---|---|
| Failure class | 93% | 93% | 100% | Jev 1.13 (TypeSafe): 69%–99%; Claude Haiku 4.5 · Claude Code: 69%–99%; Claude Sonnet 5.5 (low) · Claude Code: 78%–100% | 14 |
| Message intent | 86% | 93% | 100% | Jev 1.13 (TypeSafe): 60%–96%; Claude Haiku 4.5 · Claude Code: 69%–99%; Claude Sonnet 5.5 (low) · Claude Code: 78%–100% | 14 |
| Is it a rule? | 93% | 93% | 93% | Jev 1.13 (TypeSafe): 69%–99%; Claude Haiku 4.5 · Claude Code: 69%–99%; Claude Sonnet 5.5 (low) · Claude Code: 69%–99% | 14 |
| Context shape | 57% | 36% | 57% | Jev 1.13 (TypeSafe): 33%–79%; Claude Haiku 4.5 · Claude Code: 16%–61%; Claude Sonnet 5.5 (low) · Claude Code: 33%–79% | 14 |
4 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 (low) · Claude Code. Jev 1.13 (TypeSafe): highest Failure class 93% (95% interval 69%–99%, n 14). Lowest Context shape 57% (95% interval 33%–79%, n 14). All intervals overlap. Claude Haiku 4.5 · Claude Code: highest Failure class 93% (95% interval 69%–99%, n 14). Lowest Context shape 36% (95% interval 16%–61%, n 14). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 14 per row
14 cases per decision type
Whiskers are 95% Wilson intervals. With 14 cases a perfect score has an interval of 78% to 100%, so a decision type where every router scores 14 of 14 is at its ceiling and cannot rank them.
- Tuned set (routing-jev-vs-llm)
- Unseen holdout (square)
Gap labels, Unseen holdout vs Tuned set (routing-jev-vs-llm): Unseen holdout is x percentage points higher (+) or lower (−) than Tuned set (routing-jev-vs-llm), calculated from the two values shown; lines are the 95% Wilson interval.
| Item | Tuned set (routing-jev-vs-llm) | Unseen holdout | 95% interval | n |
|---|---|---|---|---|
| Jev 1.13 (TypeSafe) | 90% | 82% | Tuned set (routing-jev-vs-llm): 82%–95%; Unseen holdout: 70%–90% | 82 |
| Claude Haiku 4.5 · Claude Code | 89% | 79% | Tuned set (routing-jev-vs-llm): 80%–94%; Unseen holdout: 66%–87% | 82 |
| Claude Sonnet 5.5 (low) · Claude Code | 94% | 88% | Tuned set (routing-jev-vs-llm): 87%–97%; Unseen holdout: 76%–94% | 82 |
3 rows, 2 series: Tuned set (routing-jev-vs-llm), Unseen holdout. Tuned set (routing-jev-vs-llm): highest Claude Sonnet 5.5 (low) · Claude Code 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 · Claude Code 89% (95% interval 80%–94%, n 82). All intervals overlap. Unseen holdout: highest Claude Sonnet 5.5 (low) · Claude Code 88% (95% interval 76%–94%, n 56). Lowest Claude Haiku 4.5 · Claude Code 79% (95% interval 66%–87%, n 56). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 56–82 per row
Tuned set: the routing study’s 82 cases, revised against Jev answers. Holdout: 56 new cases, frozen before any router call
Whiskers are 95% Wilson intervals. The two case sets differ in mix and size, so a gap mixes a change of case set with any change in the router, and the data cannot separate them. A gap counts only when the two intervals do not overlap.
Sources: Routing on unseen holdout decisions, Routing runs: Jev router vs LLM routing
- Wall time
- Model time (API, CLI-reported)
Seconds · log scale: each gridline is 10 times the one before
| Item | Wall time | Model time (API, CLI-reported) | Median to p95 | n |
|---|---|---|---|---|
| Jev 1.13 (TypeSafe) | 0.1 s | — | Wall time: 0.1 s–0.2 s | 168 |
| Claude Haiku 4.5 · Claude Code | 9.4 s | 7.5 s | Wall time: 9.4 s–25.4 s; Model time (API, CLI-reported): 7.5 s–23.9 s | 56 |
| Claude Sonnet 5.5 (low) · Claude Code | 2.4 s | 1.5 s | Wall time: 2.4 s–3.7 s; Model time (API, CLI-reported): 1.5 s–2.4 s | 56 |
3 rows, 2 series: Wall time, Model time (API, CLI-reported). Wall time: slowest Claude Haiku 4.5 · Claude Code 9.4 s (median to p95 9.4 s–25.4 s, n 56). Fastest Jev 1.13 (TypeSafe) 0.1 s (median to p95 0.1 s–0.2 s, n 168). Not all run ranges overlap. Model time (API, CLI-reported): slowest Claude Haiku 4.5 · Claude Code 7.5 s (median to p95 7.5 s–23.9 s, n 56). Fastest Claude Sonnet 5.5 (low) · Claude Code 1.5 s (median to p95 1.5 s–2.4 s, n 56). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n 56–168 per row
Median, whisker to the 95th percentile
The routes differ. Jev: one HTTPS call to the vendor API from one Mac, all three repetitions. Claude: the Claude Code CLI from process start to exit, which adds start-up time a direct API call would not. The whisker runs from the median to the 95th percentile; it is not a confidence interval. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.
Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the highlighted row): a ratio of list-price calculations, not a measurement.
| Item | Cost | n |
|---|---|---|
| Jev 1.13 (TypeSafe) | $0.031 | 168 |
| Claude Haiku 4.5 · Claude Code | $7.13 | 56 |
| Claude Sonnet 5.5 (low) · Claude Code | $7.24 | 56 |
List-price calculation, not a run. 3 rows. Highest Claude Sonnet 5.5 (low) · Claude Code $7.24 (n 56). Lowest Jev 1.13 (TypeSafe) $0.031 (n 168).
Notesn 56–168 per row
Reported tokens per decision × list price
Calculation, not a bill. Jev: reported input tokens × $0.042 per million, output free. Claude: CLI-reported tokens × list price; the CLI wrote its prompt cache as 1-hour writes, priced at 2× input, and adds its own system prompt and tool-schema tokens. The Claude routers ran on a subscription. At the tuned-set study’s convention (every cache write at 1.25× input), Sonnet 5.5 (low) would be $4.877 here. That study shows $4.996 for it on its own cases, where the CLI reported $7.324, so its cost row and this one differ by convention and by case mix.
Sources: Routing on unseen holdout decisions, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models)
Tables
Same unseen cases, two routers: exact McNemar test
| Pair | Cases | Both right | Only first right | Only second right | Both wrong | Exact McNemar p |
|---|---|---|---|---|---|---|
| Jev 1.13 (TypeSafe) vs Claude Haiku 4.5 · Claude Code | 56 | 42 | 4 | 2 | 8 | 0.688 |
| Jev 1.13 (TypeSafe) vs Claude Sonnet 5.5 (low) · Claude Code | 56 | 42 | 4 | 7 | 3 | 0.549 |
| Claude Haiku 4.5 · Claude Code vs Claude Sonnet 5.5 (low) · Claude Code | 56 | 42 | 2 | 7 | 5 | 0.18 |
Label agreement: the author vs a second, blind labeller (GPT-6.1 Sol)
| Decision | Question | Labelled | Second label inside our acceptable set (95% Wilson interval) | Second label equals our first label (95% Wilson interval) | Set-aware kappa (calculation; adaptive author label) | Cohen’s kappa, fixed first label (calculation) |
|---|---|---|---|---|---|---|
| Failure class | failure | 14 | 14/14 (78.5% to 100.0%) | 14/14 (78.5% to 100.0%) | 1 | 1 |
| Message intent | intent | 14 | 14/14 (78.5% to 100.0%) | 14/14 (78.5% to 100.0%) | 1 | 1 |
| Is it a rule? | kind | 14 | 14/14 (78.5% to 100.0%) | 14/14 (78.5% to 100.0%) | 1 | 1 |
| Context shape | artifacts | 9 | 1/9 (2.0% to 43.5%) | 1/9 (2.0% to 43.5%) | 0 | 0 |
| Context shape | knowledge | 14 | 14/14 (78.5% to 100.0%) | 11/14 (52.4% to 92.4%) | 1 | 0.57 |
| Context shape | memories | 10 | 10/10 (72.2% to 100.0%) | 10/10 (72.2% to 100.0%) | 1 | 1 |
| Context shape | examples | 13 | 11/13 (57.8% to 95.7%) | 11/13 (57.8% to 95.7%) | 0.68 | 0.68 |
| Context shape | scope | 14 | 14/14 (78.5% to 100.0%) | 10/14 (45.4% to 88.3%) | 1 | 0.53 |
| Context shape | complexity | 14 | 14/14 (78.5% to 100.0%) | 10/14 (45.4% to 88.3%) | 1 | 0.46 |
| Context shape | turn | 4 | 4/4 (51.0% to 100.0%) | 3/4 (30.1% to 95.4%) | 1 | 0.56 |
| Context shape | transcript | 5 | 5/5 (56.6% to 100.0%) | 2/5 (11.8% to 76.9%) | 1 | 0.25 |
Tuned set vs unseen holdout, by decision group: exact rate with 95% intervals
| Router | Decision group | Tuned set | Unseen holdout | Change (points, calculation) | Intervals |
|---|---|---|---|---|---|
| Jev 1.13 (TypeSafe) | The 3 one-question types (failure class, message intent and is it a rule?) | 50/50, 100.0% (92.9% to 100.0%) | 38/42, 90.5% (77.9% to 96.2%) | −9.5 | overlap |
| Jev 1.13 (TypeSafe) | Context shape (several questions per case) | 24/32, 75.0% (57.9% to 86.7%) | 8/14, 57.1% (32.6% to 78.6%) | −17.9 | overlap |
| Claude Haiku 4.5 · Claude Code | The 3 one-question types (failure class, message intent and is it a rule?) | 49/50, 98.0% (89.5% to 99.6%) | 39/42, 92.9% (81.0% to 97.5%) | −5.1 | overlap |
| Claude Haiku 4.5 · Claude Code | Context shape (several questions per case) | 24/32, 75.0% (57.9% to 86.7%) | 5/14, 35.7% (16.3% to 61.2%) | −39.3 | overlap |
| Claude Sonnet 5.5 (low) · Claude Code | The 3 one-question types (failure class, message intent and is it a rule?) | 50/50, 100.0% (92.9% to 100.0%) | 41/42, 97.6% (87.7% to 99.6%) | −2.4 | overlap |
| Claude Sonnet 5.5 (low) · Claude Code | Context shape (several questions per case) | 27/32, 84.4% (68.2% to 93.1%) | 8/14, 57.1% (32.6% to 78.6%) | −27.2 | overlap |
Every holdout case: labels and who answered it exactly
| Case | Decision | Acceptable answers | Keys both labellers accept | Jev 1.13 (TypeSafe) | Claude Haiku 4.5 · Claude Code | Claude Sonnet 5.5 (low) · Claude Code | Jev exact in 3 repetitions |
|---|---|---|---|---|---|---|---|
| holdout-01 | Failure class | failure: transient | 1/1 | exact | exact | exact | 3/3 |
| holdout-02 | Failure class | failure: transient | 1/1 | exact | exact | exact | 3/3 |
| holdout-03 | Failure class | failure: transient | 1/1 | exact | exact | exact | 3/3 |
| holdout-04 | Failure class | failure: impossible-here | 1/1 | exact | exact | exact | 3/3 |
| holdout-05 | Failure class | failure: impossible-here | 1/1 | exact | exact | exact | 3/3 |
| holdout-06 | Failure class | failure: impossible-here | 1/1 | exact | exact | exact | 3/3 |
| holdout-07 | Failure class | failure: impossible-here | 1/1 | exact | exact | exact | 3/3 |
| holdout-08 | Failure class | failure: flaky | 1/1 | exact | exact | exact | 3/3 |
| holdout-09 | Failure class | failure: flaky | 1/1 | 0/1 keys | 0/1 keys | exact | 1/3 |
| holdout-10 | Failure class | failure: flaky | 1/1 | exact | exact | exact | 3/3 |
| holdout-11 | Failure class | failure: real | 1/1 | exact | exact | exact | 3/3 |
| holdout-12 | Failure class | failure: real | 1/1 | exact | exact | exact | 3/3 |
| holdout-13 | Failure class | failure: real | 1/1 | exact | exact | exact | 3/3 |
| holdout-14 | Failure class | failure: real | 1/1 | exact | exact | exact | 3/3 |
| holdout-15 | Message intent | intent: request | 1/1 | exact | exact | exact | 3/3 |
| holdout-16 | Message intent | intent: request | 1/1 | exact | exact | exact | 3/3 |
| holdout-17 | Message intent | intent: request | 1/1 | exact | exact | exact | 3/3 |
| holdout-18 | Message intent | intent: followup | 1/1 | exact | exact | exact | 3/3 |
| holdout-19 | Message intent | intent: followup | 1/1 | exact | exact | exact | 3/3 |
| holdout-20 | Message intent | intent: followup | 1/1 | exact | exact | exact | 3/3 |
| holdout-21 | Message intent | intent: status | 1/1 | exact | exact | exact | 3/3 |
| holdout-22 | Message intent | intent: status | 1/1 | exact | exact | exact | 3/3 |
| holdout-23 | Message intent | intent: status | 1/1 | exact | exact | exact | 3/3 |
| holdout-24 | Message intent | intent: decision | 1/1 | exact | exact | exact | 3/3 |
| holdout-25 | Message intent | intent: decision | 1/1 | 0/1 keys | exact | exact | 0/3 |
| holdout-26 | Message intent | intent: control | 1/1 | exact | exact | exact | 3/3 |
| holdout-27 | Message intent | intent: control | 1/1 | exact | exact | exact | 3/3 |
| holdout-28 | Message intent | intent: control | 1/1 | 0/1 keys | 0/1 keys | exact | 0/3 |
| holdout-29 | Is it a rule? | kind: obligation | 1/1 | exact | exact | exact | 3/3 |
| holdout-30 | Is it a rule? | kind: obligation | 1/1 | exact | exact | exact | 3/3 |
| holdout-31 | Is it a rule? | kind: obligation | 1/1 | exact | exact | exact | 3/3 |
| holdout-32 | Is it a rule? | kind: obligation | 1/1 | 0/1 keys | 0/1 keys | 0/1 keys | 0/3 |
| holdout-33 | Is it a rule? | kind: prohibition | 1/1 | exact | exact | exact | 3/3 |
| holdout-34 | Is it a rule? | kind: prohibition or obligation | 1/1 | exact | exact | exact | 3/3 |
| holdout-35 | Is it a rule? | kind: prohibition | 1/1 | exact | exact | exact | 3/3 |
| holdout-36 | Is it a rule? | kind: preference | 1/1 | exact | exact | exact | 3/3 |
| holdout-37 | Is it a rule? | kind: preference | 1/1 | exact | exact | exact | 3/3 |
| holdout-38 | Is it a rule? | kind: preference | 1/1 | exact | exact | exact | 3/3 |
| holdout-39 | Is it a rule? | kind: not-a-rule | 1/1 | exact | exact | exact | 3/3 |
| holdout-40 | Is it a rule? | kind: not-a-rule | 1/1 | exact | exact | exact | 3/3 |
| holdout-41 | Is it a rule? | kind: not-a-rule | 1/1 | exact | exact | exact | 3/3 |
| holdout-42 | Is it a rule? | kind: not-a-rule | 1/1 | exact | exact | exact | 3/3 |
| holdout-43 | Context shape | artifacts: no; knowledge: search; memories: yes; examples: yes; scope: small or medium; complexity: simple or moderate | 5/6 | exact | 5/6 keys | 5/6 keys | 3/3 |
| holdout-44 | Context shape | artifacts: no; knowledge: search; memories: yes; examples: no; scope: medium; complexity: moderate or complex | 5/6 | exact | 4/6 keys | exact | 3/3 |
| holdout-45 | Context shape | turn: follow-up; transcript: recent or full; knowledge: search; scope: small or medium; complexity: moderate | 5/5 | exact | exact | exact | 3/3 |
| holdout-46 | Context shape | turn: clarification; transcript: recent or full; knowledge: search; examples: no; scope: small or medium; complexity: moderate | 6/6 | 5/6 keys | 5/6 keys | exact | 0/3 |
| holdout-47 | Context shape | transcript: summary or recent; knowledge: search; memories: yes; examples: no; scope: medium or large; complexity: moderate or complex | 5/6 | exact | exact | 5/6 keys | 3/3 |
| holdout-48 | Context shape | artifacts: no; knowledge: none; memories: no; examples: no; scope: small; complexity: simple or moderate | 5/6 | 5/6 keys | 4/6 keys | exact | 0/3 |
| holdout-49 | Context shape | artifacts: no; knowledge: search; memories: yes; examples: yes; scope: small or medium; complexity: moderate | 5/6 | 5/6 keys | exact | exact | 0/3 |
| holdout-50 | Context shape | artifacts: no; knowledge: search; memories: yes; examples: no; scope: medium; complexity: complex | 5/6 | exact | exact | 5/6 keys | 3/3 |
| holdout-51 | Context shape | transcript: recent or full; knowledge: facts or search; memories: yes; examples: no; scope: medium; complexity: moderate | 6/6 | 3/6 keys | 2/6 keys | exact | 0/3 |
| holdout-52 | Context shape | artifacts: no; knowledge: facts or search; memories: yes; examples: yes; scope: medium or large; complexity: moderate | 6/6 | 5/6 keys | 2/6 keys | 4/6 keys | 0/3 |
| holdout-53 | Context shape | turn: new-task or follow-up; artifacts: no; knowledge: facts or search; memories: yes; examples: no; scope: small or medium; complexity: moderate | 6/7 | exact | 5/7 keys | exact | 3/3 |
| holdout-54 | Context shape | turn: follow-up; transcript: recent or full; knowledge: facts or search; examples: yes; scope: small or medium; complexity: moderate | 6/6 | exact | exact | exact | 3/3 |
| holdout-55 | Context shape | artifacts: no; knowledge: search; memories: yes; examples: no; scope: large; complexity: complex | 4/6 | exact | 5/6 keys | 5/6 keys | 3/3 |
| holdout-56 | Context shape | artifacts: no; knowledge: facts; examples: no; scope: small; complexity: simple or moderate | 4/5 | 4/5 keys | 2/5 keys | 2/5 keys | 0/3 |
Every router on the holdout
| Router | Route | Counted calls | Exact | Key accuracy | Exact, agreed keys only | Mean tokens per decision (calculation; input incl. cache / output) | USD per 1,000 (calculation) |
|---|---|---|---|---|---|---|---|
| Jev 1.13 (TypeSafe) | direct HTTPS to the vendor API with the production request body | 168 | 46 of 56 (82%, 95% interval 70% to 90%) | 113 of 125 (90%, 95% interval 84% to 94%) | 46 of 56 (82%, 95% interval 70% to 90%) | 730 / 119.3 | $0.031 |
| Claude Haiku 4.5 · Claude Code | Claude Code CLI, one call per decision, CLI default effort (extended thinking) | 56 | 44 of 56 (79%, 95% interval 66% to 87%) | 102 of 125 (82%, 95% interval 74% to 87%) | 46 of 56 (82%, 95% interval 70% to 90%) | 1,775 / 1,071 | $7.13 |
| Claude Sonnet 5.5 (low) · Claude Code | Claude Code CLI, one call per decision, effort low | 56 | 49 of 56 (88%, 95% interval 76% to 94%) | 115 of 125 (92%, 95% interval 86% to 96%) | 52 of 56 (93%, 95% interval 83% to 97%) | 1,705 / 90.5 | $7.24 |
Method
- Cases: 56 new cases, 14 per decision type, in the production case shape. The types are failure class, message intent, is-it-a-rule and context shape. Production asks a router about each case. We score only questions the rule leaves open: 125 labelled questions.
- Blindness: the author says they opened no router answer before the label freeze. File times support this order but cannot prove blindness. A script compared every case with existing cases. The author rewrote 17 close cases before the freeze.
- File times: the protocol file dates from 21:35:45 UTC on 2026-10-06. Case drafts and one uncounted second-labeller probe came first. The first counted labelling call came at 21:35:49 UTC. The freeze came at 21:37:02 UTC. The frozen file and its checksum preceded every router call.
- Second labeller: GPT-6.1 Sol through the Codex CLI, 4 counted calls. It labelled every case without the author labels. Our set accepts its label on 115 of 125 (92%, 95% interval 86% to 96%). It matches our first label on 100 of 125 (80%, 95% interval 72% to 86%). Of these questions, 24 accept two or three labels. The lowest agreement was on “artifacts”; see the caveats. Secondary scoring keeps only questions where our set accepts its label.
- Routers: Jev 1.13 used direct HTTPS with the production request body. It made 168 calls in 3 repetitions, with no retries. Claude Haiku 4.5 used CLI default effort; Claude Sonnet 5.5 used effort low. Both used the Claude Code CLI with the production system text and schema. Each made 56 counted calls, one per decision. Each route had one uncounted probe. Total calls stayed within the caps: Jev 169/200, Claude 114/120, labeller 5/5. No arm stopped early or lost cases.
- Scoring: the repository’s decision-eval runner. Exact means every scored question in a case was acceptable. Key accuracy counts each question. We use Wilson 95% intervals and exact McNemar tests on paired cases. The paired tests have no multiple-test correction. Errors count as wrong.
- Cost: reported tokens × list price per 1,000 decisions, a calculation.
Caveats
- The case author and two of the three routers are Claude models, so a same-family label bias is possible. A second labeller (GPT-6.1 Sol) labelled every case blind; the secondary scoring keeps only the keys where its label is inside our acceptable set.
- 56 cases (14 per decision type) from one author: per-type intervals are very wide.
- The routes differ: Jev ran over direct HTTPS, the Claude routers through the Claude Code CLI, which adds start-up time and tokens.
- Jev figures are its first repetition; repetitions 2 and 3 and the stability across all three are reported beside it.
- The artifacts label violates the protocol rule to leave judgement calls unscored. The author accepted only no on all 9 labelled cases. The second labeller chose no on 1 of 9. Frozen labels stay unchanged; sensitivity calculations are not new runs.
- The protocol file was created at 21:35:45.935 UTC on 2026-10-06. It followed the case drafts and an uncounted labeller probe. It preceded the first counted labelling call at 21:35:49.3 UTC, the label freeze at 21:37:02.582 UTC and every router call. It was not written before every call.
- The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.
- The three one-question case sets are near a ceiling: every router scored 12 to 14 of 14 on each. The 95% Wilson interval for 14 of 14 is 78.5% to 100%. These sets provide little room to separate routers.
- Per-question Wilson intervals treat questions as independent. Several questions share each context-shape case, so those intervals can understate uncertainty. Label-agreement intervals have the same limit.
- Set-aware kappa uses the second label as the author label when the acceptable set contains it. This adaptive calculation can inflate agreement. Only kappaStrict compares two fixed labels.
- The tuned set was revised against Jev answers. Its size and decision mix differ from the holdout. Overlapping intervals neither show nor rule out a home advantage.
- The three paired McNemar tests have no multiple-test correction. Their p values do not prove equal accuracy.
- Failure class, Message intent and Is it a rule?: every router answered at least 85% of the 14 cases exactly. These case sets are near the ceiling and barely separate the routers; context shape carries most of the differences.
- Agreement with the second labeller counts a label inside our acceptable set: 115 of 125 questions (92%). 24 of the 125 questions accept two or three labels, so agreement with our first label alone is lower: 100 of 125 (80%).
- The “artifacts” label broke our own rules. The second labeller’s answer was inside our set on 1 of 9 cases (kappa 0). We labelled “no” on every one of them. The protocol leaves out a question whose answer is a judgement call, and the production case suites accept both answers for a fresh work item. Cases where the router’s answer equalled our label: Jev 1.13 9 of 9, Haiku 4.5 4 of 9, Sonnet 5.5 (low) 7 of 9. The scoring against our labels therefore favours Jev on this question. The labels stayed frozen. With both answers accepted for “artifacts” and the same answers scored again (a calculation), exact answers are Jev 1.13 46 of 56 (82%, 95% interval 70% to 90%); Haiku 4.5 45 of 56 (80%, 95% interval 68% to 89%); Sonnet 5.5 (low) 50 of 56 (89%, 95% interval 79% to 95%). Context shape: Jev 1.13 8 of 14 (57%, 95% interval 33% to 79%); Haiku 4.5 6 of 14 (43%, 95% interval 21% to 67%); Sonnet 5.5 (low) 9 of 14 (64%, 95% interval 39% to 84%); the overall intervals still overlap.
- Label spread is narrow on some questions. One label holds at least 90% of the cases for artifacts (“no”, 9 of 9) and memories (“yes”, 9 of 10). One option has a single case for knowledge (“none”), scope (“large”) and transcript (“summary”), so those options are barely tested.
- The tuned and holdout sets differ in size and mix, so a change between them cannot be assigned to the router or to the case set. The data neither shows nor rules out a home advantage.
Sources
Routing on unseen holdout decisions
Frozen unseen routing decisions; every router call retained. Costs are calculations from recorded tokens and list prices.
Routing runs: Jev router vs LLM routing
Routing decisions recorded per case and arm.
Repricing calculation
Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.
Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free.
Anthropic list prices (Claude models)
Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “Jev vs Claude routers on unseen decisions: a blind holdout”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/routing-holdout.
Models and comparisons in this study
Write-ups on this study
Jev vs Claude routers on unseen decisions: a blind holdout test
56 new routing decisions, written blind and frozen before any router call. Jev 1.13 vs Claude Haiku 4.5 vs Sonnet 5.5: accuracy, 95% intervals, cost, speed.
More studies
All benchmarksJev vs Claude as a router: accuracy and cost
Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.
Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
96 calls. Strict passes, schema vs instructions: Haiku 18/24 vs 0/24, Sonnet 12/12 vs 12/12, GPT-6.1 Sol 12/12 vs 12/12. With intervals.
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.