• Routing
  • Jev
  • Claude Haiku
  • Claude Sonnet
  • Model Routing
  • Holdout
  • Benchmark Method

Jev vs Claude routers on unseen decisions: a blind holdout

Does a router keep its accuracy on typed routing decisions that nobody tuned against its answers?

Published · 6 charts · Download the data or a carousel

82%

95% CI 70%–90% · n = 56

46/56 · Jev 1.13 (TypeSafe): exact on unseen decisions

The answer

The holdout does not establish a router ranking. The author wrote 56 new decisions blind to router answers. Labels froze before the first router call. Jev 1.13 (TypeSafe): 46 of 56 (82%, 95% interval 70% to 90%). Claude Haiku 4.5 · Claude Code: 44 of 56 (79%, 95% interval 66% to 87%). Claude Sonnet 5.5 (low) · Claude Code: 49 of 56 (88%, 95% interval 76% to 94%). The intervals overlap, so the holdout does not rank the routers. The exact McNemar tests find no significant difference (p from 0.18 to 0.688). These paired tests have no multiple-test correction. Tuned set → holdout exact rate (calculation; counts and intervals in the chart): Jev 1.13 90% → 82%, Haiku 4.5 89% → 79%, Sonnet 5.5 (low) 94% → 88%. All 3 observed rates were lower. Each router’s tuned and holdout intervals overlap. The data does not show a clear drop for any router. Decision groups (calculation; counts and intervals in the table). On the 3 one-question types the observed exact rate changed by Jev 1.13 −9.5, Haiku 4.5 −5.1, Sonnet 5.5 (low) −2.4 points. These are calculations, not a ranking of drops. On context shape the observed exact rate changed by Jev 1.13 −17.9, Haiku 4.5 −39.3, Sonnet 5.5 (low) −27.2 points. These are calculations, not a ranking of drops. Every tuned and holdout pair of intervals overlaps. The data neither shows nor rules out a home advantage for Jev. The 3 drops range from 6.4 to 10.5 points (calculation). The sets differ in mix: context shape is 32 of 82 tuned cases and 14 of 56 holdout cases. Secondary scoring keeps only questions where our set accepts the second label. Our set accepts 115 of 125 (92%, 95% interval 86% to 96%). Our first label agrees on 100 of 125 (80%, 95% interval 72% to 86%). Jev 1.13: 46 of 56 (82%, 95% interval 70% to 90%). Haiku 4.5: 46 of 56 (82%, 95% interval 70% to 90%). Sonnet 5.5 (low): 52 of 56 (93%, 95% interval 83% to 97%). The weakest label is “artifacts”. Our set accepts the second label on 1 of 9 (11%, 95% interval 2% to 43%). Set-aware kappa is 0 (calculation). This label broke our protocol. Matches to our label: Jev 1.13: 9 of 9 (100%, 95% interval 70% to 100%). Haiku 4.5: 4 of 9 (44%, 95% interval 19% to 73%). Sonnet 5.5 (low): 7 of 9 (78%, 95% interval 45% to 94%). See the caveats. Exact by decision type, of 14 each (95% intervals in the chart): Failure class 13 to 14; Message intent 12 to 14; Is it a rule? 13; Context shape 5 to 8. Failure class, Message intent and Is it a rule? are near the ceiling for every router (85% or more), so most differences come from context shape. Cost per 1,000 decisions (calculation): Jev 1.13 $0.0307, Haiku 4.5 $7.13, Sonnet 5.5 (low) $7.24. Median time per decision: Jev 1.13: 139 ms (p50–p95 139 ms to 192 ms, n = 168). Haiku 4.5: 9.44 s (p50–p95 9.44 s to 25.41 s, n = 56). Sonnet 5.5 (low): 2.36 s (p50–p95 2.36 s to 3.66 s, n = 56). These ranges are not confidence intervals. Jev used direct HTTPS; Claude used its CLI. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.

Key numbers

79% (44/56)

Claude Haiku 4.5 · Claude Code: exact on unseen decisions

95% CI 66%–87% · n = 56

88% (49/56)

Claude Sonnet 5.5 (low) · Claude Code: exact on unseen decisions

95% CI 76%–94% · n = 56

95% (53/56)

Jev 1.13 (TypeSafe): same answers on every key in 3 repetitions

95% CI 85%–98% · n = 56

46,

Jev 1.13 (TypeSafe): exact in each of 3 repetitions

46, 47 of 56 · n = 56

92% (115/125)

Second labeller (GPT-6.1 Sol): inside our acceptable label sets

95% CI 86%–96% · n = 125

80% (100/125)

Second labeller (GPT-6.1 Sol): equal to our first label

95% CI 72%–86% · n = 125

−8.1 points

Jev 1.13 (TypeSafe): holdout minus tuned-set exact rate

n = 56

−10.5 points

Claude Haiku 4.5 · Claude Code: holdout minus tuned-set exact rate

n = 56

−6.4 points

Claude Sonnet 5.5 (low) · Claude Code: holdout minus tuned-set exact rate

n = 56

$0.0307

Jev 1.13 (TypeSafe): cost per 1,000 unseen decisions

n = 168

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 (low) · Claude Code 88% (95% interval 76%–94%, n 56). Lowest Claude Haiku 4.5 · Claude Code 79% (95% interval 66%–87%, n 56). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 56 per row

Share of the 56 holdout cases where every scored question was acceptable

Whiskers are 95% Wilson intervals. Cases written blind to every router answer, then frozen. Jev: its first of three repetitions.

Source: Routing on unseen holdout decisions

Share card (PNG)
Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 (low) · Claude Code 92% (95% interval 86%–96%, n 125). Lowest Claude Haiku 4.5 · Claude Code 82% (95% interval 74%–87%, n 125). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 125 per row

Each open question a router was asked; an unanswered question counts as wrong

Whiskers are nominal 95% Wilson intervals. Context-shape cases ask up to 8 questions each, the other decision types one. Questions in one case are not independent; these intervals do not adjust for that grouping.

Source: Routing on unseen holdout decisions

Share card (PNG)

Jev 1.13 (TypeSafe)

Failure class
Message intent
Is it a rule?
Context shape

Claude Haiku 4.5 · Claude Code

Failure class
Message intent
Is it a rule?
Context shape

Claude Sonnet 5.5 (low) · Claude Code

Failure class
Message intent
Is it a rule?
Context shape

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

4 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 (low) · Claude Code. Jev 1.13 (TypeSafe): highest Failure class 93% (95% interval 69%–99%, n 14). Lowest Context shape 57% (95% interval 33%–79%, n 14). All intervals overlap. Claude Haiku 4.5 · Claude Code: highest Failure class 93% (95% interval 69%–99%, n 14). Lowest Context shape 36% (95% interval 16%–61%, n 14). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 14 per row

14 cases per decision type

Whiskers are 95% Wilson intervals. With 14 cases a perfect score has an interval of 78% to 100%, so a decision type where every router scores 14 of 14 is at its ceiling and cannot rank them.

Source: Routing on unseen holdout decisions

Share card (PNG)
  • Tuned set (routing-jev-vs-llm)
  • Unseen holdout (square)
In chart order.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

Gap labels, Unseen holdout vs Tuned set (routing-jev-vs-llm): Unseen holdout is x percentage points higher (+) or lower (−) than Tuned set (routing-jev-vs-llm), calculated from the two values shown; lines are the 95% Wilson interval.

3 rows, 2 series: Tuned set (routing-jev-vs-llm), Unseen holdout. Tuned set (routing-jev-vs-llm): highest Claude Sonnet 5.5 (low) · Claude Code 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 · Claude Code 89% (95% interval 80%–94%, n 82). All intervals overlap. Unseen holdout: highest Claude Sonnet 5.5 (low) · Claude Code 88% (95% interval 76%–94%, n 56). Lowest Claude Haiku 4.5 · Claude Code 79% (95% interval 66%–87%, n 56). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 56–82 per row

Tuned set: the routing study’s 82 cases, revised against Jev answers. Holdout: 56 new cases, frozen before any router call

Whiskers are 95% Wilson intervals. The two case sets differ in mix and size, so a gap mixes a change of case set with any change in the router, and the data cannot separate them. A gap counts only when the two intervals do not overlap.

Sources: Routing on unseen holdout decisions, Routing runs: Jev router vs LLM routing

Share card (PNG)
  • Wall time
  • Model time (API, CLI-reported)
Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

Seconds · log scale: each gridline is 10 times the one before

3 rows, 2 series: Wall time, Model time (API, CLI-reported). Wall time: slowest Claude Haiku 4.5 · Claude Code 9.4 s (median to p95 9.4 s–25.4 s, n 56). Fastest Jev 1.13 (TypeSafe) 0.1 s (median to p95 0.1 s–0.2 s, n 168). Not all run ranges overlap. Model time (API, CLI-reported): slowest Claude Haiku 4.5 · Claude Code 7.5 s (median to p95 7.5 s–23.9 s, n 56). Fastest Claude Sonnet 5.5 (low) · Claude Code 1.5 s (median to p95 1.5 s–2.4 s, n 56). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 56–168 per row

Median, whisker to the 95th percentile

The routes differ. Jev: one HTTPS call to the vendor API from one Mac, all three repetitions. Claude: the Claude Code CLI from process start to exit, which adds start-up time a direct API call would not. The whisker runs from the median to the 95th percentile; it is not a confidence interval. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.

Source: Routing on unseen holdout decisions

Share card (PNG)
Calculation
Largest value is 240x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Sonnet 5.5 (low) · Claude Code $7.24 (n 56). Lowest Jev 1.13 (TypeSafe) $0.031 (n 168).

Notesn 56–168 per row

Reported tokens per decision × list price

Calculation, not a bill. Jev: reported input tokens × $0.042 per million, output free. Claude: CLI-reported tokens × list price; the CLI wrote its prompt cache as 1-hour writes, priced at 2× input, and adds its own system prompt and tool-schema tokens. The Claude routers ran on a subscription. At the tuned-set study’s convention (every cache write at 1.25× input), Sonnet 5.5 (low) would be $4.877 here. That study shows $4.996 for it on its own cases, where the CLI reported $7.324, so its cost row and this one differ by convention and by case mix.

Sources: Routing on unseen holdout decisions, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models)

Share card (PNG)

Tables

Same unseen cases, two routers: exact McNemar test

PairCasesBoth rightOnly first rightOnly second rightBoth wrongExact McNemar p
Jev 1.13 (TypeSafe) vs Claude Haiku 4.5 · Claude Code56424280.688
Jev 1.13 (TypeSafe) vs Claude Sonnet 5.5 (low) · Claude Code56424730.549
Claude Haiku 4.5 · Claude Code vs Claude Sonnet 5.5 (low) · Claude Code56422750.18

Label agreement: the author vs a second, blind labeller (GPT-6.1 Sol)

DecisionQuestionLabelledSecond label inside our acceptable set (95% Wilson interval)Second label equals our first label (95% Wilson interval)Set-aware kappa (calculation; adaptive author label)Cohen’s kappa, fixed first label (calculation)
Failure classfailure1414/14 (78.5% to 100.0%)14/14 (78.5% to 100.0%)11
Message intentintent1414/14 (78.5% to 100.0%)14/14 (78.5% to 100.0%)11
Is it a rule?kind1414/14 (78.5% to 100.0%)14/14 (78.5% to 100.0%)11
Context shapeartifacts91/9 (2.0% to 43.5%)1/9 (2.0% to 43.5%)00
Context shapeknowledge1414/14 (78.5% to 100.0%)11/14 (52.4% to 92.4%)10.57
Context shapememories1010/10 (72.2% to 100.0%)10/10 (72.2% to 100.0%)11
Context shapeexamples1311/13 (57.8% to 95.7%)11/13 (57.8% to 95.7%)0.680.68
Context shapescope1414/14 (78.5% to 100.0%)10/14 (45.4% to 88.3%)10.53
Context shapecomplexity1414/14 (78.5% to 100.0%)10/14 (45.4% to 88.3%)10.46
Context shapeturn44/4 (51.0% to 100.0%)3/4 (30.1% to 95.4%)10.56
Context shapetranscript55/5 (56.6% to 100.0%)2/5 (11.8% to 76.9%)10.25

Tuned set vs unseen holdout, by decision group: exact rate with 95% intervals

RouterDecision groupTuned setUnseen holdoutChange (points, calculation)Intervals
Jev 1.13 (TypeSafe)The 3 one-question types (failure class, message intent and is it a rule?)50/50, 100.0% (92.9% to 100.0%)38/42, 90.5% (77.9% to 96.2%)−9.5overlap
Jev 1.13 (TypeSafe)Context shape (several questions per case)24/32, 75.0% (57.9% to 86.7%)8/14, 57.1% (32.6% to 78.6%)−17.9overlap
Claude Haiku 4.5 · Claude CodeThe 3 one-question types (failure class, message intent and is it a rule?)49/50, 98.0% (89.5% to 99.6%)39/42, 92.9% (81.0% to 97.5%)−5.1overlap
Claude Haiku 4.5 · Claude CodeContext shape (several questions per case)24/32, 75.0% (57.9% to 86.7%)5/14, 35.7% (16.3% to 61.2%)−39.3overlap
Claude Sonnet 5.5 (low) · Claude CodeThe 3 one-question types (failure class, message intent and is it a rule?)50/50, 100.0% (92.9% to 100.0%)41/42, 97.6% (87.7% to 99.6%)−2.4overlap
Claude Sonnet 5.5 (low) · Claude CodeContext shape (several questions per case)27/32, 84.4% (68.2% to 93.1%)8/14, 57.1% (32.6% to 78.6%)−27.2overlap

Every holdout case: labels and who answered it exactly

CaseDecisionAcceptable answersKeys both labellers acceptJev 1.13 (TypeSafe)Claude Haiku 4.5 · Claude CodeClaude Sonnet 5.5 (low) · Claude CodeJev exact in 3 repetitions
holdout-01Failure classfailure: transient1/1exactexactexact3/3
holdout-02Failure classfailure: transient1/1exactexactexact3/3
holdout-03Failure classfailure: transient1/1exactexactexact3/3
holdout-04Failure classfailure: impossible-here1/1exactexactexact3/3
holdout-05Failure classfailure: impossible-here1/1exactexactexact3/3
holdout-06Failure classfailure: impossible-here1/1exactexactexact3/3
holdout-07Failure classfailure: impossible-here1/1exactexactexact3/3
holdout-08Failure classfailure: flaky1/1exactexactexact3/3
holdout-09Failure classfailure: flaky1/10/1 keys0/1 keysexact1/3
holdout-10Failure classfailure: flaky1/1exactexactexact3/3
holdout-11Failure classfailure: real1/1exactexactexact3/3
holdout-12Failure classfailure: real1/1exactexactexact3/3
holdout-13Failure classfailure: real1/1exactexactexact3/3
holdout-14Failure classfailure: real1/1exactexactexact3/3
holdout-15Message intentintent: request1/1exactexactexact3/3
holdout-16Message intentintent: request1/1exactexactexact3/3
holdout-17Message intentintent: request1/1exactexactexact3/3
holdout-18Message intentintent: followup1/1exactexactexact3/3
holdout-19Message intentintent: followup1/1exactexactexact3/3
holdout-20Message intentintent: followup1/1exactexactexact3/3
holdout-21Message intentintent: status1/1exactexactexact3/3
holdout-22Message intentintent: status1/1exactexactexact3/3
holdout-23Message intentintent: status1/1exactexactexact3/3
holdout-24Message intentintent: decision1/1exactexactexact3/3
holdout-25Message intentintent: decision1/10/1 keysexactexact0/3
holdout-26Message intentintent: control1/1exactexactexact3/3
holdout-27Message intentintent: control1/1exactexactexact3/3
holdout-28Message intentintent: control1/10/1 keys0/1 keysexact0/3
holdout-29Is it a rule?kind: obligation1/1exactexactexact3/3
holdout-30Is it a rule?kind: obligation1/1exactexactexact3/3
holdout-31Is it a rule?kind: obligation1/1exactexactexact3/3
holdout-32Is it a rule?kind: obligation1/10/1 keys0/1 keys0/1 keys0/3
holdout-33Is it a rule?kind: prohibition1/1exactexactexact3/3
holdout-34Is it a rule?kind: prohibition or obligation1/1exactexactexact3/3
holdout-35Is it a rule?kind: prohibition1/1exactexactexact3/3
holdout-36Is it a rule?kind: preference1/1exactexactexact3/3
holdout-37Is it a rule?kind: preference1/1exactexactexact3/3
holdout-38Is it a rule?kind: preference1/1exactexactexact3/3
holdout-39Is it a rule?kind: not-a-rule1/1exactexactexact3/3
holdout-40Is it a rule?kind: not-a-rule1/1exactexactexact3/3
holdout-41Is it a rule?kind: not-a-rule1/1exactexactexact3/3
holdout-42Is it a rule?kind: not-a-rule1/1exactexactexact3/3
holdout-43Context shapeartifacts: no; knowledge: search; memories: yes; examples: yes; scope: small or medium; complexity: simple or moderate5/6exact5/6 keys5/6 keys3/3
holdout-44Context shapeartifacts: no; knowledge: search; memories: yes; examples: no; scope: medium; complexity: moderate or complex5/6exact4/6 keysexact3/3
holdout-45Context shapeturn: follow-up; transcript: recent or full; knowledge: search; scope: small or medium; complexity: moderate5/5exactexactexact3/3
holdout-46Context shapeturn: clarification; transcript: recent or full; knowledge: search; examples: no; scope: small or medium; complexity: moderate6/65/6 keys5/6 keysexact0/3
holdout-47Context shapetranscript: summary or recent; knowledge: search; memories: yes; examples: no; scope: medium or large; complexity: moderate or complex5/6exactexact5/6 keys3/3
holdout-48Context shapeartifacts: no; knowledge: none; memories: no; examples: no; scope: small; complexity: simple or moderate5/65/6 keys4/6 keysexact0/3
holdout-49Context shapeartifacts: no; knowledge: search; memories: yes; examples: yes; scope: small or medium; complexity: moderate5/65/6 keysexactexact0/3
holdout-50Context shapeartifacts: no; knowledge: search; memories: yes; examples: no; scope: medium; complexity: complex5/6exactexact5/6 keys3/3
holdout-51Context shapetranscript: recent or full; knowledge: facts or search; memories: yes; examples: no; scope: medium; complexity: moderate6/63/6 keys2/6 keysexact0/3
holdout-52Context shapeartifacts: no; knowledge: facts or search; memories: yes; examples: yes; scope: medium or large; complexity: moderate6/65/6 keys2/6 keys4/6 keys0/3
holdout-53Context shapeturn: new-task or follow-up; artifacts: no; knowledge: facts or search; memories: yes; examples: no; scope: small or medium; complexity: moderate6/7exact5/7 keysexact3/3
holdout-54Context shapeturn: follow-up; transcript: recent or full; knowledge: facts or search; examples: yes; scope: small or medium; complexity: moderate6/6exactexactexact3/3
holdout-55Context shapeartifacts: no; knowledge: search; memories: yes; examples: no; scope: large; complexity: complex4/6exact5/6 keys5/6 keys3/3
holdout-56Context shapeartifacts: no; knowledge: facts; examples: no; scope: small; complexity: simple or moderate4/54/5 keys2/5 keys2/5 keys0/3

Every router on the holdout

RouterRouteCounted callsExactKey accuracyExact, agreed keys onlyMean tokens per decision (calculation; input incl. cache / output)USD per 1,000 (calculation)
Jev 1.13 (TypeSafe)direct HTTPS to the vendor API with the production request body16846 of 56 (82%, 95% interval 70% to 90%)113 of 125 (90%, 95% interval 84% to 94%)46 of 56 (82%, 95% interval 70% to 90%)730 / 119.3$0.031
Claude Haiku 4.5 · Claude CodeClaude Code CLI, one call per decision, CLI default effort (extended thinking)5644 of 56 (79%, 95% interval 66% to 87%)102 of 125 (82%, 95% interval 74% to 87%)46 of 56 (82%, 95% interval 70% to 90%)1,775 / 1,071$7.13
Claude Sonnet 5.5 (low) · Claude CodeClaude Code CLI, one call per decision, effort low5649 of 56 (88%, 95% interval 76% to 94%)115 of 125 (92%, 95% interval 86% to 96%)52 of 56 (93%, 95% interval 83% to 97%)1,705 / 90.5$7.24

Method

  1. Cases: 56 new cases, 14 per decision type, in the production case shape. The types are failure class, message intent, is-it-a-rule and context shape. Production asks a router about each case. We score only questions the rule leaves open: 125 labelled questions.
  2. Blindness: the author says they opened no router answer before the label freeze. File times support this order but cannot prove blindness. A script compared every case with existing cases. The author rewrote 17 close cases before the freeze.
  3. File times: the protocol file dates from 21:35:45 UTC on 2026-10-06. Case drafts and one uncounted second-labeller probe came first. The first counted labelling call came at 21:35:49 UTC. The freeze came at 21:37:02 UTC. The frozen file and its checksum preceded every router call.
  4. Second labeller: GPT-6.1 Sol through the Codex CLI, 4 counted calls. It labelled every case without the author labels. Our set accepts its label on 115 of 125 (92%, 95% interval 86% to 96%). It matches our first label on 100 of 125 (80%, 95% interval 72% to 86%). Of these questions, 24 accept two or three labels. The lowest agreement was on “artifacts”; see the caveats. Secondary scoring keeps only questions where our set accepts its label.
  5. Routers: Jev 1.13 used direct HTTPS with the production request body. It made 168 calls in 3 repetitions, with no retries. Claude Haiku 4.5 used CLI default effort; Claude Sonnet 5.5 used effort low. Both used the Claude Code CLI with the production system text and schema. Each made 56 counted calls, one per decision. Each route had one uncounted probe. Total calls stayed within the caps: Jev 169/200, Claude 114/120, labeller 5/5. No arm stopped early or lost cases.
  6. Scoring: the repository’s decision-eval runner. Exact means every scored question in a case was acceptable. Key accuracy counts each question. We use Wilson 95% intervals and exact McNemar tests on paired cases. The paired tests have no multiple-test correction. Errors count as wrong.
  7. Cost: reported tokens × list price per 1,000 decisions, a calculation.

Caveats

  • The case author and two of the three routers are Claude models, so a same-family label bias is possible. A second labeller (GPT-6.1 Sol) labelled every case blind; the secondary scoring keeps only the keys where its label is inside our acceptable set.
  • 56 cases (14 per decision type) from one author: per-type intervals are very wide.
  • The routes differ: Jev ran over direct HTTPS, the Claude routers through the Claude Code CLI, which adds start-up time and tokens.
  • Jev figures are its first repetition; repetitions 2 and 3 and the stability across all three are reported beside it.
  • The artifacts label violates the protocol rule to leave judgement calls unscored. The author accepted only no on all 9 labelled cases. The second labeller chose no on 1 of 9. Frozen labels stay unchanged; sensitivity calculations are not new runs.
  • The protocol file was created at 21:35:45.935 UTC on 2026-10-06. It followed the case drafts and an uncounted labeller probe. It preceded the first counted labelling call at 21:35:49.3 UTC, the label freeze at 21:37:02.582 UTC and every router call. It was not written before every call.
  • The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.
  • The three one-question case sets are near a ceiling: every router scored 12 to 14 of 14 on each. The 95% Wilson interval for 14 of 14 is 78.5% to 100%. These sets provide little room to separate routers.
  • Per-question Wilson intervals treat questions as independent. Several questions share each context-shape case, so those intervals can understate uncertainty. Label-agreement intervals have the same limit.
  • Set-aware kappa uses the second label as the author label when the acceptable set contains it. This adaptive calculation can inflate agreement. Only kappaStrict compares two fixed labels.
  • The tuned set was revised against Jev answers. Its size and decision mix differ from the holdout. Overlapping intervals neither show nor rule out a home advantage.
  • The three paired McNemar tests have no multiple-test correction. Their p values do not prove equal accuracy.
  • Failure class, Message intent and Is it a rule?: every router answered at least 85% of the 14 cases exactly. These case sets are near the ceiling and barely separate the routers; context shape carries most of the differences.
  • Agreement with the second labeller counts a label inside our acceptable set: 115 of 125 questions (92%). 24 of the 125 questions accept two or three labels, so agreement with our first label alone is lower: 100 of 125 (80%).
  • The “artifacts” label broke our own rules. The second labeller’s answer was inside our set on 1 of 9 cases (kappa 0). We labelled “no” on every one of them. The protocol leaves out a question whose answer is a judgement call, and the production case suites accept both answers for a fresh work item. Cases where the router’s answer equalled our label: Jev 1.13 9 of 9, Haiku 4.5 4 of 9, Sonnet 5.5 (low) 7 of 9. The scoring against our labels therefore favours Jev on this question. The labels stayed frozen. With both answers accepted for “artifacts” and the same answers scored again (a calculation), exact answers are Jev 1.13 46 of 56 (82%, 95% interval 70% to 90%); Haiku 4.5 45 of 56 (80%, 95% interval 68% to 89%); Sonnet 5.5 (low) 50 of 56 (89%, 95% interval 79% to 95%). Context shape: Jev 1.13 8 of 14 (57%, 95% interval 33% to 79%); Haiku 4.5 6 of 14 (43%, 95% interval 21% to 67%); Sonnet 5.5 (low) 9 of 14 (64%, 95% interval 39% to 84%); the overall intervals still overlap.
  • Label spread is narrow on some questions. One label holds at least 90% of the cases for artifacts (“no”, 9 of 9) and memories (“yes”, 9 of 10). One option has a single case for knowledge (“none”), scope (“large”) and transcript (“summary”), so those options are barely tested.
  • The tuned and holdout sets differ in size and mix, so a change between them cannot be assigned to the router or to the case set. The data neither shows nor rules out a home advantage.

Sources

  • Routing on unseen holdout decisions

    Our recorded runs ·

    Frozen unseen routing decisions; every router call retained. Costs are calculations from recorded tokens and list prices.

    Raw data: routing-holdout/results.json

  • Routing runs: Jev router vs LLM routing

    Our recorded runs ·

    Routing decisions recorded per case and arm.

    Raw data: routing/receipts.json

  • Repricing calculation

    Calculation ·

    Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.

  • Jev 1.13 list price

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free.

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Jev vs Claude routers on unseen decisions: a blind holdout”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/routing-holdout.

Models and comparisons in this study

More studies

All benchmarks
Live storyIncludes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.