• Routing
  • Jev
  • Claude Haiku
  • Claude Sonnet
  • Model Routing
  • Thought experiment

Jev vs Claude as a router: accuracy and cost

Should a small dedicated router or a general LLM make the platform’s typed routing decisions?

Published · Updated · 7 charts · Download the data or a carousel

90%

95% CI 82%–95% · n = 82

74/82 · Jev 1.13 (TypeSafe): exact decisions

Live run: 3 repeats of the same 82 decisions. 221 of 246 calls were exact (89.8%; 74, 73 and 74 of 82 per repeat). The count is shown on the 82-decision scale (89.8% of 82 is 74), the scale of the interval: repeats of one decision are not independent, so the interval is taken at n = 82, not 246.

The answer

Exact decisions: Jev 1.13 (TypeSafe) 221 of 246 live calls (90%; 74, 73 and 74 of 82 per repeat; case-level interval 82% to 95%); Claude Haiku 4.5 73 of 82 (89%, 80% to 94%); Claude Sonnet 5.5 77 of 82 (94%, 87% to 97%). The intervals overlap, so accuracy does not separate the routers here. Cost does: Jev costs $0.0337 per 1,000 decisions (a calculation from its reported input tokens) against Claude Haiku 4.5 $8.92, Claude Sonnet 5.5 $5.00. Case by case against Jev’s recorded production run (one pass), the exact McNemar test finds no difference (Claude Haiku 4.5 p = 1, Claude Sonnet 5.5 p = 0.375). Per decision, Jev is about 265x cheaper than Claude Haiku 4.5 and about 148x cheaper than Claude Sonnet 5.5. Median model time per decision through the Claude Code CLI: Claude Haiku 4.5 10.7 s, Claude Sonnet 5.5 1.6 s. Jev, called directly over HTTPS from one Mac, took a median 137 ms per call (p95 196 ms, 246 calls, wall time with the network inside it). That is a different route from the CLI, so it is not a model-against-model compute comparison. The case sets were tuned against Jev answers, which gives Jev a home advantage. Thought experiment (a calculation on 2,362 recorded calls, not a run): the same tokens cost $108.54 all on Sonnet 5.5, $161.62 under the platform policy mix (1.49x) and $105.53 with Haiku on side jobs only (2.8% less), because the main coding stages hold most of the spend.

Live story

Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.

Live story · 34 sJev vs Claude as a router: accuracy and cost

Jev vs Claude as a router: accuracy and cost

Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.

Transcript
  1. Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.
  2. Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
  3. Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
  4. Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
  5. Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
  6. Open benchmarks: intervals, sources and every failure kept.

Key numbers

89% (73/82)

Claude Haiku 4.5: exact decisions

95% CI 80%–94% · n = 82

94% (77/82)

Claude Sonnet 5.5: exact decisions

95% CI 87%–97% · n = 82

$0.0337

Jev cost per 1,000 decisions

n = 246

$161.62

Thought experiment: policy (Opus strong, Haiku ancillary) vs all Sonnet 5.5

(1.49x) · n = 2362

$105.53

Thought experiment: split (Sonnet main line, Haiku ancillary) vs all Sonnet 5.5

(0.97x) · n = 2362

93.0%

Cache-read share of recorded input (economics data)

n = 2362

Routing hub: Jev, LLM routers, the deterministic policy and gateways on one page

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Thought experiment: not a run. These values reprice recorded tokens at list prices. No model was called again.

Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 82 per row

Share of asked cases where every scored question was acceptable

Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 97% (95% interval 94%–99%, n 194). Lowest Claude Haiku 4.5 94% (95% interval 90%–97%, n 194). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 194 per row

Each open question the router was asked; an unanswered question counts as wrong

Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)

Jev 1.13 (TypeSafe)

Failure class
Message intent
Is it a rule?
Context shape

Claude Haiku 4.5

Failure class
Message intent
Is it a rule?
Context shape

Claude Sonnet 5.5

Failure class
Message intent
Is it a rule?
Context shape

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

4 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5, Claude Sonnet 5.5. Jev 1.13 (TypeSafe): highest Failure class 100% (95% interval 82%–100%, n 18). Lowest Context shape 74% (95% interval 58%–87%, n 32). All intervals overlap. Claude Haiku 4.5: highest Message intent 100% (95% interval 84%–100%, n 20). Lowest Context shape 75% (95% interval 58%–87%, n 32). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 12–32 per row3 of 4 (Jev 1.13 (TypeSafe)) at 100%: this task set cannot separate them.

A missing bar means that router was not run on that decision type. Whiskers are 95% Wilson intervals. Jev: the pooled rate of the live repeats, with the interval taken at the number of decisions of that type.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
Calculation
Largest value is 260x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).

Notesn 82–246 per row

List price × reported tokens per decision

List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
  • Wall time (CLI)
  • Wall time (direct API call)
  • Model time (API)
Claude Haiku 4.5
Claude Sonnet 5.5
Jev 1.13 (TypeSafe)

Time per decision · log scale: each gridline is 10 times the one before

3 rows, 3 series: Wall time (CLI), Wall time (direct API call), Model time (API). Wall time (CLI): slowest Claude Haiku 4.5 12.67 s (median to p95 12.67 s–34.41 s, n 82). Fastest Claude Sonnet 5.5 2.6 s (median to p95 2.6 s–4.3 s, n 82). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–246 per row

Median wall time, whisker to the 95th percentile

Whiskers run from p50 to p95. The Claude routers ran through the Claude Code CLI, so their wall time includes CLI start-up and the tool schema; one pass of 82 decisions each. Jev was called directly over HTTPS from one Mac on a home network: 246 calls in a 35-second window, client wall time with the network inside it. Its API reports no server time, so Jev has no model-time point. These are different routes: the chart shows what a caller waits per decision, not model compute time.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
Calculation
all Fable 5.1
all Opus 5.5
policy (Opus strong, Haiku ancillary)
all Sonnet 5.5
split (Sonnet main line, Haiku ancillary)
all Haiku 4.5

Hover or focus a bar for its ratio to all Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 6 rows. Highest all Fable 5.1 $370. Lowest all Haiku 4.5 $54.27.

Notes

50 benchmark runs, 2,362 model calls, repriced

Calculation, not a run: every recorded call ran on Sonnet 5.5 with routing off. Same tokens on every model; a different model or mix would take a different path.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Anthropic list prices (Claude models)

Share card (PNG)
Calculation

Haiku 4.5

act (strong)
research (strong)
verify (strong)
review (strong)
memory and onboarding (economy)
other (standard)

Sonnet 5.5

act (strong)
research (strong)
verify (strong)
review (strong)
memory and onboarding (economy)
other (standard)

Opus 5.5

act (strong)
research (strong)
verify (strong)
review (strong)
memory and onboarding (economy)
other (standard)

Fable 5.1

act (strong)
research (strong)
verify (strong)
review (strong)
memory and onboarding (economy)
other (standard)

One panel per series, all on the same axis.

List-price calculation, not a run. 6 rows, 4 series: Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1. Haiku 4.5: highest act (strong) $28.66. Lowest other (standard) $0.1. Sonnet 5.5: highest act (strong) $57.31. Lowest other (standard) $0.19.

Notes

Calculation, not a run. The tier in brackets is the routing policy tier for that stage.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Anthropic list prices (Claude models)

Share card (PNG)

Tables

Same cases, two routers (Jev: its recorded production run, one pass)

Every case once: both right, only one right, both wrong

Only the 7 discordant decisions count: 3 vs 4. Exact McNemar p = 1: no evidence of a difference.

70Both right3Only Claude Haiku 4.5 right4Only Jev 1.13 (TypeSafe) right5Both wrong
Both right
70 85%Agree and right: says nothing about which is better.
Only Claude Haiku 4.5 right
3 4%Discordant: Claude Haiku 4.5 right where Jev 1.13 (TypeSafe) was wrong.
Only Jev 1.13 (TypeSafe) right
4 5%Discordant: Jev 1.13 (TypeSafe) right where Claude Haiku 4.5 was wrong.
Both wrong
5 6%Agree and wrong: says nothing about which is better.

One square per paired decision (n = 82). Squares start in one grid and sort into the four outcomes; their order inside a quadrant carries no meaning.

Counts of cases from the table; the exact McNemar p reads only the discordant casesn = 82 cases per pair

Claude Haiku 4.5 vs Jev 1.13 (TypeSafe): 3 vs 4 discordant cases of 82, exact McNemar p = 1. Claude Sonnet 5.5 vs Jev 1.13 (TypeSafe): 4 vs 1 discordant cases of 82, exact McNemar p = 0.375.

Per-question accuracy by decision type (Jev: its recorded production run, one pass)

DecisionQuestionJev 1.13 (TypeSafe)Claude Haiku 4.5Claude Sonnet 5.5
Failure classfailure18/1817/1818/18
Message intentintent20/2020/2020/20
Is it a rule?kind12/1212/1212/12
Context shapeturn5/65/65/6
Context shapetranscript12/1613/1615/16
Context shapeartifacts4/44/44/4
Context shapeknowledge22/2422/2423/24
Context shapememories21/2121/2120/21
Context shapeexamples10/1010/1010/10
Context shapescope30/3129/3131/31
Context shapecomplexity31/3230/3231/32

Every router

One card per router: exact decisions with the 95% interval, tokens and cost

  • Jev 1.13 (TypeSafe)

    74/82exact decisions

    90% · 95% CI 82%–95% · n = 82

    Key accuracy
    184/194
    Tokens per decision (in / out)
    803 / 148
    USD per 1,000 decisions
    $0.034

    live API run 2026-10-06: 3 repeats of the same 82 decisions, 246 counted calls one at a time from one Mac. All calls: 221 of 246 exact and 552 of 582 questions; the counts shown are on the 82-decision scale of the interval. The recorded production run of 2026-10-05 scored 74 of 82. Cost is a calculation from the reported input tokens

  • Claude Haiku 4.5

    73/82exact decisions

    89% · 95% CI 80%–94% · n = 82

    Key accuracy
    183/194
    Tokens per decision (in / out)
    1,831 / 1,419
    USD per 1,000 decisions
    $8.92

    this benchmark, Claude Code CLI on a subscription account, one call per decision

  • Claude Sonnet 5.5

    77/82exact decisions

    94% · 95% CI 87%–97% · n = 82

    Key accuracy
    189/194
    Tokens per decision (in / out)
    1,785 / 107
    USD per 1,000 decisions
    $5.00

    this benchmark, Claude Code CLI on a subscription account, one call per decision

  • Clef / Clef-Flash (local)

    —exact decisions: not measured

    not measured: No local Clef server was running and installing a 6-20 GB model was out of scope for this run.

Meter: 95% Wilson interval on exact decisionsCost per 1,000 decisions: list-price calculation or provider-reported, as marked

4 routers. Jev 1.13 (TypeSafe): 74/82 exact; Claude Haiku 4.5: 73/82 exact; Claude Sonnet 5.5: 77/82 exact; Clef / Clef-Flash (local): not measured exact.

Jev live run, repeat by repeat

RunExact decisionsQuestions answered acceptablyMedian time per call (ms)
Repeat 1 (without the cold first call)74/82184/194130 ms
Repeat 273/82183/194142 ms
Repeat 374/82185/194137 ms
All 3 repeats (246 calls)221/246 = 89.8%; case-level 95% interval 81.9% to 95.0%552/582 = 94.8%; 90.8% to 97.2%137 ms
Recorded production run, 2026-10-05 (one pass, no latency recorded)74/82185/194—

Method

  1. Cases: the labelled decision suites the platform uses (failure class, message intent, is-it-a-rule, context shape), only the cases production asks a router about.
  2. Scoring: the repository’s own decision-eval runner. Exact means every scored question in a case was acceptable; key accuracy counts each question.
  3. Jev numbers come from a live run on 2026-10-06: 3 repeats of the same 82 decisions (246 counted calls), sent one at a time over HTTPS to the TypeSafe API from one Apple M3 Ultra Mac on a home network. A failed call would count as wrong; there were 0. Latency is client wall time, so the network is inside it, and the API reports no server time. Jev’s earlier recorded production run (2026-10-05, same case versions and runner) scored 74 of 82 and has no per-call latency. Cost is a calculation: reported input tokens × the published price.
  4. Claude routers ran through the Claude Code CLI with the production system text and schema, one call per decision, one pass over the 82 decisions.
  5. Stability: 73 of 82 decisions were exact in every repeat, 8 in none and 1 in some (a Context shape case, wrong in repeat 2). Jev can return different probabilities for the same request; the repeats show how much.
  6. Economics: recorded tokens of the benchmark runs repriced at list prices for each model mix (a calculation).

Caveats

  • The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
  • Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
  • Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
  • One sample per decision; production asks a second sample when confidence is low. Confidence-gated coverage is therefore not compared.
  • 82 cases in four small hand-labelled sets: intervals are wide.
  • Clef / Clef-Flash (local): not measured (No local Clef server was running and installing a 6-20 GB model was out of scope for this run).
  • Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
  • Jev ran as a direct HTTPS call; the Claude routers ran through the Claude Code CLI. These are different routes, so speed and cost compare what a caller pays per decision, not one model against the other.
  • Jev’s latency is one 35-second window from one Mac over a home network. The API reports no server time. A caller near the API would see less.

Sources

  • Routing runs: Jev router vs LLM routing

    Our recorded runs ·

    Routing decisions recorded per case and arm.

    Raw data: routing/receipts.json

  • Repricing calculation

    Calculation ·

    Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.

  • Jev 1.13 list price

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free.

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • Jev live run: 246 timed calls on the 82 routing decisions

    Our recorded runs ·

    Jev 1.13 called over HTTPS, 3 repeats of the same 82 typed decisions, one call at a time, from one Mac over a home network: client wall time, with the network inside it. The API reports no server time. Cost per 1,000 decisions is a calculation from the reported input tokens and the published price. The case sets were revised against Jev answers, so Jev has a home advantage.

    Raw data: jev-live/summary.json, jev-live/calls.json

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Jev vs Claude as a router: accuracy and cost”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/routing-jev-vs-llm.

Explainers that cite this study

Read the methods and terms in the context of these recorded results.

More write-ups that cite this study (1)

Models and comparisons in this study

More studies

All benchmarks
  • Routing
  • Jev

Jev vs Claude routers on unseen decisions: a blind holdout

Jev 1.13, Claude Haiku 4.5 and Sonnet 5.5 on 56 routing decisions written blind and frozen first: exact rate with 95% intervals, tuned vs unseen.

82% (46/56)Jev 1.13 (TypeSafe): exact on unseen decisions · n = 56

6 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • Claude Haiku

Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts

Haiku 4.5 first, then retry or escalate to Sonnet 5.5? A calculation on 64 recorded calls across 8 hard tasks: cost and time per correct answer.

46% (11/24)Strict passes, Claude Haiku 4.5 · Claude Code, eight hard tasks · n = 24

6 chartsUpdated October 7, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.