• Routing
  • Model Routing
  • Jev
  • Claude Haiku

Jev vs Claude Haiku and Sonnet as a router: tied on accuracy, 148x to 265x cheaper

82 typed routing decisions. Jev, Claude Haiku 4.5 and Sonnet 5.5 tie on accuracy, but Jev costs $0.0337 per 1,000 decisions against $5 to $8.92.

TL;DR

  • On 82 typed routing decisions, Jev 1.13 got 74 exactly right (90%; 221 of 246 live calls over 3 repeats), Claude Haiku 4.5 got 73 (89%) and Claude Sonnet 5.5 got 77 (94%).
  • The 95% intervals overlap, and an exact McNemar test finds no difference against Jev (Haiku p = 1, Sonnet p = 0.375). Accuracy does not rank these routers.
  • Cost does. Jev costs $0.0337 per 1,000 decisions. Haiku costs $8.92 and Sonnet $5.00 through the CLI. That makes Jev about 265x cheaper than Haiku and 148x cheaper than Sonnet.
  • Median model time per decision through the CLI: Haiku 10.7 s, Sonnet 1.6 s. Jev's call, a direct API call from one Mac, took a median 136.5 ms (246 calls). That is a different route, so read it as a gap between routes, not only between models.
  • Home advantage: the case sets were tuned against Jev answers. Read the tie with that in mind.

Every case, every router: /benchmarks/routing-jev-vs-llm.

What a routing decision is

Inside an agent platform, many small choices happen before any real work starts. Is this message a question or a task? Is this failure a flaky test or a real bug? Is this sentence a rule the worker should keep? How much context does the next step need?

Each of these is a typed decision: the router reads some text and fills a small schema. We tested four decision suites that production actually asks a router about:

  1. Failure class, 18 cases.
  2. Message intent, 20 cases.
  3. Is it a rule?, 12 cases.
  4. Context shape, 32 cases, with several questions each.

A decision is exact when every scored question in the case is acceptable. Key accuracy scores each question on its own.

Accuracy: a three-way tie

Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 82 per row

Share of asked cases where every scored question was acceptable

Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Live story · 34 sJev vs Claude as a router: accuracy and cost

Jev vs Claude as a router: accuracy and cost

Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.

Transcript
  1. Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.
  2. Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
  3. Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
  4. Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
  5. Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
  6. Open benchmarks: intervals, sources and every failure kept.
  • Jev 1.13 (TypeSafe): 221 of 246 live calls, 89.8% (74, 73 and 74 of 82 per repeat), interval 82% to 95%.
  • Claude Haiku 4.5: 73 of 82, 89.0%, interval 80% to 94%.
  • Claude Sonnet 5.5: 77 of 82, 93.9%, interval 87% to 97%.

Per question, the picture is the same. Jev got 552 of 582 answers right over its three repeats (184, 183 and 185 of 194), Haiku 183 of 194 and Sonnet 189.

Sonnet is ahead on the point estimate. Is the gap real? The right test for two routers on the same cases is McNemar's test, which looks only at the cases where they disagree.

  • Haiku vs Jev: both right on 70, only Haiku right on 3, only Jev right on 4, both wrong on 5. Exact p = 1.
  • Sonnet vs Jev: both right on 73, only Sonnet right on 4, only Jev right on 1, both wrong on 4. Exact p = 0.375.

Seven disagreements in one pair and five in the other are not enough to call a winner. With this sample, the test cannot tell the three routers apart.

Where the routers differ

Jev 1.13 (TypeSafe)

Failure class
Message intent
Is it a rule?
Context shape

Claude Haiku 4.5

Failure class
Message intent
Is it a rule?
Context shape

Claude Sonnet 5.5

Failure class
Message intent
Is it a rule?
Context shape

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

4 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5, Claude Sonnet 5.5. Jev 1.13 (TypeSafe): highest Failure class 100% (95% interval 82%–100%, n 18). Lowest Context shape 74% (95% interval 58%–87%, n 32). All intervals overlap. Claude Haiku 4.5: highest Message intent 100% (95% interval 84%–100%, n 20). Lowest Context shape 75% (95% interval 58%–87%, n 32). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 12–32 per row3 of 4 (Jev 1.13 (TypeSafe)) at 100%: this task set cannot separate them.

A missing bar means that router was not run on that decision type. Whiskers are 95% Wilson intervals. Jev: the pooled rate of the live repeats, with the interval taken at the number of decisions of that type.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Three of the four suites are solved by everyone. Message intent and "is it a rule?" are 100% for all three routers. Failure class is 18 of 18 for Jev and Sonnet, and 17 of 18 for Haiku.

All the action is in context shape: Jev 24, 23 and 24 of 32 in its three repeats, Haiku 24 of 32, Sonnet 27 of 32. Context shape asks several questions at once, such as how much transcript, which knowledge and how complex the turn is. The hardest single question was the transcript one: on Jev's recorded production run, 12 of 16, against Haiku 13 of 16 and Sonnet 15 of 16.

If you are building a router, that is the lesson. Simple classifications saturate fast. Multi-field judgments about context are where a router earns its keep, and where you need more cases to tell routers apart.

Cost: not close

Calculation
Largest value is 260x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).

Notesn 82–246 per row

List price × reported tokens per decision

List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions

Here the routers separate by orders of magnitude:

  • Jev: $0.0337 per 1,000 decisions. About 803 input tokens per decision, and Jev output is free at its list price.
  • Claude Sonnet 5.5: $5.00 per 1,000. About 1,785 input and 107 output tokens per decision.
  • Claude Haiku 4.5: $8.92 per 1,000. About 1,831 input and 1,419 output tokens per decision.

Haiku costs more than Sonnet here. That surprises people, but the reason is simple. Haiku ran with the CLI default extended thinking and wrote about 1,419 output tokens per decision. Sonnet ran at effort low, as production asks, and wrote about 107. Output tokens are the expensive ones.

Note that the Claude figures include the CLI's tool-schema and thinking tokens, because the CLI reports them. A direct API call would be cheaper. We estimate it would not close a gap of two orders of magnitude, but we did not measure it.

Latency

  • Wall time (CLI)
  • Wall time (direct API call)
  • Model time (API)
Claude Haiku 4.5
Claude Sonnet 5.5
Jev 1.13 (TypeSafe)

Time per decision · log scale: each gridline is 10 times the one before

3 rows, 3 series: Wall time (CLI), Wall time (direct API call), Model time (API). Wall time (CLI): slowest Claude Haiku 4.5 12.67 s (median to p95 12.67 s–34.41 s, n 82). Fastest Claude Sonnet 5.5 2.6 s (median to p95 2.6 s–4.3 s, n 82). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–246 per row

Median wall time, whisker to the 95th percentile

Whiskers run from p50 to p95. The Claude routers ran through the Claude Code CLI, so their wall time includes CLI start-up and the tool schema; one pass of 82 decisions each. Jev was called directly over HTTPS from one Mac on a home network: 246 calls in a 35-second window, client wall time with the network inside it. Its API reports no server time, so Jev has no model-time point. These are different routes: the chart shows what a caller waits per decision, not model compute time.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

  • Claude Sonnet 5.5: median model time 1.6 s, 95th percentile 2.57 s. Wall time through the CLI: median 2.6 s.
  • Claude Haiku 4.5: median model time 10.7 s, 95th percentile 32.07 s. Wall time: median 12.67 s.
  • Jev 1.13: median wall time 0.14 s (136.5 ms), 95th percentile 0.2 s (195.7 ms), over 246 calls in a 35-second window. We called its API directly from one Mac over a home network, so the network is inside that time, and the API reports no server time.

These are different routes. The Claude routers went through the CLI, which adds about a second per Sonnet decision; Jev was a direct call. The chart shows what a caller waits per decision, not model compute time. For the full delay picture, including a rule-based router at 1.42 µs per decision and the CLI share of each Claude decision, see what does a router cost you?

A router that takes 10 seconds is a problem when it sits in front of every message. Haiku's thinking default made it both the slowest and the most expensive router in this test.

So which router should you use?

On this evidence:

  • If the router runs on every message or every step, cost and latency dominate. A dedicated router like Jev ties on accuracy at a tiny fraction of the cost.
  • If a decision is rare and hard, such as a complex context judgment, Sonnet at low effort is the strongest point estimate, and it is fast.
  • Do not assume the small general model is the cheap option. Under CLI defaults, Haiku thought its way to the highest cost and the slowest time.

There is a second question: once you have a router, does routing work to different models save money? We ran that as a thought experiment on 2,362 recorded calls, in what if every call ran on Opus?. Short version: on our pipeline, the main coding stages hold most of the spend, so routing side jobs to a cheaper model saved only 2.8%.

How we measured

  • Cases. The labelled decision suites the platform uses, limited to the cases production asks a router about. 82 cases, 194 scored questions.
  • Scoring. The repository's own decision-eval runner, the same for every router.
  • Jev. A live run on 2026-10-06: the same 82 decisions, 3 repeats, 246 counted calls, one at a time over HTTPS from one Mac, scored by the same runner. Its interval is taken at the 82 cases because repeats are not independent. Its earlier recorded production run scored 74 of 82 with the same case versions. McNemar needs one pass, so it uses that recorded run.
  • Claude routers. Run through the Claude Code CLI with the production system text and schema, one call per decision. Haiku 4.5 used the CLI default extended thinking. Sonnet 5.5 used effort low.
  • Cost. List price × reported tokens per decision.
  • Statistics. 95% Wilson intervals and an exact McNemar test on paired cases.

Caveats

  • Home advantage. The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05.
  • Jev's timing is one window. 246 calls from one machine on a home network, in 35 seconds. Server load then is unknown, and a caller near the API would likely see less.
  • CLI overhead. The Claude routers ran through a CLI, which adds start-up time and tool-schema tokens a direct API call would not.
  • One sample per decision. Production asks a second sample when confidence is low, so confidence-gated coverage is not compared here.
  • Small sets. 82 cases in four hand-labelled sets give wide intervals.
  • Not measured. A local Clef / Clef-Flash router was not run, because no local server was running and installing a 6 to 20 GB model was out of scope.

Routing you can inspect

Agent uses a rules-first decision cascade for incoming work. Its records show the selected answer and fallback when a decision ran. Model-call costs can be list-price estimates; they are not verified invoices. Try Agent and inspect the decisions recorded for a task.

The data behind this post

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.