Jev vs Claude Haiku and Sonnet as a router: tied on accuracy, 148x to 265x cheaper
82 typed routing decisions. Jev, Claude Haiku 4.5 and Sonnet 5.5 tie on accuracy, but Jev costs $0.0337 per 1,000 decisions against $5 to $8.92.
TL;DR
- On 82 typed routing decisions, Jev 1.13 got 74 exactly right (90%; 221 of 246 live calls over 3 repeats), Claude Haiku 4.5 got 73 (89%) and Claude Sonnet 5.5 got 77 (94%).
- The 95% intervals overlap, and an exact McNemar test finds no difference against Jev (Haiku p = 1, Sonnet p = 0.375). Accuracy does not rank these routers.
- Cost does. Jev costs $0.0337 per 1,000 decisions. Haiku costs $8.92 and Sonnet $5.00 through the CLI. That makes Jev about 265x cheaper than Haiku and 148x cheaper than Sonnet.
- Median model time per decision through the CLI: Haiku 10.7 s, Sonnet 1.6 s. Jev's call, a direct API call from one Mac, took a median 136.5 ms (246 calls). That is a different route, so read it as a gap between routes, not only between models.
- Home advantage: the case sets were tuned against Jev answers. Read the tie with that in mind.
Every case, every router: /benchmarks/routing-jev-vs-llm.
What a routing decision is
Inside an agent platform, many small choices happen before any real work starts. Is this message a question or a task? Is this failure a flaky test or a real bug? Is this sentence a rule the worker should keep? How much context does the next step need?
Each of these is a typed decision: the router reads some text and fills a small schema. We tested four decision suites that production actually asks a router about:
- Failure class, 18 cases.
- Message intent, 20 cases.
- Is it a rule?, 12 cases.
- Context shape, 32 cases, with several questions each.
A decision is exact when every scored question in the case is acceptable. Key accuracy scores each question on its own.
Accuracy: a three-way tie
Jev vs Claude as a router: accuracy and cost
Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.
Transcript
- Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.
- Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
- Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
- Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
- Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
- Open benchmarks: intervals, sources and every failure kept.
- Jev 1.13 (TypeSafe): 221 of 246 live calls, 89.8% (74, 73 and 74 of 82 per repeat), interval 82% to 95%.
- Claude Haiku 4.5: 73 of 82, 89.0%, interval 80% to 94%.
- Claude Sonnet 5.5: 77 of 82, 93.9%, interval 87% to 97%.
Per question, the picture is the same. Jev got 552 of 582 answers right over its three repeats (184, 183 and 185 of 194), Haiku 183 of 194 and Sonnet 189.
Sonnet is ahead on the point estimate. Is the gap real? The right test for two routers on the same cases is McNemar's test, which looks only at the cases where they disagree.
- Haiku vs Jev: both right on 70, only Haiku right on 3, only Jev right on 4, both wrong on 5. Exact p = 1.
- Sonnet vs Jev: both right on 73, only Sonnet right on 4, only Jev right on 1, both wrong on 4. Exact p = 0.375.
Seven disagreements in one pair and five in the other are not enough to call a winner. With this sample, the test cannot tell the three routers apart.
Where the routers differ
Three of the four suites are solved by everyone. Message intent and "is it a rule?" are 100% for all three routers. Failure class is 18 of 18 for Jev and Sonnet, and 17 of 18 for Haiku.
All the action is in context shape: Jev 24, 23 and 24 of 32 in its three repeats, Haiku 24 of 32, Sonnet 27 of 32. Context shape asks several questions at once, such as how much transcript, which knowledge and how complex the turn is. The hardest single question was the transcript one: on Jev's recorded production run, 12 of 16, against Haiku 13 of 16 and Sonnet 15 of 16.
If you are building a router, that is the lesson. Simple classifications saturate fast. Multi-field judgments about context are where a router earns its keep, and where you need more cases to tell routers apart.
Cost: not close
Here the routers separate by orders of magnitude:
- Jev: $0.0337 per 1,000 decisions. About 803 input tokens per decision, and Jev output is free at its list price.
- Claude Sonnet 5.5: $5.00 per 1,000. About 1,785 input and 107 output tokens per decision.
- Claude Haiku 4.5: $8.92 per 1,000. About 1,831 input and 1,419 output tokens per decision.
Haiku costs more than Sonnet here. That surprises people, but the reason is simple. Haiku ran with the CLI default extended thinking and wrote about 1,419 output tokens per decision. Sonnet ran at effort low, as production asks, and wrote about 107. Output tokens are the expensive ones.
Note that the Claude figures include the CLI's tool-schema and thinking tokens, because the CLI reports them. A direct API call would be cheaper. We estimate it would not close a gap of two orders of magnitude, but we did not measure it.
Latency
- Claude Sonnet 5.5: median model time 1.6 s, 95th percentile 2.57 s. Wall time through the CLI: median 2.6 s.
- Claude Haiku 4.5: median model time 10.7 s, 95th percentile 32.07 s. Wall time: median 12.67 s.
- Jev 1.13: median wall time 0.14 s (136.5 ms), 95th percentile 0.2 s (195.7 ms), over 246 calls in a 35-second window. We called its API directly from one Mac over a home network, so the network is inside that time, and the API reports no server time.
These are different routes. The Claude routers went through the CLI, which adds about a second per Sonnet decision; Jev was a direct call. The chart shows what a caller waits per decision, not model compute time. For the full delay picture, including a rule-based router at 1.42 µs per decision and the CLI share of each Claude decision, see what does a router cost you?
A router that takes 10 seconds is a problem when it sits in front of every message. Haiku's thinking default made it both the slowest and the most expensive router in this test.
So which router should you use?
On this evidence:
- If the router runs on every message or every step, cost and latency dominate. A dedicated router like Jev ties on accuracy at a tiny fraction of the cost.
- If a decision is rare and hard, such as a complex context judgment, Sonnet at low effort is the strongest point estimate, and it is fast.
- Do not assume the small general model is the cheap option. Under CLI defaults, Haiku thought its way to the highest cost and the slowest time.
There is a second question: once you have a router, does routing work to different models save money? We ran that as a thought experiment on 2,362 recorded calls, in what if every call ran on Opus?. Short version: on our pipeline, the main coding stages hold most of the spend, so routing side jobs to a cheaper model saved only 2.8%.
How we measured
- Cases. The labelled decision suites the platform uses, limited to the cases production asks a router about. 82 cases, 194 scored questions.
- Scoring. The repository's own decision-eval runner, the same for every router.
- Jev. A live run on 2026-10-06: the same 82 decisions, 3 repeats, 246 counted calls, one at a time over HTTPS from one Mac, scored by the same runner. Its interval is taken at the 82 cases because repeats are not independent. Its earlier recorded production run scored 74 of 82 with the same case versions. McNemar needs one pass, so it uses that recorded run.
- Claude routers. Run through the Claude Code CLI with the production system text and schema, one call per decision. Haiku 4.5 used the CLI default extended thinking. Sonnet 5.5 used effort low.
- Cost. List price × reported tokens per decision.
- Statistics. 95% Wilson intervals and an exact McNemar test on paired cases.
Caveats
- Home advantage. The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05.
- Jev's timing is one window. 246 calls from one machine on a home network, in 35 seconds. Server load then is unknown, and a caller near the API would likely see less.
- CLI overhead. The Claude routers ran through a CLI, which adds start-up time and tool-schema tokens a direct API call would not.
- One sample per decision. Production asks a second sample when confidence is low, so confidence-gated coverage is not compared here.
- Small sets. 82 cases in four hand-labelled sets give wide intervals.
- Not measured. A local Clef / Clef-Flash router was not run, because no local server was running and installing a 6 to 20 GB model was out of scope.
What to read next
- Jev vs Claude Haiku vs Claude Sonnet as a router: an honest comparison
- What does a router cost you? Rules vs Jev vs an LLM router
- What if every call ran on Opus? A thought experiment
- Claude Haiku vs Sonnet vs Opus vs Fable vs Codex, head to head
- Why we count every failed attempt
Routing you can inspect
Agent uses a rules-first decision cascade for incoming work. Its records show the selected answer and fallback when a decision ran. Model-call costs can be list-price estimates; they are not verified invoices. Try Agent and inspect the decisions recorded for a task.