Jev vs Claude Haiku vs Claude Sonnet as a router: an honest comparison
Every row of our Jev, Haiku and Sonnet router comparisons: accuracy ties, Jev is far cheaper and quicker per call over its API, on a different route.
TL;DR
- Accuracy: a tie. On 82 typed routing decisions, Jev 1.13 got 221 of 246 live calls exactly right (74, 73 and 74 of 82 in its three repeats). Claude Haiku 4.5 got 73 of 82 and Claude Sonnet 5.5 77. All 95% intervals overlap. Jev's interval is 82% to 95%, taken at the 82 cases because the repeats are not independent.
- Cost: a gap of two orders of magnitude. Jev $0.0337 per 1,000 decisions (a calculation from the input tokens its API reported), Sonnet $5.00 and Haiku $8.92 (list-price calculations). That is about 148x and 265x. The gap has no interval, so the comparison pages mark it "unclear", but it is far larger than the cost of a few extra errors.
- Speed: Jev's call took a median 136.5 ms. Sonnet took 2.60 s and Haiku 12.54 s through the Claude Code CLI: about 19 times and 92 times longer (a calculation). This is a gap between routes, not only between models: Jev was a direct HTTPS call, the Claude routers went through a CLI.
- Home advantage: the case sets were tuned against Jev answers. Read Jev's accuracy with that in mind.
Pair pages: Jev vs Haiku, Jev vs Sonnet and Haiku vs Sonnet.
Why another post on this
Our first routing post told the story of the run. Since then, the routing overhead study added timing breakdowns and per-task costs, and a live run timed Jev for the first time: 246 calls over HTTPS on the same 82 decisions. The comparison pages now line up every shared metric for each pair. This post reads those pages row by row. It says what each row shows and what it does not.
How a comparison row gets a winner
The comparison pages follow fixed rules:
- A rate row has a winner only when the two 95% intervals do not overlap.
- A timing row has a winner only when one side's median sits outside the other's band (95th percentile or range), with enough runs.
- A cost row with no interval or range on either side is "unclear": the gap is stated, not tested against run-to-run variation.
- A calculation row is labelled as one.
- When the two sides ran through different routes, the page says so, and a winning row says that it compares routes, not models alone.
So "tie" means "this sample cannot separate them", not "they are equal".
Accuracy: every row is a tie
Jev vs Claude as a router: accuracy and cost
Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.
Transcript
- Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.
- Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
- Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
- Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
- Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
- Open benchmarks: intervals, sources and every failure kept.
| Metric | Jev 1.13 (live, 3 repeats) | Claude Haiku 4.5 | Claude Sonnet 5.5 |
|---|---|---|---|
| Exact decisions (n = 82) | 90% (221 of 246 calls), 82% to 95% | 89% (73), 80% to 94% | 94% (77), 87% to 97% |
| Per question (n = 194) | 95% (552 of 582 answers), 91% to 97% | 94% (183), 90% to 97% | 97% (189), 94% to 99% |
| Failure class (n = 18) | 18 in every repeat | 17 | 18 |
| Message intent (n = 20) | 20 in every repeat | 20 | 20 |
| Is it a rule? (n = 12) | 12 in every repeat | 12 | 12 |
| Context shape (n = 32) | 24, 23 and 24 | 24 | 27 |
Jev ran the same 82 decisions three times. The recorded production run of 2026-10-05 scored 74 of 82, the same as repeats 1 and 3. Of the 82 decisions, 73 were exact in every repeat, 8 in none and 1 in some. That one, a context-shape case, flipped in repeat 2.
Why the interval is 82% to 95% and not tighter: 246 calls look like a bigger sample than 82, but the repeats ask the same questions, so they are not independent draws. Pooling them as 246 would give a falsely narrow interval. We take Jev's pooled rate at 82 cases instead.
Sonnet has the highest or joint-highest point estimate on every line where anyone missed. With these sample sizes, that is not enough. An exact McNemar test on the paired cases agrees: Haiku vs Jev p = 1, Sonnet vs Jev p = 0.375 (from the routing study, on Jev's recorded production run: the test needs one pass per router).
Honest reading: all three routers are good at the three simple decision types. The only place they differ, even on point estimates, is context shape, where several questions are asked at once. If your router mostly answers multi-field context questions, Sonnet's 27 of 32 is the best result we have, and it is still inside the others' intervals.
Completion: every row is a tie
Jev returned a decision on 246 of 246 calls, and Haiku and Sonnet on 82 of 82 (95% intervals 98% to 100% and 96% to 100%). None timed out or failed to parse.
Cost: large gaps, marked "unclear"
- Jev 1.13: $0.0337 per 1,000 decisions. A calculation: 803 input tokens per decision, as its API reported, × $0.042 per million input tokens. Its output tokens are free. The recorded production run's own provider-reported cost is the same figure.
- Claude Sonnet 5.5: $5.00 per 1,000, list price × reported tokens.
- Claude Haiku 4.5: $8.92 per 1,000, list price × reported tokens.
Why "unclear" and not "Jev wins"? Each figure is one total over the calls, with no interval. Our rules do not name a winner without one. The size of the gap still matters: at 148x, run-to-run variation would have to be enormous to close it.
Haiku costs more than Sonnet because it ran with the CLI's default thinking and wrote about 1,419 output tokens per decision. Sonnet ran at effort low and wrote about 107.
The per-task view makes the gap concrete (a calculation; see what a router costs):
| Added cost per 1,000 tasks | Jev 1.13 | Sonnet 5.5 | Haiku 4.5 |
|---|---|---|---|
| Every model call routed (49.5 per task) | $1.67 | $247.30 | $441.74 |
| Only System One decisions (7 per task) | $0.24 | $34.97 | $62.47 |
Speed: Jev's call is about 19 times shorter than Sonnet's
- Jev: median 136.5 ms per call, p95 195.7 ms, fastest 100.9 ms, slowest 297.3 ms (246 calls). The first call of the run, on a fresh connection, took 224.7 ms.
- Sonnet: median 2,597 ms per decision, p95 4,298 ms. Model time 1,596 ms; CLI time 973 ms.
- Haiku: median 12,543 ms, p95 34,481 ms. Model time 10,508 ms; CLI time 1,698 ms.
Both Claude medians sit above Jev's 95th percentile, so on Jev vs Sonnet and Jev vs Haiku Jev wins the decision-time row. Haiku's median also sits above Sonnet's 95th percentile, so on Haiku vs Sonnet Sonnet wins the routing timing rows. (That page also holds hard-task rows from another study, where Sonnet wins on pass rate.)
Read the Jev row with its route in mind. The comparison pages say it too:
- Jev was a direct HTTPS call from one Mac over a home network. The network is inside its 136.5 ms. The API sends no server time, so we cannot say how much of that is the model.
- The Claude routers ran through the Claude Code CLI. About a second of each Sonnet decision is CLI start-up and tool schema, not the model.
- Taking the CLI out does not close the gap: Sonnet's model time alone, 1,596 ms at the median, is still above Jev's whole call.
- Jev's run is one 35-second window, 3 repeats of 82 requests. Server load at that time is unknown. A caller near the API would likely see less.
Per task, if every decision waits in line (a calculation and an upper bound), routing all 49.5 model calls adds up to 6.76 s with Jev, 128.6 s with Sonnet and 620.9 s with Haiku. Routing only the 7 System One decisions adds 0.96 s, 18.2 s and 87.8 s.
What is still not known
- Where Jev's time goes. The API reports no server-side time. We measured what a caller sees from one machine.
- Jev from another place or time. One Mac, one home network, one 35-second window. A server in the same region, or another hour, can differ.
- Claude through the API. Both Claude routers ran through the Claude Code CLI, which added about 1 to 1.7 s per decision and its own tokens. A direct API router would be faster and cheaper. We did not measure it.
- More cases. 82 decisions in four hand-labelled sets give wide intervals, and Jev has a home advantage on them.
The bottom line
- If you route every call: Jev ties on accuracy at a fraction of the cost, and its measured call time leaves room on a hot path. Check it from where your code runs.
- If you route a few hard context decisions: Sonnet at low effort has the best point estimates and was fast for an LLM.
- Do not default to Haiku because it is "small". Under CLI defaults it was the slowest and most expensive router in this test, with no accuracy gain.
- For decisions rules can make, use rules. They cost $0 and take microseconds: Jev vs rules.
How we measured
- Cases: 82 typed routing decisions from the platform's decision suites, 194 scored questions.
- Scoring: the repository's own decision-eval runner, the same for every router.
- Claude routers: Claude Code CLI, production system text and schema, one call per decision. Haiku 4.5 with the CLI default thinking; Sonnet 5.5 at effort low.
- Jev: a live run on 2026-10-06. The same 82 decisions, 3 repeats, 246 counted calls, one at a time, over HTTPS from one Apple M3 Ultra Mac. Time is client wall time. A failed call would have counted as wrong; there were none. The protocol was written before the first counted call.
- Statistics: 95% Wilson intervals for rates (Jev's at the 82 cases); p50 and p95 for timings; an exact McNemar test on paired cases.
Caveats
- Home advantage. The case sets and question wording were revised against Jev answers on 2026-10-04 and 2026-10-05.
- Different routes. A direct API call and a CLI are not the same thing to time or to price. The Claude CLI tokens in the cost figures include its tool schema.
- Calculations. Jev's and the Claude costs and every per-task figure are list-price calculations, not bills.
- One sample per decision for the Claude routers. Production asks a second sample when confidence is low; that is not compared here.
What to read next
- What does a router cost you? Rules vs Jev vs an LLM router
- Jev vs Claude Haiku and Sonnet as a router: tied on accuracy
- OpenRouter vs going direct
Routing you can inspect
Agent uses typed decisions like these to handle incoming work, and it records each one with its cost. Try Agent and see which choices it made for your work, and why.