• Routing
  • Model Routing
  • Jev
  • Claude Haiku

Jev vs Claude Haiku vs Claude Sonnet as a router: an honest comparison

Every row of our Jev, Haiku and Sonnet router comparisons: accuracy ties, Jev is far cheaper and quicker per call over its API, on a different route.

TL;DR

  • Accuracy: a tie. On 82 typed routing decisions, Jev 1.13 got 221 of 246 live calls exactly right (74, 73 and 74 of 82 in its three repeats). Claude Haiku 4.5 got 73 of 82 and Claude Sonnet 5.5 77. All 95% intervals overlap. Jev's interval is 82% to 95%, taken at the 82 cases because the repeats are not independent.
  • Cost: a gap of two orders of magnitude. Jev $0.0337 per 1,000 decisions (a calculation from the input tokens its API reported), Sonnet $5.00 and Haiku $8.92 (list-price calculations). That is about 148x and 265x. The gap has no interval, so the comparison pages mark it "unclear", but it is far larger than the cost of a few extra errors.
  • Speed: Jev's call took a median 136.5 ms. Sonnet took 2.60 s and Haiku 12.54 s through the Claude Code CLI: about 19 times and 92 times longer (a calculation). This is a gap between routes, not only between models: Jev was a direct HTTPS call, the Claude routers went through a CLI.
  • Home advantage: the case sets were tuned against Jev answers. Read Jev's accuracy with that in mind.

Pair pages: Jev vs Haiku, Jev vs Sonnet and Haiku vs Sonnet.

Why another post on this

Our first routing post told the story of the run. Since then, the routing overhead study added timing breakdowns and per-task costs, and a live run timed Jev for the first time: 246 calls over HTTPS on the same 82 decisions. The comparison pages now line up every shared metric for each pair. This post reads those pages row by row. It says what each row shows and what it does not.

How a comparison row gets a winner

The comparison pages follow fixed rules:

  • A rate row has a winner only when the two 95% intervals do not overlap.
  • A timing row has a winner only when one side's median sits outside the other's band (95th percentile or range), with enough runs.
  • A cost row with no interval or range on either side is "unclear": the gap is stated, not tested against run-to-run variation.
  • A calculation row is labelled as one.
  • When the two sides ran through different routes, the page says so, and a winning row says that it compares routes, not models alone.

So "tie" means "this sample cannot separate them", not "they are equal".

Accuracy: every row is a tie

Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 82 per row

Share of asked cases where every scored question was acceptable

Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Live story · 34 sJev vs Claude as a router: accuracy and cost

Jev vs Claude as a router: accuracy and cost

Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.

Transcript
  1. Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.
  2. Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
  3. Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
  4. Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
  5. Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
  6. Open benchmarks: intervals, sources and every failure kept.
MetricJev 1.13 (live, 3 repeats)Claude Haiku 4.5Claude Sonnet 5.5
Exact decisions (n = 82)90% (221 of 246 calls), 82% to 95%89% (73), 80% to 94%94% (77), 87% to 97%
Per question (n = 194)95% (552 of 582 answers), 91% to 97%94% (183), 90% to 97%97% (189), 94% to 99%
Failure class (n = 18)18 in every repeat1718
Message intent (n = 20)20 in every repeat2020
Is it a rule? (n = 12)12 in every repeat1212
Context shape (n = 32)24, 23 and 242427
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 97% (95% interval 94%–99%, n 194). Lowest Claude Haiku 4.5 94% (95% interval 90%–97%, n 194). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 194 per row

Each open question the router was asked; an unanswered question counts as wrong

Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Jev ran the same 82 decisions three times. The recorded production run of 2026-10-05 scored 74 of 82, the same as repeats 1 and 3. Of the 82 decisions, 73 were exact in every repeat, 8 in none and 1 in some. That one, a context-shape case, flipped in repeat 2.

Why the interval is 82% to 95% and not tighter: 246 calls look like a bigger sample than 82, but the repeats ask the same questions, so they are not independent draws. Pooling them as 246 would give a falsely narrow interval. We take Jev's pooled rate at 82 cases instead.

Sonnet has the highest or joint-highest point estimate on every line where anyone missed. With these sample sizes, that is not enough. An exact McNemar test on the paired cases agrees: Haiku vs Jev p = 1, Sonnet vs Jev p = 0.375 (from the routing study, on Jev's recorded production run: the test needs one pass per router).

Honest reading: all three routers are good at the three simple decision types. The only place they differ, even on point estimates, is context shape, where several questions are asked at once. If your router mostly answers multi-field context questions, Sonnet's 27 of 32 is the best result we have, and it is still inside the others' intervals.

Completion: every row is a tie

Jev returned a decision on 246 of 246 calls, and Haiku and Sonnet on 82 of 82 (95% intervals 98% to 100% and 96% to 100%). None timed out or failed to parse.

Cost: large gaps, marked "unclear"

Calculation
Largest value is 260x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).

Notesn 82–246 per row

List price × reported tokens per decision

List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions

  • Jev 1.13: $0.0337 per 1,000 decisions. A calculation: 803 input tokens per decision, as its API reported, × $0.042 per million input tokens. Its output tokens are free. The recorded production run's own provider-reported cost is the same figure.
  • Claude Sonnet 5.5: $5.00 per 1,000, list price × reported tokens.
  • Claude Haiku 4.5: $8.92 per 1,000, list price × reported tokens.

Why "unclear" and not "Jev wins"? Each figure is one total over the calls, with no interval. Our rules do not name a winner without one. The size of the gap still matters: at 148x, run-to-run variation would have to be enormous to close it.

Haiku costs more than Sonnet because it ran with the CLI's default thinking and wrote about 1,419 output tokens per decision. Sonnet ran at effort low and wrote about 107.

The per-task view makes the gap concrete (a calculation; see what a router costs):

Calculation
  • Every model call routed (49.5 per task)
  • Only System One decisions (7 per task) (square)
In chart order.
Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).

List-price calculation, not a run. 4 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $442. Lowest Deterministic routing policy (Agent, in process) $0. Only System One decisions (7 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $62.47. Lowest Deterministic routing policy (Agent, in process) $0.

Notes

Decisions per task from recorded runs × cost per decision

A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions

Added cost per 1,000 tasksJev 1.13Sonnet 5.5Haiku 4.5
Every model call routed (49.5 per task)$1.67$247.30$441.74
Only System One decisions (7 per task)$0.24$34.97$62.47

Speed: Jev's call is about 19 times shorter than Sonnet's

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

Time per decision · log scale: each gridline is 10 times the one before

4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–20000 per row

Median; whiskers = median to 95th percentile

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

  • Jev: median 136.5 ms per call, p95 195.7 ms, fastest 100.9 ms, slowest 297.3 ms (246 calls). The first call of the run, on a fresh connection, took 224.7 ms.
  • Sonnet: median 2,597 ms per decision, p95 4,298 ms. Model time 1,596 ms; CLI time 973 ms.
  • Haiku: median 12,543 ms, p95 34,481 ms. Model time 10,508 ms; CLI time 1,698 ms.

Both Claude medians sit above Jev's 95th percentile, so on Jev vs Sonnet and Jev vs Haiku Jev wins the decision-time row. Haiku's median also sits above Sonnet's 95th percentile, so on Haiku vs Sonnet Sonnet wins the routing timing rows. (That page also holds hard-task rows from another study, where Sonnet wins on pass rate.)

Read the Jev row with its route in mind. The comparison pages say it too:

  • Jev was a direct HTTPS call from one Mac over a home network. The network is inside its 136.5 ms. The API sends no server time, so we cannot say how much of that is the model.
  • The Claude routers ran through the Claude Code CLI. About a second of each Sonnet decision is CLI start-up and tool schema, not the model.
  • Taking the CLI out does not close the gap: Sonnet's model time alone, 1,596 ms at the median, is still above Jev's whole call.
  • Jev's run is one 35-second window, 3 repeats of 82 requests. Server load at that time is unknown. A caller near the API would likely see less.

Per task, if every decision waits in line (a calculation and an upper bound), routing all 49.5 model calls adds up to 6.76 s with Jev, 128.6 s with Sonnet and 620.9 s with Haiku. Routing only the 7 System One decisions adds 0.96 s, 18.2 s and 87.8 s.

What is still not known

  • Where Jev's time goes. The API reports no server-side time. We measured what a caller sees from one machine.
  • Jev from another place or time. One Mac, one home network, one 35-second window. A server in the same region, or another hour, can differ.
  • Claude through the API. Both Claude routers ran through the Claude Code CLI, which added about 1 to 1.7 s per decision and its own tokens. A direct API router would be faster and cheaper. We did not measure it.
  • More cases. 82 decisions in four hand-labelled sets give wide intervals, and Jev has a home advantage on them.

The bottom line

  • If you route every call: Jev ties on accuracy at a fraction of the cost, and its measured call time leaves room on a hot path. Check it from where your code runs.
  • If you route a few hard context decisions: Sonnet at low effort has the best point estimates and was fast for an LLM.
  • Do not default to Haiku because it is "small". Under CLI defaults it was the slowest and most expensive router in this test, with no accuracy gain.
  • For decisions rules can make, use rules. They cost $0 and take microseconds: Jev vs rules.

How we measured

  • Cases: 82 typed routing decisions from the platform's decision suites, 194 scored questions.
  • Scoring: the repository's own decision-eval runner, the same for every router.
  • Claude routers: Claude Code CLI, production system text and schema, one call per decision. Haiku 4.5 with the CLI default thinking; Sonnet 5.5 at effort low.
  • Jev: a live run on 2026-10-06. The same 82 decisions, 3 repeats, 246 counted calls, one at a time, over HTTPS from one Apple M3 Ultra Mac. Time is client wall time. A failed call would have counted as wrong; there were none. The protocol was written before the first counted call.
  • Statistics: 95% Wilson intervals for rates (Jev's at the 82 cases); p50 and p95 for timings; an exact McNemar test on paired cases.

Caveats

  • Home advantage. The case sets and question wording were revised against Jev answers on 2026-10-04 and 2026-10-05.
  • Different routes. A direct API call and a CLI are not the same thing to time or to price. The Claude CLI tokens in the cost figures include its tool schema.
  • Calculations. Jev's and the Claude costs and every per-task figure are list-price calculations, not bills.
  • One sample per decision for the Claude routers. Production asks a second sample when confidence is low; that is not compared here.

Routing you can inspect

Agent uses typed decisions like these to handle incoming work, and it records each one with its cost. Try Agent and see which choices it made for your work, and why.

The data behind this post

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.