How fast is Jev? 136.5 ms per routing decision, measured live
Jev router latency over direct HTTPS: median 136.5 ms, p95 195.7 ms, range 100.9–297.3 ms, n = 246. Claude routers used a different CLI route.
TL;DR
- Jev 1.13 took a median 136.5 ms per routing decision (p95 195.7 ms; range 100.9–297.3 ms, n = 246 calls, 0 errors). We called it live over HTTPS, one call at a time, on the 82 typed decisions of our routing study, 3 times each.
- Claude routers took seconds. Claude Sonnet 5.5 through the Claude Code CLI took a median 2.60 s (p95 4.30 s; range 1.993–5.583 s, n = 82). Claude Haiku 4.5 with its default thinking took 12.54 s (p95 34.48 s; range 5.857–51.278 s, n = 82). The routes differ: Jev was a direct HTTPS call, the Claude routers ran through a CLI.
- Accuracy counts were similar. Jev answered 74, 73 and 74 of 82 decisions exactly in the three reps. The recorded run had 74 of 82. Pooled: 221 of 246 (89.8%; approximate 95% interval 81.9% to 95.0% over 82 cases).
- Answers were stable. 73 of 82 decisions gave the same answer choices in all 3 reps (95% Wilson interval 80.4–94.1%).
- Cost stays tiny. About $0.034 per 1,000 decisions (a calculation from reported tokens and the published price, n = 246), against $7.32 for Sonnet and $8.92 for Haiku (list-price calculations, n = 82 each).
- Rules are still faster. Our pure in-process routing policy takes a median 1.42 µs (p95 2.33 µs; maximum 2,538.21 µs, n = 20,000). It excludes database reads and record writes.
How fast is Jev?
Jev 1.13 from TypeSafe answered a typed routing decision in a median 136.5 ms. The 95th percentile was 195.7 ms. The fastest call took 100.9 ms and the slowest took 297.3 ms. All 246 calls returned HTTP 200 (95% Wilson interval for completion: 98.5% to 100%). The time range is not a confidence interval.
The first call of the run took 224.7 ms. It opened a new connection (DNS, TCP and TLS). Without it, the median was 136.4 ms and the p95 was 193.2 ms (range 100.9–297.3 ms, n = 245). That call ran about 88 ms above the warm median (calculation: 224.7 − 136.4). An uncounted probe call, on its own fresh connection, took 186.5 ms. Two fresh connections are too few to size the cold cost, so read them as examples.
The three reps then gave similar times: medians of 130.3 ms, 141.7 ms and 137.1 ms, with p95s of 179.8, 195.6 and 192.2 ms. For these timing summaries, rep 1 leaves out the cold first call (n = 81; range 109.6–232.6 ms). Reps 2 and 3 each have n = 82, with ranges of 108.2–297.3 ms and 100.9–239.0 ms. These ranges overlap. The routing overhead study shows the pooled live latency. The live summary lists each rep.
The whisker on the chart runs from the median to the 95th percentile. It is not a confidence interval.
How does that compare with Claude as a router?
On decision time, the gap is more than tenfold. Each ratio below is a calculation on the medians:
| Router | Route | Median per decision | p95 | Observed range (not an interval) | n | Median ÷ Jev's median (calculation) |
|---|---|---|---|---|---|---|
| Jev 1.13 | direct HTTPS, live run | 136.5 ms | 195.7 ms | 100.9–297.3 ms | 246 | 1x |
| Claude Sonnet 5.5, effort low | Claude Code CLI | 2,597 ms | 4,298 ms | 1,993–5,583 ms | 82 | 19.0x |
| Claude Haiku 4.5, thinking on | Claude Code CLI | 12,543 ms | 34,481 ms | 5,857–51,278 ms | 82 | 91.9x |
Even the slowest Jev call (297.3 ms) took less time than the fastest Sonnet call (1,993 ms) and the fastest Haiku call (5,857 ms). That holds under any percentile rule.
The routes are not the same, so read this as two ways to deploy a router, not as two models on equal footing. Part of each Claude decision is CLI time. For Sonnet, the CLI reported a median 1,596 ms of model API time and 973 ms of CLI and harness time (n = 82 each). Their ranges were 1,060–4,757 ms and 826–3,673 ms. Medians of parts do not add up to the median of the whole. Model API time alone is still 11.7 times Jev's median wall time (calculation), and Jev's time includes the network round trip from our Mac.
On the comparison pages, the routing-overhead decision-time row puts Jev ahead. The Claude medians sit above Jev's 95th percentile, so the p50–p95 bands do not overlap. These bands are not confidence intervals. Each page names the two routes. See Jev vs Claude Sonnet 5.5 and Jev vs Claude Haiku 4.5.
Is Jev the fastest LLM router?
Jev had the lowest observed decision time of the three model routers we timed, across different routes. Jev's 95th percentile (195.7 ms) is below the Sonnet median (2,597 ms), and the Haiku median is higher still. The deterministic routing policy is faster than all three, at a median 1.42 µs (p95 2.33 µs; maximum 2,538.21 µs, n = 20,000). It makes no model call and excludes database reads and record writes. We have not timed OpenRouter's Auto Router, or a Claude router through the direct API. Until we do, "fastest" holds only for the routers on this page, and only across two different routes.
What does the delay add up to per task?
This part is a calculation, not a run. In 48 recorded bench tasks, an agent made a median 49.5 model calls per task (range 13–73). Two empty runs were excluded. Routing was off, so these are hypothetical router decisions. If a router decides before each call, and each decision waits for the one before, the added delay per task is:
- Jev: about 6.76 s (49.5 × 136.5 ms).
- Claude Sonnet 5.5 through the CLI: about 128.6 s (49.5 × 2,597 ms).
If a router decides only at the median 7 System One decision points per task (range 2–24 across the 48 tasks), the added delay drops to about 0.96 s with Jev and 18.2 s with Sonnet. The median recorded task took 10.3 minutes (range 1.32–40.79 minutes, n = 48). These calculations multiply medians; they are neither measured task delays nor upper bounds. Parallel work can hide delay, while slow decisions can add more.
Did accuracy hold in the live run?
The counts were similar. The live reps scored 74, 73 and 74 of 82 decisions exactly. Their 95% Wilson intervals were 81.9–95.0%, 80.4–94.1% and 81.9–95.0%. The recorded production run scored 74 of 82 (95% Wilson interval 81.9–95.0%). The routing accuracy chart now shows the pooled live rate. A difference of one decision is inside the spread between the live reps.
Pooled over the three reps, Jev was exactly right on 221 of 246 calls (89.8%). The reps repeat the same 82 cases, so the dataset reports an approximate 95% Wilson interval over 82 cases: 81.9% to 95.0%. It rounds the pooled rate to 74/82 before computing the interval; it is not an interval from 246 independent cases. Per question, Jev was right on 552 of 582 scored keys (94.8%; approximate 95% Wilson interval 90.8–97.2% at 194 distinct keys). The interval rounds the pooled key rate to 184/194. Questions within one decision can also be related.
The Claude routers ran once each on the same cases: Sonnet 5.5, 77 of 82 (94%, 95% interval 87% to 97%), and Haiku 4.5, 73 of 82 (89%, 80% to 94%). The intervals overlap, so accuracy does not separate the three routers. Against Jev's recorded single pass, the exact McNemar test finds no significant difference (Sonnet p = 0.375; Haiku p = 1). This does not prove equal accuracy (routing study).
A caveat comes with every Jev accuracy number: home advantage. We revised the case sets and the question wording against Jev's answers on 2026-10-04 and 2026-10-05.
How stable are Jev's answers?
- 73 of 82 decisions gave the same answer choices in all 3 reps (95% Wilson interval 80.4–94.1%). Probabilities could still change.
- Of the 82 decisions, 73 were right in all 3 reps and 8 were wrong in all 3. 1 changed correctness. 9 changed at least one answer choice; these are different measures of stability.
- Three decision types hit a ceiling: failure class (54/54 calls, 18 cases), message intent (60/60, 20 cases) and is-it-a-rule (36/36, 12 cases). Their case-level 95% Wilson intervals are 82.4–100%, 83.9–100% and 75.8–100%. These small sets cannot show perfect accuracy on new cases.
- All the misses were in the context-shape set: 71 of 96 calls exact (74%; approximate case-level 95% Wilson interval 57.9–86.7%, n = 32 cases). That interval rounds the pooled rate to 24/32. The set asks about 4.5 scored questions per call (432 questions over 96 calls, calculation), against 1 in each of the other three sets.
A wrong answer that repeats is easier to find, and to cover with a rule, than one that changes each time. That is our view, not a measured result.
What does a Jev decision cost?
Cost correction: the routing accuracy study and its cost chart still use $5.00 for Sonnet, with the five-minute cache-write price. The routing receipts report $7.324 per 1,000 in Sonnet’s CLI cost field. The underlying call receipts report one-hour writes. The calculation below uses that tier, as does the current routing-at-scale study. We omit the older cost chart and story until those assets are corrected.
- Jev: $0.0337 per 1,000 decisions, a calculation from the provider-reported total cost of the recorded run (n = 82). The live run gives the same figure as a calculation: a mean of 802.8 reported input tokens per decision (n = 246; range 475–1,415) × $0.042 per million input tokens = $0.03372 per 1,000. Output tokens are free at the published price. The live API reported tokens but no cost field.
- Claude Sonnet 5.5: $7.32 per 1,000 (corrected list-price calculation on 82 calls), 217.2 times Jev's live calculated cost (calculation).
- Claude Haiku 4.5: $8.92 per 1,000 (list-price calculation on 82 calls), 264.7 times Jev's live calculated cost (calculation). Haiku used default thinking and reported more output tokens. This comparison does not isolate the effect of thinking.
Sonnet reported 127,257 one-hour cache-write tokens and no five-minute writes. We price each call at $2/M plain input, $0.20/M cache reads, $4/M one-hour writes and $10/M output tokens. Summing these costs and dividing by 82 calls gives $7.32368 per 1,000 decisions (calculation). The recorded-input table reports the rounded cost. These are measured token counts priced in a calculation; they are not measured bills.
Assuming 49.5 router decisions per task (the median of 48 tasks), per 1,000 tasks that is $1.67 with Jev and $362.52 with Sonnet (a calculation). Claude calls used subscriptions; these list-price costs are not invoices.
What we recommend
- Use rules first. The pure policy took microseconds with $0 model cost. Production database work adds time.
- Use a small decision model where judgment pays. The measured Jev route and calculated cost may fit your latency and cost budget. Check it on unseen decisions before routing every call.
- Test a direct API route before choosing a CLI router. Sonnet's median CLI and harness time was 973 ms (range 826–3,673 ms, n = 82). We did not measure Sonnet through its direct API.
- Measure the first call separately. The two fresh-connection calls we saw took 224.7 ms and 186.5 ms; the other 245 calls took a median 136.4 ms (range 100.9–297.3 ms). Two calls are not an estimate, so time the first call on your own route.
- Measure your own route. Our client sat on a home network. A server near the API can see a different time.
How we measured
- Protocol: the file was created at 20:11:41 UTC on 2026-10-06, before the probe at 20:11:46 and counted run at 20:12:02. It was updated after the run at 20:15:33 with results and a correction to estimated times. The current file is not an untouched pre-run record.
- Cases: the 82 typed decisions of the routing study (failure class, message intent, is-it-a-rule, context shape), the same case versions, 194 scored questions.
- Calls: 3 reps × 82 = 246 counted calls to Jev 1.13 (
jev-1.13.0), one at a time. Each call went straight to the hosted API over HTTPS from one Apple Silicon Mac on a home network. The run took 35 seconds. - Timing: client wall time from before the request to after the response body was read. The API reports no server-side timing. Jev's percentiles interpolate between ranks (the run summary's rule). The Claude rows in the routing overhead study use the nearest rank. The slowest-versus-fastest comparison above does not depend on the rule.
- Scoring: as in the routing study. A decision is exact when every scored question has an acceptable answer. The repository's own decision runner gives the same counts on the live answers.
- Cost: reported input tokens × the published price (a calculation).
- Claude rows: copied from the routing overhead and routing studies, recorded on 2026-10-05 through the Claude Code CLI. We did not re-run them.
The sanitized run summary is public: /benchmarks/raw/jev-live/summary.json.
Caveats
- Different routes. Jev was a direct HTTPS call; the Claude routers ran through a CLI. This compares deployments, not models on equal footing.
- One client, one network, one 35-second window. The time is wall time from a home network, not model compute time.
- Home advantage. The cases were revised against Jev's answers.
- Small sets. 82 hand-labelled decisions give wide intervals, and the 3 reps are not independent samples of new cases.
- Not measured: OpenRouter's Auto Router and Claude through its direct API.
Read next
- The routing hub: every router, its time, its cost and what is not measured.
- What does a router cost you? Rules vs Jev vs an LLM router, per task.
- Jev vs Claude Haiku and Sonnet as a router: accuracy decision by decision.
See your own routing decisions
Disclosure: I build Agent, the product behind these benchmarks. The decision cases come from Agent's own routing decisions.
Agent records each routing decision with its reason, its time and its cost. Try Agent and see which choices it made for your work.