Calculation · dated routing observations

LLM routing: what choosing adds.

Enter your decision volume and a timing cap. Compare the model-call cost and serial waiting that the recorded routes would add.

A calculation on recorded routes, not a forecast of quality, queue throughput or an SLA. Direct HTTPS and Claude Code CLI timings include different overhead.

Example input

Your assumption. The study timed the in-process policy but did not test its accuracy on these labelled cases.

Calculation

Lowest recorded cost basis within your cap

Jev 1.13 (TypeSafe)

Jev 1.13 (TypeSafe)

Direct HTTPS from one Mac; network included

Cost/day
$0.0337
Cost/month
$0.7077
Serial hours/month: median / p95-per-call
0.7963 h / 1.1416 h

Use this as a starting point for your own trial. It does not rank quality. Direct HTTPS and CLI observations include different overhead, and future timing can differ.

Added cost and serial waiting

1,000 decisions/day × 21 days. Waiting assumes every decision runs in line. The p95 column is the sum if every call took its recorded p95, not a monthly percentile.

All router calculations
Calculations on the same dated per-decision observations
Router and routeCost/dayCost/monthSerial hours/month: median / p95-per-callRecorded cap check
Deterministic routing policy (Agent, in process)In-process rules; no model call$0$00 h / 0 hWithin observed capOnly eligible if your rules cover these decisions; accuracy was not tested.
Jev 1.13 (TypeSafe)Direct HTTPS from one Mac; network included$0.0337$0.70770.7963 h / 1.1416 hWithin observed capWithin the recorded cap; cost basis available.
Claude Sonnet 5.5 (effort low, via Claude Code)Claude Code CLI; effort low$4.996$104.91615.1492 h / 25.0717 hOver observed capThe recorded timing observation exceeds your cap.
Claude Haiku 4.5 (thinking on, via Claude Code)Claude Code CLI; thinking on$8.924$187.40473.1675 h / 201.1392 hOver observed capThe recorded timing observation exceeds your cap.
OpenRouter Auto Router / cheaper hosted inferenceHosted router / gateway; not measured in these studiesNo estimateNo estimateNo estimate / No estimateNot measuredNot measured in these studies.
Clef / Clef-Flash (local)Local router; not measured in these studiesNo estimateNo estimateNo estimate / No estimateNot measuredNot measured in these studies.

The recorded inputs

Updated 2026-10-06. Sample sizes and the transport belong to each value. Intervals describe uncertainty in the labelled study cases; they do not establish quality for your decisions.

Recorded routing inputs
Recorded timing, qualified cost basis and exact-decision accuracy
RouterRecorded median / p95Cost basis per 1,000Exact decisions, 95% interval
Deterministic routing policy (Agent, in process)In-process rules; no model call0.00142 ms / 0.00233 msRecorded · n=20000measured 2026-10-06: 1.42 µs median$0No model-call charge; host compute not priced · n=20000$0 (no model call)Not measured
Jev 1.13 (TypeSafe)Direct HTTPS from one Mac; network included136.5 ms / 195.7 msRecorded · n=246measured 2026-10-06: 137 ms median, p95 196 ms (246 calls over the API from one Mac; client wall time, no server time)$0.0337List-price calculation · 2026-09-23 · n=246$0.0337 per 1,000 (list-price calculation from input tokens; the recorded run’s own cost figure agrees)89.84% · 81.91–94.97%Recorded · effective n=82
Claude Sonnet 5.5 (effort low, via Claude Code)Claude Code CLI; effort low2,597 ms / 4,298 msRecorded · n=82recorded: 2.60 s median (82 calls)$4.996List-price calculation · 2026-09-21 · n=82$4.996 per 1,000 (list-price calculation)93.90% · 86.51–97.37%Recorded · effective n=82
Claude Haiku 4.5 (thinking on, via Claude Code)Claude Code CLI; thinking on12,543 ms / 34,481 msRecorded · n=82recorded: 12.54 s median (82 calls)$8.924List-price calculation · 2026-09-21 · n=82$8.924 per 1,000 (list-price calculation)89.02% · 80.44–94.12%Recorded · effective n=82
OpenRouter Auto Router / cheaper hosted inferenceHosted router / gateway; not measured in these studiesNot measuredNot measured: no key in the environment. The harness is ready and runs when a key is set.Not measuredNot measured: no key in the environment. The harness is ready and runs when a key is set.Not measured
Clef / Clef-Flash (local)Local router; not measured in these studiesNot measuredNot measured: no local server running.Not measuredNot measured: no local server running.Not measured

Sources and limits

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

A calculation, not a bill: the Claude calls ran on a subscription. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free). Same tokens and prices as the routing study.

  • The policy is timed in process and the LLM routers through a CLI: this compares the two ways of routing as deployed, not two models on equal footing. A direct API call would skip the CLI time (shown separately).
  • The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
  • Haiku 4.5 ran with the CLI’s default thinking, which makes it slower than Sonnet 5.5 at effort low here.
  • OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
  • CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
  • Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
  • Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
  • Jev’s timing is one 35-second window on 2026-10-06, 246 calls from one machine; server load at that time is unknown. The 246 calls are 3 repeats of 82 requests, so they are not independent draws.
  • Corrected 2026-10-07: an earlier version of this page counted the cached tokens twice (30,235). The CLI reports 17,051 input tokens including 13,184 read from the cache.
  • The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
  • Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
  • Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
  • One sample per decision; production asks a second sample when confidence is low. Confidence-gated coverage is therefore not compared.
  • 82 cases in four small hand-labelled sets: intervals are wide.
  • Clef / Clef-Flash (local): not measured (No local Clef server was running and installing a 6-20 GB model was out of scope for this run).
  • Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
  • Jev ran as a direct HTTPS call; the Claude routers ran through the Claude Code CLI. These are different routes, so speed and cost compare what a caller pays per decision, not one model against the other.
  • Jev’s latency is one 35-second window from one Mac over a home network. The API reports no server time. A caller near the API would see less.

Questions

How do I choose an LLM router?

Start with routes that have recorded cost, timing and exact-decision evidence. This tool identifies the lowest recorded cost basis within your selected observed timing cap. Trial it on your own decisions before adopting it.

Does the p95 cap guarantee response time?

No. It filters a recorded per-call p95 from one study. Future load, network and requests can differ. Multiplying that p95 is a serial-wait scenario, not the p95 of a day or month.

Why does the in-process policy have no accuracy?

These studies timed the policy but did not test its accuracy on the model routers’ labelled decisions. It is only a candidate when you state that rules cover your decisions. Zero means no model-call charge; host compute is not priced.

Are the calculated costs vendor bills?

No. Model-router costs use the study’s dated list-price calculations. CLI calls ran on subscriptions. Missing prices and timings stay unknown, and no router is called by this tool.

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.