4 studies · 8 comparisons · updated October 6, 2026

Routing: what it costs to pick a model.

A router decides which model and effort each call gets. It can be a rule in the same process, a hosted decision model such as Jev, a large model asked for its opinion, or a gateway such as OpenRouter that sends one API to many providers. Here is what each one adds in time, cost and accuracy, with the sample size behind every number.

Log scale · axis 1 µs to

Motion reduced: press Replay zoom to animate

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.

about 96 thousand×Jev 1.13 (TypeSafe) takes about 96 thousand times as long as Deterministic routing policy (Agent, in process). Calculation: ratio of the two medians (137 ms ÷ 1.42 µs). Log scale: each gridline is 10 times the one before.

4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n 82–20000 per row

Median; whiskers = median to 95th percentile

The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
  • Rules in process

    Microseconds and no tokens. They only know what the rules know.

  • Decision models and LLM routers

    They read the task. Each decision costs a model call: seconds through a CLI.

  • Gateways and cheaper inference

    One API, many providers. Closed models cost the same per token everywhere; open-weight prices vary up to 12.6x between providers.

Live stories

Drawn live in the page from the same data. Play one, or download it as a video from the player.

Live story · 51 sRouting overhead: a 1.42 µs policy vs LLM routers

Routing overhead: a 1.42 µs policy vs LLM routers

A deterministic routing policy decides in 1.42 µs (median, n = 20,000); Sonnet 5.5 as a router takes 2.60 s through the CLI. Per-task costs and delays are calculations.

Transcript
  1. Routing overhead · policy vs LLM routers. What a routing decision costs before the work starts. Time and money per decision, then per task. Jev is timed over its API, a different route from the CLI routers.
  2. The latency ladder: the in-process policy decides in 1.42 µs. Jev, a direct API call, takes 137 ms. Sonnet 5.5 as a router takes 2.60 s through the CLI, Haiku 4.5 12.5 s. Chart: Time per decision · log scale · median to p95 (n = 82–20000 each). Caveat: Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.
  3. The LLM router takes about 1.8 million times longer per decision. Jev costs $0.0337 per 1,000 decisions, provider-reported. Policy decision, median (20,000 timed): 1.42 µs (n = 20000, p95 2.33 µs). Sonnet 5.5 router ÷ policy, medians: 1.8 million× (n = 82). Jev 1.13 per 1,000 decisions (provider-reported): $0.0337 (n = 82, 137 ms median per call over its API). Caveat: OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.
  4. Of Sonnet 5.5’s 2.60 s per decision, a median 973 ms is CLI and harness time, not the model. Chart: Where an LLM router’s time goes · median per call (n = 82 each). Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
  5. Route all 49.5 model calls per task: $1.67 per 1,000 tasks with Jev, $247.30 with Sonnet 5.5. Route 7 decisions: $34.97. Chart: Added routing cost per 1,000 tasks. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
  6. If each decision waits in line, Sonnet 5.5 adds up to 129 s per task, or 18.2 s for 7 decisions. The policy adds 70.3 µs. Chart: Added routing delay per task · upper bound. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.
  7. Start-up tax for a one-word answer: Claude Code 2.53 s and 6,761 input tokens; Codex CLI 6.00 s and 17,051 input tokens, 13,184 of them read from the cache. Chart: CLI start-up tax · one-word answer · median of 5 runs (n = 5 each). Caveat: CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.
  8. Route in process where a rule is enough. Every timing, cost and gap online.
Live story · 34 sJev vs Claude as a router: accuracy and cost

Jev vs Claude as a router: accuracy and cost

Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.

Transcript
  1. Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.
  2. Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
  3. Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.
  4. Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.
  5. Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.
  6. Open benchmarks: intervals, sources and every failure kept.

Accuracy: does the router pick right?

Exact decisions and key accuracy on the same 82 routing cases, with 95% Wilson intervals. Overlapping intervals mean the sample cannot separate the routers.

Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 82 per row

Share of asked cases where every scored question was acceptable

Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Every interval overlaps every other: this chart does not order these rows.

3 rows. Highest Claude Sonnet 5.5 97% (95% interval 94%–99%, n 194). Lowest Claude Haiku 4.5 94% (95% interval 90%–97%, n 194). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 194 per row

Each open question the router was asked; an unanswered question counts as wrong

Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.

Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)

Same cases, two routers

Each case was routed by both. Cases both got right, or both got wrong, say nothing about which is better; the exact McNemar test reads only the cases where they disagree.

Paired outcomes per case, with the exact McNemar test

Every case once: both right, only one right, both wrong

Only the 7 discordant decisions count: 3 vs 4. Exact McNemar p = 1: no evidence of a difference.

70Both right3Only Claude Haiku 4.5 right4Only Jev 1.13 (TypeSafe) right5Both wrong
Both right
70 85%Agree and right: says nothing about which is better.
Only Claude Haiku 4.5 right
3 4%Discordant: Claude Haiku 4.5 right where Jev 1.13 (TypeSafe) was wrong.
Only Jev 1.13 (TypeSafe) right
4 5%Discordant: Jev 1.13 (TypeSafe) right where Claude Haiku 4.5 was wrong.
Both wrong
5 6%Agree and wrong: says nothing about which is better.

One square per paired decision (n = 82). Squares start in one grid and sort into the four outcomes; their order inside a quadrant carries no meaning.

Counts of cases from the table; the exact McNemar p reads only the discordant casesn = 82 cases per pair

Claude Haiku 4.5 vs Jev 1.13 (TypeSafe): 3 vs 4 discordant cases of 82, exact McNemar p = 1. Claude Sonnet 5.5 vs Jev 1.13 (TypeSafe): 4 vs 1 discordant cases of 82, exact McNemar p = 0.375.

Cost per 1,000 decisions

Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)

2 rows. Highest Jev 1.13 (TypeSafe) $0.034 (n 82). Lowest Deterministic routing policy (Agent, in process) $0 (n 20000).

Notesn 82–20000 per row

USD per 1,000 decisions

The policy makes no model call, so it costs nothing per decision. Jev’s figure is the cost its provider reported for the recorded production run (82 decisions). The input tokens of the live run give the same figure at the published price. The Claude routers are in the next chart: their cost is derived from list prices.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev 1.13 list price

Share card (PNG)
Calculation
Largest value is 260x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).

Notesn 82–246 per row

List price × reported tokens per decision

List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)

Per task: what routing adds

Where a routed call’s time goes, then calculations from recorded decision counts per task (not runs). The delay is an upper bound: it assumes each decision waits for the one before.

  • Model API time
  • CLI and harness time
Entrance: medians race at 7.5× real timeMotion reduced: press Replay to animateThe slowest median is 10.51 s. The clock runs at the recorded speed.
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

2 rows, 2 series: Model API time, CLI and harness time. Model API time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 10.51 s (median to p95 10.51 s–32.13 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 1.6 s (median to p95 1.6 s–2.58 s, n 82). Not all run ranges overlap. CLI and harness time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 1.7 s (median to p95 1.7 s–2.68 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 973 ms (median to p95 973 ms–1.28 s, n 82). Not all run ranges overlap.

NotesLines: median to p95 (not an interval)n = 82 per row

Median per call; whiskers = median to 95th percentile

Model API time is the API duration the CLI reports; CLI and harness time is wall time minus that, per call. Medians of the parts do not add up to the median of the whole. Jev is not split: its API reports no server time. The whisker is the median to the 95th percentile, not a confidence interval.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing

Share card (PNG)
Calculations per task: added cost, added delay and the CLI start-up tax (3)
Calculation
  • Every model call routed (49.5 per task)
  • Only System One decisions (7 per task) (square)
In chart order.
Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).

List-price calculation, not a run. 4 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $442. Lowest Deterministic routing policy (Agent, in process) $0. Only System One decisions (7 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $62.47. Lowest Deterministic routing policy (Agent, in process) $0.

Notes

Decisions per task from recorded runs × cost per decision

A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
Calculation
  • Every model call routed (49.5 per task)
  • Only System One decisions (7 per task) (square)
In chart order.
Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).

List-price calculation, not a run. 4 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 621 s. Fastest Deterministic routing policy (Agent, in process) 70.3 µs. Only System One decisions (7 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 87.8 s. Fastest Deterministic routing policy (Agent, in process) 9.9 µs.

Notes

Decisions per task × median decision time, if every decision waits in line

A calculation and an upper bound: it assumes each decision waits for the one before. Median recorded task wall time: 10.3 min. Jev’s delay uses its live median over the API from one Mac (network included); the Claude routers’ includes the CLI.

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions

Share card (PNG)
  • First output event
  • First model output
  • Total wall time
Entrance: medians race at 4.3× real timeMotion reduced: press Replay to animateThe slowest median is 6 s. The clock runs at the recorded speed.
Claude Code · Claude Haiku 4.5
Codex CLI (default model)

2 rows, 3 series: First output event, First model output, Total wall time. First output event: slowest Claude Code · Claude Haiku 4.5 563 ms (range 519 ms–726 ms, n 5). Fastest Codex CLI (default model) 489 ms (range 354 ms–1.3 s, n 5). All run ranges overlap. First model output: slowest Codex CLI (default model) 5.06 s (range 4.39 s–5.48 s, n 5). Fastest Claude Code · Claude Haiku 4.5 1.46 s (range 1.21 s–2.31 s, n 5). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 5 per row

Median of 5 runs; whiskers = fastest and slowest run

Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.

Source: Routing overhead runs: policy microbenchmark and CLI start-up

Share card (PNG)

Estimate what serial waiting costs your team

Scale routing cost and waiting to your own decision volume

Thought experiments

Thought experiment: not a run. These charts reprice recorded routing decisions and tokens at list prices for whole workloads. No router was called again.

Calculation
all Fable 5.1
all Opus 5.5
policy (Opus strong, Haiku ancillary)
all Sonnet 5.5
split (Sonnet main line, Haiku ancillary)
all Haiku 4.5

Hover or focus a bar for its ratio to all Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 6 rows. Highest all Fable 5.1 $370. Lowest all Haiku 4.5 $54.27.

Notes

50 benchmark runs, 2,362 model calls, repriced

Calculation, not a run: every recorded call ran on Sonnet 5.5 with routing off. Same tokens on every model; a different model or mix would take a different path.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Anthropic list prices (Claude models)

Share card (PNG)
The same thought experiment by pipeline stage (1)
Calculation

Haiku 4.5

act (strong)
research (strong)
verify (strong)
review (strong)
memory and onboarding (economy)
other (standard)

Sonnet 5.5

act (strong)
research (strong)
verify (strong)
review (strong)
memory and onboarding (economy)
other (standard)

Opus 5.5

act (strong)
research (strong)
verify (strong)
review (strong)
memory and onboarding (economy)
other (standard)

Fable 5.1

act (strong)
research (strong)
verify (strong)
review (strong)
memory and onboarding (economy)
other (standard)

One panel per series, all on the same axis.

List-price calculation, not a run. 6 rows, 4 series: Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1. Haiku 4.5: highest act (strong) $28.66. Lowest other (standard) $0.1. Sonnet 5.5: highest act (strong) $57.31. Lowest other (standard) $0.19.

Notes

Calculation, not a run. The tier in brackets is the routing policy tier for that stage.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Anthropic list prices (Claude models)

Share card (PNG)

Every router: measured, reported or not measured

Unknown stays unknown. A router without a measured number says why.

Routers side by side

Exact decisions with the 95% interval, tokens, cost and decision time

  • in-process rules

    Deterministic routing policy (Agent, in process)

    —no accuracy run

    Decision time
    Measuredmeasured 2026-10-06: 1.42 µs median
    Cost
    Measured$0 (no model call)
  • hosted decision model

    Jev 1.13 (TypeSafe)

    74/82exact decisions

    90% · 95% CI 82%–95% · n = 82

    Key accuracy
    184/194
    Tokens per decision (in / out)
    803 / 148
    USD per 1,000 decisions
    $0.034 Calculation
    Decision time
    Measuredmeasured 2026-10-06: 137 ms median, p95 196 ms (246 calls over the API from one Mac; client wall time, no server time)

    live API run 2026-10-06: 3 repeats of the same 82 decisions, 246 counted calls one at a time from one Mac. All calls: 221 of 246 exact and 552 of 582 questions; the counts shown are on the 82-decision scale of the interval. The recorded production run of 2026-10-05 scored 74 of 82. Cost is a calculation from the reported input tokens

  • LLM router via agent CLI

    Claude Sonnet 5.5 (effort low, via Claude Code)

    77/82exact decisions

    94% · 95% CI 87%–97% · n = 82

    Key accuracy
    189/194
    Tokens per decision (in / out)
    1,785 / 107
    USD per 1,000 decisions
    $5.00 Calculation
    Decision time
    Measuredrecorded: 2.60 s median (82 calls)

    this benchmark, Claude Code CLI on a subscription account, one call per decision

  • LLM router via agent CLI

    Claude Haiku 4.5 (thinking on, via Claude Code)

    73/82exact decisions

    89% · 95% CI 80%–94% · n = 82

    Key accuracy
    183/194
    Tokens per decision (in / out)
    1,831 / 1,419
    USD per 1,000 decisions
    $8.92 Calculation
    Decision time
    Measuredrecorded: 12.54 s median (82 calls)

    this benchmark, Claude Code CLI on a subscription account, one call per decision

  • hosted LLM router / gateway

    OpenRouter Auto Router / cheaper hosted inference

    —no accuracy run

    Decision time
    Not measuredno key in the environment. The harness is ready and runs when a key is set.
    Cost
    Not measuredno key in the environment. The harness is ready and runs when a key is set.
  • local router model

    Clef / Clef-Flash (local)

    —exact decisions: not measured

    Decision time
    Not measuredno local server running.
    Cost
    Not measuredno local server running.

    not measured: No local Clef server was running and installing a 6-20 GB model was out of scope for this run.

Meter: 95% Wilson interval on exact decisions (82 cases)Each value is measured, reported, a calculation or not measured, as marked

6 routers. Decision time and cost are marked measured, reported, a calculation or not measured; exact decisions carry their 95% interval.

Gateways and cheaper inference

Reported by OpenRouter’s public API. The gateway’s own delay is not measured: no key in the environment.

Calculation

0%per-token markup on 10 of 10 models

5.5% with the card credit feeCalculation

  • Claude Haiku 4.5
  • Claude Sonnet 5
  • Claude Sonnet 5.5
  • Claude Opus 4.8
  • Claude Opus 5
  • Claude Opus 5.5
  • Claude Fable 5.1
  • GPT-6 Luna
  • Gemini 3.8 Flash
  • Gemini 3.5 Flash
  • per-token markup over the first-party list price (input and output)
  • With the 5.5% card credit fee (input) (calculation, hollow)

List-price calculation, not a run. 10 rows, 3 series: Per-token markup (input), Per-token markup (output), With the 5.5% card credit fee (input). Per-token markup (input): all at 0%. Per-token markup (output): all at 0%.

Notes

Per-token markup, and the markup after the Standard credit-purchase fee (calculation)

A calculation: OpenRouter list price ÷ first-party list price − 1, then × (1 + 5.5%) for credits bought by card on the Standard plan (purchases large enough that the minimum fee does not apply). Fees as published on 2026-10-06; first-party prices from the vendors’ list prices.

Sources: OpenRouter public API: models and provider endpoints (snapshot), OpenRouter pricing and fees, Anthropic list prices (Claude models), OpenAI list prices, Google Gemini list prices

Share card (PNG)
Calculation
ModelSpread
  1. DeepSeek V4 Flash 0423n = 15 providers
  2. DeepSeek V4 Pro 0423n = 15 providers
  3. gpt-oss-120bn = 20 providers
  4. Llama 3.3 70B Instructn = 10 providers
  5. GLM 5.3n = 32 providers
  6. Kimi K3n = 19 providers
  7. Llama 4 Maverickn = 3 providers
  8. 15 more models: the same price at every providerClaude Haiku 4.5, Claude Sonnet 5, Claude Sonnet 5.5, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Fable 5.1, GPT-6 Sol, GPT-6 Luna, GPT-6 Astra, GPT-5.5, Gemini 3.8 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.1 Pro Preview
    1x

List-price calculation, not a run. 22 rows. Highest DeepSeek V4 Flash 0423 12.6x (n 15). Lowest Gemini 3.1 Pro Preview 1x (n 2).

Notesn 2–32 per row

Standard tier, blended price (3 input : 1 output), most expensive provider ÷ cheapest provider

A calculation on prices reported by OpenRouter’s public API, snapshot 2026-10-06. n = providers with a standard-tier endpoint. 1x means every provider charges the same. Highlighted: a spread of 4x or more.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)

Every provider, priced by model Price a workload via the cheapest provider

Router and gateway pages

  • Router · Agent

    Deterministic routing policy

    Agent’s rule-based routing: an in-process policy picks the model and effort for each call from the task stage and signals. No model call, so no token cost.

    • Routing calls that returned a decision100% (20000/20000)n=20000
    • Time to make one routing decision1.42 µsn=20000
    • Cost per 1,000 routing decisions: no model call vs provider-reported$0.00n=20000
  • Router · TypeSafe

    Jev 1.13

    A small routing model from TypeSafe that answers typed routing decisions over an API. Priced on input tokens only. Measured here as a router (accuracy, live time per call and cost per 1,000 decisions) and in the System One arena.

    • Who decides right? Accuracy on 1,000+ checkable decisions77% (805/1048)n=1048
    • Where each model is strong: Games42% (81/192)n=192
    • Where each model is strong: Logic and thought experiments74% (116/156)n=156
  • Gateway · OpenRouter

    OpenRouter

    A gateway that routes one API to many inference providers. Its per-token price is compared with first-party list prices; its fee is charged when credits are bought.

    • Claude Haiku 4.5: OpenRouter vs Anthropic list price (Input)$1.00
    • Claude Haiku 4.5: OpenRouter vs Anthropic list price (Output)$5.00
    • Claude Sonnet 5: OpenRouter vs Anthropic list price (Input)$2.00

Head to head

All comparisons

A side wins a row only when the 95% intervals or run ranges do not overlap. Price rows never name a winner. The bar shows wins, ties and unclear rows.

The studies

Live story
  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

Live storyIncludes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Live story
  • Inference
  • Providers

Inference provider index: 27 models, 52 providers

Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.

265endpoints, 52 providers, 27 models · Provider endpoints in the snapshot · n = 27

34 chartsUpdated October 6, 2026

Live story
  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.