OpenRouter vs going direct: what the gateway really costs
OpenRouter charged the vendor's per-token price on 10 of 10 models we checked. The cost is a 5.5% credit fee. What we know, and what is still unmeasured.
TL;DR
- We read OpenRouter's public, keyless API on 2026-10-06: 265 endpoints from 52 providers for 27 models. Every price here is reported by OpenRouter, not measured by us.
- For 10 of 10 models with a first-party list price in our price table, OpenRouter's per-token price was the same as the vendor's. Claude, GPT-6 Luna and Gemini: same input price, same output price.
- The gateway charges when you buy credits, not per request: 5.5% on the Standard plan by card ($0.80 minimum), 8% on Business and 5% by crypto (OpenRouter's pricing page, fetched 2026-10-06). So the effective markup on a card purchase is +5.5%.
- 15 of 15 closed models (Claude, GPT, Gemini) with 2 or more providers had one standard-tier price at every provider: Anthropic, Amazon Bedrock, Google Vertex, Azure and the others. Price differences come only from named tiers (flex, priority, fast) and regional endpoints.
- Not measured: the gateway's delay. The keyless API returned latency for 0 of 265 endpoints, and we had no key in the environment to time calls. The harness is ready and runs when a key is present.
Every price and endpoint: /benchmarks/inference-provider-index.
Same model, different price: 265 provider endpoints compared
Reported by OpenRouter’s public API on 2026-10-06: open-weight prices vary up to 12.6x across providers; closed models have one standard price, and the gateway adds a 5.5% credit fee.
Transcript
- Provider index · third-party-reported · 2026-10-06. Same model, different price. 265 endpoints, 52 providers, 27 models. Prices as OpenRouter’s public API reports them.
- 265 endpoints, 52 providers, 27 models. The keyless API returned latency for 0 of them, so speed is not compared. Provider endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Endpoints with a latency figure: 0 of 265. Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
- Open-weight models vary most: DeepSeek V4 Flash 12.6×, DeepSeek V4 Pro 11.2×. All 15 closed models with 2+ providers: one standard price (1×). Chart: Priciest ÷ cheapest standard provider · blended price (n = 2–32 each). Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
- One model, 10 providers: Llama 3.3 70B input costs $0.10 per million tokens at DeepInfra (fp8) and $1.04 at Together. Chart: Llama 3.3 70B · USD per million tokens · one row per provider. Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
- Gateway markup: for 10 of 10 models OpenRouter charges the vendor’s per-token price. The cost is a 5.5% fee when you buy credits. Models at the first-party per-token price: 10 of 10. Credit-purchase fee, Standard plan by card: 5.5% ($0.80 minimum by card). Effective markup per token after the fee: +5.5% (calculation). Caveat: First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.
- Example: Sonnet 5.5 costs $2.00 in and $10.00 out per million tokens on both. A price is not a measurement: no winner. Chart: Claude Sonnet 5.5 · OpenRouter vs Anthropic list price · USD per million. Caveat: Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.
- Prices change often: every endpoint, tier and source date online.
The question
A gateway gives you one API and one bill for many models and providers. What does that cost you compared with calling the vendor directly? Three parts can add cost:
- A per-token markup: the gateway lists a higher price than the vendor.
- A fee: the gateway charges on top, per request or per purchase.
- Delay: one more network hop before the provider starts work.
We can answer the first two from public data. The third needs live calls, and we did not make them.
Part 1: the per-token price is the same
For each model that has a first-party list price in our product price table, we divided OpenRouter's list price by the vendor's and subtracted 1. The result was 0% on input and 0% on output for all 10 models:
- Anthropic: Claude Haiku 4.5, Sonnet 5, Sonnet 5.5, Opus 4.8, Opus 5, Opus 5.5 and Fable 5.1.
- OpenAI: GPT-6 Luna.
- Google: Gemini 3.8 Flash and Gemini 3.5 Flash.
Two examples side by side:
Claude Sonnet 5.5: $2.00 in and $10.00 out per million tokens on both OpenRouter and Anthropic.
Gemini 3.8 Flash: $0.75 in and $3.75 out on both OpenRouter and Google AI Studio.
The comparison pages hold every row: OpenRouter vs Anthropic (14 rows, all the same price), OpenRouter vs OpenAI and OpenRouter vs Google AI Studio. A reported price has no interval, so these pages name no winner. They state the gap, and here the gap is zero.
Part 2: the fee is on credits, not on tokens
OpenRouter's pricing page says that inference is billed at the provider's list price and that the fee is charged when you buy credits:
- Standard plan, card: 5.5%, with a $0.80 minimum per purchase.
- Business plan: 8%.
- Crypto: 5%.
So on the Standard plan, $100 of tokens costs you $105.50 of card payments (our arithmetic). The $0.80 minimum matters for small top-ups: 5.5% of a purchase is less than $0.80 below about $14.55, so small purchases pay a higher percentage (our arithmetic).
That is why the markup chart has a third series. The per-token markup is 0%. With the card fee, the effective markup is +5.5% for every model.
Part 3: does the provider matter for closed models?
The spread is the most expensive standard-tier provider's blended price divided by the cheapest one's (3 input tokens to 1 output token, a calculation). For all 15 closed models with 2 or more providers, the spread was 1.0x: every standard-tier provider listed the same price.
Claude Opus 5.5 is a typical case: $4.00 in, $20.00 out and $0.20 for a cache read at Amazon Bedrock, Anthropic, Azure, Claude Platform on AWS and Google Vertex alike. The pairwise pages agree: Anthropic vs Amazon Bedrock, Amazon Bedrock vs Google Vertex, Amazon Bedrock vs Azure, Google AI Studio vs Google Vertex and OpenAI vs Azure.
So for a closed model, you pick a provider for reasons other than the standard list price: your cloud contract, data residency, rate limits, a named tier or a regional endpoint. Those reasons can be real. The list price is not one of them.
Open-weight models are a different story. Their prices differ by up to 12.6x between providers. That is the subject of the cheapest place to run open models.
What we do not know
We would rather show a gap than fill it:
- Gateway delay. The time to first token and the total time through OpenRouter vs direct are not measured: no key in the environment. A live harness is ready. It sends nothing without a key, and it has a hard spending cap.
- Billed cost per call. Also not measured, for the same reason. List price × tokens is a calculation, not a bill.
- Provider latency and throughput. OpenRouter's keyless API returned these for 0 of 265 endpoints. We cannot say whether a provider is faster or slower.
- Price drift. The first-party prices in our table were verified before the snapshot date. A vendor price change in between would show as a markup. None did.
- Prices outside OpenRouter. A provider's price on its own site can differ from its price through OpenRouter. We compared OpenRouter's listing with the vendor's published list price only.
So should you use a gateway?
On price alone, from this snapshot:
- For closed models, a gateway costs the credit fee, about 5.5% on the Standard plan, and nothing per token. If one API, one bill and access to many providers save you more than 5.5% of your token spend in engineering time, the gateway pays for itself. If you spend a lot on one vendor, going direct saves the fee.
- For open-weight models, the provider you pick matters far more than the gateway fee. A 5.5% fee is small next to a 6x to 12x spread between providers.
- Measure the delay before you commit. It is the one cost we could not see. If every call in an agent loop adds a hop, a small delay adds up. We will publish the timing when we run it.
How we measured
- Snapshot: OpenRouter's public, keyless API on 2026-10-06: the model list and the endpoint list for 27 curated models (28 requests). Prices, context, quantization and uptime are copied as the API reported them.
- First-party prices: the vendors' published list prices as recorded in the product price table.
- Markup: OpenRouter list price ÷ first-party list price − 1. Effective markup adds the 5.5% card fee.
- Tier: read from the endpoint tag. Flex, priority, fast, ultrafast and batch are named tiers; a region suffix is regional; anything else is standard. This is a heuristic.
Caveats
- Third-party-reported. Every price comes from OpenRouter's API or pricing page on 2026-10-06. Prices change often; check before you rely on one.
- Snapshot, not a series. One day's prices. We did not track how often they change.
- No timing. Without a key, nothing here says whether the gateway is slower.
- Fees by plan. Enterprise or negotiated terms can differ from the published fees.
What to read next
- The cheapest place to run open models right now
- What does a router cost you? Rules vs Jev vs an LLM router
- How to estimate your AI coding bill
See what your calls really cost
Agent records the model, the route, the tokens and the cost of every call it makes. Try Agent and see where your spend goes.