Explainer · Inference gateway
Inference gateways and OpenRouter, explained
Definition
An inference gateway is a service that puts many model providers behind one API: you send a request for a model to the gateway, and it forwards the request to one of the providers that serve that model, then returns the answer and one bill. OpenRouter is the best-known public example. A gateway gives you one key, one schema and fallback between providers; it can add a fee and a network hop, and it routes between providers of a model, not between models for a task.
Agent team · · 4 min read · Every number is from the public studies
Interactive
One request, 10 providers: how a gateway routes Llama 3.3 70B Instruct
The gateway forwards your request to one provider that serves the model and returns the answer and one bill. Prices are per million tokens.
+ 4 more providers in the Table view.
| Provider | Quantization | Input | Output |
|---|---|---|---|
| DeepInfra | fp8 | $0.1 | $0.32 |
| Novita | bf16 | $0.14 | $0.4 |
| AkashML | fp8 | $0.2 | $0.52 |
| Parasail | fp8 | $0.22 | $0.5 |
| SambaNova | — | $0.45 | $0.9 |
| Groq | — | $0.59 | $0.79 |
| CoreWeave | fp16 | $0.71 | $0.71 |
| Google Vertex | — | $0.72 | $0.72 |
| Cloudflare | fp8 | $0.29 | $2.25 |
| Together | — | $1.04 | $1.04 |
10 providers list Llama 3.3 70B Instruct in this snapshot; the first 6 by blended price are drawn. Lowest listed input price among them: DeepInfra (fp8) at $0.1 per million input tokens. Gateway per-token markup over first-party prices: 0% on 10 of 10. Credit fee: 5.5% ($0.80 minimum by card).
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose. Route shown: DeepInfra (fp8 weights). A tag such as fp8 is the provider's stated quantization.
Source: Inference provider index
What a gateway does
Without a gateway, each provider needs its own account, key, SDK quirks and invoice. With a gateway:
- One API for many models. Claude, GPT, Gemini and open-weight models through one request format.
- Provider choice per model. An open-weight model such as gpt-oss-120b or a DeepSeek model is served by many providers at different prices, precisions and context lengths. The gateway can pick one by price, speed or uptime, or follow your preference.
- Fallback. When one provider fails, the gateway can retry another.
- One bill. You buy credits once and spend them across providers.
A gateway is not an LLM router in the task sense. A router decides which model should handle a request; a gateway decides who serves the model you asked for. See what is an LLM router.
What does the gateway add to the price?
We took a snapshot of OpenRouter's public, keyless API on 2026-10-06: 265 endpoints from 52 providers for 27 models. Every price below is third-party-reported, not measured by a run.
- For 10 of 10 models with a first-party list price, OpenRouter's per-token price was the same as the vendor's price.
- The cost of the gateway is the 5.5% credit-purchase fee on the Standard plan (card purchases, third-party-reported). So the effective markup is +5.5%, a calculation.
In short: per token you pay the vendor's price; the fee is where the gateway earns.
Providers differ, mostly for open-weight models
For closed models, every standard-tier provider charged one price: 15 of 15 closed models (Claude, GPT, Gemini) with 2 or more providers had a single standard price. Their price differences come from named tiers (flex, priority, fast) and regional endpoints.
Open-weight models are different. Many companies host the same weights, and they compete on price:
The most expensive standard-tier provider charged 12.6x the cheapest for DeepSeek V4 Flash 0423 (15 providers), 11.2x for DeepSeek V4 Pro 0423, 6.9x for gpt-oss-120b (20 providers) and 6.7x for Llama 3.3 70B Instruct. These spreads are a calculation on a 3:1 input-to-output blended price.
A low price is not always a like-for-like offer: the cheapest endpoints often report lower precision (fp8, fp4) or a shorter context. See quantization and cheap inference.
What is not measured yet
- Gateway delay. Time to first token and total time through the gateway vs direct are not measured: no key in the environment. A live harness is ready and sends nothing without a key.
- Provider latency and throughput. The keyless API returned a latency figure for 0 of 265 endpoints.
- Billed cost per call through the gateway: not measured.
Until those runs exist, a claim that a gateway is faster or slower than a direct call has no data behind it here.
When to use a gateway
- Use one when you want many models through one key, when you use open-weight models and want to shop across hosts, or when provider fallback matters more than a few percent of cost.
- Go direct when you use one or two vendors at volume, when you need the vendor's own features first (new parameters, batch pricing, enterprise terms), or when a 5.5% fee is larger than the convenience is worth.
- Either way, check the endpoint: provider, precision, context length and tier, not only the model name.
Browse every provider's position per model at /providers, or price a workload through a provider in the AI cost calculator.
Frequently asked questions
Does OpenRouter charge more than the official API?
Per token, not in our snapshot: for 10 of 10 models with a first-party list price, OpenRouter's price equalled the vendor's. The cost is the 5.5% credit-purchase fee on the Standard plan, which makes the effective markup +5.5% (third-party-reported, 2026-10-06).
Is OpenRouter slower than calling the provider directly?
We have not measured it. The gateway's own delay needs a keyed run, and no key was in the environment. Our pages say "not measured" until that run exists.
Why do providers charge different prices for the same model?
Mostly for open-weight models, which many companies can host. They differ in hardware, precision, context length and margin. For closed models in our snapshot, every standard-tier provider charged the same price.
Is a gateway the same as an LLM router?
No. A gateway forwards a request for one model to a provider that serves it. An LLM router chooses which model, and which effort, a task should use.
Watch the data
Same model, different price: 265 provider endpoints compared
Reported by OpenRouter’s public API on 2026-10-06: open-weight prices vary up to 12.6x across providers; closed models have one standard price, and the gateway adds a 5.5% credit fee.
Transcript
- Provider index · third-party-reported · 2026-10-06. Same model, different price. 265 endpoints, 52 providers, 27 models. Prices as OpenRouter’s public API reports them.
- 265 endpoints, 52 providers, 27 models. The keyless API returned latency for 0 of them, so speed is not compared. Provider endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Endpoints with a latency figure: 0 of 265. Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
- Open-weight models vary most: DeepSeek V4 Flash 12.6×, DeepSeek V4 Pro 11.2×. All 15 closed models with 2+ providers: one standard price (1×). Chart: Priciest ÷ cheapest standard provider · blended price (n = 2–32 each). Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
- One model, 10 providers: Llama 3.3 70B input costs $0.10 per million tokens at DeepInfra (fp8) and $1.04 at Together. Chart: Llama 3.3 70B · USD per million tokens · one row per provider. Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
- Gateway markup: for 10 of 10 models OpenRouter charges the vendor’s per-token price. The cost is a 5.5% fee when you buy credits. Models at the first-party per-token price: 10 of 10. Credit-purchase fee, Standard plan by card: 5.5% ($0.80 minimum by card). Effective markup per token after the fee: +5.5% (calculation). Caveat: First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.
- Example: Sonnet 5.5 costs $2.00 in and $10.00 out per million tokens on both. A price is not a measurement: no winner. Chart: Claude Sonnet 5.5 · OpenRouter vs Anthropic list price · USD per million. Caveat: Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.
- Prices change often: every endpoint, tier and source date online.