Explainer · Inference gateway

Inference gateways and OpenRouter, explained

Definition

An inference gateway is a service that puts many model providers behind one API: you send a request for a model to the gateway, and it forwards the request to one of the providers that serve that model, then returns the answer and one bill. OpenRouter is the best-known public example. A gateway gives you one key, one schema and fallback between providers; it can add a fee and a network hop, and it routes between providers of a model, not between models for a task.

Agent team · · 4 min read · Every number is from the public studies

Interactive

One request, 10 providers: how a gateway routes Llama 3.3 70B Instruct

The gateway forwards your request to one provider that serves the model and returns the answer and one bill. Prices are per million tokens.

Reported prices
Route to

+ 4 more providers in the Table view.

10 providers list Llama 3.3 70B Instruct in this snapshot; the first 6 by blended price are drawn. Lowest listed input price among them: DeepInfra (fp8) at $0.1 per million input tokens. Gateway per-token markup over first-party prices: 0% on 10 of 10. Credit fee: 5.5% ($0.80 minimum by card).

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose. Route shown: DeepInfra (fp8 weights). A tag such as fp8 is the provider's stated quantization.

Source: Inference provider index

What a gateway does

Without a gateway, each provider needs its own account, key, SDK quirks and invoice. With a gateway:

  • One API for many models. Claude, GPT, Gemini and open-weight models through one request format.
  • Provider choice per model. An open-weight model such as gpt-oss-120b or a DeepSeek model is served by many providers at different prices, precisions and context lengths. The gateway can pick one by price, speed or uptime, or follow your preference.
  • Fallback. When one provider fails, the gateway can retry another.
  • One bill. You buy credits once and spend them across providers.

A gateway is not an LLM router in the task sense. A router decides which model should handle a request; a gateway decides who serves the model you asked for. See what is an LLM router.

What does the gateway add to the price?

We took a snapshot of OpenRouter's public, keyless API on 2026-10-06: 265 endpoints from 52 providers for 27 models. Every price below is third-party-reported, not measured by a run.

Calculation

0%per-token markup on 10 of 10 models

5.5% with the card credit feeCalculation

  • Claude Haiku 4.5
  • Claude Sonnet 5
  • Claude Sonnet 5.5
  • Claude Opus 4.8
  • Claude Opus 5
  • Claude Opus 5.5
  • Claude Fable 5.1
  • GPT-6 Luna
  • Gemini 3.8 Flash
  • Gemini 3.5 Flash
  • per-token markup over the first-party list price (input and output)
  • With the 5.5% card credit fee (input) (calculation, hollow)

List-price calculation, not a run. 10 rows, 3 series: Per-token markup (input), Per-token markup (output), With the 5.5% card credit fee (input). Per-token markup (input): all at 0%. Per-token markup (output): all at 0%.

Notes

Per-token markup, and the markup after the Standard credit-purchase fee (calculation)

A calculation: OpenRouter list price ÷ first-party list price − 1, then × (1 + 5.5%) for credits bought by card on the Standard plan (purchases large enough that the minimum fee does not apply). Fees as published on 2026-10-06; first-party prices from the vendors’ list prices.

Sources: OpenRouter public API: models and provider endpoints (snapshot), OpenRouter pricing and fees, Anthropic list prices (Claude models), OpenAI list prices, Google Gemini list prices

  • For 10 of 10 models with a first-party list price, OpenRouter's per-token price was the same as the vendor's price.
  • The cost of the gateway is the 5.5% credit-purchase fee on the Standard plan (card purchases, third-party-reported). So the effective markup is +5.5%, a calculation.

In short: per token you pay the vendor's price; the fee is where the gateway earns.

Providers differ, mostly for open-weight models

For closed models, every standard-tier provider charged one price: 15 of 15 closed models (Claude, GPT, Gemini) with 2 or more providers had a single standard price. Their price differences come from named tiers (flex, priority, fast) and regional endpoints.

Open-weight models are different. Many companies host the same weights, and they compete on price:

Calculation
ModelSpread
  1. DeepSeek V4 Flash 0423n = 15 providers
  2. DeepSeek V4 Pro 0423n = 15 providers
  3. gpt-oss-120bn = 20 providers
  4. Llama 3.3 70B Instructn = 10 providers
  5. GLM 5.3n = 32 providers
  6. Kimi K3n = 19 providers
  7. Llama 4 Maverickn = 3 providers
  8. 15 more models: the same price at every providerClaude Haiku 4.5, Claude Sonnet 5, Claude Sonnet 5.5, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Fable 5.1, GPT-6 Sol, GPT-6 Luna, GPT-6 Astra, GPT-5.5, Gemini 3.8 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.1 Pro Preview
    1x

List-price calculation, not a run. 22 rows. Highest DeepSeek V4 Flash 0423 12.6x (n 15). Lowest Gemini 3.1 Pro Preview 1x (n 2).

Notesn 2–32 per row

Standard tier, blended price (3 input : 1 output), most expensive provider ÷ cheapest provider

A calculation on prices reported by OpenRouter’s public API, snapshot 2026-10-06. n = providers with a standard-tier endpoint. 1x means every provider charges the same. Highlighted: a spread of 4x or more.

Source: OpenRouter public API: models and provider endpoints (snapshot)

The most expensive standard-tier provider charged 12.6x the cheapest for DeepSeek V4 Flash 0423 (15 providers), 11.2x for DeepSeek V4 Pro 0423, 6.9x for gpt-oss-120b (20 providers) and 6.7x for Llama 3.3 70B Instruct. These spreads are a calculation on a 3:1 input-to-output blended price.

A low price is not always a like-for-like offer: the cheapest endpoints often report lower precision (fp8, fp4) or a shorter context. See quantization and cheap inference.

What is not measured yet

  • Gateway delay. Time to first token and total time through the gateway vs direct are not measured: no key in the environment. A live harness is ready and sends nothing without a key.
  • Provider latency and throughput. The keyless API returned a latency figure for 0 of 265 endpoints.
  • Billed cost per call through the gateway: not measured.

Until those runs exist, a claim that a gateway is faster or slower than a direct call has no data behind it here.

When to use a gateway

  • Use one when you want many models through one key, when you use open-weight models and want to shop across hosts, or when provider fallback matters more than a few percent of cost.
  • Go direct when you use one or two vendors at volume, when you need the vendor's own features first (new parameters, batch pricing, enterprise terms), or when a 5.5% fee is larger than the convenience is worth.
  • Either way, check the endpoint: provider, precision, context length and tier, not only the model name.

Browse every provider's position per model at /providers, or price a workload through a provider in the AI cost calculator.

Frequently asked questions

Does OpenRouter charge more than the official API?

Per token, not in our snapshot: for 10 of 10 models with a first-party list price, OpenRouter's price equalled the vendor's. The cost is the 5.5% credit-purchase fee on the Standard plan, which makes the effective markup +5.5% (third-party-reported, 2026-10-06).

Is OpenRouter slower than calling the provider directly?

We have not measured it. The gateway's own delay needs a keyed run, and no key was in the environment. Our pages say "not measured" until that run exists.

Why do providers charge different prices for the same model?

Mostly for open-weight models, which many companies can host. They differ in hardware, precision, context length and margin. For closed models in our snapshot, every standard-tier provider charged the same price.

Is a gateway the same as an LLM router?

No. A gateway forwards a request for one model to a provider that serves it. An LLM router chooses which model, and which effort, a task should use.

Watch the data

Live story · 42 sSame model, different price: 265 provider endpoints compared

Same model, different price: 265 provider endpoints compared

Reported by OpenRouter’s public API on 2026-10-06: open-weight prices vary up to 12.6x across providers; closed models have one standard price, and the gateway adds a 5.5% credit fee.

Transcript
  1. Provider index · third-party-reported · 2026-10-06. Same model, different price. 265 endpoints, 52 providers, 27 models. Prices as OpenRouter’s public API reports them.
  2. 265 endpoints, 52 providers, 27 models. The keyless API returned latency for 0 of them, so speed is not compared. Provider endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Endpoints with a latency figure: 0 of 265. Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
  3. Open-weight models vary most: DeepSeek V4 Flash 12.6×, DeepSeek V4 Pro 11.2×. All 15 closed models with 2+ providers: one standard price (1×). Chart: Priciest ÷ cheapest standard provider · blended price (n = 2–32 each). Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
  4. One model, 10 providers: Llama 3.3 70B input costs $0.10 per million tokens at DeepInfra (fp8) and $1.04 at Together. Chart: Llama 3.3 70B · USD per million tokens · one row per provider. Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
  5. Gateway markup: for 10 of 10 models OpenRouter charges the vendor’s per-token price. The cost is a 5.5% fee when you buy credits. Models at the first-party per-token price: 10 of 10. Credit-purchase fee, Standard plan by card: 5.5% ($0.80 minimum by card). Effective markup per token after the fee: +5.5% (calculation). Caveat: First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.
  6. Example: Sonnet 5.5 costs $2.00 in and $10.00 out per million tokens on both. A price is not a measurement: no winner. Chart: Claude Sonnet 5.5 · OpenRouter vs Anthropic list price · USD per million. Caveat: Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.
  7. Prices change often: every endpoint, tier and source date online.

The data behind this explainer

  • Inference
  • Providers

Inference provider index: 27 models, 52 providers

Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.

265endpoints, 52 providers, 27 models · Provider endpoints in the snapshot · n = 27

34 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.