• OpenRouter
  • Inference
  • Gateway
  • LLM pricing

OpenRouter vs going direct: what the gateway really costs

OpenRouter charged the vendor's per-token price on 10 of 10 models we checked. The cost is a 5.5% credit fee. What we know, and what is still unmeasured.

TL;DR

  • We read OpenRouter's public, keyless API on 2026-10-06: 265 endpoints from 52 providers for 27 models. Every price here is reported by OpenRouter, not measured by us.
  • For 10 of 10 models with a first-party list price in our price table, OpenRouter's per-token price was the same as the vendor's. Claude, GPT-6 Luna and Gemini: same input price, same output price.
  • The gateway charges when you buy credits, not per request: 5.5% on the Standard plan by card ($0.80 minimum), 8% on Business and 5% by crypto (OpenRouter's pricing page, fetched 2026-10-06). So the effective markup on a card purchase is +5.5%.
  • 15 of 15 closed models (Claude, GPT, Gemini) with 2 or more providers had one standard-tier price at every provider: Anthropic, Amazon Bedrock, Google Vertex, Azure and the others. Price differences come only from named tiers (flex, priority, fast) and regional endpoints.
  • Not measured: the gateway's delay. The keyless API returned latency for 0 of 265 endpoints, and we had no key in the environment to time calls. The harness is ready and runs when a key is present.

Every price and endpoint: /benchmarks/inference-provider-index.

Live story · 42 sSame model, different price: 265 provider endpoints compared

Same model, different price: 265 provider endpoints compared

Reported by OpenRouter’s public API on 2026-10-06: open-weight prices vary up to 12.6x across providers; closed models have one standard price, and the gateway adds a 5.5% credit fee.

Transcript
  1. Provider index · third-party-reported · 2026-10-06. Same model, different price. 265 endpoints, 52 providers, 27 models. Prices as OpenRouter’s public API reports them.
  2. 265 endpoints, 52 providers, 27 models. The keyless API returned latency for 0 of them, so speed is not compared. Provider endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Endpoints with a latency figure: 0 of 265. Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
  3. Open-weight models vary most: DeepSeek V4 Flash 12.6×, DeepSeek V4 Pro 11.2×. All 15 closed models with 2+ providers: one standard price (1×). Chart: Priciest ÷ cheapest standard provider · blended price (n = 2–32 each). Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
  4. One model, 10 providers: Llama 3.3 70B input costs $0.10 per million tokens at DeepInfra (fp8) and $1.04 at Together. Chart: Llama 3.3 70B · USD per million tokens · one row per provider. Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
  5. Gateway markup: for 10 of 10 models OpenRouter charges the vendor’s per-token price. The cost is a 5.5% fee when you buy credits. Models at the first-party per-token price: 10 of 10. Credit-purchase fee, Standard plan by card: 5.5% ($0.80 minimum by card). Effective markup per token after the fee: +5.5% (calculation). Caveat: First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.
  6. Example: Sonnet 5.5 costs $2.00 in and $10.00 out per million tokens on both. A price is not a measurement: no winner. Chart: Claude Sonnet 5.5 · OpenRouter vs Anthropic list price · USD per million. Caveat: Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.
  7. Prices change often: every endpoint, tier and source date online.

The question

A gateway gives you one API and one bill for many models and providers. What does that cost you compared with calling the vendor directly? Three parts can add cost:

  1. A per-token markup: the gateway lists a higher price than the vendor.
  2. A fee: the gateway charges on top, per request or per purchase.
  3. Delay: one more network hop before the provider starts work.

We can answer the first two from public data. The third needs live calls, and we did not make them.

Part 1: the per-token price is the same

Calculation

0%per-token markup on 10 of 10 models

5.5% with the card credit feeCalculation

  • Claude Haiku 4.5
  • Claude Sonnet 5
  • Claude Sonnet 5.5
  • Claude Opus 4.8
  • Claude Opus 5
  • Claude Opus 5.5
  • Claude Fable 5.1
  • GPT-6 Luna
  • Gemini 3.8 Flash
  • Gemini 3.5 Flash
  • per-token markup over the first-party list price (input and output)
  • With the 5.5% card credit fee (input) (calculation, hollow)

List-price calculation, not a run. 10 rows, 3 series: Per-token markup (input), Per-token markup (output), With the 5.5% card credit fee (input). Per-token markup (input): all at 0%. Per-token markup (output): all at 0%.

Notes

Per-token markup, and the markup after the Standard credit-purchase fee (calculation)

A calculation: OpenRouter list price ÷ first-party list price − 1, then × (1 + 5.5%) for credits bought by card on the Standard plan (purchases large enough that the minimum fee does not apply). Fees as published on 2026-10-06; first-party prices from the vendors’ list prices.

Sources: OpenRouter public API: models and provider endpoints (snapshot), OpenRouter pricing and fees, Anthropic list prices (Claude models), OpenAI list prices, Google Gemini list prices

For each model that has a first-party list price in our product price table, we divided OpenRouter's list price by the vendor's and subtracted 1. The result was 0% on input and 0% on output for all 10 models:

  • Anthropic: Claude Haiku 4.5, Sonnet 5, Sonnet 5.5, Opus 4.8, Opus 5, Opus 5.5 and Fable 5.1.
  • OpenAI: GPT-6 Luna.
  • Google: Gemini 3.8 Flash and Gemini 3.5 Flash.

Two examples side by side:

Reported

OpenRouter

Anthropicfirst-party list price

Same price on all 2 price types
  • InputOpenRouter: $2.00 same priceAnthropic: $2.00
  • OutputOpenRouter: $10.00 same priceAnthropic: $10.00

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $2.00. Output: all at $10.00.

Notes

USD per million tokens, before any credit-purchase fee

OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.

Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)

Claude Sonnet 5.5: $2.00 in and $10.00 out per million tokens on both OpenRouter and Anthropic.

Reported

OpenRouter

Google AI Studiofirst-party list price

Same price on all 2 price types
  • InputOpenRouter: $0.75 same priceGoogle AI Studio: $0.75
  • OutputOpenRouter: $3.75 same priceGoogle AI Studio: $3.75

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $0.75. Output: all at $3.75.

Notes

USD per million tokens, before any credit-purchase fee

OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Google price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.

Sources: OpenRouter public API: models and provider endpoints (snapshot), Google Gemini list prices

Gemini 3.8 Flash: $0.75 in and $3.75 out on both OpenRouter and Google AI Studio.

The comparison pages hold every row: OpenRouter vs Anthropic (14 rows, all the same price), OpenRouter vs OpenAI and OpenRouter vs Google AI Studio. A reported price has no interval, so these pages name no winner. They state the gap, and here the gap is zero.

Part 2: the fee is on credits, not on tokens

OpenRouter's pricing page says that inference is billed at the provider's list price and that the fee is charged when you buy credits:

  • Standard plan, card: 5.5%, with a $0.80 minimum per purchase.
  • Business plan: 8%.
  • Crypto: 5%.

So on the Standard plan, $100 of tokens costs you $105.50 of card payments (our arithmetic). The $0.80 minimum matters for small top-ups: 5.5% of a purchase is less than $0.80 below about $14.55, so small purchases pay a higher percentage (our arithmetic).

That is why the markup chart has a third series. The per-token markup is 0%. With the card fee, the effective markup is +5.5% for every model.

Part 3: does the provider matter for closed models?

Calculation
ModelSpread
  1. DeepSeek V4 Flash 0423n = 15 providers
  2. DeepSeek V4 Pro 0423n = 15 providers
  3. gpt-oss-120bn = 20 providers
  4. Llama 3.3 70B Instructn = 10 providers
  5. GLM 5.3n = 32 providers
  6. Kimi K3n = 19 providers
  7. Llama 4 Maverickn = 3 providers
  8. 15 more models: the same price at every providerClaude Haiku 4.5, Claude Sonnet 5, Claude Sonnet 5.5, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Fable 5.1, GPT-6 Sol, GPT-6 Luna, GPT-6 Astra, GPT-5.5, Gemini 3.8 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.1 Pro Preview
    1x

List-price calculation, not a run. 22 rows. Highest DeepSeek V4 Flash 0423 12.6x (n 15). Lowest Gemini 3.1 Pro Preview 1x (n 2).

Notesn 2–32 per row

Standard tier, blended price (3 input : 1 output), most expensive provider ÷ cheapest provider

A calculation on prices reported by OpenRouter’s public API, snapshot 2026-10-06. n = providers with a standard-tier endpoint. 1x means every provider charges the same. Highlighted: a spread of 4x or more.

Source: OpenRouter public API: models and provider endpoints (snapshot)

The spread is the most expensive standard-tier provider's blended price divided by the cheapest one's (3 input tokens to 1 output token, a calculation). For all 15 closed models with 2 or more providers, the spread was 1.0x: every standard-tier provider listed the same price.

Reported
  1. Inputall 5 at $4.00

  2. Outputall 5 at $20.00

  3. Cache readall 5 at $0.20

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $4.00. Output: all at $20.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Claude Opus 5.5 is a typical case: $4.00 in, $20.00 out and $0.20 for a cache read at Amazon Bedrock, Anthropic, Azure, Claude Platform on AWS and Google Vertex alike. The pairwise pages agree: Anthropic vs Amazon Bedrock, Amazon Bedrock vs Google Vertex, Amazon Bedrock vs Azure, Google AI Studio vs Google Vertex and OpenAI vs Azure.

So for a closed model, you pick a provider for reasons other than the standard list price: your cloud contract, data residency, rate limits, a named tier or a regional endpoint. Those reasons can be real. The list price is not one of them.

Open-weight models are a different story. Their prices differ by up to 12.6x between providers. That is the subject of the cheapest place to run open models.

What we do not know

We would rather show a gap than fill it:

  • Gateway delay. The time to first token and the total time through OpenRouter vs direct are not measured: no key in the environment. A live harness is ready. It sends nothing without a key, and it has a hard spending cap.
  • Billed cost per call. Also not measured, for the same reason. List price × tokens is a calculation, not a bill.
  • Provider latency and throughput. OpenRouter's keyless API returned these for 0 of 265 endpoints. We cannot say whether a provider is faster or slower.
  • Price drift. The first-party prices in our table were verified before the snapshot date. A vendor price change in between would show as a markup. None did.
  • Prices outside OpenRouter. A provider's price on its own site can differ from its price through OpenRouter. We compared OpenRouter's listing with the vendor's published list price only.

So should you use a gateway?

On price alone, from this snapshot:

  • For closed models, a gateway costs the credit fee, about 5.5% on the Standard plan, and nothing per token. If one API, one bill and access to many providers save you more than 5.5% of your token spend in engineering time, the gateway pays for itself. If you spend a lot on one vendor, going direct saves the fee.
  • For open-weight models, the provider you pick matters far more than the gateway fee. A 5.5% fee is small next to a 6x to 12x spread between providers.
  • Measure the delay before you commit. It is the one cost we could not see. If every call in an agent loop adds a hop, a small delay adds up. We will publish the timing when we run it.

How we measured

  • Snapshot: OpenRouter's public, keyless API on 2026-10-06: the model list and the endpoint list for 27 curated models (28 requests). Prices, context, quantization and uptime are copied as the API reported them.
  • First-party prices: the vendors' published list prices as recorded in the product price table.
  • Markup: OpenRouter list price ÷ first-party list price − 1. Effective markup adds the 5.5% card fee.
  • Tier: read from the endpoint tag. Flex, priority, fast, ultrafast and batch are named tiers; a region suffix is regional; anything else is standard. This is a heuristic.

Caveats

  • Third-party-reported. Every price comes from OpenRouter's API or pricing page on 2026-10-06. Prices change often; check before you rely on one.
  • Snapshot, not a series. One day's prices. We did not track how often they change.
  • No timing. Without a key, nothing here says whether the gateway is slower.
  • Fees by plan. Enterprise or negotiated terms can differ from the published fees.

See what your calls really cost

Agent records the model, the route, the tokens and the cost of every call it makes. Try Agent and see where your spend goes.

The data behind this post

  • Inference
  • Providers

Inference provider index: 27 models, 52 providers

Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.

265endpoints, 52 providers, 27 models · Provider endpoints in the snapshot · n = 27

34 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.