Explainer · Inference provider

What is an inference provider? Bedrock, Vertex, Azure and first-party prices

Definition

Inference provider is a company that sells API access to an AI model: the vendor, a cloud, a specialist host or a gateway. In our 2026-10-06 snapshot, Bedrock listed Claude Sonnet 5.5 at the same price as the Anthropic API: $2 input and $10 output per million tokens. For seven open-weight models, standard-tier price spreads ranged from 1.7x to 12.6x (a calculation). The largest spread covered 15 providers.

Agent team · · 5 min read · Every number is from the public studies

Reported
  1. Inputall 5 at $2.00

  2. Outputall 5 at $10.00

  3. Cache readall 5 at $0.20

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $2.00. Output: all at $10.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Kinds of inference provider

Four kinds sell model access:

  • The vendor. Anthropic sells Claude API access at its first-party list price.
  • A cloud. Amazon, Google and Microsoft own Bedrock, Vertex and Azure.
  • A specialist host. DeepInfra, Together and Cerebras host open-weight models.
  • A gateway. OpenRouter forwards requests to these providers and sends one bill. See inference gateways.

Our provider index covers OpenRouter's public API snapshot on 2026-10-06: 265 endpoints, 52 providers and 27 curated models. Every price is third-party-reported. These full listing counts have no sampling interval. They do not cover the whole market. We ran no provider calls. These listings are not measured bills.

Closed models: the same list price across clouds

Claude Sonnet 5.5 has five providers in the snapshot (n = 5):

Reported
  1. Inputall 5 at $2.00

  2. Outputall 5 at $10.00

  3. Cache readall 5 at $0.20

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $2.00. Output: all at $10.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Amazon Bedrock, Anthropic, Azure, Claude Platform on AWS and Google Vertex each listed $2 input, $10 output and $0.20 cache read per million tokens. Across providers, 15 of 15 closed models had matching standard-tier input and output prices for each model. This covers 7 Claude, 4 GPT and 4 Gemini models with 2 or more providers. Prices differ between models. In the Anthropic vs Amazon Bedrock comparison, all 21 shared prices tie.

Here, standard price does not separate closed-model providers. Quotas, regions, data terms and features may. We did not measure them.

Prices still differ by named tier. The 15 closed models have 118 endpoints. We read tiers from endpoint tags (a heuristic): 49 standard, 41 regional, 12 flex, 8 priority, 7 fast and 1 ultrafast. Against the standard price of the same model (a calculation on the endpoint table):

  • Regional endpoints at Bedrock, Vertex and Azure list 10% more in 41 of 41 cases.
  • Flex endpoints list half the price in 12 of 12 cases.
  • Priority endpoints list 1.8x in 8 of 8 cases.
  • Fast and ultrafast endpoints list 2x to 6x in 8 of 8 cases.

For Claude Sonnet 5.5, three regional endpoints (Vertex twice, Azure once) listed $2.20 input and $11 output.

OpenRouter, our price source, adds a fee, not a per-token markup. For 10 of 10 models with a first-party list price, its per-token price equalled the vendor's. OpenRouter charges 5.5% (minimum $0.80) when you buy credits by card on its Standard plan (third-party-reported). For a large purchase, the effective markup is +5.5% (a calculation). First-party prices predate the snapshot. An intervening vendor price change would show as a markup.

Calculation

0%per-token markup on 10 of 10 models

5.5% with the card credit feeCalculation

  • Claude Haiku 4.5
  • Claude Sonnet 5
  • Claude Sonnet 5.5
  • Claude Opus 4.8
  • Claude Opus 5
  • Claude Opus 5.5
  • Claude Fable 5.1
  • GPT-6 Luna
  • Gemini 3.8 Flash
  • Gemini 3.5 Flash
  • per-token markup over the first-party list price (input and output)
  • With the 5.5% card credit fee (input) (calculation, hollow)

List-price calculation, not a run. 10 rows, 3 series: Per-token markup (input), Per-token markup (output), With the 5.5% card credit fee (input). Per-token markup (input): all at 0%. Per-token markup (output): all at 0%.

Notes

Per-token markup, and the markup after the Standard credit-purchase fee (calculation)

A calculation: OpenRouter list price ÷ first-party list price − 1, then × (1 + 5.5%) for credits bought by card on the Standard plan (purchases large enough that the minimum fee does not apply). Fees as published on 2026-10-06; first-party prices from the vendors’ list prices.

Sources: OpenRouter public API: models and provider endpoints (snapshot), OpenRouter pricing and fees, Anthropic list prices (Claude models), OpenAI list prices, Google Gemini list prices

Open-weight models: the price spread is wide

The chart divides the highest standard-tier price by the lowest for each model. Blended price is (3 × input + output) ÷ 4, a 3:1 mix of input and output tokens (a calculation):

Calculation
ModelSpread
  1. DeepSeek V4 Flash 0423n = 15 providers
  2. DeepSeek V4 Pro 0423n = 15 providers
  3. gpt-oss-120bn = 20 providers
  4. Llama 3.3 70B Instructn = 10 providers
  5. GLM 5.3n = 32 providers
  6. Kimi K3n = 19 providers
  7. Llama 4 Maverickn = 3 providers
  8. 15 more models: the same price at every providerClaude Haiku 4.5, Claude Sonnet 5, Claude Sonnet 5.5, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Fable 5.1, GPT-6 Sol, GPT-6 Luna, GPT-6 Astra, GPT-5.5, Gemini 3.8 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.1 Pro Preview
    1x

List-price calculation, not a run. 22 rows. Highest DeepSeek V4 Flash 0423 12.6x (n 15). Lowest Gemini 3.1 Pro Preview 1x (n 2).

Notesn 2–32 per row

Standard tier, blended price (3 input : 1 output), most expensive provider ÷ cheapest provider

A calculation on prices reported by OpenRouter’s public API, snapshot 2026-10-06. n = providers with a standard-tier endpoint. 1x means every provider charges the same. Highlighted: a spread of 4x or more.

Source: OpenRouter public API: models and provider endpoints (snapshot)

All 15 closed models sit at 1x. Of the 7 open-weight models with 2 or more providers, 5 had a spread of 4x or more. Each ratio is a reported-price calculation, not a measured cost. Here n counts providers with a standard-tier endpoint:

  • DeepSeek V4 Flash 0423: 12.6x (n = 15).
  • DeepSeek V4 Pro 0423: 11.2x (n = 15).
  • gpt-oss-120b: 6.9x (n = 20).
  • Llama 3.3 70B Instruct: 6.7x (n = 10).
  • GLM 5.3: 4.7x (n = 32).

Kimi K3 (1.8x, n = 19) and Llama 4 Maverick (1.7x, n = 3) had smaller spreads. For Llama 3.3 70B Instruct, DeepInfra (fp8) listed the lowest blended price, $0.10 input and $0.32 output, and Together listed the highest, $1.04 for both.

Low price does not make endpoints equal. For 6 of the 7 open-weight models, the lowest-priced provider reports fp8 or fp4 precision (a count from the models table). See quantization and cheap inference. Clouds sell open weights too: Bedrock listed gpt-oss-120b at $0.15 input and $0.60 output, and Vertex at $0.090 and $0.36, a gap of 1.7x on both (a calculation across n = 2 providers). The Bedrock vs Vertex comparison marks both rows unclear and names no winner, because a list price has no interval.

What this snapshot does not measure

  • Speed. The API returned a latency figure for 0 of 265 endpoints. We make no speed claim. Without an OpenRouter key, we measured neither gateway delay nor billed call cost.
  • Quality by provider. We ran no provider tasks. None of the 118 closed-model endpoints reports its precision, so we cannot check what precision a cloud runs.
  • Quotas, regions, data terms and features. We did not measure them.
  • Prices on each provider's own site. Providers may list different prices on their sites. Prices change often.

A checklist for choosing a provider

  1. Pick the model first. Then compare standard-tier prices for that exact model.
  2. Check the tier and region of the endpoint. A regional endpoint listed 10% more in our snapshot.
  3. For open-weight models, check precision (fp4, fp8 or fp16) and context length before comparing prices.
  4. Read the provider's own price page on the day you buy. Check quotas, regions, data terms and features yourself.
  5. Run your own tasks on each candidate. Record time and pass rate.
  6. Turn a list price into a monthly bill with the AI cost calculator.

Frequently asked questions

Is Claude cheaper on AWS Bedrock than on the Anthropic API?

Not in our snapshot. For Claude Sonnet 5.5, both listed $2 input and $10 output per million tokens (n = 5 providers). Bedrock also listed 10 regional Claude endpoints, each 10% above its standard price (a calculation). OpenRouter reported these prices, so check Bedrock's own page before you buy.

Is Claude on Vertex AI the same model?

It has the same model name and listed the same standard price as Anthropic: $2 and $10 for Sonnet 5.5. We ran no task through Vertex, so we have no pass rate, answer or speed comparison (0 runs). Treat "same model" as a listing, not as a tested result.

Why do open models cost different amounts at different providers?

These listings do not show why prices differ. For 6 of the 7 open-weight models, the lowest-priced provider reports fp8 or fp4 precision. Llama 3.3 70B Instruct spans 6.7x across 10 providers (a calculation). We did not test whether precision changes quality.

Which provider is fastest?

We do not know. The API returned a latency figure for 0 of 265 endpoints, and we ran no calls through any provider. Time each candidate on your own task.

Watch the data

Live story · 42 sSame model, different price: 265 provider endpoints compared

Same model, different price: 265 provider endpoints compared

Reported by OpenRouter’s public API on 2026-10-06: open-weight prices vary up to 12.6x across providers; closed models have one standard price, and the gateway adds a 5.5% credit fee.

Transcript
  1. Provider index · third-party-reported · 2026-10-06. Same model, different price. 265 endpoints, 52 providers, 27 models. Prices as OpenRouter’s public API reports them.
  2. 265 endpoints, 52 providers, 27 models. The keyless API returned latency for 0 of them, so speed is not compared. Provider endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Endpoints with a latency figure: 0 of 265. Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
  3. Open-weight models vary most: DeepSeek V4 Flash 12.6×, DeepSeek V4 Pro 11.2×. All 15 closed models with 2+ providers: one standard price (1×). Chart: Priciest ÷ cheapest standard provider · blended price (n = 2–32 each). Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
  4. One model, 10 providers: Llama 3.3 70B input costs $0.10 per million tokens at DeepInfra (fp8) and $1.04 at Together. Chart: Llama 3.3 70B · USD per million tokens · one row per provider. Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
  5. Gateway markup: for 10 of 10 models OpenRouter charges the vendor’s per-token price. The cost is a 5.5% fee when you buy credits. Models at the first-party per-token price: 10 of 10. Credit-purchase fee, Standard plan by card: 5.5% ($0.80 minimum by card). Effective markup per token after the fee: +5.5% (calculation). Caveat: First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.
  6. Example: Sonnet 5.5 costs $2.00 in and $10.00 out per million tokens on both. A price is not a measurement: no winner. Chart: Claude Sonnet 5.5 · OpenRouter vs Anthropic list price · USD per million. Caveat: Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.
  7. Prices change often: every endpoint, tier and source date online.

The data behind this explainer

  • Inference
  • Providers

Inference provider index: 27 models, 52 providers

Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.

265endpoints, 52 providers, 27 models · Provider endpoints in the snapshot · n = 27

34 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.