52 providers · 265 endpoints · snapshot October 6, 2026

Inference providers, priced side by side.

The same model often costs very different amounts from one provider to the next. Each dot below is one provider; the diamond is the model maker’s own list price. Prices are reported by OpenRouter’s public API; we did not measure them, and the API reported no latency.

Reported
Calculation: (3 × input + output) ÷ 4

22 models · Select a provider to pin it on every row.

ModelSpread
  1. DeepSeek V4 Flash 042315 providers · StreamLake $0.0525 → Cloudflare $0.66
    Spread 12.6x: highest ÷ lowest Blended 3:1 price (calculation)
  2. DeepSeek V4 Pro 042315 providers · StreamLake $0.261 → Reka $2.92
    Spread 11.2x: highest ÷ lowest Blended 3:1 price (calculation)
  3. gpt-oss-120b20 providers · CoreWeave $0.065 → Cerebras $0.45
    Spread 6.9x: highest ÷ lowest Blended 3:1 price (calculation)
  4. Llama 3.3 70B Instruct10 providers · DeepInfra $0.155 → Together $1.04
    Spread 6.7x: highest ÷ lowest Blended 3:1 price (calculation)
  5. GLM 5.332 providers · Novita $0.645 → Relace $3.02
    Spread 4.7x: highest ÷ lowest Blended 3:1 price (calculation)
  6. Kimi K319 providers · Relace $3.87 → Alibaba $6.90
    Spread 1.8x: highest ÷ lowest Blended 3:1 price (calculation)
  7. Llama 4 Maverick3 providers · DigitalOcean $0.3037 → Parasail $0.5125
    Spread 1.7x: highest ÷ lowest Blended 3:1 price (calculation)
  8. 15 more models: one blended 3:1 price at every providerClaude Haiku 4.5, Claude Sonnet 5, Claude Sonnet 5.5, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Fable 5.1, GPT-6 Sol, GPT-6 Luna, GPT-6 Astra, GPT-5.5, Gemini 3.8 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.1 Pro Preview
    1x
  • one provider (blended: a calculation, hollow)
  • first-party list price
  • Spread: highest ÷ lowest price of the row (calculation)

Third-party reported values. 163 listings, 4 series: Blended 3:1, Input, Output, Cache read. Blended 3:1: highest Claude Fable 5.1 · Amazon Bedrock $20.00 (n 4). Lowest DeepSeek V4 Flash 0423 · StreamLake (fp8) $0.053 (n 15). Input: highest Claude Fable 5.1 · Amazon Bedrock $10.00 (n 4). Lowest DeepSeek V4 Flash 0423 · Relace (fp4) $0.012 (n 15).

Notesn 2–32 per row

22 models with two or more standard-tier providers, widest spread first. USD per million tokens, snapshot 2026-10-06.

Prices reported by OpenRouter’s public API (third-party-reported, not measured by Agent). One dot per provider: its cheapest standard-tier endpoint. The diamond is the first-party list price. Spread: highest ÷ lowest price of the row. A parenthesis names the quantization the provider reported. n = providers with a standard-tier endpoint.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Gateways: what OpenRouter adds

A gateway resells other providers’ endpoints. Its per-token price matches the first-party list price; the fee comes when you buy credits.

Calculation

0%per-token markup on 10 of 10 models

5.5% with the card credit feeCalculation

  • Claude Haiku 4.5
  • Claude Sonnet 5
  • Claude Sonnet 5.5
  • Claude Opus 4.8
  • Claude Opus 5
  • Claude Opus 5.5
  • Claude Fable 5.1
  • GPT-6 Luna
  • Gemini 3.8 Flash
  • Gemini 3.5 Flash
  • per-token markup over the first-party list price (input and output)
  • With the 5.5% card credit fee (input) (calculation, hollow)

List-price calculation, not a run. 10 rows, 3 series: Per-token markup (input), Per-token markup (output), With the 5.5% card credit fee (input). Per-token markup (input): all at 0%. Per-token markup (output): all at 0%.

Notes

Per-token markup, and the markup after the Standard credit-purchase fee (calculation)

A calculation: OpenRouter list price ÷ first-party list price − 1, then × (1 + 5.5%) for credits bought by card on the Standard plan (purchases large enough that the minimum fee does not apply). Fees as published on 2026-10-06; first-party prices from the vendors’ list prices.

Sources: OpenRouter public API: models and provider endpoints (snapshot), OpenRouter pricing and fees, Anthropic list prices (Claude models), OpenAI list prices, Google Gemini list prices

Providers with a page

The gateways, clouds and first-party APIs most people compare. The bar counts the models where its price is the lowest, tied for the lowest, or above the lowest: a calculation on reported prices (3 input : 1 output tokens).

20 providers

  • Inference provider · OpenRouter

    OpenRouter

    A gateway that routes one API to many inference providers. Its per-token price is compared with first-party list prices; its fee is charged when credits are bought.

    A gateway: it resells other providers’ endpoints. See its markup.

    3 comparisons

  • Inference provider · Anthropic

    Anthropic

    Anthropic’s own API: the first-party list price of the Claude models, and Anthropic’s endpoint as listed on OpenRouter.

    7 models in the index · median 1x the cheapest

    Counted over 7 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 100% (median, reported)5 comparisons

  • Inference provider · OpenAI

    OpenAI

    OpenAI’s own API: the first-party list price of GPT models, and OpenAI’s endpoints as listed on OpenRouter.

    4 models in the index · median 1x the cheapest

    Counted over 4 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 100% (median, reported)2 comparisons

  • Inference provider · Google

    Google AI Studio

    Google’s Gemini API (AI Studio): the first-party list price of Gemini models, and its endpoints as listed on OpenRouter.

    4 models in the index · median 1x the cheapest

    Counted over 4 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 100% (median, reported)2 comparisons

  • Inference provider · Google

    Google Vertex AI

    Google Cloud’s model platform. An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    14 models in the index · median 1x the cheapest

    Counted over 13 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 100% (median, reported)16 comparisons

  • Inference provider · Amazon

    Amazon Bedrock

    Amazon Web Services’ model platform. An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    13 models in the index · median 1x the cheapest

    Counted over 8 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 100% (median, reported)14 comparisons

  • Inference provider · Microsoft

    Azure

    Microsoft’s cloud model platform. An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    13 models in the index · median 1x the cheapest

    Counted over 11 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 100% (median, reported)5 comparisons

  • Inference provider · Anthropic

    Claude Platform on AWS

    Anthropic’s Claude platform hosted on AWS. An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    5 models in the index · median 1x the cheapest

    Counted over 5 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 100% (median, reported)4 comparisons

  • Inference provider · Groq

    Groq

    An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    2 models in the index · median 4.1x the cheapest

    Counted over 2 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 100% (median, reported)12 comparisons

  • Inference provider · Together AI

    Together AI

    An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    4 models in the index · median 3.7x the cheapest

    Counted over 4 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 95% (median, reported)13 comparisons

  • Inference provider · Fireworks AI

    Fireworks AI

    An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    2 models in the index · median 2.4x the cheapest

    Counted over 2 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 99% (median, reported)8 comparisons

  • Inference provider · DeepInfra

    DeepInfra

    An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    6 models in the index · median 1.5x the cheapest

    Counted over 6 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 99% (median, reported)13 comparisons

  • Inference provider · Cerebras

    Cerebras

    An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    1 model in the index · median 6.9x the cheapest

    Counted over 1 model with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 100% (median, reported)11 comparisons

  • Inference provider · SambaNova

    SambaNova

    An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    2 models in the index · median 4.4x the cheapest

    Counted over 2 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 99% (median, reported)12 comparisons

  • Inference provider · Nebius

    Nebius

    An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    2 models in the index · median 3.7x the cheapest

    Counted over 2 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 97% (median, reported)13 comparisons

  • Inference provider · Parasail

    Parasail

    An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    7 models in the index · median 3.3x the cheapest

    Counted over 7 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 100% (median, reported)13 comparisons

  • Inference provider · Novita AI

    Novita AI

    An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    6 models in the index · median 1.5x the cheapest

    Counted over 6 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 98% (median, reported)13 comparisons

  • Inference provider · Baseten

    Baseten

    An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    3 models in the index · median 3.1x the cheapest

    Counted over 3 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 98% (median, reported)13 comparisons

  • Inference provider · Cloudflare

    Cloudflare Workers AI

    An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    4 models in the index · median 5.4x the cheapest

    Counted over 4 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 98% (median, reported)11 comparisons

  • Inference provider · SiliconFlow

    SiliconFlow

    An inference provider listed on OpenRouter. Its prices here are the ones OpenRouter’s public API reported for its endpoints.

    4 models in the index · median 3.6x the cheapest

    Counted over 4 models with 2 or more standard-tier providers · blended 3:1, a calculation

    Uptime 99% (median, reported)13 comparisons

One model, every provider

Input, output and cache-read prices on their own strips. Point at a provider to link its three prices; select it to pin it on every chart of this page. The Table view sorts any column.

Price per provider: DeepSeek V4 Flash 0423

Reported by OpenRouter’s public API, snapshot October 6, 2026. Third-party-reported prices, not measured by Agent.

Reported

15 providers with a standard-tier endpoint · cheapest StreamLake at $0.0525 blended, priciest Cloudflare at $0.66 (12.6x) · showing 15 of 15 rows

  1. InputRelace $0.012 → Cloudflare $0.44 · 36.7x

  2. OutputStreamLake $0.084 → OpenInference $1.41 · 16.8x

  3. Cache readStreamLake $0.0084 → Parasail $0.07 · 8.3x

  • one provider (reported price, USD per million tokens, log scale per strip)

Blended price and “vs cheapest” are calculations on the reported prices (3 input : 1 output tokens), against the cheapest standard-tier provider. A lower price can come with lower precision (fp4, fp8) or a shorter context. Uptime is what the API reported for the last day; the API reported no latency or throughput.

Provider vs provider

All comparisons

Price rows never name a winner: a reported list price has no interval, so a gap is stated, not ranked.

Where every provider’s price sits

52 providers listed for the 27 models in the index. “Sole cheapest” counts models where it alone has the lowest standard-tier blended price and at least one other provider serves the model. “Tied cheapest” counts models where it shares the lowest price with other providers.

Calculation
  • Sole cheapest
  • Tied cheapest
  • Above cheapest
  1. Google Vertex
  2. Azure
  3. Amazon Bedrock
  4. Anthropic
  5. Parasail
  6. Alibaba
  7. DeepInfra
  8. DigitalOcean
  9. Novita
  10. Claude Platform on AWS
  11. OpenAI
  12. Google AI Studio
  13. AkashML
  14. Cloudflare
  15. Relace
  16. SiliconFlow
  17. Together
  18. Mistral
  19. BaseTen
  20. AtlasCloud
  21. Baidu
  22. GMICloud
  23. Phala
  24. Venice
  25. Fireworks
  26. Wafer
  27. InferenceNet
  28. Sail Research
  29. CoreWeave
  30. Crusoe
  31. Decart
  32. Groq
  33. Makora
  34. Mancer 2
  35. Modal
  36. Morph
  37. Nebius
  38. Reka
  39. SambaNova
  40. StreamLake
  41. xAI
  42. Cerebras
  43. Chutes
  44. DekaLLM
  45. Friendli
  46. Inceptron
  47. Mara
  48. Moonshot AI
  49. NextBit
  50. OpenInference
  51. PrimeIntellect
  52. Z.AI

List-price calculation, not a run. 52 providers, 3 series: Sole cheapest, Tied cheapest, Above cheapest. Sole cheapest: highest StreamLake 2. Lowest Z.AI 0. Tied cheapest: highest Google Vertex 11. Lowest Z.AI 0.

Notes

Models with two or more standard-tier providers: sole cheapest, tied for cheapest, or above the cheapest (blended 3:1)

Counts of a calculation on reported prices (blended 3 input : 1 output). Not a score: a price says nothing about speed or quality.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Reported by OpenRouter’s public API, snapshot October 6, 2026. Uptime is as reported, not measured by Agent.

How these numbers were made

The rules

  1. One snapshot of OpenRouter’s public, keyless API on October 6, 2026. Each price is copied as the API reported it. Agent measured none of them.
  2. One dot per provider and model: its cheapest standard-tier endpoint. The Table views list every endpoint and tier.
  3. Blended price = (3 × input + output) ÷ 4. Spread = highest ÷ lowest blended price of one model. Both are calculations on the reported prices.
  4. Providers at the same price are tied, not ranked. A price has no interval, so no page names a winner on price.

The limits

  • Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
  • A provider listed on OpenRouter is reached through OpenRouter; its price there may differ from the price on the provider’s own site.
  • The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
  • No latency or throughput figure: the keyless API returned none. A cheaper provider is not shown to be slower or faster.
  • First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.
  • Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.

The full method Download the data (JSON)Every chart point (CSV)

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.