• Inference
  • Providers
  • Open Weight
  • LLM pricing

The cheapest place to run open models right now (October 2026 snapshot)

DeepSeek V4, gpt-oss-120b, Llama, Kimi K3 and GLM 5.3 across 52 providers. Prices differ up to 12.6x. Snapshot 2026-10-06, with the caveats.

TL;DR

  • Prices from OpenRouter's public API on 2026-10-06. They are reported, not measured by us, and they change often.
  • Open-weight models differ a lot by provider. The most expensive standard-tier provider charged this much more than the cheapest (blended 3 input : 1 output, our calculation): DeepSeek V4 Flash 12.6x, DeepSeek V4 Pro 11.2x, gpt-oss-120b 6.9x, Llama 3.3 70B 6.7x, GLM 5.3 4.7x.
  • Closed models do not: all 15 had one standard-tier price everywhere.
  • The cheapest endpoint often runs lower precision (fp4 or fp8) or a shorter context. A cheaper endpoint is not the same product.
  • We have no latency or quality data for these providers. The keyless API returned latency for 0 of 265 endpoints.

Every endpoint, with its quantization, context and uptime: /benchmarks/inference-provider-index.

How big is the spread?

Calculation
ModelSpread
  1. DeepSeek V4 Flash 0423n = 15 providers
  2. DeepSeek V4 Pro 0423n = 15 providers
  3. gpt-oss-120bn = 20 providers
  4. Llama 3.3 70B Instructn = 10 providers
  5. GLM 5.3n = 32 providers
  6. Kimi K3n = 19 providers
  7. Llama 4 Maverickn = 3 providers
  8. 15 more models: the same price at every providerClaude Haiku 4.5, Claude Sonnet 5, Claude Sonnet 5.5, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Fable 5.1, GPT-6 Sol, GPT-6 Luna, GPT-6 Astra, GPT-5.5, Gemini 3.8 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.1 Pro Preview
    1x

List-price calculation, not a run. 22 rows. Highest DeepSeek V4 Flash 0423 12.6x (n 15). Lowest Gemini 3.1 Pro Preview 1x (n 2).

Notesn 2–32 per row

Standard tier, blended price (3 input : 1 output), most expensive provider ÷ cheapest provider

A calculation on prices reported by OpenRouter’s public API, snapshot 2026-10-06. n = providers with a standard-tier endpoint. 1x means every provider charges the same. Highlighted: a spread of 4x or more.

Source: OpenRouter public API: models and provider endpoints (snapshot)

For each model, we took every provider's cheapest standard-tier endpoint, blended its price at 3 input tokens to 1 output token, and divided the most expensive by the cheapest. A spread of 1x means every provider charges the same.

  • DeepSeek V4 Flash 0423: 12.6x across 15 providers.
  • DeepSeek V4 Pro 0423: 11.2x across 15 providers.
  • gpt-oss-120b: 6.9x across 20 providers.
  • Llama 3.3 70B Instruct: 6.7x across 10 providers.
  • GLM 5.3: 4.7x across 32 providers.
  • Kimi K3: 1.8x across 19 providers.
  • Llama 4 Maverick: 1.7x across 3 providers.

Every closed model (Claude, GPT and Gemini) sits at 1.0x. For those, see OpenRouter vs going direct.

Model by model

Each chart shows one bar group per provider, sorted from the lowest to the highest blended price. A parenthesis names the quantization the provider reported. No parenthesis means the provider did not report one, which does not mean full precision. All prices are USD per million tokens.

DeepSeek V4 Flash

Reported
  1. InputRelace $0.012 → Cloudflare $0.44 · 36.7x

  2. OutputStreamLake $0.084 → OpenInference $1.41 · 16.8x

  3. Cache readStreamLake $0.0084 → Parasail $0.07 · 8.3x

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 15 rows, 3 series: Input, Output, Cache read. Input: highest Cloudflare $0.44. Lowest Relace (fp4) $0.012. Output: highest OpenInference (fp4) $1.41. Lowest StreamLake (fp8) $0.084.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

  • Cheapest blended: StreamLake (fp8) at $0.042 in / $0.084 out.
  • Next: DeepInfra (fp8) at $0.09 / $0.18, GMICloud (fp8) at $0.091 / $0.182, Venice at $0.0966 / $0.1925.
  • Most expensive: Cloudflare at $0.44 / $1.32.
  • Watch the mix: Relace (fp4) charges only $0.012 for input but $1.28 for output. For a job that reads a lot and writes little, it can beat the blended ranking.

DeepSeek V4 Pro

Reported
  1. InputRelace $0.2067 → NextBit $1.74 · 8.4x

  2. OutputStreamLake $0.4176 → Reka $9.00 · 21.6x

  3. Cache readStreamLake $0.0174 → Venice $0.33 · 19.0x

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 15 rows, 3 series: Input, Output, Cache read. Input: highest NextBit (fp8) $1.74. Lowest Relace (fp4) $0.21. Output: highest Reka $9.00. Lowest StreamLake (fp8) $0.42.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

  • Cheapest blended: StreamLake (fp8) at $0.2088 / $0.4176, far below the next group. GMICloud (fp8) is at $0.957 / $1.914 and Parasail (fp8) at $0.45 / $3.48.
  • Most expensive: Reka at $0.90 / $9.00.
  • 10 of the 15 providers list between $1.044 and $1.74 for input.

gpt-oss-120b

Reported
  1. Input2 tie at $0.03 → Cerebras $0.35 · 11.7x

  2. Output2 tie at $0.17 → SambaNova $0.95 · 5.6x

  3. Cache readDigitalOcean $0.012 → Cerebras $0.35 · 29.2x

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 20 rows, 3 series: Input, Output, Cache read. Input: highest Cerebras (fp16) $0.35. Lowest DekaLLM (bf16) $0.03. Output: highest SambaNova $0.95. Lowest DeepInfra (bf16) $0.17.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

  • Cheapest blended: CoreWeave (fp4) at $0.03 / $0.17. Close behind are DekaLLM (bf16) at $0.03 / $0.18 and DeepInfra (bf16) at $0.037 / $0.17.
  • Most expensive: Cerebras (fp16) at $0.35 / $0.75.
  • The precision trade: the cheapest endpoint reports fp4. DekaLLM and DeepInfra report bf16 at almost the same price. If precision matters to you, the cheapest bf16 endpoints cost about the same here.
  • Pairs: Groq vs Cerebras ($0.15 vs $0.35 input), DeepInfra vs Parasail and Together vs Fireworks.

Llama 3.3 70B Instruct

Reported
  1. InputDeepInfra $0.10 → Together $1.04 · 10.4x

  2. OutputDeepInfra $0.32 → Cloudflare $2.25 · 7.0x

  3. Cache readAkashML $0.10 → CoreWeave $0.71 · 7.1x

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 10 rows, 3 series: Input, Output, Cache read. Input: highest Together $1.04. Lowest DeepInfra (fp8) $0.1. Output: highest Cloudflare (fp8) $2.25. Lowest DeepInfra (fp8) $0.32.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

  • Cheapest blended: DeepInfra (fp8) at $0.10 / $0.32, then Novita (bf16) at $0.135 / $0.40.
  • Most expensive: Together at $1.04 / $1.04.
  • Check the context: Novita's endpoint reports a 12,288-token context and Cloudflare's 24,000. Most others report about 131,000. A short context can make a low price useless for long prompts.

GLM 5.3

Reported
  1. InputRelace $0.03 → 14 tie at $1.40 · 46.7x

  2. OutputNovita $1.32 → Relace $12.00 · 9.1x

  3. Cache readRelace $0.03 → 11 tie at $0.26 · 8.7x

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 32 rows, 3 series: Input, Output, Cache read. Input: highest AtlasCloud (fp8) $1.40. Lowest Relace $0.03. Output: highest Relace $12.00. Lowest Novita (fp8) $1.32.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

  • Cheapest blended: Novita (fp8) at $0.42 / $1.32.
  • Most expensive blended: Relace at $0.03 / $12.00. It has the lowest input price of all and the highest output price.
  • A crowd at one price: 14 providers list exactly $1.40 / $4.40.

Kimi K3 and Llama 4 Maverick

Reported
  1. InputRelace $0.83 → Alibaba $3.45 · 4.2x

  2. OutputPhala $9.75 → Alibaba $17.25 · 1.8x

  3. Cache readPhala $0.195 → AkashML $1.30 · 6.7x

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 19 rows, 3 series: Input, Output, Cache read. Input: highest Alibaba $3.45. Lowest Relace (fp4) $0.83. Output: highest Alibaba $17.25. Lowest Phala $9.75.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Kimi K3 is tight: all 19 providers sit between $0.83 and $3.45 for input and $9.75 and $17.25 for output. The cheapest blended is Relace (fp4) at $0.83 / $13.00; the most expensive is Alibaba at $3.45 / $17.25. Llama 4 Maverick has only three standard-tier providers: DigitalOcean at $0.1875 / $0.6525, Novita (fp8) at $0.27 / $0.85 and Parasail (fp8) at $0.35 / $1.00.

Five rules before you switch providers

  1. Use your own token mix. We rank by a 3:1 blend. An agent that reads large contexts and writes short patches is input-heavy, and the ranking can change. The Relace rows for DeepSeek V4 Flash and GLM 5.3 show how much.
  2. Read the quantization. fp4 and fp8 endpoints are often the cheapest. Lower precision can change output quality. We did not measure quality on any of these endpoints.
  3. Read the context length. A 12,288-token context is not a 131,072-token context at a discount.
  4. Look at cache-read prices. Repeated prompts are often billed as cache reads. On gpt-oss-120b, cache reads range from $0.012 (DigitalOcean) to $0.35 (Cerebras). Several providers list none.
  5. Refetch before you decide. This is one day's snapshot. Prices change often.

What we do not know

  • Speed. The keyless API returned latency and throughput for 0 of 265 endpoints. A cheaper provider is not shown to be slower or faster.
  • Quality. We did not run any task on these endpoints. Two endpoints with the same model name can differ in precision, context and serving settings.
  • Direct prices. These are prices through OpenRouter. A provider's own site can list a different price.

How we measured

  • Snapshot: OpenRouter's public, keyless API on 2026-10-06: the model list and the endpoint list for 27 curated models (28 requests). 265 endpoints from 52 providers.
  • Per-model charts: the standard tier only, one bar group per provider, its cheapest standard endpoint. Flex, priority, fast and regional endpoints are left out because they are priced differently on purpose; the endpoint table on the study page lists them.
  • Blended price = (3 × input + output) ÷ 4. Spread = most expensive ÷ cheapest blended price. Both are calculations on reported prices.
  • Comparison pages between providers name no winner. A reported price has no interval, so each gap is stated, not ranked.

Caveats

  • Third-party-reported prices, snapshot 2026-10-06. Not a measurement by Agent.
  • Tier detection is a heuristic read from the endpoint tag.
  • Curated list. 27 models, not every model on the gateway.

Know which provider did the work

Agent records the model, the provider route, the tokens and the cost of each call. Try Agent and compare your own spend by route.

The data behind this post

  • Inference
  • Providers

Inference provider index: 27 models, 52 providers

Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.

265endpoints, 52 providers, 27 models · Provider endpoints in the snapshot · n = 27

34 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.