Prices from OpenRouter's public API on 2026-10-06. They are reported, not measured by us, and they change often.
Open-weight models differ a lot by provider. The most expensive standard-tier provider charged this much more than the cheapest (blended 3 input : 1 output, our calculation): DeepSeek V4 Flash 12.6x, DeepSeek V4 Pro 11.2x, gpt-oss-120b 6.9x, Llama 3.3 70B 6.7x, GLM 5.3 4.7x.
Closed models do not: all 15 had one standard-tier price everywhere.
The cheapest endpoint often runs lower precision (fp4 or fp8) or a shorter context. A cheaper endpoint is not the same product.
We have no latency or quality data for these providers. The keyless API returned latency for 0 of 265 endpoints.
15 more models: the same price at every providerClaude Haiku 4.5, Claude Sonnet 5, Claude Sonnet 5.5, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Fable 5.1, GPT-6 Sol, GPT-6 Luna, GPT-6 Astra, GPT-5.5, Gemini 3.8 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.1 Pro Preview
= 1x
How much more the priciest provider charges than the cheapest: data
Item
Price spread
n
DeepSeek V4 Flash 0423
12.6x
15
DeepSeek V4 Pro 0423
11.2x
15
gpt-oss-120b
6.9x
20
Llama 3.3 70B Instruct
6.7x
10
GLM 5.3
4.7x
32
Kimi K3
1.8x
19
Llama 4 Maverick
1.7x
3
Claude Haiku 4.5
1x
4
Claude Sonnet 5
1x
5
Claude Sonnet 5.5
1x
5
Claude Opus 4.8
1x
5
Claude Opus 5
1x
5
Claude Opus 5.5
1x
5
Claude Fable 5.1
1x
4
GPT-6 Sol
1x
2
GPT-6 Luna
1x
2
GPT-6 Astra
1x
2
GPT-5.5
1x
2
Gemini 3.8 Flash
1x
2
Gemini 3.5 Flash
1x
2
Gemini 3.5 Flash Lite
1x
2
Gemini 3.1 Pro Preview
1x
2
List-price calculation, not a run. 22 rows. Highest DeepSeek V4 Flash 0423 12.6x (n 15). Lowest Gemini 3.1 Pro Preview 1x (n 2).
Notesn 2–32 per row
Standard tier, blended price (3 input : 1 output), most expensive provider ÷ cheapest provider
List-price calculation, not a run. 22 rows. Highest DeepSeek V4 Flash 0423 12.6x (n 15). Lowest Gemini 3.1 Pro Preview 1x (n 2).
A calculation on prices reported by OpenRouter’s public API, snapshot 2026-10-06. n = providers with a standard-tier endpoint. 1x means every provider charges the same. Highlighted: a spread of 4x or more.
For each model, we took every provider's cheapest standard-tier endpoint, blended its price at 3 input tokens to 1 output token, and divided the most expensive by the cheapest. A spread of 1x means every provider charges the same.
DeepSeek V4 Flash 0423: 12.6x across 15 providers.
Each chart shows one bar group per provider, sorted from the lowest to the highest blended price. A parenthesis names the quantization the provider reported. No parenthesis means the provider did not report one, which does not mean full precision. All prices are USD per million tokens.
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Cheapest blended: StreamLake (fp8) at $0.042 in / $0.084 out.
Next: DeepInfra (fp8) at $0.09 / $0.18, GMICloud (fp8) at $0.091 / $0.182, Venice at $0.0966 / $0.1925.
Most expensive: Cloudflare at $0.44 / $1.32.
Watch the mix: Relace (fp4) charges only $0.012 for input but $1.28 for output. For a job that reads a lot and writes little, it can beat the blended ranking.
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Cheapest blended: StreamLake (fp8) at $0.2088 / $0.4176, far below the next group. GMICloud (fp8) is at $0.957 / $1.914 and Parasail (fp8) at $0.45 / $3.48.
Most expensive: Reka at $0.90 / $9.00.
10 of the 15 providers list between $1.044 and $1.74 for input.
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Cheapest blended: CoreWeave (fp4) at $0.03 / $0.17. Close behind are DekaLLM (bf16) at $0.03 / $0.18 and DeepInfra (bf16) at $0.037 / $0.17.
Most expensive: Cerebras (fp16) at $0.35 / $0.75.
The precision trade: the cheapest endpoint reports fp4. DekaLLM and DeepInfra report bf16 at almost the same price. If precision matters to you, the cheapest bf16 endpoints cost about the same here.
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Cheapest blended: DeepInfra (fp8) at $0.10 / $0.32, then Novita (bf16) at $0.135 / $0.40.
Most expensive: Together at $1.04 / $1.04.
Check the context: Novita's endpoint reports a 12,288-token context and Cloudflare's 24,000. Most others report about 131,000. A short context can make a low price useless for long prompts.
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Kimi K3 is tight: all 19 providers sit between $0.83 and $3.45 for input and $9.75 and $17.25 for output. The cheapest blended is Relace (fp4) at $0.83 / $13.00; the most expensive is Alibaba at $3.45 / $17.25. Llama 4 Maverick has only three standard-tier providers: DigitalOcean at $0.1875 / $0.6525, Novita (fp8) at $0.27 / $0.85 and Parasail (fp8) at $0.35 / $1.00.
Five rules before you switch providers
Use your own token mix. We rank by a 3:1 blend. An agent that reads large contexts and writes short patches is input-heavy, and the ranking can change. The Relace rows for DeepSeek V4 Flash and GLM 5.3 show how much.
Read the quantization. fp4 and fp8 endpoints are often the cheapest. Lower precision can change output quality. We did not measure quality on any of these endpoints.
Read the context length. A 12,288-token context is not a 131,072-token context at a discount.
Look at cache-read prices. Repeated prompts are often billed as cache reads. On gpt-oss-120b, cache reads range from $0.012 (DigitalOcean) to $0.35 (Cerebras). Several providers list none.
Refetch before you decide. This is one day's snapshot. Prices change often.
What we do not know
Speed. The keyless API returned latency and throughput for 0 of 265 endpoints. A cheaper provider is not shown to be slower or faster.
Quality. We did not run any task on these endpoints. Two endpoints with the same model name can differ in precision, context and serving settings.
Direct prices. These are prices through OpenRouter. A provider's own site can list a different price.
How we measured
Snapshot: OpenRouter's public, keyless API on 2026-10-06: the model list and the endpoint list for 27 curated models (28 requests). 265 endpoints from 52 providers.
Per-model charts: the standard tier only, one bar group per provider, its cheapest standard endpoint. Flex, priority, fast and regional endpoints are left out because they are priced differently on purpose; the endpoint table on the study page lists them.
Blended price = (3 × input + output) ÷ 4. Spread = most expensive ÷ cheapest blended price. Both are calculations on reported prices.
Comparison pages between providers name no winner. A reported price has no interval, so each gap is stated, not ranked.
Caveats
Third-party-reported prices, snapshot 2026-10-06. Not a measurement by Agent.
Tier detection is a heuristic read from the endpoint tag.
Curated list. 27 models, not every model on the gateway.
OpenRouter charged the vendor's per-token price on 10 of 10 models we checked. The cost is a 5.5% credit fee. What we know, and what is still unmeasured.
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
20 min read
Turn the numbers into shipped work.
Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.