Explainer · Open-weight model
Open-weight models: why the same model has many prices
Definition
Open-weight models are AI models whose trained weights are public, so anyone who accepts the license can download and run them. Many companies host the same open model. Our snapshot shows a price spread up to 12.6x (calculation, 15 standard-tier providers for DeepSeek V4 Flash 0423). All 15 closed models with 2 or more providers had one standard-tier price across those providers. This applies to our snapshot, not every closed model.
Agent team · · 5 min read · Every number is from the public studies
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| StreamLake (fp8) | $0.042 | $0.084 | $0.0084 |
| DeepInfra (fp8) | $0.09 | $0.18 | $0.018 |
| GMICloud (fp8) | $0.091 | $0.18 | $0.018 |
| Venice | $0.097 | $0.19 | $0.02 |
| DigitalOcean | $0.098 | $0.2 | $0.02 |
| Alibaba (fp8) | $0.13 | $0.27 | $0.027 |
| SiliconFlow (fp8) | $0.13 | $0.28 | $0.028 |
| AtlasCloud (fp4) | $0.14 | $0.28 | $0.028 |
| Baidu (fp8) | $0.14 | $0.28 | $0.028 |
| Novita (fp8) | $0.14 | $0.28 | $0.028 |
| Parasail (fp8) | $0.14 | $0.28 | $0.07 |
| Mancer 2 (fp8) | $0.19 | $0.5 | — |
| Relace (fp4) | $0.012 | $1.28 | $0.012 |
| OpenInference (fp4) | $0.013 | $1.41 | $0.013 |
| Cloudflare | $0.44 | $1.32 | $0.014 |
Third-party reported values. 15 rows, 3 series: Input, Output, Cache read. Input: highest Cloudflare $0.44. Lowest Relace (fp4) $0.012. Output: highest OpenInference (fp4) $1.41. Lowest StreamLake (fp8) $0.084.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
What open-weight means
The weights are the numbers a model learned in training. An open-weight maker publishes them. A closed-model maker keeps them and sells access to a service that runs them.
Open-weight is not the same as open-source. Open source means the source code is public under a license that lets you use, change and share it. An open-weight release may come without the training data or code, and its license may limit use.
Closed models still run in more than one place. In the provider index, Claude Sonnet 5.5 had 5 standard-tier providers. They were Anthropic, Amazon Bedrock, Azure, Claude Platform on AWS and Google Vertex. All five standard-tier endpoints listed $2.00 per million input tokens and $10.00 per million output tokens.
One open model, many prices
The provider index is one snapshot of OpenRouter's public API on 2026-10-06. It listed 265 endpoints from 52 providers for 27 curated models. This is a selected model set, not the whole market. The prices are third-party-reported, not measured by us. They apply through OpenRouter and may differ from each provider's direct prices. All prices below are USD per million tokens, before gateway credit-purchase fees.
We select each provider's cheapest standard-tier endpoint by blended price. Tier labels come from endpoint tags; this is a heuristic. Spread = highest blended price ÷ lowest blended price among the selected standard-tier endpoints. The blend is 3 input tokens to 1 output token. The spread is a calculation on reported prices.
The study names four open-weight models with a wide spread. Each spread below is a calculation on reported prices:
- DeepSeek V4 Flash 0423: 12.6x across 15 providers. Cloudflare was the most expensive and StreamLake (fp8) the cheapest.
- DeepSeek V4 Pro 0423: 11.2x across 15 providers.
- gpt-oss-120b: 6.9x across 20 providers.
- Llama 3.3 70B Instruct: 6.7x across 10 providers.
Three more models also differ: GLM 5.3 at 4.7x (32 providers), Kimi K3 at 1.8x (19) and Llama 4 Maverick at 1.7x (3).
The provider count does not set the spread alone. The 15 closed models with 2 or more providers (Claude, GPT and Gemini) all sat at 1x. These are snapshot price ranges, not 95% intervals from model calls. A price has no interval, so we name no winner. We state the gap.
The same model at 20 providers
For gpt-oss-120b, input prices ran from $0.03 per million tokens (CoreWeave and DekaLLM) to $0.35 (Cerebras). Output prices ran from $0.17 to $0.95. The blended price is a calculation: (3 × input + output) ÷ 4. For CoreWeave that is (3 × 0.03 + 0.17) ÷ 4 = $0.065. For Cerebras it is (3 × 0.35 + 0.75) ÷ 4 = $0.45. The ratio is 6.9x (calculation).
What a low price can hide
A price belongs to an endpoint, not to a model. Four things change what you buy.
- Precision: some providers report the precision they serve. For gpt-oss-120b, 12 of 20 did: fp4 at 5, bf16 at 4, fp8 at 2 and fp16 at 1. CoreWeave, the cheapest of the 20, reported fp4 ($0.065 blended, calculation). Cerebras, the most expensive, reported fp16 ($0.45 blended, calculation). The precision labels do not explain the price order: the next three cheapest (DekaLLM, DeepInfra and AkashML) reported bf16. No label does not mean full precision.
- Context length. Of 10 Llama 3.3 70B Instruct providers, 6 list 131,072 tokens. Novita lists 12,288 tokens at $0.135 input and $0.40 output. AkashML lists 131,072 at $0.20 and $0.52, so Novita has the lower price and the shorter context. A short context does not always mean a low price: Cloudflare lists 24,000 and had the second-highest blended price of the 10.
- Input and output mix. The blend assumes 3 input tokens for each output token. SambaNova lists gpt-oss-120b at $0.14 for input but $0.95 for output, the highest output price of the 20. Rank by your own mix.
- Quality and speed. We measured neither. The keyless API gave a latency figure for 0 of 265 endpoints. A cheaper provider is not shown to be slower, faster, better or worse.
Llama 3.3 70B Instruct runs from $0.10 input and $0.32 output at DeepInfra (fp8) to $1.04 and $1.04 at Together. The calculation is (3 × 0.10 + 0.32) ÷ 4 = $0.155 against (3 × 1.04 + 1.04) ÷ 4 = $1.04, a 6.7x spread.
How to choose a provider
- Test your own tasks on two or three providers. Send each the same prompt. Keep every call in the count. Read each pass rate with its n and 95% interval (how to read AI benchmarks honestly).
- Price your own token mix in the AI cost calculator.
- Compare standard with standard. Providers price flex, priority, fast and regional endpoints differently on purpose.
- Read the endpoint, not the model name: precision, context and tier. See quantization and cheap inference.
- Check the price again before you rely on it. Prices change often. Some providers have a page, such as DeepInfra.
- Decide if you want one key for many providers. Inference gateways and OpenRouter explains what a gateway costs.
The prices for DeepSeek V4, gpt-oss-120b, Llama, Kimi K3 and GLM 5.3 are in the cheapest place to run open models.
Frequently asked questions
What is the difference between open-weight and open-source?
Open-weight means the trained weights are public. Open source also requires a license that lets you use, change and share the code. Training data and code may stay private. Read the license.
Why is the same open model cheaper at some providers?
We see the gap but not its cause. The snapshot reports precision, context, tier and price differences. We did not test hardware, speed or margin.
Is the cheapest provider the same quality?
We did not measure quality or speed. Price alone does not make endpoints equal. Test your own tasks before you switch.
Are open-weight models cheaper than Claude or GPT?
Some have lower reported token prices, but we did not compare quality. Sonnet 5.5's blended price is $4.00 at all 5 standard-tier providers (calculation: (3 × 2.00 + 10.00) ÷ 4). gpt-oss-120b ranges from $0.065 to $0.45 at 20 providers (calculation). A closed model can be cheap too: GPT-6 Luna lists $0.10 input and $0.50 output at 2 standard-tier providers.
Watch the data
Same model, different price: 265 provider endpoints compared
Reported by OpenRouter’s public API on 2026-10-06: open-weight prices vary up to 12.6x across providers; closed models have one standard price, and the gateway adds a 5.5% credit fee.
Transcript
- Provider index · third-party-reported · 2026-10-06. Same model, different price. 265 endpoints, 52 providers, 27 models. Prices as OpenRouter’s public API reports them.
- 265 endpoints, 52 providers, 27 models. The keyless API returned latency for 0 of them, so speed is not compared. Provider endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Endpoints with a latency figure: 0 of 265. Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
- Open-weight models vary most: DeepSeek V4 Flash 12.6×, DeepSeek V4 Pro 11.2×. All 15 closed models with 2+ providers: one standard price (1×). Chart: Priciest ÷ cheapest standard provider · blended price (n = 2–32 each). Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
- One model, 10 providers: Llama 3.3 70B input costs $0.10 per million tokens at DeepInfra (fp8) and $1.04 at Together. Chart: Llama 3.3 70B · USD per million tokens · one row per provider. Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
- Gateway markup: for 10 of 10 models OpenRouter charges the vendor’s per-token price. The cost is a 5.5% fee when you buy credits. Models at the first-party per-token price: 10 of 10. Credit-purchase fee, Standard plan by card: 5.5% ($0.80 minimum by card). Effective markup per token after the fee: +5.5% (calculation). Caveat: First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.
- Example: Sonnet 5.5 costs $2.00 in and $10.00 out per million tokens on both. A price is not a measurement: no winner. Chart: Claude Sonnet 5.5 · OpenRouter vs Anthropic list price · USD per million. Caveat: Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.
- Prices change often: every endpoint, tier and source date online.