Explainer · Quantization
Quantization and cheap inference, explained
Definition
Quantization is the practice of storing and running a model's weights (and sometimes its activations and cache) in fewer bits, such as 8-bit (fp8) or 4-bit (fp4) numbers instead of 16-bit (bf16 or fp16). A quantized model needs less memory and less compute per token, so a provider can serve it faster or cheaper on the same hardware. The trade is accuracy: a lower precision can change the model's outputs, by a little or a lot, depending on the model, the method and the task.
Agent team · · 4 min read · Every number is from the public studies
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| Relace (fp4) | $0.83 | $13.00 | $0.45 |
| Phala | $1.95 | $9.75 | $0.2 |
| Sail Research (fp4) | $0.84 | $13.50 | $0.3 |
| Decart (mxfp4) | $2.01 | $10.05 | $0.2 |
| InferenceNet (fp4) | $0.95 | $14.00 | $0.31 |
| Wafer | $0.95 | $14.00 | $0.4 |
| Morph (fp8) | $1.27 | $13.30 | $0.28 |
| Makora | $1.53 | $12.75 | $0.2 |
| AkashML (fp4) | $1.30 | $14.00 | $1.30 |
| DigitalOcean | $2.55 | $12.95 | $0.26 |
| Together | $2.70 | $13.50 | $0.27 |
| DeepInfra (mxfp4) | $2.85 | $14.25 | $0.28 |
| BaseTen (fp8) | $3.00 | $15.00 | $0.3 |
| Chutes (mxfp4) | $3.00 | $15.00 | $0.3 |
| Fireworks | $3.00 | $15.00 | $0.3 |
| Modal (mxfp4) | $3.00 | $15.00 | $0.3 |
| Moonshot AI (mxfp4) | $3.00 | $15.00 | $0.3 |
| Parasail (fp4) | $3.00 | $15.00 | $0.3 |
| Alibaba | $3.45 | $17.25 | $0.34 |
Third-party reported values. 19 rows, 3 series: Input, Output, Cache read. Input: highest Alibaba $3.45. Lowest Relace (fp4) $0.83. Output: highest Alibaba $17.25. Lowest Phala $9.75.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Why bits matter for price
A model with tens of billions of parameters must hold every weight in memory and move it through the processor for every token. At 16 bits per weight that is two bytes each; at 8 bits, one; at 4 bits, half a byte. Fewer bytes per weight means:
- a model fits on fewer or smaller accelerators,
- more requests share one machine,
- each token moves less data, which is often the bottleneck.
All of that lowers the provider's cost per token, and in a competitive market, the price.
Quantization in real price lists
Open-weight models are hosted by many providers, and some of them report the precision they serve. In our snapshot of OpenRouter's public API (2026-10-06, third-party-reported), counted from the endpoint table of the provider index:
- 183 of 265 endpoints did not report a precision.
- 41 reported fp8, 25 fp4, 5 mxfp4 and 3 nvfp4.
- 6 reported bf16 and 2 fp16.
Here is gpt-oss-120b, one bar per provider's cheapest standard endpoint, with the reported precision in brackets:
The cheapest standard provider reported fp4; the most expensive reported fp16. The spread between them is 6.9x on a blended price (a calculation). Precision is not the only reason for the gap, since hardware, speed and margin differ too, but it is the first thing to check.
DeepSeek V4 Flash 0423 shows a second pattern:
- The cheapest standard provider on a blended price reported fp8.
- Two fp4 endpoints had the lowest input price but a much higher output price, so they are cheap only for input-heavy work.
- The most expensive provider reported no precision and listed a 384,000-token context, where most others listed about one million.
The spread here is 12.6x. A price comparison that ignores precision, context and the input-to-output mix compares different products.
What quantization changes in quality
The effect depends on the method:
- fp8 is close to the 16-bit model on many tasks, and is now a common serving format.
- fp4 and other 4-bit formats save more and risk more. Newer 4-bit formats (mxfp4, nvfp4) keep scale factors per small block of weights to limit the loss.
- Some models ship quantized. A model trained or released at a given precision may lose nothing when served at it. Check what the model's own authors publish.
We have not measured output quality per provider or per precision. The provider index is a price study; quality and latency through each endpoint need keyed runs. Until then, treat precision as a reason to test, not as a verdict.
How to buy cheap inference safely
- Check the endpoint, not only the model name: precision, context length, tier and region.
- Price your own token mix. A provider that is cheap on input and expensive on output can lose on a chat workload. The AI cost calculator prices a mix through any provider.
- Run your evaluation on the endpoint you will use. A model's published score was measured at its own precision, on its own serving stack.
- Pin the provider when quality matters, so a gateway does not move you to a cheaper, lower-precision endpoint.
- Watch for silent changes. Providers change hardware and formats; re-run your checks after a price drop.
Frequently asked questions
Does quantization reduce model quality?
It can. 8-bit serving is often close to the 16-bit model; 4-bit saves more and risks more, depending on the model and the method. We have not measured quality per precision; test on your own tasks with the endpoint you plan to use.
What is the difference between fp8 and fp4?
They are number formats with 8 and 4 bits per value. fp4 needs half the memory of fp8 for the weights, which lowers cost, but it represents each value less precisely.
Why is the same open model so much cheaper at some providers?
Precision is one reason: in our snapshot the cheapest gpt-oss-120b provider reported fp4 and the most expensive reported fp16. Hardware, context length, speed and margin are others. Closed models showed no such spread: every standard-tier provider charged one price.
How do I know which precision a provider uses?
Look at the endpoint details your gateway or provider publishes. In our snapshot, 183 of 265 endpoints did not report a precision, so the absence of a label does not mean full precision.
Watch the data
Same model, different price: 265 provider endpoints compared
Reported by OpenRouter’s public API on 2026-10-06: open-weight prices vary up to 12.6x across providers; closed models have one standard price, and the gateway adds a 5.5% credit fee.
Transcript
- Provider index · third-party-reported · 2026-10-06. Same model, different price. 265 endpoints, 52 providers, 27 models. Prices as OpenRouter’s public API reports them.
- 265 endpoints, 52 providers, 27 models. The keyless API returned latency for 0 of them, so speed is not compared. Provider endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Endpoints with a latency figure: 0 of 265. Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
- Open-weight models vary most: DeepSeek V4 Flash 12.6×, DeepSeek V4 Pro 11.2×. All 15 closed models with 2+ providers: one standard price (1×). Chart: Priciest ÷ cheapest standard provider · blended price (n = 2–32 each). Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
- One model, 10 providers: Llama 3.3 70B input costs $0.10 per million tokens at DeepInfra (fp8) and $1.04 at Together. Chart: Llama 3.3 70B · USD per million tokens · one row per provider. Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
- Gateway markup: for 10 of 10 models OpenRouter charges the vendor’s per-token price. The cost is a 5.5% fee when you buy credits. Models at the first-party per-token price: 10 of 10. Credit-purchase fee, Standard plan by card: 5.5% ($0.80 minimum by card). Effective markup per token after the fee: +5.5% (calculation). Caveat: First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.
- Example: Sonnet 5.5 costs $2.00 in and $10.00 out per million tokens on both. A price is not a measurement: no winner. Chart: Claude Sonnet 5.5 · OpenRouter vs Anthropic list price · USD per million. Caveat: Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.
- Prices change often: every endpoint, tier and source date online.