Inference provider index: 27 models, 52 providers
For the same model, how much do inference providers differ in price, and what does a gateway such as OpenRouter add over the first-party list price?
Published · 34 charts · Download the data or a carousel
265
The answer
OpenRouter’s public API listed 265 endpoints from 52 providers for 27 models on 2026-10-06. For 10 of the 10 models with a first-party list price, OpenRouter’s per-token price was the same as the vendor’s; the cost of the gateway is the 5.5% credit-purchase fee (Standard plan, third-party-reported), so the effective markup is +5.5%. 15 of 15 closed models (Claude, GPT, Gemini) with 2 or more providers had one price across every standard-tier provider; their price differences come from named tiers (flex, priority, fast) and regional endpoints. Open-weight models differ by provider: DeepSeek V4 Flash 0423 12.6x (15 providers), DeepSeek V4 Pro 0423 11.2x (15 providers), gpt-oss-120b 6.9x (20 providers), Llama 3.3 70B Instruct 6.7x (10 providers) between the most expensive and the cheapest standard-tier provider (a calculation on reported prices). The cheapest endpoints often report lower precision or a shorter context. Latency and throughput were not in the keyless API (0 of 265 endpoints), and the gateway’s own delay is not measured: no key in the environment.
Live story
Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.
Same model, different price: 265 provider endpoints compared
Reported by OpenRouter’s public API on 2026-10-06: open-weight prices vary up to 12.6x across providers; closed models have one standard price, and the gateway adds a 5.5% credit fee.
Transcript
- Provider index · third-party-reported · 2026-10-06. Same model, different price. 265 endpoints, 52 providers, 27 models. Prices as OpenRouter’s public API reports them.
- 265 endpoints, 52 providers, 27 models. The keyless API returned latency for 0 of them, so speed is not compared. Provider endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Endpoints with a latency figure: 0 of 265. Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
- Open-weight models vary most: DeepSeek V4 Flash 12.6×, DeepSeek V4 Pro 11.2×. All 15 closed models with 2+ providers: one standard price (1×). Chart: Priciest ÷ cheapest standard provider · blended price (n = 2–32 each). Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
- One model, 10 providers: Llama 3.3 70B input costs $0.10 per million tokens at DeepInfra (fp8) and $1.04 at Together. Chart: Llama 3.3 70B · USD per million tokens · one row per provider. Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
- Gateway markup: for 10 of 10 models OpenRouter charges the vendor’s per-token price. The cost is a 5.5% fee when you buy credits. Models at the first-party per-token price: 10 of 10. Credit-purchase fee, Standard plan by card: 5.5% ($0.80 minimum by card). Effective markup per token after the fee: +5.5% (calculation). Caveat: First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.
- Example: Sonnet 5.5 costs $2.00 in and $10.00 out per million tokens on both. A price is not a measurement: no winner. Chart: Claude Sonnet 5.5 · OpenRouter vs Anthropic list price · USD per million. Caveat: Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.
- Prices change often: every endpoint, tier and source date online.
Key numbers
12.6x
Largest standard-tier price spread (DeepSeek V4 Flash 0423)
(most expensive: Cloudflare; cheapest: StreamLake (fp8)) · n = 15
10 of 10
Models where OpenRouter’s per-token price equals the first-party list price
n = 10
5.5%
Credit-purchase fee on OpenRouter’s Standard plan (third-party-reported)
($0.80 minimum by card)
0 of 265
Endpoints with a latency figure in the keyless API
n = 265
Compare providers for one model
Pick a model, then sort or filter its providers. The study charts below show the same standard-tier prices; this table also lists every named tier and regional endpoint.
Price per provider: Claude Haiku 4.5
Reported by OpenRouter’s public API, snapshot October 6, 2026. Third-party-reported prices, not measured by Agent.
4 providers with a standard-tier endpoint · all 4 providers report the same price ($2.00 blended) · showing 4 of 4 rows
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
- first-party list price
| vs cheapest | ||||||||
|---|---|---|---|---|---|---|---|---|
| Amazon Bedrock | not reported | 200k | $1.00 | $5.00 | $0.10 | $2.00 | 100% | 1x |
| Anthropic | not reported | 200k | $1.00 | $5.00 | $0.10 | $2.00 | 100% | 1x |
| Azure | not reported | 200k | $1.00 | $5.00 | $0.10 | $2.00 | 100% | 1x |
| Google Vertex | not reported | 200k | $1.00 | $5.00 | $0.10 | $2.00 | 100% | 1x |
Blended price and “vs cheapest” are calculations on the reported prices (3 input : 1 output tokens), against the cheapest standard-tier provider. A lower price can come with lower precision (fp4, fp8) or a shorter context. Uptime is what the API reported for the last day; the API reported no latency or throughput.
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
- DeepSeek V4 Flash 0423
- DeepSeek V4 Pro 0423
- gpt-oss-120b
- Llama 3.3 70B Instruct
- GLM 5.3
- Kimi K3
- Llama 4 Maverick
- 15 more models: the same price at every provider1x
| Item | Price spread | n |
|---|---|---|
| DeepSeek V4 Flash 0423 | 12.6x | 15 |
| DeepSeek V4 Pro 0423 | 11.2x | 15 |
| gpt-oss-120b | 6.9x | 20 |
| Llama 3.3 70B Instruct | 6.7x | 10 |
| GLM 5.3 | 4.7x | 32 |
| Kimi K3 | 1.8x | 19 |
| Llama 4 Maverick | 1.7x | 3 |
| Claude Haiku 4.5 | 1x | 4 |
| Claude Sonnet 5 | 1x | 5 |
| Claude Sonnet 5.5 | 1x | 5 |
| Claude Opus 4.8 | 1x | 5 |
| Claude Opus 5 | 1x | 5 |
| Claude Opus 5.5 | 1x | 5 |
| Claude Fable 5.1 | 1x | 4 |
| GPT-6 Sol | 1x | 2 |
| GPT-6 Luna | 1x | 2 |
| GPT-6 Astra | 1x | 2 |
| GPT-5.5 | 1x | 2 |
| Gemini 3.8 Flash | 1x | 2 |
| Gemini 3.5 Flash | 1x | 2 |
| Gemini 3.5 Flash Lite | 1x | 2 |
| Gemini 3.1 Pro Preview | 1x | 2 |
List-price calculation, not a run. 22 rows. Highest DeepSeek V4 Flash 0423 12.6x (n 15). Lowest Gemini 3.1 Pro Preview 1x (n 2).
Notesn 2–32 per row
Standard tier, blended price (3 input : 1 output), most expensive provider ÷ cheapest provider
A calculation on prices reported by OpenRouter’s public API, snapshot 2026-10-06. n = providers with a standard-tier endpoint. 1x means every provider charges the same. Highlighted: a spread of 4x or more.
Source: OpenRouter public API: models and provider endpoints (snapshot)
0%per-token markup on 10 of 10 models
5.5% with the card credit feeCalculation
- Claude Haiku 4.5
- Claude Sonnet 5
- Claude Sonnet 5.5
- Claude Opus 4.8
- Claude Opus 5
- Claude Opus 5.5
- Claude Fable 5.1
- GPT-6 Luna
- Gemini 3.8 Flash
- Gemini 3.5 Flash
- per-token markup over the first-party list price (input and output)
- With the 5.5% card credit fee (input) (calculation, hollow)
| Item | Per-token markup (input) | Per-token markup (output) | With the 5.5% card credit fee (input) |
|---|---|---|---|
| Claude Haiku 4.5 | 0% | 0% | 5.5% |
| Claude Sonnet 5 | 0% | 0% | 5.5% |
| Claude Sonnet 5.5 | 0% | 0% | 5.5% |
| Claude Opus 4.8 | 0% | 0% | 5.5% |
| Claude Opus 5 | 0% | 0% | 5.5% |
| Claude Opus 5.5 | 0% | 0% | 5.5% |
| Claude Fable 5.1 | 0% | 0% | 5.5% |
| GPT-6 Luna | 0% | 0% | 5.5% |
| Gemini 3.8 Flash | 0% | 0% | 5.5% |
| Gemini 3.5 Flash | 0% | 0% | 5.5% |
List-price calculation, not a run. 10 rows, 3 series: Per-token markup (input), Per-token markup (output), With the 5.5% card credit fee (input). Per-token markup (input): all at 0%. Per-token markup (output): all at 0%.
Notes
Per-token markup, and the markup after the Standard credit-purchase fee (calculation)
A calculation: OpenRouter list price ÷ first-party list price − 1, then × (1 + 5.5%) for credits bought by card on the Standard plan (purchases large enough that the minimum fee does not apply). Fees as published on 2026-10-06; first-party prices from the vendors’ list prices.
Sources: OpenRouter public API: models and provider endpoints (snapshot), OpenRouter pricing and fees, Anthropic list prices (Claude models), OpenAI list prices, Google Gemini list prices
Per-model price charts (32)
OpenRouter
Anthropicfirst-party list price
Same price on all 2 price types- InputOpenRouter: $1.00 same priceAnthropic: $1.00
- OutputOpenRouter: $5.00 same priceAnthropic: $5.00
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output |
|---|---|---|
| OpenRouter | $1.00 | $5.00 |
| Anthropic (first-party list price) | $1.00 | $5.00 |
Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $1.00. Output: all at $5.00.
Notes
USD per million tokens, before any credit-purchase fee
OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.
Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)
OpenRouter
Anthropicfirst-party list price
Same price on all 2 price types- InputOpenRouter: $2.00 same priceAnthropic: $2.00
- OutputOpenRouter: $10.00 same priceAnthropic: $10.00
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output |
|---|---|---|
| OpenRouter | $2.00 | $10.00 |
| Anthropic (first-party list price) | $2.00 | $10.00 |
Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $2.00. Output: all at $10.00.
Notes
USD per million tokens, before any credit-purchase fee
OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.
Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)
OpenRouter
Anthropicfirst-party list price
Same price on all 2 price types- InputOpenRouter: $2.00 same priceAnthropic: $2.00
- OutputOpenRouter: $10.00 same priceAnthropic: $10.00
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output |
|---|---|---|
| OpenRouter | $2.00 | $10.00 |
| Anthropic (first-party list price) | $2.00 | $10.00 |
Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $2.00. Output: all at $10.00.
Notes
USD per million tokens, before any credit-purchase fee
OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.
Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)
OpenRouter
Anthropicfirst-party list price
Same price on all 2 price types- InputOpenRouter: $5.00 same priceAnthropic: $5.00
- OutputOpenRouter: $25.00 same priceAnthropic: $25.00
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output |
|---|---|---|
| OpenRouter | $5.00 | $25.00 |
| Anthropic (first-party list price) | $5.00 | $25.00 |
Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $5.00. Output: all at $25.00.
Notes
USD per million tokens, before any credit-purchase fee
OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.
Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)
OpenRouter
Anthropicfirst-party list price
Same price on all 2 price types- InputOpenRouter: $5.00 same priceAnthropic: $5.00
- OutputOpenRouter: $25.00 same priceAnthropic: $25.00
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output |
|---|---|---|
| OpenRouter | $5.00 | $25.00 |
| Anthropic (first-party list price) | $5.00 | $25.00 |
Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $5.00. Output: all at $25.00.
Notes
USD per million tokens, before any credit-purchase fee
OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.
Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)
OpenRouter
Anthropicfirst-party list price
Same price on all 2 price types- InputOpenRouter: $4.00 same priceAnthropic: $4.00
- OutputOpenRouter: $20.00 same priceAnthropic: $20.00
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output |
|---|---|---|
| OpenRouter | $4.00 | $20.00 |
| Anthropic (first-party list price) | $4.00 | $20.00 |
Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $4.00. Output: all at $20.00.
Notes
USD per million tokens, before any credit-purchase fee
OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.
Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)
OpenRouter
Anthropicfirst-party list price
Same price on all 2 price types- InputOpenRouter: $10.00 same priceAnthropic: $10.00
- OutputOpenRouter: $50.00 same priceAnthropic: $50.00
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output |
|---|---|---|
| OpenRouter | $10.00 | $50.00 |
| Anthropic (first-party list price) | $10.00 | $50.00 |
Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $10.00. Output: all at $50.00.
Notes
USD per million tokens, before any credit-purchase fee
OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.
Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)
OpenRouter
OpenAIfirst-party list price
Same price on all 2 price types- InputOpenRouter: $0.10 same priceOpenAI: $0.10
- OutputOpenRouter: $0.50 same priceOpenAI: $0.50
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output |
|---|---|---|
| OpenRouter | $0.1 | $0.5 |
| OpenAI (first-party list price) | $0.1 | $0.5 |
Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $0.1. Output: all at $0.5.
Notes
USD per million tokens, before any credit-purchase fee
OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. OpenAI price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.
Sources: OpenRouter public API: models and provider endpoints (snapshot), OpenAI list prices
OpenRouter
Google AI Studiofirst-party list price
Same price on all 2 price types- InputOpenRouter: $0.75 same priceGoogle AI Studio: $0.75
- OutputOpenRouter: $3.75 same priceGoogle AI Studio: $3.75
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output |
|---|---|---|
| OpenRouter | $0.75 | $3.75 |
| Google AI Studio (first-party list price) | $0.75 | $3.75 |
Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $0.75. Output: all at $3.75.
Notes
USD per million tokens, before any credit-purchase fee
OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Google price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.
Sources: OpenRouter public API: models and provider endpoints (snapshot), Google Gemini list prices
OpenRouter
Google AI Studiofirst-party list price
Same price on all 2 price types- InputOpenRouter: $1.50 same priceGoogle AI Studio: $1.50
- OutputOpenRouter: $9.00 same priceGoogle AI Studio: $9.00
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output |
|---|---|---|
| OpenRouter | $1.50 | $9.00 |
| Google AI Studio (first-party list price) | $1.50 | $9.00 |
Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $1.50. Output: all at $9.00.
Notes
USD per million tokens, before any credit-purchase fee
OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Google price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.
Sources: OpenRouter public API: models and provider endpoints (snapshot), Google Gemini list prices
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| Amazon Bedrock | $1.00 | $5.00 | $0.1 |
| Anthropic | $1.00 | $5.00 | $0.1 |
| Azure | $1.00 | $5.00 | $0.1 |
| Google Vertex | $1.00 | $5.00 | $0.1 |
Third-party reported values. 4 rows, 3 series: Input, Output, Cache read. Input: all at $1.00. Output: all at $5.00.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| Amazon Bedrock | $2.00 | $10.00 | $0.2 |
| Anthropic | $2.00 | $10.00 | $0.2 |
| Azure | $2.00 | $10.00 | $0.2 |
| Claude Platform on AWS | $2.00 | $10.00 | $0.2 |
| Google Vertex | $2.00 | $10.00 | $0.2 |
Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $2.00. Output: all at $10.00.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| Amazon Bedrock | $2.00 | $10.00 | $0.2 |
| Anthropic | $2.00 | $10.00 | $0.2 |
| Azure | $2.00 | $10.00 | $0.2 |
| Claude Platform on AWS | $2.00 | $10.00 | $0.2 |
| Google Vertex | $2.00 | $10.00 | $0.2 |
Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $2.00. Output: all at $10.00.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| Amazon Bedrock | $5.00 | $25.00 | $0.5 |
| Anthropic | $5.00 | $25.00 | $0.5 |
| Azure | $5.00 | $25.00 | $0.5 |
| Claude Platform on AWS | $5.00 | $25.00 | $0.5 |
| Google Vertex | $5.00 | $25.00 | $0.5 |
Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $5.00. Output: all at $25.00.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| Amazon Bedrock | $5.00 | $25.00 | $0.5 |
| Anthropic | $5.00 | $25.00 | $0.5 |
| Azure | $5.00 | $25.00 | $0.5 |
| Claude Platform on AWS | $5.00 | $25.00 | $0.5 |
| Google Vertex | $5.00 | $25.00 | $0.5 |
Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $5.00. Output: all at $25.00.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| Amazon Bedrock | $4.00 | $20.00 | $0.2 |
| Anthropic | $4.00 | $20.00 | $0.2 |
| Azure | $4.00 | $20.00 | $0.2 |
| Claude Platform on AWS | $4.00 | $20.00 | $0.2 |
| Google Vertex | $4.00 | $20.00 | $0.2 |
Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $4.00. Output: all at $20.00.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| Amazon Bedrock | $10.00 | $50.00 | $0.25 |
| Anthropic | $10.00 | $50.00 | $0.25 |
| Azure | $10.00 | $50.00 | $0.25 |
| Google Vertex | $10.00 | $50.00 | $0.25 |
Third-party reported values. 4 rows, 3 series: Input, Output, Cache read. Input: all at $10.00. Output: all at $50.00.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Azure
OpenAI
Same price on all 3 price types- InputAzure: $2.00 same priceOpenAI: $2.00
- OutputAzure: $10.00 same priceOpenAI: $10.00
- Cache readAzure: $0.20 same priceOpenAI: $0.20
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output | Cache read |
|---|---|---|---|
| Azure | $2.00 | $10.00 | $0.2 |
| OpenAI | $2.00 | $10.00 | $0.2 |
Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $2.00. Output: all at $10.00.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Azure
OpenAI
Same price on all 3 price types- InputAzure: $0.10 same priceOpenAI: $0.10
- OutputAzure: $0.50 same priceOpenAI: $0.50
- Cache readAzure: $0.01 same priceOpenAI: $0.01
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output | Cache read |
|---|---|---|---|
| Azure | $0.1 | $0.5 | $0.01 |
| OpenAI | $0.1 | $0.5 | $0.01 |
Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $0.1. Output: all at $0.5.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Azure
OpenAI
Same price on all 3 price types- InputAzure: $10.00 same priceOpenAI: $10.00
- OutputAzure: $50.00 same priceOpenAI: $50.00
- Cache readAzure: $1.00 same priceOpenAI: $1.00
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output | Cache read |
|---|---|---|---|
| Azure | $10.00 | $50.00 | $1.00 |
| OpenAI | $10.00 | $50.00 | $1.00 |
Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $10.00. Output: all at $50.00.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Azure
OpenAI
Same price on all 3 price types- InputAzure: $5.00 same priceOpenAI: $5.00
- OutputAzure: $30.00 same priceOpenAI: $30.00
- Cache readAzure: $0.50 same priceOpenAI: $0.50
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output | Cache read |
|---|---|---|---|
| Azure | $5.00 | $30.00 | $0.5 |
| OpenAI | $5.00 | $30.00 | $0.5 |
Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $5.00. Output: all at $30.00.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| CoreWeave (fp4) | $0.03 | $0.17 | $0.03 |
| DekaLLM (bf16) | $0.03 | $0.18 | $0.03 |
| DeepInfra (bf16) | $0.037 | $0.17 | — |
| AkashML (bf16) | $0.037 | $0.19 | $0.037 |
| Mancer 2 (fp8) | $0.045 | $0.25 | — |
| Crusoe (bf16) | $0.05 | $0.25 | $0.05 |
| Novita (fp4) | $0.05 | $0.25 | — |
| DigitalOcean | $0.06 | $0.42 | $0.012 |
| Google Vertex | $0.09 | $0.36 | — |
| BaseTen (fp4) | $0.1 | $0.5 | $0.1 |
| Amazon Bedrock | $0.15 | $0.6 | — |
| Groq | $0.15 | $0.6 | $0.075 |
| Nebius (fp4) | $0.15 | $0.6 | — |
| Phala | $0.15 | $0.6 | — |
| SiliconFlow (fp8) | $0.15 | $0.6 | $0.075 |
| Together | $0.15 | $0.6 | — |
| Parasail (fp4) | $0.1 | $0.75 | $0.055 |
| Mara | $0.15 | $0.75 | — |
| SambaNova | $0.14 | $0.95 | — |
| Cerebras (fp16) | $0.35 | $0.75 | $0.35 |
Third-party reported values. 20 rows, 3 series: Input, Output, Cache read. Input: highest Cerebras (fp16) $0.35. Lowest DekaLLM (bf16) $0.03. Output: highest SambaNova $0.95. Lowest DeepInfra (bf16) $0.17.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Google AI Studio
Google Vertex
Same price on all 3 price types- InputGoogle AI Studio: $0.75 same priceGoogle Vertex: $0.75
- OutputGoogle AI Studio: $3.75 same priceGoogle Vertex: $3.75
- Cache readGoogle AI Studio: $0.075 same priceGoogle Vertex: $0.075
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output | Cache read |
|---|---|---|---|
| Google AI Studio | $0.75 | $3.75 | $0.075 |
| Google Vertex | $0.75 | $3.75 | $0.075 |
Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $0.75. Output: all at $3.75.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Google AI Studio
Google Vertex
Same price on all 3 price types- InputGoogle AI Studio: $1.50 same priceGoogle Vertex: $1.50
- OutputGoogle AI Studio: $9.00 same priceGoogle Vertex: $9.00
- Cache readGoogle AI Studio: $0.15 same priceGoogle Vertex: $0.15
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output | Cache read |
|---|---|---|---|
| Google AI Studio | $1.50 | $9.00 | $0.15 |
| Google Vertex | $1.50 | $9.00 | $0.15 |
Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $1.50. Output: all at $9.00.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Google AI Studio
Google Vertex
Same price on all 3 price types- InputGoogle AI Studio: $0.30 same priceGoogle Vertex: $0.30
- OutputGoogle AI Studio: $2.50 same priceGoogle Vertex: $2.50
- Cache readGoogle AI Studio: $0.03 same priceGoogle Vertex: $0.03
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output | Cache read |
|---|---|---|---|
| Google AI Studio | $0.3 | $2.50 | $0.03 |
| Google Vertex | $0.3 | $2.50 | $0.03 |
Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $0.3. Output: all at $2.50.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Google AI Studio
Google Vertex
Same price on all 3 price types- InputGoogle AI Studio: $2.00 same priceGoogle Vertex: $2.00
- OutputGoogle AI Studio: $12.00 same priceGoogle Vertex: $12.00
- Cache readGoogle AI Studio: $0.20 same priceGoogle Vertex: $0.20
USD per million tokens, from $0. Equal prices draw one dot.
| Item | Input | Output | Cache read |
|---|---|---|---|
| Google AI Studio | $2.00 | $12.00 | $0.2 |
| Google Vertex | $2.00 | $12.00 | $0.2 |
Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $2.00. Output: all at $12.00.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Input
Output
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output |
|---|---|---|
| DigitalOcean | $0.19 | $0.65 |
| Novita (fp8) | $0.27 | $0.85 |
| Parasail (fp8) | $0.35 | $1.00 |
Third-party reported values. 3 rows, 2 series: Input, Output. Input: highest Parasail (fp8) $0.35. Lowest DigitalOcean $0.19. Output: highest Parasail (fp8) $1.00. Lowest DigitalOcean $0.65.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| DeepInfra (fp8) | $0.1 | $0.32 | — |
| Novita (bf16) | $0.14 | $0.4 | — |
| AkashML (fp8) | $0.2 | $0.52 | $0.1 |
| Parasail (fp8) | $0.22 | $0.5 | $0.11 |
| SambaNova | $0.45 | $0.9 | — |
| Groq | $0.59 | $0.79 | $0.29 |
| CoreWeave (fp16) | $0.71 | $0.71 | $0.71 |
| Google Vertex | $0.72 | $0.72 | — |
| Cloudflare (fp8) | $0.29 | $2.25 | — |
| Together | $1.04 | $1.04 | — |
Third-party reported values. 10 rows, 3 series: Input, Output, Cache read. Input: highest Together $1.04. Lowest DeepInfra (fp8) $0.1. Output: highest Cloudflare (fp8) $2.25. Lowest DeepInfra (fp8) $0.32.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| StreamLake (fp8) | $0.21 | $0.42 | $0.017 |
| GMICloud (fp8) | $0.96 | $1.91 | $0.08 |
| Relace (fp4) | $0.21 | $4.20 | $0.21 |
| Parasail (fp8) | $0.45 | $3.48 | $0.1 |
| DigitalOcean | $1.04 | $2.09 | $0.21 |
| Cloudflare | $1.15 | $2.55 | $0.2 |
| DeepInfra (fp8) | $1.30 | $2.60 | $0.1 |
| Alibaba (fp8) | $1.42 | $2.83 | $0.12 |
| SiliconFlow (fp8) | $1.50 | $3.13 | $0.14 |
| Novita (fp8) | $1.60 | $3.20 | $0.14 |
| Venice | $1.65 | $3.30 | $0.33 |
| AtlasCloud (fp4) | $1.68 | $3.38 | $0.13 |
| Baidu (fp8) | $1.69 | $3.38 | $0.14 |
| NextBit (fp8) | $1.74 | $3.48 | $0.14 |
| Reka | $0.9 | $9.00 | $0.18 |
Third-party reported values. 15 rows, 3 series: Input, Output, Cache read. Input: highest NextBit (fp8) $1.74. Lowest Relace (fp4) $0.21. Output: highest Reka $9.00. Lowest StreamLake (fp8) $0.42.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| StreamLake (fp8) | $0.042 | $0.084 | $0.0084 |
| DeepInfra (fp8) | $0.09 | $0.18 | $0.018 |
| GMICloud (fp8) | $0.091 | $0.18 | $0.018 |
| Venice | $0.097 | $0.19 | $0.02 |
| DigitalOcean | $0.098 | $0.2 | $0.02 |
| Alibaba (fp8) | $0.13 | $0.27 | $0.027 |
| SiliconFlow (fp8) | $0.13 | $0.28 | $0.028 |
| AtlasCloud (fp4) | $0.14 | $0.28 | $0.028 |
| Baidu (fp8) | $0.14 | $0.28 | $0.028 |
| Novita (fp8) | $0.14 | $0.28 | $0.028 |
| Parasail (fp8) | $0.14 | $0.28 | $0.07 |
| Mancer 2 (fp8) | $0.19 | $0.5 | — |
| Relace (fp4) | $0.012 | $1.28 | $0.012 |
| OpenInference (fp4) | $0.013 | $1.41 | $0.013 |
| Cloudflare | $0.44 | $1.32 | $0.014 |
Third-party reported values. 15 rows, 3 series: Input, Output, Cache read. Input: highest Cloudflare $0.44. Lowest Relace (fp4) $0.012. Output: highest OpenInference (fp4) $1.41. Lowest StreamLake (fp8) $0.084.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| Relace (fp4) | $0.83 | $13.00 | $0.45 |
| Phala | $1.95 | $9.75 | $0.2 |
| Sail Research (fp4) | $0.84 | $13.50 | $0.3 |
| Decart (mxfp4) | $2.01 | $10.05 | $0.2 |
| InferenceNet (fp4) | $0.95 | $14.00 | $0.31 |
| Wafer | $0.95 | $14.00 | $0.4 |
| Morph (fp8) | $1.27 | $13.30 | $0.28 |
| Makora | $1.53 | $12.75 | $0.2 |
| AkashML (fp4) | $1.30 | $14.00 | $1.30 |
| DigitalOcean | $2.55 | $12.95 | $0.26 |
| Together | $2.70 | $13.50 | $0.27 |
| DeepInfra (mxfp4) | $2.85 | $14.25 | $0.28 |
| BaseTen (fp8) | $3.00 | $15.00 | $0.3 |
| Chutes (mxfp4) | $3.00 | $15.00 | $0.3 |
| Fireworks | $3.00 | $15.00 | $0.3 |
| Modal (mxfp4) | $3.00 | $15.00 | $0.3 |
| Moonshot AI (mxfp4) | $3.00 | $15.00 | $0.3 |
| Parasail (fp4) | $3.00 | $15.00 | $0.3 |
| Alibaba | $3.45 | $17.25 | $0.34 |
Third-party reported values. 19 rows, 3 series: Input, Output, Cache read. Input: highest Alibaba $3.45. Lowest Relace (fp4) $0.83. Output: highest Alibaba $17.25. Lowest Phala $9.75.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Input
Output
Cache read
- one provider (reported price, USD per million tokens, log scale per strip)
| Item | Input | Output | Cache read |
|---|---|---|---|
| Novita (fp8) | $0.42 | $1.32 | $0.078 |
| Reka | $0.17 | $3.00 | $0.17 |
| Sail Research (fp8) | $0.2 | $3.40 | $0.15 |
| Morph (fp8) | $0.18 | $3.55 | $0.14 |
| DeepInfra (fp4) | $0.56 | $2.50 | $0.13 |
| SiliconFlow (fp8) | $0.7 | $2.20 | $0.13 |
| InferenceNet | $0.14 | $4.40 | $0.07 |
| Makora (fp4) | $0.18 | $4.40 | $0.19 |
| AkashML (fp8) | $0.19 | $4.40 | $0.19 |
| Phala | $0.84 | $2.64 | $0.16 |
| Inceptron (fp4) | $0.6 | $3.39 | $0.2 |
| DigitalOcean | $0.91 | $2.86 | $0.17 |
| GMICloud (fp8) | $0.98 | $3.08 | $0.18 |
| Alibaba | $1.19 | $3.74 | $0.24 |
| Decart (fp4) | $1.19 | $3.74 | $0.2 |
| Wafer | $0.15 | $7.00 | $0.14 |
| Friendli | $1.26 | $3.96 | $0.23 |
| AtlasCloud (fp8) | $1.40 | $4.40 | $0.26 |
| Baidu (fp8) | $1.40 | $4.40 | $0.26 |
| BaseTen (fp4) | $1.40 | $4.40 | $0.14 |
| Cloudflare | $1.40 | $4.40 | $0.26 |
| Crusoe (fp4) | $1.40 | $4.40 | $0.26 |
| Fireworks | $1.40 | $4.40 | $0.26 |
| Mistral (nvfp4) | $1.40 | $4.40 | $0.14 |
| Modal | $1.40 | $4.40 | $0.26 |
| Nebius (fp4) | $1.40 | $4.40 | — |
| Parasail (fp8) | $1.40 | $4.40 | $0.26 |
| PrimeIntellect | $1.40 | $4.40 | $0.26 |
| Together | $1.40 | $4.40 | $0.26 |
| Venice | $1.40 | $4.40 | $0.26 |
| Z.AI (fp8) | $1.40 | $4.40 | $0.26 |
| Relace | $0.03 | $12.00 | $0.03 |
Third-party reported values. 32 rows, 3 series: Input, Output, Cache read. Input: highest AtlasCloud (fp8) $1.40. Lowest Relace $0.03. Output: highest Relace $12.00. Lowest Novita (fp8) $1.32.
Notes
Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06
Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.
Source: OpenRouter public API: models and provider endpoints (snapshot)
Tables
Provider index per model (reported by OpenRouter’s public API, snapshot 2026-10-06)
| Model | Providers | Endpoints (all tiers) | OpenRouter list price, in / out (USD per M) | Cheapest standard provider | Most expensive standard provider | Spread (blended) | First-party list price, in / out |
|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | 4 | 8 | $1.00 / $5.00 | Amazon Bedrock: $1.00 / $5.00 | Google Vertex: $1.00 / $5.00 | 1.0x | Anthropic: $1.00 / $5.00 |
| Claude Sonnet 5 | 5 | 10 | $2.00 / $10.00 | Amazon Bedrock: $2.00 / $10.00 | Google Vertex: $2.00 / $10.00 | 1.0x | Anthropic: $2.00 / $10.00 |
| Claude Sonnet 5.5 | 5 | 8 | $2.00 / $10.00 | Amazon Bedrock: $2.00 / $10.00 | Google Vertex: $2.00 / $10.00 | 1.0x | Anthropic: $2.00 / $10.00 |
| Claude Opus 4.8 | 5 | 11 | $5.00 / $25.00 | Amazon Bedrock: $5.00 / $25.00 | Google Vertex: $5.00 / $25.00 | 1.0x | Anthropic: $5.00 / $25.00 |
| Claude Opus 5 | 5 | 11 | $5.00 / $25.00 | Amazon Bedrock: $5.00 / $25.00 | Google Vertex: $5.00 / $25.00 | 1.0x | Anthropic: $5.00 / $25.00 |
| Claude Opus 5.5 | 5 | 11 | $4.00 / $20.00 | Amazon Bedrock: $4.00 / $20.00 | Google Vertex: $4.00 / $20.00 | 1.0x | Anthropic: $4.00 / $20.00 |
| Claude Fable 5.1 | 4 | 4 | $10.00 / $50.00 | Amazon Bedrock: $10.00 / $50.00 | Google Vertex: $10.00 / $50.00 | 1.0x | Anthropic: $10.00 / $50.00 |
| GPT-6 Sol | 3 | 7 | $2.00 / $10.00 | Azure: $2.00 / $10.00 | OpenAI: $2.00 / $10.00 | 1.0x | not in the price table |
| GPT-6 Luna | 3 | 7 | $0.10 / $0.50 | Azure: $0.10 / $0.50 | OpenAI: $0.10 / $0.50 | 1.0x | OpenAI: $0.10 / $0.50 |
| GPT-6 Astra | 3 | 7 | $10.00 / $50.00 | Azure: $10.00 / $50.00 | OpenAI: $10.00 / $50.00 | 1.0x | not in the price table |
| GPT-5.5 | 3 | 7 | $5.00 / $30.00 | Azure: $5.00 / $30.00 | OpenAI: $5.00 / $30.00 | 1.0x | not in the price table |
| gpt-oss-120b | 20 | 23 | $0.037 / $0.17 | CoreWeave (fp4): $0.03 / $0.17 | Cerebras (fp16): $0.35 / $0.75 | 6.9x | not in the price table |
| Gemini 3.8 Flash | 2 | 6 | $0.75 / $3.75 | Google AI Studio: $0.75 / $3.75 | Google Vertex: $0.75 / $3.75 | 1.0x | Google: $0.75 / $3.75 |
| Gemini 3.5 Flash | 2 | 7 | $1.50 / $9.00 | Google AI Studio: $1.50 / $9.00 | Google Vertex: $1.50 / $9.00 | 1.0x | Google: $1.50 / $9.00 |
| Gemini 3.5 Flash Lite | 2 | 8 | $0.30 / $2.50 | Google AI Studio: $0.30 / $2.50 | Google Vertex: $0.30 / $2.50 | 1.0x | not in the price table |
| Gemini 3.1 Pro Preview | 2 | 6 | $2.00 / $12.00 | Google AI Studio: $2.00 / $12.00 | Google Vertex: $2.00 / $12.00 | 1.0x | not in the price table |
| Llama 4 Maverick | 4 | 4 | $0.19 / $0.65 | DigitalOcean: $0.19 / $0.65 | Parasail (fp8): $0.35 / $1.00 | 1.7x | not in the price table |
| Llama 3.3 70B Instruct | 10 | 11 | $0.22 / $0.50 | DeepInfra (fp8): $0.10 / $0.32 | Together: $1.04 / $1.04 | 6.7x | not in the price table |
| DeepSeek V4 Pro 0423 | 16 | 16 | $0.21 / $0.42 | StreamLake (fp8): $0.21 / $0.42 | Reka: $0.90 / $9.00 | 11.2x | not in the price table |
| DeepSeek V4 Flash 0423 | 16 | 16 | $0.012 / $1.28 | StreamLake (fp8): $0.042 / $0.084 | Cloudflare: $0.44 / $1.32 | 12.6x | not in the price table |
| Qwen3.8 Max (0902) | 1 | 1 | $2.00 / $6.00 | Alibaba: $2.00 / $6.00 | only one provider | n/a | not in the price table |
| Qwen3.8 Flash | 1 | 1 | $0.15 / $0.47 | Alibaba: $0.15 / $0.47 | only one provider | n/a | not in the price table |
| Mistral Medium 3.5 | 1 | 3 | $1.50 / $7.50 | Mistral: $1.50 / $7.50 | only one provider | n/a | not in the price table |
| Mistral Large 3 2512 | 1 | 2 | $0.50 / $1.50 | Mistral: $0.50 / $1.50 | only one provider | n/a | not in the price table |
| Grok 4.7 | 1 | 5 | $2.00 / $6.00 | xAI: $2.00 / $6.00 | only one provider | n/a | not in the price table |
| Kimi K3 | 20 | 24 | $0.95 / $14.00 | Relace (fp4): $0.83 / $13.00 | Alibaba: $3.45 / $17.25 | 1.8x | not in the price table |
| GLM 5.3 | 32 | 41 | $0.07 / $7.00 | Novita (fp8): $0.42 / $1.32 | Relace: $0.03 / $12.00 | 4.7x | not in the price table |
Gateway overhead: what is and is not measured
| Metric | Status |
|---|---|
| Per-token price through OpenRouter vs first-party | Reported list prices, snapshot 2026-10-06 |
| Credit-purchase fee | 5.5% on Standard (third-party-reported, 2026-10-06) |
| Provider latency and throughput (last 30 min) | unknown: the keyless API returned none for 265 endpoints |
| Time to first token and total time through the gateway vs direct | not measured: no key in the environment. The live harness sends nothing without an OpenRouter key. |
| Billed cost per call through the gateway | not measured: no key in the environment. The live harness sends nothing without an OpenRouter key. |
Every endpoint OpenRouter listed (reported by OpenRouter’s public API, snapshot 2026-10-06)
| Model | Provider | Tier | Quantization | Context (tokens) | Input (USD per M) | Output (USD per M) | Cache read (USD per M) | Uptime, last day (%) |
|---|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | Azure | standard | not reported | 1,000,000 | $10.00 | $50.00 | $0.25 | 99% |
| Claude Fable 5.1 | Anthropic | standard | not reported | 1,000,000 | $10.00 | $50.00 | $0.25 | 100% |
| Claude Fable 5.1 | Amazon Bedrock | standard | not reported | 1,000,000 | $10.00 | $50.00 | $0.25 | — |
| Claude Fable 5.1 | Google Vertex | standard | not reported | 1,000,000 | $10.00 | $50.00 | $0.25 | 100% |
| Claude Haiku 4.5 | Azure | standard | not reported | 200,000 | $1.00 | $5.00 | $0.1 | 100% |
| Claude Haiku 4.5 | Amazon Bedrock | standard | not reported | 200,000 | $1.00 | $5.00 | $0.1 | 100% |
| Claude Haiku 4.5 | Google Vertex | standard | not reported | 200,000 | $1.00 | $5.00 | $0.1 | 100% |
| Claude Haiku 4.5 | Anthropic | standard | not reported | 200,000 | $1.00 | $5.00 | $0.1 | 100% |
| Claude Haiku 4.5 | Amazon Bedrock | regional | not reported | 200,000 | $1.10 | $5.50 | $0.11 | 100% |
| Claude Haiku 4.5 | Google Vertex | regional | not reported | 200,000 | $1.10 | $5.50 | $0.11 | 100% |
| Claude Haiku 4.5 | Amazon Bedrock | regional | not reported | 200,000 | $1.10 | $5.50 | $0.11 | 100% |
| Claude Haiku 4.5 | Google Vertex | regional | not reported | 200,000 | $1.10 | $5.50 | $0.11 | 100% |
| Claude Opus 4.8 | Claude Platform on AWS | standard | not reported | 1,000,000 | $5.00 | $25.00 | $0.5 | 100% |
| Claude Opus 4.8 | Azure | standard | not reported | 1,000,000 | $5.00 | $25.00 | $0.5 | 98% |
| Claude Opus 4.8 | Amazon Bedrock | standard | not reported | 1,000,000 | $5.00 | $25.00 | $0.5 | 95% |
| Claude Opus 4.8 | Google Vertex | standard | not reported | 1,000,000 | $5.00 | $25.00 | $0.5 | 100% |
| Claude Opus 4.8 | Anthropic | standard | not reported | 1,000,000 | $5.00 | $25.00 | $0.5 | 100% |
| Claude Opus 4.8 | Azure | regional | not reported | 1,000,000 | $5.50 | $27.50 | $0.55 | — |
| Claude Opus 4.8 | Amazon Bedrock | regional | not reported | 1,000,000 | $5.50 | $27.50 | $0.55 | 100% |
| Claude Opus 4.8 | Google Vertex | regional | not reported | 1,000,000 | $5.50 | $27.50 | $0.55 | — |
| Claude Opus 4.8 | Amazon Bedrock | regional | not reported | 1,000,000 | $5.50 | $27.50 | $0.55 | 100% |
| Claude Opus 4.8 | Google Vertex | regional | not reported | 1,000,000 | $5.50 | $27.50 | $0.55 | 100% |
| Claude Opus 4.8 | Anthropic | fast | not reported | 1,000,000 | $10.00 | $50.00 | $1.00 | 100% |
| Claude Opus 5 | Claude Platform on AWS | standard | not reported | 1,000,000 | $5.00 | $25.00 | $0.5 | 99% |
| Claude Opus 5 | Google Vertex | standard | not reported | 1,000,000 | $5.00 | $25.00 | $0.5 | 100% |
| Claude Opus 5 | Amazon Bedrock | standard | not reported | 1,000,000 | $5.00 | $25.00 | $0.5 | 100% |
| Claude Opus 5 | Azure | standard | not reported | 1,000,000 | $5.00 | $25.00 | $0.5 | 100% |
| Claude Opus 5 | Anthropic | standard | not reported | 1,000,000 | $5.00 | $25.00 | $0.5 | 100% |
| Claude Opus 5 | Amazon Bedrock | regional | not reported | 1,000,000 | $5.50 | $27.50 | $0.55 | 100% |
| Claude Opus 5 | Azure | regional | not reported | 1,000,000 | $5.50 | $27.50 | $0.55 | — |
| Claude Opus 5 | Google Vertex | regional | not reported | 1,000,000 | $5.50 | $27.50 | $0.55 | 100% |
| Claude Opus 5 | Google Vertex | regional | not reported | 1,000,000 | $5.50 | $27.50 | $0.55 | 100% |
| Claude Opus 5 | Amazon Bedrock | regional | not reported | 1,000,000 | $5.50 | $27.50 | $0.55 | 100% |
| Claude Opus 5 | Anthropic | fast | not reported | 1,000,000 | $10.00 | $50.00 | $1.00 | 100% |
| Claude Opus 5.5 | Amazon Bedrock | standard | not reported | 1,000,000 | $4.00 | $20.00 | $0.2 | 97% |
| Claude Opus 5.5 | Azure | standard | not reported | 1,000,000 | $4.00 | $20.00 | $0.2 | 100% |
| Claude Opus 5.5 | Google Vertex | standard | not reported | 1,000,000 | $4.00 | $20.00 | $0.2 | 100% |
| Claude Opus 5.5 | Claude Platform on AWS | standard | not reported | 1,000,000 | $4.00 | $20.00 | $0.2 | 100% |
| Claude Opus 5.5 | Anthropic | standard | not reported | 1,000,000 | $4.00 | $20.00 | $0.2 | 100% |
| Claude Opus 5.5 | Amazon Bedrock | regional | not reported | 1,000,000 | $4.40 | $22.00 | $0.22 | 100% |
| Claude Opus 5.5 | Amazon Bedrock | regional | not reported | 1,000,000 | $4.40 | $22.00 | $0.22 | 100% |
| Claude Opus 5.5 | Azure | regional | not reported | 1,000,000 | $4.40 | $22.00 | $0.22 | 100% |
| Claude Opus 5.5 | Google Vertex | regional | not reported | 1,000,000 | $4.40 | $22.00 | $0.22 | 100% |
| Claude Opus 5.5 | Google Vertex | regional | not reported | 1,000,000 | $4.40 | $22.00 | $0.22 | 100% |
| Claude Opus 5.5 | Anthropic | fast | not reported | 1,000,000 | $8.00 | $40.00 | $0.4 | 100% |
| Claude Sonnet 5 | Claude Platform on AWS | standard | not reported | 1,000,000 | $2.00 | $10.00 | $0.2 | 100% |
| Claude Sonnet 5 | Azure | standard | not reported | 1,000,000 | $2.00 | $10.00 | $0.2 | 100% |
| Claude Sonnet 5 | Google Vertex | standard | not reported | 1,000,000 | $2.00 | $10.00 | $0.2 | 100% |
| Claude Sonnet 5 | Amazon Bedrock | standard | not reported | 1,000,000 | $2.00 | $10.00 | $0.2 | 100% |
| Claude Sonnet 5 | Anthropic | standard | not reported | 1,000,000 | $2.00 | $10.00 | $0.2 | 100% |
| Claude Sonnet 5 | Amazon Bedrock | regional | not reported | 1,000,000 | $2.20 | $11.00 | $0.22 | 100% |
| Claude Sonnet 5 | Azure | regional | not reported | 1,000,000 | $2.20 | $11.00 | $0.22 | — |
| Claude Sonnet 5 | Google Vertex | regional | not reported | 1,000,000 | $2.20 | $11.00 | $0.22 | 100% |
| Claude Sonnet 5 | Google Vertex | regional | not reported | 1,000,000 | $2.20 | $11.00 | $0.22 | 100% |
| Claude Sonnet 5 | Amazon Bedrock | regional | not reported | 1,000,000 | $2.20 | $11.00 | $0.22 | 100% |
| Claude Sonnet 5.5 | Google Vertex | standard | not reported | 1,000,000 | $2.00 | $10.00 | $0.2 | 100% |
| Claude Sonnet 5.5 | Amazon Bedrock | standard | not reported | 1,000,000 | $2.00 | $10.00 | $0.2 | 100% |
| Claude Sonnet 5.5 | Azure | standard | not reported | 1,000,000 | $2.00 | $10.00 | $0.2 | 100% |
| Claude Sonnet 5.5 | Claude Platform on AWS | standard | not reported | 1,000,000 | $2.00 | $10.00 | $0.2 | 100% |
| Claude Sonnet 5.5 | Anthropic | standard | not reported | 1,000,000 | $2.00 | $10.00 | $0.2 | 100% |
| Claude Sonnet 5.5 | Google Vertex | regional | not reported | 1,000,000 | $2.20 | $11.00 | $0.22 | 100% |
| Claude Sonnet 5.5 | Azure | regional | not reported | 1,000,000 | $2.20 | $11.00 | $0.22 | 100% |
| Claude Sonnet 5.5 | Google Vertex | regional | not reported | 1,000,000 | $2.20 | $11.00 | $0.22 | 100% |
| DeepSeek V4 Flash 0423 | Relace | standard | fp4 | 1,048,576 | $0.012 | $1.28 | $0.012 | 100% |
| DeepSeek V4 Flash 0423 | OpenInference | standard | fp4 | 1,048,576 | $0.013 | $1.41 | $0.013 | 99% |
| DeepSeek V4 Flash 0423 | StreamLake | standard | fp8 | 1,024,000 | $0.042 | $0.084 | $0.0084 | 99% |
| DeepSeek V4 Flash 0423 | DeepInfra | standard | fp8 | 1,048,576 | $0.09 | $0.18 | $0.018 | 100% |
| DeepSeek V4 Flash 0423 | GMICloud | standard | fp8 | 1,048,575 | $0.091 | $0.18 | $0.018 | 100% |
| DeepSeek V4 Flash 0423 | Venice | standard | not reported | 1,000,000 | $0.097 | $0.19 | $0.02 | 96% |
| DeepSeek V4 Flash 0423 | DigitalOcean | standard | not reported | 1,048,576 | $0.098 | $0.2 | $0.02 | 100% |
| DeepSeek V4 Flash 0423 | SiliconFlow | standard | fp8 | 1,048,576 | $0.13 | $0.28 | $0.028 | 100% |
| DeepSeek V4 Flash 0423 | Alibaba | standard | fp8 | 1,000,000 | $0.13 | $0.27 | $0.027 | 99% |
| DeepSeek V4 Flash 0423 | Baidu | standard | fp8 | 1,048,576 | $0.14 | $0.28 | $0.028 | 99% |
| DeepSeek V4 Flash 0423 | Novita | standard | fp8 | 1,048,576 | $0.14 | $0.28 | $0.028 | 100% |
| DeepSeek V4 Flash 0423 | AtlasCloud | standard | fp4 | 1,048,576 | $0.14 | $0.28 | $0.028 | 99% |
| DeepSeek V4 Flash 0423 | Parasail | standard | fp8 | 1,048,576 | $0.14 | $0.28 | $0.07 | 100% |
| DeepSeek V4 Flash 0423 | Mancer 2 | standard | fp8 | 1,048,576 | $0.19 | $0.5 | — | 95% |
| DeepSeek V4 Flash 0423 | Azure | regional | not reported | 1,048,576 | $0.21 | $0.56 | $0.031 | 95% |
| DeepSeek V4 Flash 0423 | Cloudflare | standard | not reported | 384,000 | $0.44 | $1.32 | $0.014 | 98% |
| DeepSeek V4 Pro 0423 | Relace | standard | fp4 | 1,048,576 | $0.21 | $4.20 | $0.21 | 100% |
| DeepSeek V4 Pro 0423 | StreamLake | standard | fp8 | 1,024,000 | $0.21 | $0.42 | $0.017 | 99% |
| DeepSeek V4 Pro 0423 | Parasail | standard | fp8 | 1,048,576 | $0.45 | $3.48 | $0.1 | 98% |
| DeepSeek V4 Pro 0423 | Reka | standard | not reported | 1,048,576 | $0.9 | $9.00 | $0.18 | 99% |
| DeepSeek V4 Pro 0423 | GMICloud | standard | fp8 | 1,048,576 | $0.96 | $1.91 | $0.08 | 97% |
| DeepSeek V4 Pro 0423 | DigitalOcean | standard | not reported | 1,048,576 | $1.04 | $2.09 | $0.21 | 100% |
| DeepSeek V4 Pro 0423 | Cloudflare | standard | not reported | 1,048,576 | $1.15 | $2.55 | $0.2 | 98% |
| DeepSeek V4 Pro 0423 | DeepInfra | standard | fp8 | 1,048,576 | $1.30 | $2.60 | $0.1 | 100% |
| DeepSeek V4 Pro 0423 | Alibaba | standard | fp8 | 1,000,000 | $1.42 | $2.83 | $0.12 | 90% |
| DeepSeek V4 Pro 0423 | SiliconFlow | standard | fp8 | 1,048,576 | $1.50 | $3.13 | $0.14 | 99% |
| DeepSeek V4 Pro 0423 | Novita | standard | fp8 | 1,048,576 | $1.60 | $3.20 | $0.14 | 100% |
| DeepSeek V4 Pro 0423 | Venice | standard | not reported | 1,000,000 | $1.65 | $3.30 | $0.33 | 97% |
| DeepSeek V4 Pro 0423 | AtlasCloud | standard | fp4 | 1,048,576 | $1.68 | $3.38 | $0.13 | 99% |
| DeepSeek V4 Pro 0423 | Baidu | standard | fp8 | 1,048,576 | $1.69 | $3.38 | $0.14 | 100% |
| DeepSeek V4 Pro 0423 | NextBit | standard | fp8 | 1,048,576 | $1.74 | $3.48 | $0.14 | 99% |
| DeepSeek V4 Pro 0423 | Azure | regional | not reported | 1,048,576 | $1.91 | $3.83 | $0.16 | 99% |
| Gemini 3.1 Pro Preview | Google Vertex | flex | not reported | 1,048,576 | $1.00 | $6.00 | $0.1 | 93% |
| Gemini 3.1 Pro Preview | Google AI Studio | flex | not reported | 1,048,576 | $1.00 | $6.00 | $0.1 | 100% |
| Gemini 3.1 Pro Preview | Google Vertex | standard | not reported | 1,048,576 | $2.00 | $12.00 | $0.2 | 98% |
| Gemini 3.1 Pro Preview | Google AI Studio | standard | not reported | 1,048,576 | $2.00 | $12.00 | $0.2 | 100% |
| Gemini 3.1 Pro Preview | Google Vertex | priority | not reported | 1,048,576 | $3.60 | $21.60 | $0.36 | 100% |
| Gemini 3.1 Pro Preview | Google AI Studio | priority | not reported | 1,048,576 | $3.60 | $21.60 | $0.36 | 99% |
| Gemini 3.5 Flash | Google Vertex | flex | not reported | 1,048,576 | $0.75 | $4.50 | $0.075 | 99% |
| Gemini 3.5 Flash | Google AI Studio | flex | not reported | 1,048,576 | $0.75 | $4.50 | $0.075 | 100% |
| Gemini 3.5 Flash | Google Vertex | standard | not reported | 1,048,576 | $1.50 | $9.00 | $0.15 | 99% |
| Gemini 3.5 Flash | Google AI Studio | standard | not reported | 1,048,576 | $1.50 | $9.00 | $0.15 | 100% |
| Gemini 3.5 Flash | Google Vertex | regional | not reported | 1,048,576 | $1.65 | $9.90 | $0.17 | — |
| Gemini 3.5 Flash | Google Vertex | priority | not reported | 1,048,576 | $2.70 | $16.20 | $0.27 | 100% |
| Gemini 3.5 Flash | Google AI Studio | priority | not reported | 1,048,576 | $2.70 | $16.20 | $0.27 | 100% |
| Gemini 3.5 Flash Lite | Google Vertex | flex | not reported | 1,048,576 | $0.15 | $1.25 | $0.015 | 100% |
| Gemini 3.5 Flash Lite | Google AI Studio | flex | not reported | 1,048,576 | $0.15 | $1.25 | $0.015 | 100% |
| Gemini 3.5 Flash Lite | Google AI Studio | standard | not reported | 1,048,576 | $0.3 | $2.50 | $0.03 | 100% |
| Gemini 3.5 Flash Lite | Google Vertex | standard | not reported | 1,048,576 | $0.3 | $2.50 | $0.03 | 100% |
| Gemini 3.5 Flash Lite | Google Vertex | regional | not reported | 1,048,576 | $0.33 | $2.75 | $0.033 | 100% |
| Gemini 3.5 Flash Lite | Google Vertex | regional | not reported | 1,048,576 | $0.33 | $2.75 | $0.033 | 100% |
| Gemini 3.5 Flash Lite | Google Vertex | priority | not reported | 1,048,576 | $0.54 | $4.50 | $0.054 | 100% |
| Gemini 3.5 Flash Lite | Google AI Studio | priority | not reported | 1,048,576 | $0.54 | $4.50 | $0.054 | 100% |
| Gemini 3.8 Flash | Google AI Studio | flex | not reported | 1,048,576 | $0.38 | $1.88 | $0.037 | 100% |
| Gemini 3.8 Flash | Google Vertex | flex | not reported | 1,048,576 | $0.38 | $1.88 | $0.037 | 99% |
| Gemini 3.8 Flash | Google AI Studio | standard | not reported | 1,048,576 | $0.75 | $3.75 | $0.075 | 100% |
| Gemini 3.8 Flash | Google Vertex | standard | not reported | 1,048,576 | $0.75 | $3.75 | $0.075 | 97% |
| Gemini 3.8 Flash | Google AI Studio | priority | not reported | 1,048,576 | $1.35 | $6.75 | $0.14 | 100% |
| Gemini 3.8 Flash | Google Vertex | priority | not reported | 1,048,576 | $1.35 | $6.75 | $0.14 | 100% |
| GLM 5.3 | Relace | standard | not reported | 1,048,576 | $0.03 | $12.00 | $0.03 | 100% |
| GLM 5.3 | Wafer | regional | not reported | 1,048,576 | $0.07 | $7.00 | $0.065 | 100% |
| GLM 5.3 | InferenceNet | standard | not reported | 1,048,576 | $0.14 | $4.40 | $0.07 | 100% |
| GLM 5.3 | Wafer | standard | not reported | 1,048,576 | $0.15 | $7.00 | $0.14 | 100% |
| GLM 5.3 | Reka | standard | not reported | 262,144 | $0.17 | $3.00 | $0.17 | 100% |
| GLM 5.3 | Morph | standard | fp8 | 1,048,576 | $0.18 | $3.55 | $0.14 | 99% |
| GLM 5.3 | Makora | standard | fp4 | 980,000 | $0.18 | $4.40 | $0.19 | 96% |
| GLM 5.3 | AkashML | standard | fp8 | 1,048,576 | $0.19 | $4.40 | $0.19 | 100% |
| GLM 5.3 | Sail Research | regional | fp8 | 1,048,576 | $0.2 | $3.40 | $0.15 | 99% |
| GLM 5.3 | Sail Research | standard | fp8 | 1,048,576 | $0.2 | $3.40 | $0.15 | 99% |
| GLM 5.3 | Novita | standard | fp8 | 1,048,576 | $0.42 | $1.32 | $0.078 | 97% |
| GLM 5.3 | DeepInfra | standard | fp4 | 1,048,576 | $0.56 | $2.50 | $0.13 | 97% |
| GLM 5.3 | Inceptron | standard | fp4 | 1,048,576 | $0.6 | $3.39 | $0.2 | 98% |
| GLM 5.3 | SiliconFlow | standard | fp8 | 1,048,576 | $0.7 | $2.20 | $0.13 | 100% |
| GLM 5.3 | Phala | standard | not reported | 1,048,576 | $0.84 | $2.64 | $0.16 | 99% |
| GLM 5.3 | DigitalOcean | standard | not reported | 1,048,576 | $0.91 | $2.86 | $0.17 | 100% |
| GLM 5.3 | GMICloud | standard | fp8 | 1,048,576 | $0.98 | $3.08 | $0.18 | 99% |
| GLM 5.3 | Alibaba | standard | not reported | 1,000,000 | $1.19 | $3.74 | $0.24 | 100% |
| GLM 5.3 | Decart | standard | fp4 | 1,048,576 | $1.19 | $3.74 | $0.2 | 100% |
| GLM 5.3 | Friendli | standard | not reported | 1,048,576 | $1.26 | $3.96 | $0.23 | 100% |
| GLM 5.3 | Mistral | standard | nvfp4 | 1,048,576 | $1.40 | $4.40 | $0.14 | 100% |
| GLM 5.3 | Baidu | standard | fp8 | 1,048,576 | $1.40 | $4.40 | $0.26 | 100% |
| GLM 5.3 | BaseTen | standard | fp4 | 1,048,576 | $1.40 | $4.40 | $0.14 | 98% |
| GLM 5.3 | Mistral | standard | nvfp4 | 1,048,576 | $1.40 | $4.40 | $0.14 | 99% |
| GLM 5.3 | Nebius | standard | fp4 | 1,024,000 | $1.40 | $4.40 | — | 96% |
| GLM 5.3 | Crusoe | standard | fp4 | 1,048,576 | $1.40 | $4.40 | $0.26 | 98% |
| GLM 5.3 | PrimeIntellect | standard | not reported | 1,048,576 | $1.40 | $4.40 | $0.26 | 99% |
| GLM 5.3 | Venice | standard | not reported | 1,000,000 | $1.40 | $4.40 | $0.26 | 99% |
| GLM 5.3 | Together | standard | not reported | 1,048,575 | $1.40 | $4.40 | $0.26 | 96% |
| GLM 5.3 | Parasail | standard | fp8 | 1,048,576 | $1.40 | $4.40 | $0.26 | 100% |
| GLM 5.3 | Modal | standard | not reported | 1,048,576 | $1.40 | $4.40 | $0.26 | 98% |
| GLM 5.3 | BaseTen | standard | fp4 | 1,048,576 | $1.40 | $4.40 | $0.14 | 95% |
| GLM 5.3 | Fireworks | standard | not reported | 1,048,576 | $1.40 | $4.40 | $0.26 | 100% |
| GLM 5.3 | Cloudflare | standard | not reported | 1,048,576 | $1.40 | $4.40 | $0.26 | 98% |
| GLM 5.3 | AtlasCloud | standard | fp8 | 1,048,576 | $1.40 | $4.40 | $0.26 | 100% |
| GLM 5.3 | Z.AI | standard | fp8 | 1,048,576 | $1.40 | $4.40 | $0.26 | 100% |
| GLM 5.3 | Mistral | standard | nvfp4 | 1,048,576 | $1.54 | $4.84 | $0.15 | 100% |
| GLM 5.3 | Fireworks | fast | not reported | 1,048,576 | $2.10 | $6.60 | $0.39 | 100% |
| GLM 5.3 | BaseTen | fast | fp8 | 1,048,576 | $2.10 | $6.60 | $0.21 | 100% |
| GLM 5.3 | BaseTen | fast | fp8 | 1,048,576 | $2.10 | $6.60 | $0.21 | 99% |
| GLM 5.3 | Alibaba | fast | not reported | 1,000,000 | $2.80 | $8.80 | $0.56 | 100% |
| GPT-5.5 | OpenAI | flex | not reported | 1,050,000 | $2.50 | $15.00 | $0.25 | 100% |
| GPT-5.5 | Azure | standard | not reported | 1,050,000 | $5.00 | $30.00 | $0.5 | 100% |
| GPT-5.5 | OpenAI | standard | not reported | 1,050,000 | $5.00 | $30.00 | $0.5 | 100% |
| GPT-5.5 | Azure | regional | not reported | 1,050,000 | $5.50 | $33.00 | $0.55 | 100% |
| GPT-5.5 | Azure | regional | not reported | 1,050,000 | $5.50 | $33.00 | $0.55 | 100% |
| GPT-5.5 | Amazon Bedrock | regional | not reported | 1,050,000 | $5.50 | $33.00 | $0.55 | — |
| GPT-5.5 | OpenAI | fast | not reported | 1,050,000 | $12.50 | $75.00 | $1.25 | 100% |
| GPT-6 Astra | OpenAI | flex | not reported | 1,050,000 | $5.00 | $25.00 | $0.5 | 100% |
| GPT-6 Astra | Azure | standard | not reported | 1,050,000 | $10.00 | $50.00 | $1.00 | 100% |
| GPT-6 Astra | OpenAI | standard | not reported | 1,050,000 | $10.00 | $50.00 | $1.00 | 100% |
| GPT-6 Astra | Amazon Bedrock | regional | not reported | 1,050,000 | $11.00 | $55.00 | $1.10 | — |
| GPT-6 Astra | Azure | regional | not reported | 1,050,000 | $11.00 | $55.00 | $1.10 | 100% |
| GPT-6 Astra | OpenAI | fast | not reported | 1,050,000 | $20.00 | $100 | $2.00 | 100% |
| GPT-6 Astra | OpenAI | ultrafast | not reported | 1,050,000 | $60.00 | $300 | $6.00 | 100% |
| GPT-6 Luna | OpenAI | flex | not reported | 1,050,000 | $0.05 | $0.25 | $0.005 | 98% |
| GPT-6 Luna | OpenAI | standard | not reported | 1,050,000 | $0.1 | $0.5 | $0.01 | 100% |
| GPT-6 Luna | Azure | standard | not reported | 1,050,000 | $0.1 | $0.5 | $0.01 | 100% |
| GPT-6 Luna | Azure | regional | not reported | 1,050,000 | $0.11 | $0.55 | $0.011 | 100% |
| GPT-6 Luna | Azure | regional | not reported | 1,050,000 | $0.11 | $0.55 | $0.011 | 100% |
| GPT-6 Luna | Amazon Bedrock | regional | not reported | 1,050,000 | $0.11 | $0.55 | $0.011 | 99% |
| GPT-6 Luna | OpenAI | fast | not reported | 1,050,000 | $0.2 | $1.00 | $0.02 | 100% |
| GPT-6 Sol | OpenAI | flex | not reported | 1,050,000 | $1.00 | $5.00 | $0.1 | 100% |
| GPT-6 Sol | OpenAI | standard | not reported | 1,050,000 | $2.00 | $10.00 | $0.2 | 100% |
| GPT-6 Sol | Azure | standard | not reported | 1,050,000 | $2.00 | $10.00 | $0.2 | 100% |
| GPT-6 Sol | Azure | regional | not reported | 1,050,000 | $2.20 | $11.00 | $0.22 | 100% |
| GPT-6 Sol | Azure | regional | not reported | 1,050,000 | $2.20 | $11.00 | $0.22 | 100% |
| GPT-6 Sol | Amazon Bedrock | regional | not reported | 1,050,000 | $2.20 | $11.00 | $0.22 | 100% |
| GPT-6 Sol | OpenAI | fast | not reported | 1,050,000 | $4.00 | $20.00 | $0.4 | 100% |
| gpt-oss-120b | CoreWeave | standard | fp4 | 131,072 | $0.03 | $0.17 | $0.03 | 99% |
| gpt-oss-120b | DekaLLM | standard | bf16 | 131,072 | $0.03 | $0.18 | $0.03 | 100% |
| gpt-oss-120b | DeepInfra | standard | bf16 | 131,072 | $0.037 | $0.17 | — | 99% |
| gpt-oss-120b | AkashML | standard | bf16 | 131,072 | $0.037 | $0.19 | $0.037 | 100% |
| gpt-oss-120b | Mancer 2 | standard | fp8 | 131,072 | $0.045 | $0.25 | — | 99% |
| gpt-oss-120b | Crusoe | standard | bf16 | 131,072 | $0.05 | $0.25 | $0.05 | 100% |
| gpt-oss-120b | Novita | standard | fp4 | 131,072 | $0.05 | $0.25 | — | 99% |
| gpt-oss-120b | DigitalOcean | standard | not reported | 128,000 | $0.06 | $0.42 | $0.012 | 100% |
| gpt-oss-120b | Google Vertex | standard | not reported | 131,072 | $0.09 | $0.36 | — | 67% |
| gpt-oss-120b | BaseTen | standard | fp4 | 128,072 | $0.1 | $0.5 | $0.1 | 100% |
| gpt-oss-120b | BaseTen | standard | fp4 | 128,072 | $0.1 | $0.5 | $0.1 | 100% |
| gpt-oss-120b | Parasail | standard | fp4 | 131,072 | $0.1 | $0.75 | $0.055 | 100% |
| gpt-oss-120b | SambaNova | standard | not reported | 131,072 | $0.14 | $0.95 | — | 100% |
| gpt-oss-120b | Amazon Bedrock | regional | not reported | 131,072 | $0.15 | $0.6 | — | 100% |
| gpt-oss-120b | Nebius | standard | fp4 | 131,072 | $0.15 | $0.6 | — | 97% |
| gpt-oss-120b | Amazon Bedrock | standard | not reported | 131,072 | $0.15 | $0.6 | — | 99% |
| gpt-oss-120b | DeepInfra | standard | bf16 | 131,072 | $0.15 | $0.6 | — | 100% |
| gpt-oss-120b | SiliconFlow | standard | fp8 | 131,072 | $0.15 | $0.6 | $0.075 | 83% |
| gpt-oss-120b | Phala | standard | not reported | 131,072 | $0.15 | $0.6 | — | 99% |
| gpt-oss-120b | Together | standard | not reported | 131,072 | $0.15 | $0.6 | — | 87% |
| gpt-oss-120b | Groq | standard | not reported | 131,072 | $0.15 | $0.6 | $0.075 | 99% |
| gpt-oss-120b | Mara | standard | not reported | 131,072 | $0.15 | $0.75 | — | 97% |
| gpt-oss-120b | Cerebras | standard | fp16 | 131,072 | $0.35 | $0.75 | $0.35 | 100% |
| Grok 4.7 | xAI | standard | not reported | 500,000 | $2.00 | $6.00 | $0.5 | 99% |
| Grok 4.7 | xAI | standard | not reported | 500,000 | $2.00 | $6.00 | $0.5 | 99% |
| Grok 4.7 | xAI | regional | not reported | 500,000 | $2.20 | $6.60 | $0.55 | 100% |
| Grok 4.7 | xAI | priority | not reported | 500,000 | $4.00 | $12.00 | $1.00 | 100% |
| Grok 4.7 | xAI | priority | not reported | 500,000 | $4.00 | $12.00 | $1.00 | 99% |
| Kimi K3 | Relace | standard | fp4 | 1,048,576 | $0.83 | $13.00 | $0.45 | 100% |
| Kimi K3 | Sail Research | standard | fp4 | 1,048,576 | $0.84 | $13.50 | $0.3 | 100% |
| Kimi K3 | InferenceNet | standard | fp4 | 1,048,576 | $0.95 | $14.00 | $0.31 | 100% |
| Kimi K3 | Wafer | standard | not reported | 1,048,576 | $0.95 | $14.00 | $0.4 | 99% |
| Kimi K3 | Morph | standard | fp8 | 1,048,576 | $1.27 | $13.30 | $0.28 | 100% |
| Kimi K3 | AkashML | standard | fp4 | 1,048,576 | $1.30 | $14.00 | $1.30 | 98% |
| Kimi K3 | Makora | standard | not reported | 1,048,576 | $1.53 | $12.75 | $0.2 | 97% |
| Kimi K3 | Phala | standard | not reported | 1,048,576 | $1.95 | $9.75 | $0.2 | 97% |
| Kimi K3 | Decart | standard | mxfp4 | 1,048,576 | $2.01 | $10.05 | $0.2 | 87% |
| Kimi K3 | DigitalOcean | standard | not reported | 1,048,576 | $2.55 | $12.95 | $0.26 | 100% |
| Kimi K3 | Together | standard | not reported | 1,048,576 | $2.70 | $13.50 | $0.27 | 99% |
| Kimi K3 | Wafer | regional | not reported | 1,048,576 | $2.80 | $14.00 | $0.3 | 99% |
| Kimi K3 | DeepInfra | standard | mxfp4 | 1,048,576 | $2.85 | $14.25 | $0.28 | 99% |
| Kimi K3 | Amazon Bedrock | regional | not reported | 1,048,576 | $3.00 | $15.00 | $0.3 | 93% |
| Kimi K3 | Chutes | standard | mxfp4 | 1,048,576 | $3.00 | $15.00 | $0.3 | 97% |
| Kimi K3 | Parasail | standard | fp4 | 1,048,576 | $3.00 | $15.00 | $0.3 | 98% |
| Kimi K3 | Modal | standard | mxfp4 | 1,048,576 | $3.00 | $15.00 | $0.3 | 98% |
| Kimi K3 | Fireworks | standard | not reported | 1,048,576 | $3.00 | $15.00 | $0.3 | 99% |
| Kimi K3 | BaseTen | standard | fp8 | 1,048,576 | $3.00 | $15.00 | $0.3 | 98% |
| Kimi K3 | Moonshot AI | standard | mxfp4 | 1,048,576 | $3.00 | $15.00 | $0.3 | 100% |
| Kimi K3 | Alibaba | standard | not reported | 1,048,576 | $3.45 | $17.25 | $0.34 | 99% |
| Kimi K3 | InferenceNet | fast | fp4 | 250,000 | $3.50 | $15.00 | $0.45 | 100% |
| Kimi K3 | Fireworks | regional | not reported | 1,048,576 | $4.50 | $22.50 | $0.45 | 99% |
| Kimi K3 | Fireworks | fast | not reported | 1,048,576 | $4.50 | $22.50 | $0.45 | 97% |
| Llama 3.3 70B Instruct | DeepInfra | standard | fp8 | 131,072 | $0.1 | $0.32 | — | 98% |
| Llama 3.3 70B Instruct | Novita | standard | bf16 | 12,288 | $0.14 | $0.4 | — | 98% |
| Llama 3.3 70B Instruct | AkashML | standard | fp8 | 131,072 | $0.2 | $0.52 | $0.1 | 99% |
| Llama 3.3 70B Instruct | Parasail | standard | fp8 | 131,072 | $0.22 | $0.5 | $0.11 | 100% |
| Llama 3.3 70B Instruct | Cloudflare | standard | fp8 | 24,000 | $0.29 | $2.25 | — | 99% |
| Llama 3.3 70B Instruct | SambaNova | standard | not reported | 131,072 | $0.45 | $0.9 | — | 99% |
| Llama 3.3 70B Instruct | Groq | standard | not reported | 131,072 | $0.59 | $0.79 | $0.29 | 100% |
| Llama 3.3 70B Instruct | CoreWeave | standard | fp16 | 128,000 | $0.71 | $0.71 | $0.71 | 98% |
| Llama 3.3 70B Instruct | Google Vertex | regional | not reported | 128,000 | $0.72 | $0.72 | — | — |
| Llama 3.3 70B Instruct | Google Vertex | standard | not reported | 128,000 | $0.72 | $0.72 | — | — |
| Llama 3.3 70B Instruct | Together | standard | not reported | 131,072 | $1.04 | $1.04 | — | 93% |
| Llama 4 Maverick | DigitalOcean | standard | not reported | 128,000 | $0.19 | $0.65 | — | 99% |
| Llama 4 Maverick | Novita | standard | fp8 | 1,048,576 | $0.27 | $0.85 | — | 98% |
| Llama 4 Maverick | Parasail | standard | fp8 | 524,288 | $0.35 | $1.00 | $0.17 | 100% |
| Llama 4 Maverick | Google Vertex | regional | not reported | 524,288 | $0.35 | $1.15 | — | — |
| Mistral Large 3 2512 | Mistral | standard | not reported | 262,144 | $0.5 | $1.50 | $0.05 | 100% |
| Mistral Large 3 2512 | Mistral | regional | not reported | 262,144 | $0.55 | $1.65 | $0.055 | 100% |
| Mistral Medium 3.5 | Mistral | standard | not reported | 262,144 | $1.50 | $7.50 | — | 100% |
| Mistral Medium 3.5 | Mistral | standard | not reported | 262,144 | $1.50 | $7.50 | — | 100% |
| Mistral Medium 3.5 | Mistral | regional | not reported | 262,144 | $1.65 | $8.25 | — | 100% |
| Qwen3.8 Flash | Alibaba | standard | not reported | 1,000,000 | $0.15 | $0.47 | $0.016 | 99% |
| Qwen3.8 Max (0902) | Alibaba | standard | not reported | 1,000,000 | $2.00 | $6.00 | $0.25 | 100% |
Method
- One snapshot of OpenRouter’s public, keyless API on 2026-10-06: the model list and the endpoint list of 27 curated models (28 requests). Prices, context, quantization and uptime are copied as the API reported them; a field it did not return stays empty.
- Tier: read from the endpoint tag suffix. Flex, priority, fast, ultrafast and batch are named tiers; a region suffix (us, eu, europe, a cloud region) is regional; anything else is standard. This is a heuristic.
- Per-model charts show the standard tier only, with one bar per provider: its cheapest standard endpoint. The endpoint table lists every endpoint in every tier.
- Blended price = (3 × input + output) ÷ 4, a 3:1 input:output token mix. Spread = most expensive ÷ cheapest blended price among standard-tier providers. Both are calculations on reported prices.
- Markup = OpenRouter list price ÷ first-party list price − 1. First-party prices are the vendors’ published list prices as recorded in the product price table. The effective markup adds the credit-purchase fee for card purchases on the Standard plan.
- Gateway delay (time to first token, total time, billed cost per call): not measured: no key in the environment. A live harness is ready; it sends nothing without a key and has a hard spending cap.
Caveats
- Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
- A provider listed on OpenRouter is reached through OpenRouter; its price there may differ from the price on the provider’s own site.
- The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
- No latency or throughput figure: the keyless API returned none. A cheaper provider is not shown to be slower or faster.
- First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.
- Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.
Sources
OpenRouter public API: models and provider endpoints (snapshot)
Prices, context, quantization and uptime per provider endpoint as reported by OpenRouter’s public, keyless API on 2026-10-06. Third-party-reported, not measured by Agent. Latency and throughput were not returned.
OpenRouter states that inference is billed at the provider list price and that its fee is charged when credits are bought (5.5% on Standard by card, $0.80 minimum; 8% on Business; 5% by crypto). Page fetched 2026-10-06.
Anthropic list prices (Claude models)
Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.
Token prices as listed by the vendor on 2026-10-03.
Gemini 3.x Flash prices as listed by the vendor on 2026-09-21. The vendor announced a doubling from 2027-01-01.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “Inference provider index: 27 models, 52 providers”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/inference-provider-index.
Explainers that cite this study
Read the methods and terms in the context of these recorded results.
More comparisons based on this study (90)
These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.
- Amazon Bedrock vs Baseten
- Amazon Bedrock vs Cerebras
- Amazon Bedrock vs Claude Platform on AWS
- Amazon Bedrock vs DeepInfra
- Amazon Bedrock vs Groq
- Amazon Bedrock vs Nebius
- Amazon Bedrock vs Novita AI
- Amazon Bedrock vs Parasail
- Amazon Bedrock vs SambaNova
- Amazon Bedrock vs SiliconFlow
- Amazon Bedrock vs Together AI
- Anthropic vs Amazon Bedrock
- Anthropic vs Azure
- Anthropic vs Claude Platform on AWS
- Anthropic vs Google Vertex AI
- Azure vs Claude Platform on AWS
- Baseten vs Cloudflare Workers AI
- Baseten vs SiliconFlow
- Cerebras vs Baseten
- Cerebras vs Nebius
- Cerebras vs Novita AI
- Cerebras vs Parasail
- Cerebras vs SambaNova
- Cerebras vs SiliconFlow
- Cloudflare Workers AI vs SiliconFlow
- DeepInfra vs Baseten
- DeepInfra vs Cerebras
- DeepInfra vs Cloudflare Workers AI
- DeepInfra vs Nebius
- DeepInfra vs Novita AI
- DeepInfra vs Parasail
- DeepInfra vs SambaNova
- DeepInfra vs SiliconFlow
- Fireworks AI vs Baseten
- Fireworks AI vs Cloudflare Workers AI
- Fireworks AI vs DeepInfra
- Fireworks AI vs Nebius
- Fireworks AI vs Novita AI
- Fireworks AI vs Parasail
- Fireworks AI vs SiliconFlow
- Google AI Studio vs Google Vertex AI
- Google Vertex AI vs Azure
- Google Vertex AI vs Baseten
- Google Vertex AI vs Cerebras
- Google Vertex AI vs Claude Platform on AWS
- Google Vertex AI vs Cloudflare Workers AI
- Google Vertex AI vs DeepInfra
- Google Vertex AI vs Groq
- Google Vertex AI vs Nebius
- Google Vertex AI vs Novita AI
- Google Vertex AI vs Parasail
- Google Vertex AI vs SambaNova
- Google Vertex AI vs SiliconFlow
- Google Vertex AI vs Together AI
- Groq vs Baseten
- Groq vs Cloudflare Workers AI
- Groq vs DeepInfra
- Groq vs Nebius
- Groq vs Novita AI
- Groq vs Parasail
- Groq vs SambaNova
- Groq vs SiliconFlow
- Nebius vs Baseten
- Nebius vs Cloudflare Workers AI
- Nebius vs Novita AI
- Nebius vs Parasail
- Nebius vs SiliconFlow
- Novita AI vs Baseten
- Novita AI vs Cloudflare Workers AI
- Novita AI vs SiliconFlow
- OpenAI vs Azure
- Parasail vs Baseten
- Parasail vs Cloudflare Workers AI
- Parasail vs Novita AI
- Parasail vs SiliconFlow
- SambaNova vs Baseten
- SambaNova vs Cloudflare Workers AI
- SambaNova vs Nebius
- SambaNova vs Novita AI
- SambaNova vs Parasail
- SambaNova vs SiliconFlow
- Together AI vs Baseten
- Together AI vs Cerebras
- Together AI vs Cloudflare Workers AI
- Together AI vs DeepInfra
- Together AI vs Nebius
- Together AI vs Novita AI
- Together AI vs Parasail
- Together AI vs SambaNova
- Together AI vs SiliconFlow
Models and comparisons in this study
Write-ups on this study
AI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
Claude Fable 5.1 vs Opus 5.5 vs Sonnet 5.5: speed, tokens and price tested
24 of 24: Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 each passed every hard task. Fable cost 3.3x Opus and 6.5x Sonnet per pass (list-price calculation).
LLM API pricing comparison, October 2026: Claude vs GPT vs Gemini per million tokens
21 LLM API prices per million tokens, October 2026: Claude, GPT, Gemini and more. Blended prices span 100x. Price per token is not price per task.
Most AI model comparisons are ties: 44 of 1,196 rows show a clear gap
44 of 1,196 AI model comparison rows show a gap under our overlap rules. Most gaps are timing rows. No effort row separates quality.
The cheapest way to run an AI coding agent: 7 levers from measured runs
7 levers that may cut an AI coding agent's bill, sized from our data: prompt cache 3.9x, Fable/Sonnet cost per pass 6.5x, and 5 more. List-price calculations.
AI coding benchmarks roundup, October 2026: sixteen studies, every number in one place
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
OpenRouter vs going direct: what the gateway really costs
OpenRouter charged the vendor's per-token price on 10 of 10 models we checked. The cost is a 5.5% credit fee. What we know, and what is still unmeasured.
The cheapest place to run open models right now (October 2026 snapshot)
DeepSeek V4, gpt-oss-120b, Llama, Kimi K3 and GLM 5.3 across 52 providers. Prices differ up to 12.6x. Snapshot 2026-10-06, with the caveats.
More studies
All benchmarksDoes a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
96 calls. Strict passes, schema vs instructions: Haiku 18/24 vs 0/24, Sonnet 12/12 vs 12/12, GPT-6.1 Sol 12/12 vs 12/12. With intervals.
Does a new Claude Code session reuse the prompt cache of an earlier one?
30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.