• Inference
  • Providers
  • OpenRouter
  • Pricing
  • Gateway
  • Open Weight
  • Third Party Reported

Inference provider index: 27 models, 52 providers

For the same model, how much do inference providers differ in price, and what does a gateway such as OpenRouter add over the first-party list price?

Published · 34 charts · Download the data or a carousel

265

n = 27

Provider endpoints in the snapshot · endpoints, 52 providers, 27 models

The answer

OpenRouter’s public API listed 265 endpoints from 52 providers for 27 models on 2026-10-06. For 10 of the 10 models with a first-party list price, OpenRouter’s per-token price was the same as the vendor’s; the cost of the gateway is the 5.5% credit-purchase fee (Standard plan, third-party-reported), so the effective markup is +5.5%. 15 of 15 closed models (Claude, GPT, Gemini) with 2 or more providers had one price across every standard-tier provider; their price differences come from named tiers (flex, priority, fast) and regional endpoints. Open-weight models differ by provider: DeepSeek V4 Flash 0423 12.6x (15 providers), DeepSeek V4 Pro 0423 11.2x (15 providers), gpt-oss-120b 6.9x (20 providers), Llama 3.3 70B Instruct 6.7x (10 providers) between the most expensive and the cheapest standard-tier provider (a calculation on reported prices). The cheapest endpoints often report lower precision or a shorter context. Latency and throughput were not in the keyless API (0 of 265 endpoints), and the gateway’s own delay is not measured: no key in the environment.

Live story

Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.

Live story · 42 sSame model, different price: 265 provider endpoints compared

Same model, different price: 265 provider endpoints compared

Reported by OpenRouter’s public API on 2026-10-06: open-weight prices vary up to 12.6x across providers; closed models have one standard price, and the gateway adds a 5.5% credit fee.

Transcript
  1. Provider index · third-party-reported · 2026-10-06. Same model, different price. 265 endpoints, 52 providers, 27 models. Prices as OpenRouter’s public API reports them.
  2. 265 endpoints, 52 providers, 27 models. The keyless API returned latency for 0 of them, so speed is not compared. Provider endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Endpoints with a latency figure: 0 of 265. Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
  3. Open-weight models vary most: DeepSeek V4 Flash 12.6×, DeepSeek V4 Pro 11.2×. All 15 closed models with 2+ providers: one standard price (1×). Chart: Priciest ÷ cheapest standard provider · blended price (n = 2–32 each). Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
  4. One model, 10 providers: Llama 3.3 70B input costs $0.10 per million tokens at DeepInfra (fp8) and $1.04 at Together. Chart: Llama 3.3 70B · USD per million tokens · one row per provider. Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
  5. Gateway markup: for 10 of 10 models OpenRouter charges the vendor’s per-token price. The cost is a 5.5% fee when you buy credits. Models at the first-party per-token price: 10 of 10. Credit-purchase fee, Standard plan by card: 5.5% ($0.80 minimum by card). Effective markup per token after the fee: +5.5% (calculation). Caveat: First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.
  6. Example: Sonnet 5.5 costs $2.00 in and $10.00 out per million tokens on both. A price is not a measurement: no winner. Chart: Claude Sonnet 5.5 · OpenRouter vs Anthropic list price · USD per million. Caveat: Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.
  7. Prices change often: every endpoint, tier and source date online.

Key numbers

12.6x

Largest standard-tier price spread (DeepSeek V4 Flash 0423)

(most expensive: Cloudflare; cheapest: StreamLake (fp8)) · n = 15

10 of 10

Models where OpenRouter’s per-token price equals the first-party list price

n = 10

5.5%

Credit-purchase fee on OpenRouter’s Standard plan (third-party-reported)

($0.80 minimum by card)

0 of 265

Endpoints with a latency figure in the keyless API

n = 265

Compare providers for one model

Pick a model, then sort or filter its providers. The study charts below show the same standard-tier prices; this table also lists every named tier and regional endpoint.

Price per provider: Claude Haiku 4.5

Reported by OpenRouter’s public API, snapshot October 6, 2026. Third-party-reported prices, not measured by Agent.

Reported

4 providers with a standard-tier endpoint · all 4 providers report the same price ($2.00 blended) · showing 4 of 4 rows

  1. Inputall 4 at $1.00

  2. Outputall 4 at $5.00

  3. Cache readall 4 at $0.10

  • one provider (reported price, USD per million tokens, log scale per strip)
  • first-party list price

Blended price and “vs cheapest” are calculations on the reported prices (3 input : 1 output tokens), against the cheapest standard-tier provider. A lower price can come with lower precision (fp4, fp8) or a shorter context. Uptime is what the API reported for the last day; the API reported no latency or throughput.

Every provider Price a workload via a provider

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Calculation
ModelSpread
  1. DeepSeek V4 Flash 0423n = 15 providers
  2. DeepSeek V4 Pro 0423n = 15 providers
  3. gpt-oss-120bn = 20 providers
  4. Llama 3.3 70B Instructn = 10 providers
  5. GLM 5.3n = 32 providers
  6. Kimi K3n = 19 providers
  7. Llama 4 Maverickn = 3 providers
  8. 15 more models: the same price at every providerClaude Haiku 4.5, Claude Sonnet 5, Claude Sonnet 5.5, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Fable 5.1, GPT-6 Sol, GPT-6 Luna, GPT-6 Astra, GPT-5.5, Gemini 3.8 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.1 Pro Preview
    1x

List-price calculation, not a run. 22 rows. Highest DeepSeek V4 Flash 0423 12.6x (n 15). Lowest Gemini 3.1 Pro Preview 1x (n 2).

Notesn 2–32 per row

Standard tier, blended price (3 input : 1 output), most expensive provider ÷ cheapest provider

A calculation on prices reported by OpenRouter’s public API, snapshot 2026-10-06. n = providers with a standard-tier endpoint. 1x means every provider charges the same. Highlighted: a spread of 4x or more.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Calculation

0%per-token markup on 10 of 10 models

5.5% with the card credit feeCalculation

  • Claude Haiku 4.5
  • Claude Sonnet 5
  • Claude Sonnet 5.5
  • Claude Opus 4.8
  • Claude Opus 5
  • Claude Opus 5.5
  • Claude Fable 5.1
  • GPT-6 Luna
  • Gemini 3.8 Flash
  • Gemini 3.5 Flash
  • per-token markup over the first-party list price (input and output)
  • With the 5.5% card credit fee (input) (calculation, hollow)

List-price calculation, not a run. 10 rows, 3 series: Per-token markup (input), Per-token markup (output), With the 5.5% card credit fee (input). Per-token markup (input): all at 0%. Per-token markup (output): all at 0%.

Notes

Per-token markup, and the markup after the Standard credit-purchase fee (calculation)

A calculation: OpenRouter list price ÷ first-party list price − 1, then × (1 + 5.5%) for credits bought by card on the Standard plan (purchases large enough that the minimum fee does not apply). Fees as published on 2026-10-06; first-party prices from the vendors’ list prices.

Sources: OpenRouter public API: models and provider endpoints (snapshot), OpenRouter pricing and fees, Anthropic list prices (Claude models), OpenAI list prices, Google Gemini list prices

Share card (PNG)
Per-model price charts (32)
Reported

OpenRouter

Anthropicfirst-party list price

Same price on all 2 price types
  • InputOpenRouter: $1.00 same priceAnthropic: $1.00
  • OutputOpenRouter: $5.00 same priceAnthropic: $5.00

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $1.00. Output: all at $5.00.

Notes

USD per million tokens, before any credit-purchase fee

OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.

Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)

Share card (PNG)
Reported

OpenRouter

Anthropicfirst-party list price

Same price on all 2 price types
  • InputOpenRouter: $2.00 same priceAnthropic: $2.00
  • OutputOpenRouter: $10.00 same priceAnthropic: $10.00

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $2.00. Output: all at $10.00.

Notes

USD per million tokens, before any credit-purchase fee

OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.

Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)

Share card (PNG)
Reported

OpenRouter

Anthropicfirst-party list price

Same price on all 2 price types
  • InputOpenRouter: $2.00 same priceAnthropic: $2.00
  • OutputOpenRouter: $10.00 same priceAnthropic: $10.00

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $2.00. Output: all at $10.00.

Notes

USD per million tokens, before any credit-purchase fee

OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.

Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)

Share card (PNG)
Reported

OpenRouter

Anthropicfirst-party list price

Same price on all 2 price types
  • InputOpenRouter: $5.00 same priceAnthropic: $5.00
  • OutputOpenRouter: $25.00 same priceAnthropic: $25.00

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $5.00. Output: all at $25.00.

Notes

USD per million tokens, before any credit-purchase fee

OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.

Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)

Share card (PNG)
Reported

OpenRouter

Anthropicfirst-party list price

Same price on all 2 price types
  • InputOpenRouter: $5.00 same priceAnthropic: $5.00
  • OutputOpenRouter: $25.00 same priceAnthropic: $25.00

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $5.00. Output: all at $25.00.

Notes

USD per million tokens, before any credit-purchase fee

OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.

Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)

Share card (PNG)
Reported

OpenRouter

Anthropicfirst-party list price

Same price on all 2 price types
  • InputOpenRouter: $4.00 same priceAnthropic: $4.00
  • OutputOpenRouter: $20.00 same priceAnthropic: $20.00

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $4.00. Output: all at $20.00.

Notes

USD per million tokens, before any credit-purchase fee

OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.

Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)

Share card (PNG)
Reported

OpenRouter

Anthropicfirst-party list price

Same price on all 2 price types
  • InputOpenRouter: $10.00 same priceAnthropic: $10.00
  • OutputOpenRouter: $50.00 same priceAnthropic: $50.00

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $10.00. Output: all at $50.00.

Notes

USD per million tokens, before any credit-purchase fee

OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Anthropic price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.

Sources: OpenRouter public API: models and provider endpoints (snapshot), Anthropic list prices (Claude models)

Share card (PNG)
Reported

OpenRouter

OpenAIfirst-party list price

Same price on all 2 price types
  • InputOpenRouter: $0.10 same priceOpenAI: $0.10
  • OutputOpenRouter: $0.50 same priceOpenAI: $0.50

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $0.1. Output: all at $0.5.

Notes

USD per million tokens, before any credit-purchase fee

OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. OpenAI price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.

Sources: OpenRouter public API: models and provider endpoints (snapshot), OpenAI list prices

Share card (PNG)
Reported

OpenRouter

Google AI Studiofirst-party list price

Same price on all 2 price types
  • InputOpenRouter: $0.75 same priceGoogle AI Studio: $0.75
  • OutputOpenRouter: $3.75 same priceGoogle AI Studio: $3.75

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $0.75. Output: all at $3.75.

Notes

USD per million tokens, before any credit-purchase fee

OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Google price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.

Sources: OpenRouter public API: models and provider endpoints (snapshot), Google Gemini list prices

Share card (PNG)
Reported

OpenRouter

Google AI Studiofirst-party list price

Same price on all 2 price types
  • InputOpenRouter: $1.50 same priceGoogle AI Studio: $1.50
  • OutputOpenRouter: $9.00 same priceGoogle AI Studio: $9.00

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 2 series: Input, Output. Input: all at $1.50. Output: all at $9.00.

Notes

USD per million tokens, before any credit-purchase fee

OpenRouter price reported by OpenRouter’s public API, snapshot 2026-10-06. Google price from the vendor’s published list price as recorded in the product price table. OpenRouter charges a 5.5% fee when credits are bought, not per request; the markup chart shows the price with that fee.

Sources: OpenRouter public API: models and provider endpoints (snapshot), Google Gemini list prices

Share card (PNG)
Reported
  1. Inputall 4 at $1.00

  2. Outputall 4 at $5.00

  3. Cache readall 4 at $0.10

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 4 rows, 3 series: Input, Output, Cache read. Input: all at $1.00. Output: all at $5.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported
  1. Inputall 5 at $2.00

  2. Outputall 5 at $10.00

  3. Cache readall 5 at $0.20

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $2.00. Output: all at $10.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported
  1. Inputall 5 at $2.00

  2. Outputall 5 at $10.00

  3. Cache readall 5 at $0.20

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $2.00. Output: all at $10.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported
  1. Inputall 5 at $5.00

  2. Outputall 5 at $25.00

  3. Cache readall 5 at $0.50

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $5.00. Output: all at $25.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported
  1. Inputall 5 at $5.00

  2. Outputall 5 at $25.00

  3. Cache readall 5 at $0.50

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $5.00. Output: all at $25.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported
  1. Inputall 5 at $4.00

  2. Outputall 5 at $20.00

  3. Cache readall 5 at $0.20

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $4.00. Output: all at $20.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported
  1. Inputall 4 at $10.00

  2. Outputall 4 at $50.00

  3. Cache readall 4 at $0.25

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 4 rows, 3 series: Input, Output, Cache read. Input: all at $10.00. Output: all at $50.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported

Azure

OpenAI

Same price on all 3 price types
  • InputAzure: $2.00 same priceOpenAI: $2.00
  • OutputAzure: $10.00 same priceOpenAI: $10.00
  • Cache readAzure: $0.20 same priceOpenAI: $0.20

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $2.00. Output: all at $10.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported

Azure

OpenAI

Same price on all 3 price types
  • InputAzure: $0.10 same priceOpenAI: $0.10
  • OutputAzure: $0.50 same priceOpenAI: $0.50
  • Cache readAzure: $0.01 same priceOpenAI: $0.01

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $0.1. Output: all at $0.5.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported

Azure

OpenAI

Same price on all 3 price types
  • InputAzure: $10.00 same priceOpenAI: $10.00
  • OutputAzure: $50.00 same priceOpenAI: $50.00
  • Cache readAzure: $1.00 same priceOpenAI: $1.00

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $10.00. Output: all at $50.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported

Azure

OpenAI

Same price on all 3 price types
  • InputAzure: $5.00 same priceOpenAI: $5.00
  • OutputAzure: $30.00 same priceOpenAI: $30.00
  • Cache readAzure: $0.50 same priceOpenAI: $0.50

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $5.00. Output: all at $30.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported
  1. Input2 tie at $0.03 → Cerebras $0.35 · 11.7x

  2. Output2 tie at $0.17 → SambaNova $0.95 · 5.6x

  3. Cache readDigitalOcean $0.012 → Cerebras $0.35 · 29.2x

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 20 rows, 3 series: Input, Output, Cache read. Input: highest Cerebras (fp16) $0.35. Lowest DekaLLM (bf16) $0.03. Output: highest SambaNova $0.95. Lowest DeepInfra (bf16) $0.17.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported

Google AI Studio

Google Vertex

Same price on all 3 price types
  • InputGoogle AI Studio: $0.75 same priceGoogle Vertex: $0.75
  • OutputGoogle AI Studio: $3.75 same priceGoogle Vertex: $3.75
  • Cache readGoogle AI Studio: $0.075 same priceGoogle Vertex: $0.075

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $0.75. Output: all at $3.75.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported

Google AI Studio

Google Vertex

Same price on all 3 price types
  • InputGoogle AI Studio: $1.50 same priceGoogle Vertex: $1.50
  • OutputGoogle AI Studio: $9.00 same priceGoogle Vertex: $9.00
  • Cache readGoogle AI Studio: $0.15 same priceGoogle Vertex: $0.15

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $1.50. Output: all at $9.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported

Google AI Studio

Google Vertex

Same price on all 3 price types
  • InputGoogle AI Studio: $0.30 same priceGoogle Vertex: $0.30
  • OutputGoogle AI Studio: $2.50 same priceGoogle Vertex: $2.50
  • Cache readGoogle AI Studio: $0.03 same priceGoogle Vertex: $0.03

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $0.3. Output: all at $2.50.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported

Google AI Studio

Google Vertex

Same price on all 3 price types
  • InputGoogle AI Studio: $2.00 same priceGoogle Vertex: $2.00
  • OutputGoogle AI Studio: $12.00 same priceGoogle Vertex: $12.00
  • Cache readGoogle AI Studio: $0.20 same priceGoogle Vertex: $0.20

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $2.00. Output: all at $12.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported
  1. InputDigitalOcean $0.1875 → Parasail $0.35 · 1.9x

  2. OutputDigitalOcean $0.6525 → Parasail $1.00 · 1.5x

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 3 rows, 2 series: Input, Output. Input: highest Parasail (fp8) $0.35. Lowest DigitalOcean $0.19. Output: highest Parasail (fp8) $1.00. Lowest DigitalOcean $0.65.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported
  1. InputDeepInfra $0.10 → Together $1.04 · 10.4x

  2. OutputDeepInfra $0.32 → Cloudflare $2.25 · 7.0x

  3. Cache readAkashML $0.10 → CoreWeave $0.71 · 7.1x

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 10 rows, 3 series: Input, Output, Cache read. Input: highest Together $1.04. Lowest DeepInfra (fp8) $0.1. Output: highest Cloudflare (fp8) $2.25. Lowest DeepInfra (fp8) $0.32.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported
  1. InputRelace $0.2067 → NextBit $1.74 · 8.4x

  2. OutputStreamLake $0.4176 → Reka $9.00 · 21.6x

  3. Cache readStreamLake $0.0174 → Venice $0.33 · 19.0x

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 15 rows, 3 series: Input, Output, Cache read. Input: highest NextBit (fp8) $1.74. Lowest Relace (fp4) $0.21. Output: highest Reka $9.00. Lowest StreamLake (fp8) $0.42.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported
  1. InputRelace $0.012 → Cloudflare $0.44 · 36.7x

  2. OutputStreamLake $0.084 → OpenInference $1.41 · 16.8x

  3. Cache readStreamLake $0.0084 → Parasail $0.07 · 8.3x

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 15 rows, 3 series: Input, Output, Cache read. Input: highest Cloudflare $0.44. Lowest Relace (fp4) $0.012. Output: highest OpenInference (fp4) $1.41. Lowest StreamLake (fp8) $0.084.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported
  1. InputRelace $0.83 → Alibaba $3.45 · 4.2x

  2. OutputPhala $9.75 → Alibaba $17.25 · 1.8x

  3. Cache readPhala $0.195 → AkashML $1.30 · 6.7x

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 19 rows, 3 series: Input, Output, Cache read. Input: highest Alibaba $3.45. Lowest Relace (fp4) $0.83. Output: highest Alibaba $17.25. Lowest Phala $9.75.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)
Reported
  1. InputRelace $0.03 → 14 tie at $1.40 · 46.7x

  2. OutputNovita $1.32 → Relace $12.00 · 9.1x

  3. Cache readRelace $0.03 → 11 tie at $0.26 · 8.7x

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 32 rows, 3 series: Input, Output, Cache read. Input: highest AtlasCloud (fp8) $1.40. Lowest Relace $0.03. Output: highest Relace $12.00. Lowest Novita (fp8) $1.32.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Share card (PNG)

Tables

Provider index per model (reported by OpenRouter’s public API, snapshot 2026-10-06)

ModelProvidersEndpoints (all tiers)OpenRouter list price, in / out (USD per M)Cheapest standard providerMost expensive standard providerSpread (blended)First-party list price, in / out
Claude Haiku 4.548$1.00 / $5.00Amazon Bedrock: $1.00 / $5.00Google Vertex: $1.00 / $5.001.0xAnthropic: $1.00 / $5.00
Claude Sonnet 5510$2.00 / $10.00Amazon Bedrock: $2.00 / $10.00Google Vertex: $2.00 / $10.001.0xAnthropic: $2.00 / $10.00
Claude Sonnet 5.558$2.00 / $10.00Amazon Bedrock: $2.00 / $10.00Google Vertex: $2.00 / $10.001.0xAnthropic: $2.00 / $10.00
Claude Opus 4.8511$5.00 / $25.00Amazon Bedrock: $5.00 / $25.00Google Vertex: $5.00 / $25.001.0xAnthropic: $5.00 / $25.00
Claude Opus 5511$5.00 / $25.00Amazon Bedrock: $5.00 / $25.00Google Vertex: $5.00 / $25.001.0xAnthropic: $5.00 / $25.00
Claude Opus 5.5511$4.00 / $20.00Amazon Bedrock: $4.00 / $20.00Google Vertex: $4.00 / $20.001.0xAnthropic: $4.00 / $20.00
Claude Fable 5.144$10.00 / $50.00Amazon Bedrock: $10.00 / $50.00Google Vertex: $10.00 / $50.001.0xAnthropic: $10.00 / $50.00
GPT-6 Sol37$2.00 / $10.00Azure: $2.00 / $10.00OpenAI: $2.00 / $10.001.0xnot in the price table
GPT-6 Luna37$0.10 / $0.50Azure: $0.10 / $0.50OpenAI: $0.10 / $0.501.0xOpenAI: $0.10 / $0.50
GPT-6 Astra37$10.00 / $50.00Azure: $10.00 / $50.00OpenAI: $10.00 / $50.001.0xnot in the price table
GPT-5.537$5.00 / $30.00Azure: $5.00 / $30.00OpenAI: $5.00 / $30.001.0xnot in the price table
gpt-oss-120b2023$0.037 / $0.17CoreWeave (fp4): $0.03 / $0.17Cerebras (fp16): $0.35 / $0.756.9xnot in the price table
Gemini 3.8 Flash26$0.75 / $3.75Google AI Studio: $0.75 / $3.75Google Vertex: $0.75 / $3.751.0xGoogle: $0.75 / $3.75
Gemini 3.5 Flash27$1.50 / $9.00Google AI Studio: $1.50 / $9.00Google Vertex: $1.50 / $9.001.0xGoogle: $1.50 / $9.00
Gemini 3.5 Flash Lite28$0.30 / $2.50Google AI Studio: $0.30 / $2.50Google Vertex: $0.30 / $2.501.0xnot in the price table
Gemini 3.1 Pro Preview26$2.00 / $12.00Google AI Studio: $2.00 / $12.00Google Vertex: $2.00 / $12.001.0xnot in the price table
Llama 4 Maverick44$0.19 / $0.65DigitalOcean: $0.19 / $0.65Parasail (fp8): $0.35 / $1.001.7xnot in the price table
Llama 3.3 70B Instruct1011$0.22 / $0.50DeepInfra (fp8): $0.10 / $0.32Together: $1.04 / $1.046.7xnot in the price table
DeepSeek V4 Pro 04231616$0.21 / $0.42StreamLake (fp8): $0.21 / $0.42Reka: $0.90 / $9.0011.2xnot in the price table
DeepSeek V4 Flash 04231616$0.012 / $1.28StreamLake (fp8): $0.042 / $0.084Cloudflare: $0.44 / $1.3212.6xnot in the price table
Qwen3.8 Max (0902)11$2.00 / $6.00Alibaba: $2.00 / $6.00only one providern/anot in the price table
Qwen3.8 Flash11$0.15 / $0.47Alibaba: $0.15 / $0.47only one providern/anot in the price table
Mistral Medium 3.513$1.50 / $7.50Mistral: $1.50 / $7.50only one providern/anot in the price table
Mistral Large 3 251212$0.50 / $1.50Mistral: $0.50 / $1.50only one providern/anot in the price table
Grok 4.715$2.00 / $6.00xAI: $2.00 / $6.00only one providern/anot in the price table
Kimi K32024$0.95 / $14.00Relace (fp4): $0.83 / $13.00Alibaba: $3.45 / $17.251.8xnot in the price table
GLM 5.33241$0.07 / $7.00Novita (fp8): $0.42 / $1.32Relace: $0.03 / $12.004.7xnot in the price table

Gateway overhead: what is and is not measured

MetricStatus
Per-token price through OpenRouter vs first-partyReported list prices, snapshot 2026-10-06
Credit-purchase fee5.5% on Standard (third-party-reported, 2026-10-06)
Provider latency and throughput (last 30 min)unknown: the keyless API returned none for 265 endpoints
Time to first token and total time through the gateway vs directnot measured: no key in the environment. The live harness sends nothing without an OpenRouter key.
Billed cost per call through the gatewaynot measured: no key in the environment. The live harness sends nothing without an OpenRouter key.

Every endpoint OpenRouter listed (reported by OpenRouter’s public API, snapshot 2026-10-06)

ModelProviderTierQuantizationContext (tokens)Input (USD per M)Output (USD per M)Cache read (USD per M)Uptime, last day (%)
Claude Fable 5.1Azurestandardnot reported1,000,000$10.00$50.00$0.2599%
Claude Fable 5.1Anthropicstandardnot reported1,000,000$10.00$50.00$0.25100%
Claude Fable 5.1Amazon Bedrockstandardnot reported1,000,000$10.00$50.00$0.25—
Claude Fable 5.1Google Vertexstandardnot reported1,000,000$10.00$50.00$0.25100%
Claude Haiku 4.5Azurestandardnot reported200,000$1.00$5.00$0.1100%
Claude Haiku 4.5Amazon Bedrockstandardnot reported200,000$1.00$5.00$0.1100%
Claude Haiku 4.5Google Vertexstandardnot reported200,000$1.00$5.00$0.1100%
Claude Haiku 4.5Anthropicstandardnot reported200,000$1.00$5.00$0.1100%
Claude Haiku 4.5Amazon Bedrockregionalnot reported200,000$1.10$5.50$0.11100%
Claude Haiku 4.5Google Vertexregionalnot reported200,000$1.10$5.50$0.11100%
Claude Haiku 4.5Amazon Bedrockregionalnot reported200,000$1.10$5.50$0.11100%
Claude Haiku 4.5Google Vertexregionalnot reported200,000$1.10$5.50$0.11100%
Claude Opus 4.8Claude Platform on AWSstandardnot reported1,000,000$5.00$25.00$0.5100%
Claude Opus 4.8Azurestandardnot reported1,000,000$5.00$25.00$0.598%
Claude Opus 4.8Amazon Bedrockstandardnot reported1,000,000$5.00$25.00$0.595%
Claude Opus 4.8Google Vertexstandardnot reported1,000,000$5.00$25.00$0.5100%
Claude Opus 4.8Anthropicstandardnot reported1,000,000$5.00$25.00$0.5100%
Claude Opus 4.8Azureregionalnot reported1,000,000$5.50$27.50$0.55—
Claude Opus 4.8Amazon Bedrockregionalnot reported1,000,000$5.50$27.50$0.55100%
Claude Opus 4.8Google Vertexregionalnot reported1,000,000$5.50$27.50$0.55—
Claude Opus 4.8Amazon Bedrockregionalnot reported1,000,000$5.50$27.50$0.55100%
Claude Opus 4.8Google Vertexregionalnot reported1,000,000$5.50$27.50$0.55100%
Claude Opus 4.8Anthropicfastnot reported1,000,000$10.00$50.00$1.00100%
Claude Opus 5Claude Platform on AWSstandardnot reported1,000,000$5.00$25.00$0.599%
Claude Opus 5Google Vertexstandardnot reported1,000,000$5.00$25.00$0.5100%
Claude Opus 5Amazon Bedrockstandardnot reported1,000,000$5.00$25.00$0.5100%
Claude Opus 5Azurestandardnot reported1,000,000$5.00$25.00$0.5100%
Claude Opus 5Anthropicstandardnot reported1,000,000$5.00$25.00$0.5100%
Claude Opus 5Amazon Bedrockregionalnot reported1,000,000$5.50$27.50$0.55100%
Claude Opus 5Azureregionalnot reported1,000,000$5.50$27.50$0.55—
Claude Opus 5Google Vertexregionalnot reported1,000,000$5.50$27.50$0.55100%
Claude Opus 5Google Vertexregionalnot reported1,000,000$5.50$27.50$0.55100%
Claude Opus 5Amazon Bedrockregionalnot reported1,000,000$5.50$27.50$0.55100%
Claude Opus 5Anthropicfastnot reported1,000,000$10.00$50.00$1.00100%
Claude Opus 5.5Amazon Bedrockstandardnot reported1,000,000$4.00$20.00$0.297%
Claude Opus 5.5Azurestandardnot reported1,000,000$4.00$20.00$0.2100%
Claude Opus 5.5Google Vertexstandardnot reported1,000,000$4.00$20.00$0.2100%
Claude Opus 5.5Claude Platform on AWSstandardnot reported1,000,000$4.00$20.00$0.2100%
Claude Opus 5.5Anthropicstandardnot reported1,000,000$4.00$20.00$0.2100%
Claude Opus 5.5Amazon Bedrockregionalnot reported1,000,000$4.40$22.00$0.22100%
Claude Opus 5.5Amazon Bedrockregionalnot reported1,000,000$4.40$22.00$0.22100%
Claude Opus 5.5Azureregionalnot reported1,000,000$4.40$22.00$0.22100%
Claude Opus 5.5Google Vertexregionalnot reported1,000,000$4.40$22.00$0.22100%
Claude Opus 5.5Google Vertexregionalnot reported1,000,000$4.40$22.00$0.22100%
Claude Opus 5.5Anthropicfastnot reported1,000,000$8.00$40.00$0.4100%
Claude Sonnet 5Claude Platform on AWSstandardnot reported1,000,000$2.00$10.00$0.2100%
Claude Sonnet 5Azurestandardnot reported1,000,000$2.00$10.00$0.2100%
Claude Sonnet 5Google Vertexstandardnot reported1,000,000$2.00$10.00$0.2100%
Claude Sonnet 5Amazon Bedrockstandardnot reported1,000,000$2.00$10.00$0.2100%
Claude Sonnet 5Anthropicstandardnot reported1,000,000$2.00$10.00$0.2100%
Claude Sonnet 5Amazon Bedrockregionalnot reported1,000,000$2.20$11.00$0.22100%
Claude Sonnet 5Azureregionalnot reported1,000,000$2.20$11.00$0.22—
Claude Sonnet 5Google Vertexregionalnot reported1,000,000$2.20$11.00$0.22100%
Claude Sonnet 5Google Vertexregionalnot reported1,000,000$2.20$11.00$0.22100%
Claude Sonnet 5Amazon Bedrockregionalnot reported1,000,000$2.20$11.00$0.22100%
Claude Sonnet 5.5Google Vertexstandardnot reported1,000,000$2.00$10.00$0.2100%
Claude Sonnet 5.5Amazon Bedrockstandardnot reported1,000,000$2.00$10.00$0.2100%
Claude Sonnet 5.5Azurestandardnot reported1,000,000$2.00$10.00$0.2100%
Claude Sonnet 5.5Claude Platform on AWSstandardnot reported1,000,000$2.00$10.00$0.2100%
Claude Sonnet 5.5Anthropicstandardnot reported1,000,000$2.00$10.00$0.2100%
Claude Sonnet 5.5Google Vertexregionalnot reported1,000,000$2.20$11.00$0.22100%
Claude Sonnet 5.5Azureregionalnot reported1,000,000$2.20$11.00$0.22100%
Claude Sonnet 5.5Google Vertexregionalnot reported1,000,000$2.20$11.00$0.22100%
DeepSeek V4 Flash 0423Relacestandardfp41,048,576$0.012$1.28$0.012100%
DeepSeek V4 Flash 0423OpenInferencestandardfp41,048,576$0.013$1.41$0.01399%
DeepSeek V4 Flash 0423StreamLakestandardfp81,024,000$0.042$0.084$0.008499%
DeepSeek V4 Flash 0423DeepInfrastandardfp81,048,576$0.09$0.18$0.018100%
DeepSeek V4 Flash 0423GMICloudstandardfp81,048,575$0.091$0.18$0.018100%
DeepSeek V4 Flash 0423Venicestandardnot reported1,000,000$0.097$0.19$0.0296%
DeepSeek V4 Flash 0423DigitalOceanstandardnot reported1,048,576$0.098$0.2$0.02100%
DeepSeek V4 Flash 0423SiliconFlowstandardfp81,048,576$0.13$0.28$0.028100%
DeepSeek V4 Flash 0423Alibabastandardfp81,000,000$0.13$0.27$0.02799%
DeepSeek V4 Flash 0423Baidustandardfp81,048,576$0.14$0.28$0.02899%
DeepSeek V4 Flash 0423Novitastandardfp81,048,576$0.14$0.28$0.028100%
DeepSeek V4 Flash 0423AtlasCloudstandardfp41,048,576$0.14$0.28$0.02899%
DeepSeek V4 Flash 0423Parasailstandardfp81,048,576$0.14$0.28$0.07100%
DeepSeek V4 Flash 0423Mancer 2standardfp81,048,576$0.19$0.5—95%
DeepSeek V4 Flash 0423Azureregionalnot reported1,048,576$0.21$0.56$0.03195%
DeepSeek V4 Flash 0423Cloudflarestandardnot reported384,000$0.44$1.32$0.01498%
DeepSeek V4 Pro 0423Relacestandardfp41,048,576$0.21$4.20$0.21100%
DeepSeek V4 Pro 0423StreamLakestandardfp81,024,000$0.21$0.42$0.01799%
DeepSeek V4 Pro 0423Parasailstandardfp81,048,576$0.45$3.48$0.198%
DeepSeek V4 Pro 0423Rekastandardnot reported1,048,576$0.9$9.00$0.1899%
DeepSeek V4 Pro 0423GMICloudstandardfp81,048,576$0.96$1.91$0.0897%
DeepSeek V4 Pro 0423DigitalOceanstandardnot reported1,048,576$1.04$2.09$0.21100%
DeepSeek V4 Pro 0423Cloudflarestandardnot reported1,048,576$1.15$2.55$0.298%
DeepSeek V4 Pro 0423DeepInfrastandardfp81,048,576$1.30$2.60$0.1100%
DeepSeek V4 Pro 0423Alibabastandardfp81,000,000$1.42$2.83$0.1290%
DeepSeek V4 Pro 0423SiliconFlowstandardfp81,048,576$1.50$3.13$0.1499%
DeepSeek V4 Pro 0423Novitastandardfp81,048,576$1.60$3.20$0.14100%
DeepSeek V4 Pro 0423Venicestandardnot reported1,000,000$1.65$3.30$0.3397%
DeepSeek V4 Pro 0423AtlasCloudstandardfp41,048,576$1.68$3.38$0.1399%
DeepSeek V4 Pro 0423Baidustandardfp81,048,576$1.69$3.38$0.14100%
DeepSeek V4 Pro 0423NextBitstandardfp81,048,576$1.74$3.48$0.1499%
DeepSeek V4 Pro 0423Azureregionalnot reported1,048,576$1.91$3.83$0.1699%
Gemini 3.1 Pro PreviewGoogle Vertexflexnot reported1,048,576$1.00$6.00$0.193%
Gemini 3.1 Pro PreviewGoogle AI Studioflexnot reported1,048,576$1.00$6.00$0.1100%
Gemini 3.1 Pro PreviewGoogle Vertexstandardnot reported1,048,576$2.00$12.00$0.298%
Gemini 3.1 Pro PreviewGoogle AI Studiostandardnot reported1,048,576$2.00$12.00$0.2100%
Gemini 3.1 Pro PreviewGoogle Vertexprioritynot reported1,048,576$3.60$21.60$0.36100%
Gemini 3.1 Pro PreviewGoogle AI Studioprioritynot reported1,048,576$3.60$21.60$0.3699%
Gemini 3.5 FlashGoogle Vertexflexnot reported1,048,576$0.75$4.50$0.07599%
Gemini 3.5 FlashGoogle AI Studioflexnot reported1,048,576$0.75$4.50$0.075100%
Gemini 3.5 FlashGoogle Vertexstandardnot reported1,048,576$1.50$9.00$0.1599%
Gemini 3.5 FlashGoogle AI Studiostandardnot reported1,048,576$1.50$9.00$0.15100%
Gemini 3.5 FlashGoogle Vertexregionalnot reported1,048,576$1.65$9.90$0.17—
Gemini 3.5 FlashGoogle Vertexprioritynot reported1,048,576$2.70$16.20$0.27100%
Gemini 3.5 FlashGoogle AI Studioprioritynot reported1,048,576$2.70$16.20$0.27100%
Gemini 3.5 Flash LiteGoogle Vertexflexnot reported1,048,576$0.15$1.25$0.015100%
Gemini 3.5 Flash LiteGoogle AI Studioflexnot reported1,048,576$0.15$1.25$0.015100%
Gemini 3.5 Flash LiteGoogle AI Studiostandardnot reported1,048,576$0.3$2.50$0.03100%
Gemini 3.5 Flash LiteGoogle Vertexstandardnot reported1,048,576$0.3$2.50$0.03100%
Gemini 3.5 Flash LiteGoogle Vertexregionalnot reported1,048,576$0.33$2.75$0.033100%
Gemini 3.5 Flash LiteGoogle Vertexregionalnot reported1,048,576$0.33$2.75$0.033100%
Gemini 3.5 Flash LiteGoogle Vertexprioritynot reported1,048,576$0.54$4.50$0.054100%
Gemini 3.5 Flash LiteGoogle AI Studioprioritynot reported1,048,576$0.54$4.50$0.054100%
Gemini 3.8 FlashGoogle AI Studioflexnot reported1,048,576$0.38$1.88$0.037100%
Gemini 3.8 FlashGoogle Vertexflexnot reported1,048,576$0.38$1.88$0.03799%
Gemini 3.8 FlashGoogle AI Studiostandardnot reported1,048,576$0.75$3.75$0.075100%
Gemini 3.8 FlashGoogle Vertexstandardnot reported1,048,576$0.75$3.75$0.07597%
Gemini 3.8 FlashGoogle AI Studioprioritynot reported1,048,576$1.35$6.75$0.14100%
Gemini 3.8 FlashGoogle Vertexprioritynot reported1,048,576$1.35$6.75$0.14100%
GLM 5.3Relacestandardnot reported1,048,576$0.03$12.00$0.03100%
GLM 5.3Waferregionalnot reported1,048,576$0.07$7.00$0.065100%
GLM 5.3InferenceNetstandardnot reported1,048,576$0.14$4.40$0.07100%
GLM 5.3Waferstandardnot reported1,048,576$0.15$7.00$0.14100%
GLM 5.3Rekastandardnot reported262,144$0.17$3.00$0.17100%
GLM 5.3Morphstandardfp81,048,576$0.18$3.55$0.1499%
GLM 5.3Makorastandardfp4980,000$0.18$4.40$0.1996%
GLM 5.3AkashMLstandardfp81,048,576$0.19$4.40$0.19100%
GLM 5.3Sail Researchregionalfp81,048,576$0.2$3.40$0.1599%
GLM 5.3Sail Researchstandardfp81,048,576$0.2$3.40$0.1599%
GLM 5.3Novitastandardfp81,048,576$0.42$1.32$0.07897%
GLM 5.3DeepInfrastandardfp41,048,576$0.56$2.50$0.1397%
GLM 5.3Inceptronstandardfp41,048,576$0.6$3.39$0.298%
GLM 5.3SiliconFlowstandardfp81,048,576$0.7$2.20$0.13100%
GLM 5.3Phalastandardnot reported1,048,576$0.84$2.64$0.1699%
GLM 5.3DigitalOceanstandardnot reported1,048,576$0.91$2.86$0.17100%
GLM 5.3GMICloudstandardfp81,048,576$0.98$3.08$0.1899%
GLM 5.3Alibabastandardnot reported1,000,000$1.19$3.74$0.24100%
GLM 5.3Decartstandardfp41,048,576$1.19$3.74$0.2100%
GLM 5.3Friendlistandardnot reported1,048,576$1.26$3.96$0.23100%
GLM 5.3Mistralstandardnvfp41,048,576$1.40$4.40$0.14100%
GLM 5.3Baidustandardfp81,048,576$1.40$4.40$0.26100%
GLM 5.3BaseTenstandardfp41,048,576$1.40$4.40$0.1498%
GLM 5.3Mistralstandardnvfp41,048,576$1.40$4.40$0.1499%
GLM 5.3Nebiusstandardfp41,024,000$1.40$4.40—96%
GLM 5.3Crusoestandardfp41,048,576$1.40$4.40$0.2698%
GLM 5.3PrimeIntellectstandardnot reported1,048,576$1.40$4.40$0.2699%
GLM 5.3Venicestandardnot reported1,000,000$1.40$4.40$0.2699%
GLM 5.3Togetherstandardnot reported1,048,575$1.40$4.40$0.2696%
GLM 5.3Parasailstandardfp81,048,576$1.40$4.40$0.26100%
GLM 5.3Modalstandardnot reported1,048,576$1.40$4.40$0.2698%
GLM 5.3BaseTenstandardfp41,048,576$1.40$4.40$0.1495%
GLM 5.3Fireworksstandardnot reported1,048,576$1.40$4.40$0.26100%
GLM 5.3Cloudflarestandardnot reported1,048,576$1.40$4.40$0.2698%
GLM 5.3AtlasCloudstandardfp81,048,576$1.40$4.40$0.26100%
GLM 5.3Z.AIstandardfp81,048,576$1.40$4.40$0.26100%
GLM 5.3Mistralstandardnvfp41,048,576$1.54$4.84$0.15100%
GLM 5.3Fireworksfastnot reported1,048,576$2.10$6.60$0.39100%
GLM 5.3BaseTenfastfp81,048,576$2.10$6.60$0.21100%
GLM 5.3BaseTenfastfp81,048,576$2.10$6.60$0.2199%
GLM 5.3Alibabafastnot reported1,000,000$2.80$8.80$0.56100%
GPT-5.5OpenAIflexnot reported1,050,000$2.50$15.00$0.25100%
GPT-5.5Azurestandardnot reported1,050,000$5.00$30.00$0.5100%
GPT-5.5OpenAIstandardnot reported1,050,000$5.00$30.00$0.5100%
GPT-5.5Azureregionalnot reported1,050,000$5.50$33.00$0.55100%
GPT-5.5Azureregionalnot reported1,050,000$5.50$33.00$0.55100%
GPT-5.5Amazon Bedrockregionalnot reported1,050,000$5.50$33.00$0.55—
GPT-5.5OpenAIfastnot reported1,050,000$12.50$75.00$1.25100%
GPT-6 AstraOpenAIflexnot reported1,050,000$5.00$25.00$0.5100%
GPT-6 AstraAzurestandardnot reported1,050,000$10.00$50.00$1.00100%
GPT-6 AstraOpenAIstandardnot reported1,050,000$10.00$50.00$1.00100%
GPT-6 AstraAmazon Bedrockregionalnot reported1,050,000$11.00$55.00$1.10—
GPT-6 AstraAzureregionalnot reported1,050,000$11.00$55.00$1.10100%
GPT-6 AstraOpenAIfastnot reported1,050,000$20.00$100$2.00100%
GPT-6 AstraOpenAIultrafastnot reported1,050,000$60.00$300$6.00100%
GPT-6 LunaOpenAIflexnot reported1,050,000$0.05$0.25$0.00598%
GPT-6 LunaOpenAIstandardnot reported1,050,000$0.1$0.5$0.01100%
GPT-6 LunaAzurestandardnot reported1,050,000$0.1$0.5$0.01100%
GPT-6 LunaAzureregionalnot reported1,050,000$0.11$0.55$0.011100%
GPT-6 LunaAzureregionalnot reported1,050,000$0.11$0.55$0.011100%
GPT-6 LunaAmazon Bedrockregionalnot reported1,050,000$0.11$0.55$0.01199%
GPT-6 LunaOpenAIfastnot reported1,050,000$0.2$1.00$0.02100%
GPT-6 SolOpenAIflexnot reported1,050,000$1.00$5.00$0.1100%
GPT-6 SolOpenAIstandardnot reported1,050,000$2.00$10.00$0.2100%
GPT-6 SolAzurestandardnot reported1,050,000$2.00$10.00$0.2100%
GPT-6 SolAzureregionalnot reported1,050,000$2.20$11.00$0.22100%
GPT-6 SolAzureregionalnot reported1,050,000$2.20$11.00$0.22100%
GPT-6 SolAmazon Bedrockregionalnot reported1,050,000$2.20$11.00$0.22100%
GPT-6 SolOpenAIfastnot reported1,050,000$4.00$20.00$0.4100%
gpt-oss-120bCoreWeavestandardfp4131,072$0.03$0.17$0.0399%
gpt-oss-120bDekaLLMstandardbf16131,072$0.03$0.18$0.03100%
gpt-oss-120bDeepInfrastandardbf16131,072$0.037$0.17—99%
gpt-oss-120bAkashMLstandardbf16131,072$0.037$0.19$0.037100%
gpt-oss-120bMancer 2standardfp8131,072$0.045$0.25—99%
gpt-oss-120bCrusoestandardbf16131,072$0.05$0.25$0.05100%
gpt-oss-120bNovitastandardfp4131,072$0.05$0.25—99%
gpt-oss-120bDigitalOceanstandardnot reported128,000$0.06$0.42$0.012100%
gpt-oss-120bGoogle Vertexstandardnot reported131,072$0.09$0.36—67%
gpt-oss-120bBaseTenstandardfp4128,072$0.1$0.5$0.1100%
gpt-oss-120bBaseTenstandardfp4128,072$0.1$0.5$0.1100%
gpt-oss-120bParasailstandardfp4131,072$0.1$0.75$0.055100%
gpt-oss-120bSambaNovastandardnot reported131,072$0.14$0.95—100%
gpt-oss-120bAmazon Bedrockregionalnot reported131,072$0.15$0.6—100%
gpt-oss-120bNebiusstandardfp4131,072$0.15$0.6—97%
gpt-oss-120bAmazon Bedrockstandardnot reported131,072$0.15$0.6—99%
gpt-oss-120bDeepInfrastandardbf16131,072$0.15$0.6—100%
gpt-oss-120bSiliconFlowstandardfp8131,072$0.15$0.6$0.07583%
gpt-oss-120bPhalastandardnot reported131,072$0.15$0.6—99%
gpt-oss-120bTogetherstandardnot reported131,072$0.15$0.6—87%
gpt-oss-120bGroqstandardnot reported131,072$0.15$0.6$0.07599%
gpt-oss-120bMarastandardnot reported131,072$0.15$0.75—97%
gpt-oss-120bCerebrasstandardfp16131,072$0.35$0.75$0.35100%
Grok 4.7xAIstandardnot reported500,000$2.00$6.00$0.599%
Grok 4.7xAIstandardnot reported500,000$2.00$6.00$0.599%
Grok 4.7xAIregionalnot reported500,000$2.20$6.60$0.55100%
Grok 4.7xAIprioritynot reported500,000$4.00$12.00$1.00100%
Grok 4.7xAIprioritynot reported500,000$4.00$12.00$1.0099%
Kimi K3Relacestandardfp41,048,576$0.83$13.00$0.45100%
Kimi K3Sail Researchstandardfp41,048,576$0.84$13.50$0.3100%
Kimi K3InferenceNetstandardfp41,048,576$0.95$14.00$0.31100%
Kimi K3Waferstandardnot reported1,048,576$0.95$14.00$0.499%
Kimi K3Morphstandardfp81,048,576$1.27$13.30$0.28100%
Kimi K3AkashMLstandardfp41,048,576$1.30$14.00$1.3098%
Kimi K3Makorastandardnot reported1,048,576$1.53$12.75$0.297%
Kimi K3Phalastandardnot reported1,048,576$1.95$9.75$0.297%
Kimi K3Decartstandardmxfp41,048,576$2.01$10.05$0.287%
Kimi K3DigitalOceanstandardnot reported1,048,576$2.55$12.95$0.26100%
Kimi K3Togetherstandardnot reported1,048,576$2.70$13.50$0.2799%
Kimi K3Waferregionalnot reported1,048,576$2.80$14.00$0.399%
Kimi K3DeepInfrastandardmxfp41,048,576$2.85$14.25$0.2899%
Kimi K3Amazon Bedrockregionalnot reported1,048,576$3.00$15.00$0.393%
Kimi K3Chutesstandardmxfp41,048,576$3.00$15.00$0.397%
Kimi K3Parasailstandardfp41,048,576$3.00$15.00$0.398%
Kimi K3Modalstandardmxfp41,048,576$3.00$15.00$0.398%
Kimi K3Fireworksstandardnot reported1,048,576$3.00$15.00$0.399%
Kimi K3BaseTenstandardfp81,048,576$3.00$15.00$0.398%
Kimi K3Moonshot AIstandardmxfp41,048,576$3.00$15.00$0.3100%
Kimi K3Alibabastandardnot reported1,048,576$3.45$17.25$0.3499%
Kimi K3InferenceNetfastfp4250,000$3.50$15.00$0.45100%
Kimi K3Fireworksregionalnot reported1,048,576$4.50$22.50$0.4599%
Kimi K3Fireworksfastnot reported1,048,576$4.50$22.50$0.4597%
Llama 3.3 70B InstructDeepInfrastandardfp8131,072$0.1$0.32—98%
Llama 3.3 70B InstructNovitastandardbf1612,288$0.14$0.4—98%
Llama 3.3 70B InstructAkashMLstandardfp8131,072$0.2$0.52$0.199%
Llama 3.3 70B InstructParasailstandardfp8131,072$0.22$0.5$0.11100%
Llama 3.3 70B InstructCloudflarestandardfp824,000$0.29$2.25—99%
Llama 3.3 70B InstructSambaNovastandardnot reported131,072$0.45$0.9—99%
Llama 3.3 70B InstructGroqstandardnot reported131,072$0.59$0.79$0.29100%
Llama 3.3 70B InstructCoreWeavestandardfp16128,000$0.71$0.71$0.7198%
Llama 3.3 70B InstructGoogle Vertexregionalnot reported128,000$0.72$0.72——
Llama 3.3 70B InstructGoogle Vertexstandardnot reported128,000$0.72$0.72——
Llama 3.3 70B InstructTogetherstandardnot reported131,072$1.04$1.04—93%
Llama 4 MaverickDigitalOceanstandardnot reported128,000$0.19$0.65—99%
Llama 4 MaverickNovitastandardfp81,048,576$0.27$0.85—98%
Llama 4 MaverickParasailstandardfp8524,288$0.35$1.00$0.17100%
Llama 4 MaverickGoogle Vertexregionalnot reported524,288$0.35$1.15——
Mistral Large 3 2512Mistralstandardnot reported262,144$0.5$1.50$0.05100%
Mistral Large 3 2512Mistralregionalnot reported262,144$0.55$1.65$0.055100%
Mistral Medium 3.5Mistralstandardnot reported262,144$1.50$7.50—100%
Mistral Medium 3.5Mistralstandardnot reported262,144$1.50$7.50—100%
Mistral Medium 3.5Mistralregionalnot reported262,144$1.65$8.25—100%
Qwen3.8 FlashAlibabastandardnot reported1,000,000$0.15$0.47$0.01699%
Qwen3.8 Max (0902)Alibabastandardnot reported1,000,000$2.00$6.00$0.25100%

Method

  1. One snapshot of OpenRouter’s public, keyless API on 2026-10-06: the model list and the endpoint list of 27 curated models (28 requests). Prices, context, quantization and uptime are copied as the API reported them; a field it did not return stays empty.
  2. Tier: read from the endpoint tag suffix. Flex, priority, fast, ultrafast and batch are named tiers; a region suffix (us, eu, europe, a cloud region) is regional; anything else is standard. This is a heuristic.
  3. Per-model charts show the standard tier only, with one bar per provider: its cheapest standard endpoint. The endpoint table lists every endpoint in every tier.
  4. Blended price = (3 × input + output) ÷ 4, a 3:1 input:output token mix. Spread = most expensive ÷ cheapest blended price among standard-tier providers. Both are calculations on reported prices.
  5. Markup = OpenRouter list price ÷ first-party list price − 1. First-party prices are the vendors’ published list prices as recorded in the product price table. The effective markup adds the credit-purchase fee for card purchases on the Standard plan.
  6. Gateway delay (time to first token, total time, billed cost per call): not measured: no key in the environment. A live harness is ready; it sends nothing without a key and has a hard spending cap.

Caveats

  • Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
  • A provider listed on OpenRouter is reached through OpenRouter; its price there may differ from the price on the provider’s own site.
  • The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
  • No latency or throughput figure: the keyless API returned none. A cheaper provider is not shown to be slower or faster.
  • First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.
  • Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.

Sources

  • OpenRouter public API: models and provider endpoints (snapshot)

    Vendor price list ·

    Prices, context, quantization and uptime per provider endpoint as reported by OpenRouter’s public, keyless API on 2026-10-06. Third-party-reported, not measured by Agent. Latency and throughput were not returned.

    Raw data: provider-index/index.json

  • OpenRouter pricing and fees

    Vendor price list ·

    OpenRouter states that inference is billed at the provider list price and that its fee is charged when credits are bought (5.5% on Standard by card, $0.80 minimum; 8% on Business; 5% by crypto). Page fetched 2026-10-06.

    Raw data: provider-index/index.json

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • OpenAI list prices

    Vendor price list ·

    Token prices as listed by the vendor on 2026-10-03.

  • Google Gemini list prices

    Vendor price list ·

    Gemini 3.x Flash prices as listed by the vendor on 2026-09-21. The vendor announced a doubling from 2027-01-01.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “Inference provider index: 27 models, 52 providers”, updated October 6, 2026, https://agent.sasid.ai/benchmarks/inference-provider-index.

Explainers that cite this study

Read the methods and terms in the context of these recorded results.

More comparisons based on this study (90)

These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.

Models and comparisons in this study

More studies

All benchmarks
  • Prompt Caching
  • Cache Reuse

Does a new Claude Code session reuse the prompt cache of an earlier one?

30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.

0of 2 (95% interval 0% to 66%) · Later sessions with at least 50% of turn-1 input cached, A: new folder each time · n = 2

4 chartsUpdated October 7, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.