• LLM pricing
  • Providers
  • Third Party Reported
  • OpenRouter

LLM API pricing comparison, October 2026: Claude vs GPT vs Gemini per million tokens

21 LLM API prices per million tokens, October 2026: Claude, GPT, Gemini and more. Blended prices span 100x. Price per token is not price per task.

TL;DR

  • Prices span 100x. Across 21 models, the blended price runs from $0.20 per million tokens (GPT-6 Luna) to $20.00 (Claude Fable 5.1 and GPT-6 Astra). Blended is a 3:1 input to output mix (a calculation). OpenRouter's public API reported 20 prices on 2026-10-06. The 21st is OpenAI's list price for GPT-6.1 Sol.
  • Standard prices match inside the gateway. All 15 Claude, GPT and Gemini models with multiple providers list matching standard-tier input and output prices through OpenRouter.
  • Price per token is not price per task. Claude Haiku 4.5 lists half the input and output prices of Sonnet 5.5 (calculation). On eight hard tasks it cost $0.0672 per strict pass; Sonnet 5.5 cost $0.01435 (n = 24 each, calculation). Haiku passed 11/24 (95% interval 28% to 65%); Sonnet passed 24/24 (86% to 100%).
  • Most rows have no quality data. We ran our hard-task test on 5 of the 21 rows. For the other 16, including every Gemini, Grok, Mistral and Qwen row, this post gives prices only.
Live story · 42 sSame model, different price: 265 provider endpoints compared

Same model, different price: 265 provider endpoints compared

Reported by OpenRouter’s public API on 2026-10-06: open-weight prices vary up to 12.6x across providers; closed models have one standard price, and the gateway adds a 5.5% credit fee.

Transcript
  1. Provider index · third-party-reported · 2026-10-06. Same model, different price. 265 endpoints, 52 providers, 27 models. Prices as OpenRouter’s public API reports them.
  2. 265 endpoints, 52 providers, 27 models. The keyless API returned latency for 0 of them, so speed is not compared. Provider endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Endpoints with a latency figure: 0 of 265. Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
  3. Open-weight models vary most: DeepSeek V4 Flash 12.6×, DeepSeek V4 Pro 11.2×. All 15 closed models with 2+ providers: one standard price (1×). Chart: Priciest ÷ cheapest standard provider · blended price (n = 2–32 each). Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
  4. One model, 10 providers: Llama 3.3 70B input costs $0.10 per million tokens at DeepInfra (fp8) and $1.04 at Together. Chart: Llama 3.3 70B · USD per million tokens · one row per provider. Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.
  5. Gateway markup: for 10 of 10 models OpenRouter charges the vendor’s per-token price. The cost is a 5.5% fee when you buy credits. Models at the first-party per-token price: 10 of 10. Credit-purchase fee, Standard plan by card: 5.5% ($0.80 minimum by card). Effective markup per token after the fee: +5.5% (calculation). Caveat: First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.
  6. Example: Sonnet 5.5 costs $2.00 in and $10.00 out per million tokens on both. A price is not a measurement: no winner. Chart: Claude Sonnet 5.5 · OpenRouter vs Anthropic list price · USD per million. Caveat: Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.
  7. Prices change often: every endpoint, tier and source date online.

The price table

All prices are USD per million tokens, reported by OpenRouter's public API, snapshot 2026-10-06 (third-party-reported), standard tier. We sorted by blended price, (3 × input + output) ÷ 4, a calculation. The GPT-6.1 Sol row is OpenAI's own list price, verified 2026-10-03. It is not GPT-6 Sol.

ModelInputOutputCache readBlended 3:1
GPT-6 Luna$0.10$0.50$0.01$0.20
Qwen3.8 Flash$0.15$0.47$0.016$0.23
Mistral Large 3 2512$0.50$1.50$0.05$0.75
Gemini 3.5 Flash Lite$0.30$2.50$0.03$0.85
Gemini 3.8 Flash$0.75$3.75$0.075$1.50
Claude Haiku 4.5$1.00$5.00$0.10$2.00
Grok 4.7$2.00$6.00$0.50$3.00
Mistral Medium 3.5$1.50$7.50not reported$3.00
Qwen3.8 Max (0902)$2.00$6.00$0.25$3.00
Gemini 3.5 Flash$1.50$9.00$0.15$3.375
Claude Sonnet 5$2.00$10.00$0.20$4.00
Claude Sonnet 5.5$2.00$10.00$0.20$4.00
GPT-6 Sol$2.00$10.00$0.20$4.00
GPT-6.1 Sol (OpenAI list)$2.00$10.00$0.10$4.00
Gemini 3.1 Pro Preview$2.00$12.00$0.20$4.50
Claude Opus 5.5$4.00$20.00under re-check$8.00
Claude Opus 4.8$5.00$25.00$0.50$10.00
Claude Opus 5$5.00$25.00$0.50$10.00
GPT-5.5$5.00$30.00$0.50$11.25
Claude Fable 5.1$10.00$50.00$0.25$20.00
GPT-6 Astra$10.00$50.00$1.00$20.00
  • A higher version is not always dearer. Opus 5.5 lists $4 / $20, 20% below Opus 5 and Opus 4.8 at $5 / $25 (calculation).

Same price at every provider

The 15 Claude, GPT and Gemini models with multiple providers have matching standard-tier input and output prices through OpenRouter. Each has 2 to 5 standard-tier providers. The spread (dearest ÷ cheapest blended price) is 1.0x (calculation). This does not compare the providers’ direct prices. See all providers. Open-weight models differ by up to 12.6x (study).

Reported
  1. Inputall 5 at $2.00

  2. Outputall 5 at $10.00

  3. Cache readall 5 at $0.20

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $2.00. Output: all at $10.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Reported

Google AI Studio

Google Vertex

Same price on all 3 price types
  • InputGoogle AI Studio: $0.75 same priceGoogle Vertex: $0.75
  • OutputGoogle AI Studio: $3.75 same priceGoogle Vertex: $3.75
  • Cache readGoogle AI Studio: $0.075 same priceGoogle Vertex: $0.075

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $0.75. Output: all at $3.75.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

  • The snapshot shows no input or output price markup on matched rows. For 10 of 10 matched models, OpenRouter lists the recorded vendor input and output prices. The first-party lists predate the snapshot. Standard-plan card credit purchases add 5.5%, with a $0.80 minimum fee (third-party-reported). The minimum can make the effective percentage higher on small purchases. Five Claude, GPT and Gemini snapshot rows have no vendor list price in our source table. Compare: OpenRouter vs Anthropic, OpenAI and Google AI Studio.
  • Tiers and regions have their own prices. Of 130 endpoints for these 20 models, 44 regional ones list 1.1x the standard price and 12 flex ones list 0.5x. 18 fast, priority or ultrafast ones list 1.8x to 6.0x (calculations; we infer the tier from the endpoint tag).

Cache reads, the second price

Among the 15 Claude, GPT and Gemini snapshot rows, Opus 5.5 is under re-check. Of the remaining 14, 13 list cache reads at one tenth of the input price. Fable 5.1 lists one fortieth (calculations). Reported cache-read prices across the table run from $0.01 (GPT-6 Luna) to $1.00 (GPT-6 Astra).

Reported

Azure

OpenAI

Same price on all 3 price types
  • InputAzure: $0.10 same priceOpenAI: $0.10
  • OutputAzure: $0.50 same priceOpenAI: $0.50
  • Cache readAzure: $0.01 same priceOpenAI: $0.01

USD per million tokens, from $0. Equal prices draw one dot.

Third-party reported values. 2 rows, 3 series: Input, Output, Cache read. Input: all at $0.1. Output: all at $0.5.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Four rows list $2 in and $10 out. GPT-6.1 Sol lists a $0.10 cache read and the other three list $0.20, so matching input and output prices do not imply matching total costs. The same tokens, recorded on 33 SWE-bench tasks, cost $52.42 at GPT-6.1 Sol prices and $87.23 at Sonnet 5.5 prices (calculation, not a run). In that record, 94.0% of input tokens were cache reads.

The snapshot has no cache-write prices. The repricing calculation uses the recorded first-party rates: Anthropic one-hour writes cost twice the input price. It prices non-Anthropic writes as plain input. Thus both cache-read prices and the write assumption affect this comparison.

Price per token is not price per task

A price per token hides how many tokens a task uses and how often the model fails.

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

  • Haiku 4.5 vs Sonnet 5.5: Haiku cost $0.0672 per strict pass, and Sonnet cost $0.01435 (calculation, n = 24 each). Haiku passed 11 of 24 (95% interval 28% to 65%); Sonnet passed 24 of 24 (86% to 100%). The intervals do not overlap, so Sonnet is ahead on pass rate here. Five of Haiku’s strict failures had correct answers in the wrong format; eight had wrong answers. Lenient grading gives Haiku 16/24 (95% interval 47% to 82%). A failed call still enters the cost calculation. Haiku reported a median 5,064 output tokens per call (range 1,899 to 9,321). Sonnet reported 1,050 (range 176 to 3,895; n = 24 each). These are observed ranges, not confidence intervals.
  • Same pass rate, different price: Sonnet 5.5, Opus 5.5 and Fable 5.1 each passed 24 of 24 (95% interval 86% to 100%), so the set hits a ceiling. Cost per pass still differs: Sonnet $0.01435, Fable $0.09331. Fable lists 5x Sonnet's price and cost 6.5x per pass (calculation).
Calculation
Largest value is 310x the smallest; Log shows the small bars.
Claude Fable 5.1
Claude Opus 5
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol
Claude Haiku 4.5
Gemini 3.x Flash
Jev 1.13 (router)

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 8 rows. Highest Claude Fable 5.1 $12.85. Lowest Jev 1.13 (router) $0.042.

Notes

Cost per resolved SWE-bench instance if 162.9M input and 1.8M output tokens had been billed at each model's list price

Calculation, not a run: tokens recorded by Agent on claude-sonnet-5-5 (33 attempts, 25 resolved) times list prices effective 2026-09-21. Another model would use a different number of tokens and resolve a different set. Jev is a routing model and cannot do this work; its bar is a price floor only.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models), Google Gemini list prices, OpenAI list prices, Jev 1.13 list price

The repricing chart points the other way. The same recorded tokens put Haiku at half of Sonnet per resolved task ($1.745 against $3.489; calculation). That calculation assumes every model uses the same tokens and resolves the same 25 of 33 tasks. The hard-task run shows that assumption can fail. Read the chart as price sensitivity only.

How to compare on your own work

  1. Record one task's tokens: plain input, cache reads, cache writes and output.
  2. Price each kind at your tier's list price on the day you decide.
  3. Run the same checks on every model. Count each failed call.
  4. Divide total cost by passes.

The AI cost calculator prices your tokens at each model's list price, with and without prompt caching. It is arithmetic, not a benchmark.

How we measured

  • Prices: one snapshot of OpenRouter's public, keyless API on 2026-10-06: 265 endpoints, 52 providers, 27 models. The table shows 20 of them, plus the GPT-6.1 Sol list price. The other 7 are open-weight (open-model post).
  • Calculations: the 100x range is the highest blended price ÷ the lowest across the 21 rows. Cost per strict pass is list price × reported tokens for every call, failures included, ÷ strict passes (hard tasks, five tasks). Repricing uses the tokens recorded on 33 SWE-bench tasks (study). The calls ran through Claude Code and Codex CLI on subscriptions. Their token counts include each CLI’s own context; these are route-and-model results, not bare API cost measurements. None of this is a bill. This post made no model calls. The price snapshot extract, hard-task receipts and SWE-bench token records are public.

Caveats

  • Snapshot of 2026-10-06. Prices change often. Our Google price source records a notice that Gemini 3.x Flash prices double from 2027-01-01 (recorded 2026-09-21). We did not recheck that notice for this post.
  • Third-party-reported. OpenRouter's price can differ from a vendor's own site. First-party lists date from 2026-09-21 (Anthropic, Google) and 2026-10-03 (OpenAI).
  • One mix, one set. Your own token mix can change the order. The hard set has eight tasks. Haiku ran at default effort and reported a median 4,556 reasoning tokens per call. The observed range was 1,452 to 8,569 (n = 24; not a confidence interval). These tokens are part of its reported output, not an extra token category to bill.
  • Opus 5.5 cache-read price. The stored snapshot and calculation table use $0.20 per million tokens; that rate is under re-check. The Opus 5.5 bars in both cost charts use this provisional rate. Treat those cost bars as provisional too.
  • Quality data covers 5 rows. Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol ran the eight hard tasks. A cheaper row is not shown to be better or worse. GPT-6 Luna ran only in a latency study, with 3 to 5 calls per cell: too few to rank quality.

Price the task, not the token

Try Agent. It records the model, tokens and cost of each call.

The data behind this post

  • Inference
  • Providers

Inference provider index: 27 models, 52 providers

Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.

265endpoints, 52 providers, 27 models · Provider endpoints in the snapshot · n = 27

34 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.