Explainer · Price per million tokens

LLM pricing per million tokens, explained with real token counts

Definition

Price per million tokens is an LLM vendor’s dollar rate per 1,000,000 tokens. Input, output and cached tokens have separate rates. Multiply each token count by its rate, divide by 1,000,000, then add the results. Our stored Sonnet 5.5 rates per million tokens are $2 for input, $10 for output and $0.20 for cache reads.

Agent team · · 5 min read · Every number is from the public studies

Calculation
$87.23Sum of the 4 parts

Parts sorted by value, largest first

  1. Cache writes (1 h)
  2. Cache reads
  3. Output
  4. Uncached input

Shares are calculated from the values shown.

List-price calculation, not a run. 4 rows. Highest Cache writes (1 h) $38.99. Lowest Uncached input $0.01.

Notes

Recorded tokens at Sonnet 5.5 list price, by token kind

List-price calculation on recorded tokens: 153.1M cache reads, 9.7M cache writes, 1.8M output, 3.2k uncached input. An agent loop re-reads its context on every call, so cache reads dominate the token count; by price the largest part is cache writes (1 h).

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)

Input vs output token price: the four meters

Our run used four token meters. Stored Claude Sonnet 5.5 list prices per million (effective 2026-09-21):

  • Uncached input, $2. The vendor bills these at the uncached rate.
  • Cache write, $4. The vendor stores tokens for later calls. A one-hour write costs twice input (calculation: 2 × $2).
  • Cache read, $0.20. The vendor serves stored tokens at 10% of input (calculation).
  • Output, $10. The model writes these tokens.

Output costs 5 times input for all 7 repriced text models (calculation: output price ÷ input price). Our runs do not show why vendors chose those prices.

A worked bill from recorded tokens

Our repricing study records Agent’s token ledger for 33 SWE-bench Verified attempts. It resolved 25 of 33 (76%, 95% Wilson interval 59% to 87%). It contains 162.9M input tokens, 94.0% of them cache reads, and 1.76M output tokens. At Sonnet 5.5 list prices that costs $87.23. Calculation: subscription calls, with no invoice.

The ledger excludes compaction calls. The platform estimate includes them: $92.64.

Calculation
$87.23Sum of the 4 parts

Parts sorted by value, largest first

  1. Cache writes (1 h)
  2. Cache reads
  3. Output
  4. Uncached input

Shares are calculated from the values shown.

List-price calculation, not a run. 4 rows. Highest Cache writes (1 h) $38.99. Lowest Uncached input $0.01.

Notes

Recorded tokens at Sonnet 5.5 list price, by token kind

List-price calculation on recorded tokens: 153.1M cache reads, 9.7M cache writes, 1.8M output, 3.2k uncached input. An agent loop re-reads its context on every call, so cache reads dominate the token count; by price the largest part is cache writes (1 h).

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)

  • Cache writes: $38.99 for 9.7M tokens, 45% of the bill.
  • Cache reads: $30.62 for 153.1M tokens, 35%.
  • Output: $17.61 for 1.76M tokens, 20%.
  • Uncached input: $0.01 for 3.2k tokens.

Costs and shares are calculations across all 33 attempts, failures included. Output was 1.1% of all tokens but 20% of the calculated bill. Cache reads were 93% of all tokens but 35% of the bill. The $2 input price applied to only 3.2k tokens.

Across all 164.6M tokens, the bill is $0.53 per million tokens (calculation: $87.23 ÷ 164.6 million tokens × 1 million). Cheap cache reads put it well below the $2 input price. Per resolved instance it was $3.49 (calculation: $87.23 ÷ 25).

Without caching, these tokens would cost $343.33, 3.9 times the bill (calculation: all 162.9M input tokens at $2 per million, plus output). See prompt caching, explained.

The same tokens at other list prices

Same tokens, other models’ list prices: a calculation, not a run. Other models would use different tokens and resolve different instances. This shows price sensitivity, not rankings.

Priced asInput / output, $ per millionAll 33 attempts
Claude Haiku 4.5$1 / $5$43.61
GPT-6.1 Sol$2 / $10$52.42
Claude Sonnet 5.5 (ran the work)$2 / $10$87.23
Claude Opus 5.5$4 / $20$143.83
Claude Fable 5.1$10 / $50$321.31
Calculation
Largest value is 310x the smallest; Log shows the small bars.
Claude Fable 5.1
Claude Opus 5
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol
Claude Haiku 4.5
Gemini 3.x Flash
Jev 1.13 (router)

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 8 rows. Highest Claude Fable 5.1 $12.85. Lowest Jev 1.13 (router) $0.042.

Notes

Cost per resolved SWE-bench instance if 162.9M input and 1.8M output tokens had been billed at each model's list price

Calculation, not a run: tokens recorded by Agent on claude-sonnet-5-5 (33 attempts, 25 resolved) times list prices effective 2026-09-21. Another model would use a different number of tokens and resolve a different set. Jev is a routing model and cannot do this work; its bar is a price floor only.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models), Google Gemini list prices, OpenAI list prices, Jev 1.13 list price

Both Sol and Sonnet list $2 and $10, yet these tokens cost $52.42 at Sol prices versus $87.23 at Sonnet prices. Sol cache reads cost $0.10, half Sonnet’s $0.20. Other vendors’ cache writes use plain input prices. Compare cache prices before headline prices.

Opus 5.5 uses the stored $0.20 cache-read rate per million. These dated assumptions are not live quotes. The Jev router bar is a price floor; it cannot do this coding work.

Blended prices, tiers and gateway fees

A blended price folds the meters into one number. Our provider index mixes 3 input tokens with 1 output token: $4.00 per million for Sonnet 5.5 (calculation: (3 × $2 + $10) ÷ 4).

Tier and region. Flex, priority, fast and regional endpoints have separate prices. Our Sonnet 5.5 snapshot has 3 regional endpoints on Google Vertex and Azure: $2.20 input and $11 output, 10% above standard (calculation: ($2.20 ÷ $2 − 1) × 100%).

Provider. All 5 standard-tier Sonnet 5.5 providers, including Anthropic, reported $2 input, $10 output and $0.20 cache reads. Across 15 closed models with two or more providers, reported standard-tier input and output prices matched. One snapshot, not measured bills. Open-weight prices differ: DeepSeek V4 Flash 0423’s highest price among 15 standard-tier providers is 12.6 times its lowest (calculation on blended prices). Low prices can mean lower precision or a shorter context.

Reported
  1. Inputall 5 at $2.00

  2. Outputall 5 at $10.00

  3. Cache readall 5 at $0.20

  • one provider (reported price, USD per million tokens, log scale per strip)

Third-party reported values. 5 rows, 3 series: Input, Output, Cache read. Input: all at $2.00. Output: all at $10.00.

Notes

Standard tier, one bar per provider (its cheapest standard endpoint); reported by OpenRouter’s public API, snapshot 2026-10-06

Prices reported by OpenRouter’s public API, snapshot 2026-10-06; third-party-reported, not measured by Agent. Sorted from the lowest to the highest blended price (3 input : 1 output). A parenthesis names the quantization the provider reported; lower precision or a shorter context can explain a lower price, so check the endpoint table. Flex, priority, fast and regional endpoints are left out here because they are priced differently on purpose.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Gateway fee. OpenRouter matched vendor per-token prices for 10 of 10 models with first-party prices. Standard credit purchases carry a 5.5% fee, minimum $0.80 by card. Buying $87.23 in credits would add $4.80 (calculation: 5.5% × $87.23). Assumption: one card purchase, not a per-call fee. See inference gateways.

OpenRouter’s public API reported endpoint prices (third-party snapshot, 2026-10-06). Its pricing and FAQ snapshot supplies the credit fee. Endpoint tags supply tier labels; this can misclassify endpoints. Prices change often.

How to estimate your own bill

  1. Count tokens by meter: uncached input, cache writes, cache reads and output. A "tokens per task" guess hides this split.
  2. Multiply each count by its rate per million. Divide each result by 1,000,000, then add them.
  3. Divide by resolved tasks. Failed attempts cost tokens too.
  4. Check the price at your provider and tier.
cost USD = (uncached input tokens × input rate
          + cache-read tokens × cache-read rate
          + cache-write tokens × write rate
          + output tokens × output rate) / 1,000,000

Each rate is USD per million tokens.

Our mix pools 33 attempts, two campaigns, one model and one task set; yours will differ. Try the AI cost calculator, or read how to estimate your AI coding bill.

Frequently asked questions

Why are output tokens more expensive than input tokens?

Output costs 5 times input for all 7 repriced text models (calculation). Sonnet 5.5 lists $10 output and $2 input per million. Our runs do not show why vendors chose those prices. Output was 1.1% of the tokens and 20% of the bill (calculation, n = 33 attempts).

What is a blended price?

It mixes input and output prices at a fixed ratio. Our provider index uses 3 input tokens to 1 output token: $4.00 per million for Claude Sonnet 5.5 (calculation). It does not predict bills: our 33-attempt ledger, with 94.0% cached input, cost $0.53 per million tokens (calculation).

How much does caching cut a token bill?

Our input was 94.0% cache reads. Sonnet 5.5 cost $87.23 and would be $343.33 without caching (calculation, n = 33 attempts). In 3 five-turn Claude Code sessions, caching cut calculated Sonnet 5.5 cost 50%: $0.1350 against $0.2698. Calculation over 15 recorded turns, not invoices or predictions for other sessions. Turn 1 cost more with caching: one-hour writes cost twice input.

Is the price per token the same at every provider?

For 15 closed models with two or more providers, reported standard-tier input and output prices matched. Open-weight prices differed: the spread was 12.6 times for DeepSeek V4 Flash 0423 across 15 standard-tier providers (calculation on blended prices). OpenRouter snapshot: 2026-10-06. Prices change often.

Watch the data

Live story · 33 sWhat if every call ran on Opus? Repricing real agent tokens

What if every call ran on Opus? Repricing real agent tokens

A calculation, not a run: Agent's recorded SWE-bench tokens cost $87.23 at Sonnet 5.5 prices, $143.83 at Opus 5.5 and $343.33 without caching.

Transcript
  1. Thought experiment · recorded tokens × list prices. What if every call ran on Opus? Agent recorded every token on 33 SWE-bench attempts. We repriced them.
  2. 162.9M input tokens, 94.0% of them read from the prompt cache. Input tokens recorded: 162.9M (n = 33). Output tokens recorded: 1.8M (n = 33). Input served from cache: 94.0% (n = 33). Caveat: Recorded costs are list-price estimates for subscription calls; no invoice backs them.
  3. Same tokens on Opus 5.5: $143.83 instead of $87.23, 1.65× the bill. On Haiku 4.5: $43.61. At Haiku 4.5 prices: $43.61 (n = 33). At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  4. Per resolved instance: $3.49 on Sonnet 5.5, $5.75 on Opus 5.5, $12.85 on Fable 5.1. Chart: Thought experiment: the same tokens at other list prices. Calculation, not a run. Caveat: Different models use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.
  5. In this calculation caching matters more than the model: without it, Sonnet would cost $343.33, 3.9× the recorded $87.23. Chart: Thought experiment: what prompt caching saved. Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  6. Price sensitivity, not predictions. Every repricing labelled as a calculation.

The data behind this explainer

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

  • Inference
  • Providers

Inference provider index: 27 models, 52 providers

Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.

265endpoints, 52 providers, 27 models · Provider endpoints in the snapshot · n = 27

34 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.