Explainer · Price per million tokens
LLM pricing per million tokens, explained with real token counts
Definition
Price per million tokens is an LLM vendor’s dollar rate per 1,000,000 tokens. Input, output and cached tokens have separate rates. Multiply each token count by its rate, divide by 1,000,000, then add the results. Our stored Sonnet 5.5 rates per million tokens are $2 for input, $10 for output and $0.20 for cache reads.
Agent team · · 5 min read · Every number is from the public studies
Parts sorted by value, largest first
- Cache writes (1 h)
- Cache reads
- Output
- Uncached input
Shares are calculated from the values shown.
| Item | Cost |
|---|---|
| Cache writes (1 h) | $38.99 |
| Cache reads | $30.62 |
| Output | $17.61 |
| Uncached input | $0.01 |
List-price calculation, not a run. 4 rows. Highest Cache writes (1 h) $38.99. Lowest Uncached input $0.01.
Notes
Recorded tokens at Sonnet 5.5 list price, by token kind
List-price calculation on recorded tokens: 153.1M cache reads, 9.7M cache writes, 1.8M output, 3.2k uncached input. An agent loop re-reads its context on every call, so cache reads dominate the token count; by price the largest part is cache writes (1 h).
Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)
Input vs output token price: the four meters
Our run used four token meters. Stored Claude Sonnet 5.5 list prices per million (effective 2026-09-21):
- Uncached input, $2. The vendor bills these at the uncached rate.
- Cache write, $4. The vendor stores tokens for later calls. A one-hour write costs twice input (calculation: 2 × $2).
- Cache read, $0.20. The vendor serves stored tokens at 10% of input (calculation).
- Output, $10. The model writes these tokens.
Output costs 5 times input for all 7 repriced text models (calculation: output price ÷ input price). Our runs do not show why vendors chose those prices.
A worked bill from recorded tokens
Our repricing study records Agent’s token ledger for 33 SWE-bench Verified attempts. It resolved 25 of 33 (76%, 95% Wilson interval 59% to 87%). It contains 162.9M input tokens, 94.0% of them cache reads, and 1.76M output tokens. At Sonnet 5.5 list prices that costs $87.23. Calculation: subscription calls, with no invoice.
The ledger excludes compaction calls. The platform estimate includes them: $92.64.
- Cache writes: $38.99 for 9.7M tokens, 45% of the bill.
- Cache reads: $30.62 for 153.1M tokens, 35%.
- Output: $17.61 for 1.76M tokens, 20%.
- Uncached input: $0.01 for 3.2k tokens.
Costs and shares are calculations across all 33 attempts, failures included. Output was 1.1% of all tokens but 20% of the calculated bill. Cache reads were 93% of all tokens but 35% of the bill. The $2 input price applied to only 3.2k tokens.
Across all 164.6M tokens, the bill is $0.53 per million tokens (calculation: $87.23 ÷ 164.6 million tokens × 1 million). Cheap cache reads put it well below the $2 input price. Per resolved instance it was $3.49 (calculation: $87.23 ÷ 25).
Without caching, these tokens would cost $343.33, 3.9 times the bill (calculation: all 162.9M input tokens at $2 per million, plus output). See prompt caching, explained.
The same tokens at other list prices
Same tokens, other models’ list prices: a calculation, not a run. Other models would use different tokens and resolve different instances. This shows price sensitivity, not rankings.
| Priced as | Input / output, $ per million | All 33 attempts |
|---|---|---|
| Claude Haiku 4.5 | $1 / $5 | $43.61 |
| GPT-6.1 Sol | $2 / $10 | $52.42 |
| Claude Sonnet 5.5 (ran the work) | $2 / $10 | $87.23 |
| Claude Opus 5.5 | $4 / $20 | $143.83 |
| Claude Fable 5.1 | $10 / $50 | $321.31 |
Both Sol and Sonnet list $2 and $10, yet these tokens cost $52.42 at Sol prices versus $87.23 at Sonnet prices. Sol cache reads cost $0.10, half Sonnet’s $0.20. Other vendors’ cache writes use plain input prices. Compare cache prices before headline prices.
Opus 5.5 uses the stored $0.20 cache-read rate per million. These dated assumptions are not live quotes. The Jev router bar is a price floor; it cannot do this coding work.
Blended prices, tiers and gateway fees
A blended price folds the meters into one number. Our provider index mixes 3 input tokens with 1 output token: $4.00 per million for Sonnet 5.5 (calculation: (3 × $2 + $10) ÷ 4).
Tier and region. Flex, priority, fast and regional endpoints have separate prices. Our Sonnet 5.5 snapshot has 3 regional endpoints on Google Vertex and Azure: $2.20 input and $11 output, 10% above standard (calculation: ($2.20 ÷ $2 − 1) × 100%).
Provider. All 5 standard-tier Sonnet 5.5 providers, including Anthropic, reported $2 input, $10 output and $0.20 cache reads. Across 15 closed models with two or more providers, reported standard-tier input and output prices matched. One snapshot, not measured bills. Open-weight prices differ: DeepSeek V4 Flash 0423’s highest price among 15 standard-tier providers is 12.6 times its lowest (calculation on blended prices). Low prices can mean lower precision or a shorter context.
Gateway fee. OpenRouter matched vendor per-token prices for 10 of 10 models with first-party prices. Standard credit purchases carry a 5.5% fee, minimum $0.80 by card. Buying $87.23 in credits would add $4.80 (calculation: 5.5% × $87.23). Assumption: one card purchase, not a per-call fee. See inference gateways.
OpenRouter’s public API reported endpoint prices (third-party snapshot, 2026-10-06). Its pricing and FAQ snapshot supplies the credit fee. Endpoint tags supply tier labels; this can misclassify endpoints. Prices change often.
How to estimate your own bill
- Count tokens by meter: uncached input, cache writes, cache reads and output. A "tokens per task" guess hides this split.
- Multiply each count by its rate per million. Divide each result by 1,000,000, then add them.
- Divide by resolved tasks. Failed attempts cost tokens too.
- Check the price at your provider and tier.
cost USD = (uncached input tokens × input rate
+ cache-read tokens × cache-read rate
+ cache-write tokens × write rate
+ output tokens × output rate) / 1,000,000
Each rate is USD per million tokens.Our mix pools 33 attempts, two campaigns, one model and one task set; yours will differ. Try the AI cost calculator, or read how to estimate your AI coding bill.
Frequently asked questions
Why are output tokens more expensive than input tokens?
Output costs 5 times input for all 7 repriced text models (calculation). Sonnet 5.5 lists $10 output and $2 input per million. Our runs do not show why vendors chose those prices. Output was 1.1% of the tokens and 20% of the bill (calculation, n = 33 attempts).
What is a blended price?
It mixes input and output prices at a fixed ratio. Our provider index uses 3 input tokens to 1 output token: $4.00 per million for Claude Sonnet 5.5 (calculation). It does not predict bills: our 33-attempt ledger, with 94.0% cached input, cost $0.53 per million tokens (calculation).
How much does caching cut a token bill?
Our input was 94.0% cache reads. Sonnet 5.5 cost $87.23 and would be $343.33 without caching (calculation, n = 33 attempts). In 3 five-turn Claude Code sessions, caching cut calculated Sonnet 5.5 cost 50%: $0.1350 against $0.2698. Calculation over 15 recorded turns, not invoices or predictions for other sessions. Turn 1 cost more with caching: one-hour writes cost twice input.
Is the price per token the same at every provider?
For 15 closed models with two or more providers, reported standard-tier input and output prices matched. Open-weight prices differed: the spread was 12.6 times for DeepSeek V4 Flash 0423 across 15 standard-tier providers (calculation on blended prices). OpenRouter snapshot: 2026-10-06. Prices change often.
Watch the data
What if every call ran on Opus? Repricing real agent tokens
A calculation, not a run: Agent's recorded SWE-bench tokens cost $87.23 at Sonnet 5.5 prices, $143.83 at Opus 5.5 and $343.33 without caching.
Transcript
- Thought experiment · recorded tokens × list prices. What if every call ran on Opus? Agent recorded every token on 33 SWE-bench attempts. We repriced them.
- 162.9M input tokens, 94.0% of them read from the prompt cache. Input tokens recorded: 162.9M (n = 33). Output tokens recorded: 1.8M (n = 33). Input served from cache: 94.0% (n = 33). Caveat: Recorded costs are list-price estimates for subscription calls; no invoice backs them.
- Same tokens on Opus 5.5: $143.83 instead of $87.23, 1.65× the bill. On Haiku 4.5: $43.61. At Haiku 4.5 prices: $43.61 (n = 33). At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
- Per resolved instance: $3.49 on Sonnet 5.5, $5.75 on Opus 5.5, $12.85 on Fable 5.1. Chart: Thought experiment: the same tokens at other list prices. Calculation, not a run. Caveat: Different models use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.
- In this calculation caching matters more than the model: without it, Sonnet would cost $343.33, 3.9× the recorded $87.23. Chart: Thought experiment: what prompt caching saved. Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
- Price sensitivity, not predictions. Every repricing labelled as a calculation.