How to estimate your AI coding bill from real token mixes
Estimate AI coding costs from recorded token mixes: an agent task at $2.64, a hard call at $0.014, a routing decision at $0.005. Calculations, limits stated.
TL;DR
- An AI coding bill is token mix × list price × volume ÷ success rate. Most estimates go wrong on the token mix, not the price.
- One recorded agent task (our SWE-bench run) used about 4.6M cache-read tokens, 0.29M cache-write tokens and 53k output tokens. At Claude Sonnet 5.5 list prices that is $2.64 per attempt and $3.49 per resolved task.
- In that mix, cache writes were 45% of the cost, cache reads 35% and output 20%. Plain input was almost nothing.
- Without prompt caching, the same work would cost about 3.9 times as much at Sonnet prices.
- A single hard one-turn call in Claude Code cost $0.0143 on Sonnet and $0.0308 on Haiku 4.5 (our calculation). A cheaper price per token can mean a more expensive call.
- Every figure here is a calculation on recorded tokens. The calls ran on subscriptions; none of this is an invoice.
Try the numbers yourself: /tools/ai-cost-calculator.
Why estimates go wrong
Most people estimate an AI bill like this: "a task is about 10,000 tokens, Sonnet costs $2 per million, so a task costs two cents." Real agent work breaks that estimate three ways:
- Agents re-read their context. Each turn sends the whole conversation again. A task with 50 turns sends the early context 50 times.
- Tokens have four prices, not one. Plain input, cache reads, cache writes and output each have their own price. Cache reads can cost a tenth of plain input. Output can cost five times plain input.
- Failures cost money too. A failed attempt uses tokens and produces nothing you can ship.
So a good estimate starts from a recorded token mix, not a guess. Below we apply the calculator's token-price formula to five recorded mixes. The calculator has four presets. It does not include every example or divide by a success rate for you.
Match this guide to the calculator
Choose vendor list price to use the study's dated price snapshot. Its prices can differ from today's vendor prices. The presets and controls are:
| Calculator preset | This guide | What it does |
|---|---|---|
| SWE-bench-style coding task | Mix A | Uses a rounded mix derived from all 33 attempts. At the snapshot Sonnet price, one attempt is $2.6434. |
| Short task through a coding CLI | Mix B | Uses 685 uncached input, 1,401 cache reads and 107 output tokens. Cache writes were not split out. |
| Short Q&A call (API) | Additional API example | Uses 279 input and 243 output tokens from six calls. It is not the Codex CLI mix. |
| Routing decision | Mix D | Uses 1,785 input and 107 output tokens. Cache reads were not split out, so this conservative formula gives $4.64 per 1,000 at the Sonnet snapshot price. The study's $5.00 is its recorded cost estimate, not this preset's result. |
Mix C and Mix E have no preset. Use Custom token mix when you have all four counts; output tokens alone cannot reproduce a full call cost. Set Tasks per month for volume and choose the model to compare prices. The page also shows a no-cache estimate. Divide by your success rate separately. The page does not have a success-rate control.
Open Mix A for 200 attempts at the Sonnet snapshot price.
Step 1: pick the unit of work
Decide what you pay for:
- An agent task: one issue taken from start to a delivered fix.
- A single call: one prompt, one answer, for example a review comment or a code snippet.
- A decision: one tiny call that picks a route, a label or a next step.
These differ by three orders of magnitude, so do not mix them.
Step 2: get the token mix
Split the tokens into four buckets: uncached input, cache reads, cache writes and output. Here are the recorded mixes we use.
Mix A: one agent task (SWE-bench Verified)
Our agent attempted 33 SWE-bench Verified instances. Across all 33 it recorded 153.1M cache reads, 9.7M one-hour cache writes, 1.8M output and 3.2k uncached input tokens. Per attempt (our division by 33), that is about:
- 4.64M cache reads (our calculation: 153.1M / 33)
- 0.29M cache writes
- 53k output
- about 100 uncached input
94.0% of the input came from the cache. That is normal for an agent loop.
Mix B: one short call in Claude Code
From our five-task head-to-head, Sonnet 5.5 in Claude Code used a mean 1,401 cache-read tokens and 685 other input tokens per call, and a median 107 output tokens. Most of that input is the CLI's own context, not the prompt.
Mix C: one hard call in Claude Code
From the hard head-to-head, a median call wrote 1,050 output tokens on Sonnet 5.5 and 5,064 on Haiku 4.5, of which 4,556 were reported reasoning tokens.
Mix D: one routing decision
From the router study: a decision through the Claude Code CLI includes its tool schema and thinking tokens. We report it directly as a cost per 1,000 decisions (below).
Mix E: one call through Codex CLI
For a one-line request, the Codex CLI sent 19,551 input tokens. The OpenAI API sent 17. Use this mix to price the CLI's own context. We cover it in Claude Code vs Codex CLI: the hidden context tax.
Step 3: apply list prices per bucket
The formula:
cost = uncached input × input price + cache reads × cache-read price + cache writes × cache-write price + output × output price
List prices per million tokens, as used in our studies:
| Model | Input | Cache read | Output |
|---|---|---|---|
| Claude Haiku 4.5 | $1 | $0.10 | $5 |
| Claude Sonnet 5.5 | $2 | $0.20 | $10 |
| Claude Opus 5.5 | $4 | $0.20 | $20 |
| Claude Fable 5.1 | $10 | $0.25 | $50 |
| GPT-6.1 Sol | $2 | $0.10 | $10 |
Anthropic one-hour cache writes cost twice the input price. For OpenAI, we price writes as plain input.
Mix A at Sonnet prices: 4.64M × $0.20 + 0.29M × $4 + 53k × $10 ≈ $0.93 + $1.18 + $0.53 = about $2.64. The published per-attempt figure, from the unrounded totals, is $2.643.
Over all 33 attempts the dollars split like this: cache writes $38.99, cache reads $30.62, output $17.61, uncached input $0.01. The biggest bucket by token count, cache reads, is not the biggest by price. Cache writes are.
The same mix at other list prices (per attempt, a calculation): Haiku 4.5 $1.32, GPT-6.1 Sol $1.59, Sonnet 5.5 $2.64, Opus 5.5 $4.36, Fable 5.1 $9.74. A different model would use different tokens, so read these as price sensitivity, not as a forecast for that model.
Step 4: divide by the success rate
You pay for every attempt, but you only keep the ones that succeed. So:
cost per success = cost per attempt ÷ success rate
Our agent resolved 25 of 33 attempts. At Sonnet prices, $2.643 per attempt becomes $3.49 per resolved task.
This step can reverse a ranking. On the hard head-to-head:
Haiku 4.5 has half of Sonnet's list price. But at the CLI default it wrote about five times the output tokens, and it passed 11 of 24 calls. Its cost per call was about $0.0308 (our calculation: $0.0672 per pass × 11 passes ÷ 24 calls). Sonnet's was $0.0143, and it passed all 24. Per strict pass: $0.0143 on Sonnet against $0.0672 on Haiku.
Step 5: multiply by volume
Now scale. Some examples at list price (our calculations):
- 100 agent tasks a month on Mix A: about $264 on Sonnet 5.5 and $436 on Opus 5.5. At our success rate, about 76 of them would be resolved.
- 10,000 hard one-turn calls on Sonnet in Claude Code: about $143.
- 1,000 routing decisions: $5.00 on Sonnet 5.5, $8.92 on Haiku 4.5 (it thinks longer under the CLI default) and $0.0337 on Jev 1.13, a dedicated router.
Step 6: add the overhead you do not see
Four costs often stay out of the first estimate:
- CLI context. For the 19,551-token Codex CLI one-liner, the context alone costs $0.002 to $0.039 per call at GPT-6.1 Sol prices, depending on the cache.
- Pipeline steps. Our agent's notional figure was $2.81 per attempt, not $2.64, because it also counts context-compaction calls. Planning, verification and review are real work and use real tokens.
- Gateway fees and provider choice. A gateway can add a fee on top of list price. On 2026-10-06, OpenRouter listed the vendor's own per-token price for 10 of 10 models we checked and charged 5.5% when credits are bought by card (third-party-reported): OpenRouter vs going direct. For open-weight models, the provider matters more: prices for the same model differed by up to 12.6x (cheapest open-model providers).
- Retries and caps. In our coding calibration, the latest build cost $11.06 for three real pull requests, and one earlier capped run stopped at its cap before delivery. A cap protects the budget, but a stopped attempt still costs.
Step 7: check the caching assumption
At Sonnet prices, caching cut the bill 3.9x, more than the Sonnet-to-Opus gap (1.6x). Both are our calculations: Mix A over 33 attempts costs $87.23 with caching and $343.33 without at Sonnet prices ($343.33 / $87.23 ≈ 3.9), and $143.83 with caching at Opus prices ($143.83 / $87.23 ≈ 1.6). Without caching, Sonnet would cost more than Opus with caching ($143.83).
Before you trust any estimate, confirm that your tool actually hits the cache. Agent loops depend on it.
A worked estimate
Say a team plans 200 agent tasks a month, like Mix A, on Sonnet 5.5, with a 75% success rate. Our calculation:
- Cost per attempt: $2.64.
- Monthly attempts: 200 × $2.643 = $529.
- Cost per success: $2.64 ÷ 0.75 = $3.52.
- With caching lost: about 3.9 × $529 ≈ $2,080.
The gap between line 2 and line 4 is the main risk in the budget. The model choice comes second.
How the numbers were made
- Token mixes are sums or medians of recorded tokens from our studies. Per-attempt and per-call figures are our arithmetic on those values.
- Prices are list prices from our price table, effective 2026-09-21 (OpenAI 2026-10-03).
- All costs are calculations. The underlying calls ran on subscriptions.
Caveats
- Your mix is not our mix. Repository size, task length and tool use change the token counts a lot. Record your own mix when you can.
- Another model uses other tokens. Repricing one model's tokens at another's price shows sensitivity, not that model's real cost.
- Prices change. Check the cache-read and cache-write prices first; they move agent bills the most.
- Small samples behind the short-call and hard-call mixes (15 to 24 calls each).
What to read next
- How much does prompt caching actually save?
- What if every call ran on Opus?
- What one resolved SWE-bench task really costs
- Claude Sonnet vs Opus: when is Opus worth the price?
- OpenRouter vs going direct: what the gateway costs
Get your real token mix
Agent shows recorded usage and cost for task work. A provider may omit token buckets or a reliable cost; treat those values as unknown. Try Agent, inspect the receipts, and use a complete recorded mix when you have one.