• LLM pricing
  • Cost Calculator
  • Prompt Caching
  • Tokens

How to estimate your AI coding bill from real token mixes

Estimate AI coding costs from recorded token mixes: an agent task at $2.64, a hard call at $0.014, a routing decision at $0.005. Calculations, limits stated.

TL;DR

  • An AI coding bill is token mix × list price × volume ÷ success rate. Most estimates go wrong on the token mix, not the price.
  • One recorded agent task (our SWE-bench run) used about 4.6M cache-read tokens, 0.29M cache-write tokens and 53k output tokens. At Claude Sonnet 5.5 list prices that is $2.64 per attempt and $3.49 per resolved task.
  • In that mix, cache writes were 45% of the cost, cache reads 35% and output 20%. Plain input was almost nothing.
  • Without prompt caching, the same work would cost about 3.9 times as much at Sonnet prices.
  • A single hard one-turn call in Claude Code cost $0.0143 on Sonnet and $0.0308 on Haiku 4.5 (our calculation). A cheaper price per token can mean a more expensive call.
  • Every figure here is a calculation on recorded tokens. The calls ran on subscriptions; none of this is an invoice.

Try the numbers yourself: /tools/ai-cost-calculator.

Why estimates go wrong

Most people estimate an AI bill like this: "a task is about 10,000 tokens, Sonnet costs $2 per million, so a task costs two cents." Real agent work breaks that estimate three ways:

  1. Agents re-read their context. Each turn sends the whole conversation again. A task with 50 turns sends the early context 50 times.
  2. Tokens have four prices, not one. Plain input, cache reads, cache writes and output each have their own price. Cache reads can cost a tenth of plain input. Output can cost five times plain input.
  3. Failures cost money too. A failed attempt uses tokens and produces nothing you can ship.

So a good estimate starts from a recorded token mix, not a guess. Below we apply the calculator's token-price formula to five recorded mixes. The calculator has four presets. It does not include every example or divide by a success rate for you.

Match this guide to the calculator

Choose vendor list price to use the study's dated price snapshot. Its prices can differ from today's vendor prices. The presets and controls are:

Calculator presetThis guideWhat it does
SWE-bench-style coding taskMix AUses a rounded mix derived from all 33 attempts. At the snapshot Sonnet price, one attempt is $2.6434.
Short task through a coding CLIMix BUses 685 uncached input, 1,401 cache reads and 107 output tokens. Cache writes were not split out.
Short Q&A call (API)Additional API exampleUses 279 input and 243 output tokens from six calls. It is not the Codex CLI mix.
Routing decisionMix DUses 1,785 input and 107 output tokens. Cache reads were not split out, so this conservative formula gives $4.64 per 1,000 at the Sonnet snapshot price. The study's $5.00 is its recorded cost estimate, not this preset's result.

Mix C and Mix E have no preset. Use Custom token mix when you have all four counts; output tokens alone cannot reproduce a full call cost. Set Tasks per month for volume and choose the model to compare prices. The page also shows a no-cache estimate. Divide by your success rate separately. The page does not have a success-rate control.

Open Mix A for 200 attempts at the Sonnet snapshot price.

Step 1: pick the unit of work

Decide what you pay for:

  • An agent task: one issue taken from start to a delivered fix.
  • A single call: one prompt, one answer, for example a review comment or a code snippet.
  • A decision: one tiny call that picks a route, a label or a next step.

These differ by three orders of magnitude, so do not mix them.

Step 2: get the token mix

Split the tokens into four buckets: uncached input, cache reads, cache writes and output. Here are the recorded mixes we use.

Mix A: one agent task (SWE-bench Verified)

Our agent attempted 33 SWE-bench Verified instances. Across all 33 it recorded 153.1M cache reads, 9.7M one-hour cache writes, 1.8M output and 3.2k uncached input tokens. Per attempt (our division by 33), that is about:

  • 4.64M cache reads (our calculation: 153.1M / 33)
  • 0.29M cache writes
  • 53k output
  • about 100 uncached input

94.0% of the input came from the cache. That is normal for an agent loop.

Mix B: one short call in Claude Code

From our five-task head-to-head, Sonnet 5.5 in Claude Code used a mean 1,401 cache-read tokens and 685 other input tokens per call, and a median 107 output tokens. Most of that input is the CLI's own context, not the prompt.

Mix C: one hard call in Claude Code

From the hard head-to-head, a median call wrote 1,050 output tokens on Sonnet 5.5 and 5,064 on Haiku 4.5, of which 4,556 were reported reasoning tokens.

Mix D: one routing decision

From the router study: a decision through the Claude Code CLI includes its tool schema and thinking tokens. We report it directly as a cost per 1,000 decisions (below).

Mix E: one call through Codex CLI

For a one-line request, the Codex CLI sent 19,551 input tokens. The OpenAI API sent 17. Use this mix to price the CLI's own context. We cover it in Claude Code vs Codex CLI: the hidden context tax.

Step 3: apply list prices per bucket

The formula:

cost = uncached input × input price + cache reads × cache-read price + cache writes × cache-write price + output × output price

List prices per million tokens, as used in our studies:

ModelInputCache readOutput
Claude Haiku 4.5$1$0.10$5
Claude Sonnet 5.5$2$0.20$10
Claude Opus 5.5$4$0.20$20
Claude Fable 5.1$10$0.25$50
GPT-6.1 Sol$2$0.10$10

Anthropic one-hour cache writes cost twice the input price. For OpenAI, we price writes as plain input.

Mix A at Sonnet prices: 4.64M × $0.20 + 0.29M × $4 + 53k × $10 ≈ $0.93 + $1.18 + $0.53 = about $2.64. The published per-attempt figure, from the unrounded totals, is $2.643.

Calculation
$87.23Sum of the 4 parts

Parts sorted by value, largest first

  1. Cache writes (1 h)
  2. Cache reads
  3. Output
  4. Uncached input

Shares are calculated from the values shown.

List-price calculation, not a run. 4 rows. Highest Cache writes (1 h) $38.99. Lowest Uncached input $0.01.

Notes

Recorded tokens at Sonnet 5.5 list price, by token kind

List-price calculation on recorded tokens: 153.1M cache reads, 9.7M cache writes, 1.8M output, 3.2k uncached input. An agent loop re-reads its context on every call, so cache reads dominate the token count; by price the largest part is cache writes (1 h).

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)

Over all 33 attempts the dollars split like this: cache writes $38.99, cache reads $30.62, output $17.61, uncached input $0.01. The biggest bucket by token count, cache reads, is not the biggest by price. Cache writes are.

The same mix at other list prices (per attempt, a calculation): Haiku 4.5 $1.32, GPT-6.1 Sol $1.59, Sonnet 5.5 $2.64, Opus 5.5 $4.36, Fable 5.1 $9.74. A different model would use different tokens, so read these as price sensitivity, not as a forecast for that model.

Step 4: divide by the success rate

You pay for every attempt, but you only keep the ones that succeed. So:

cost per success = cost per attempt ÷ success rate

Our agent resolved 25 of 33 attempts. At Sonnet prices, $2.643 per attempt becomes $3.49 per resolved task.

This step can reverse a ranking. On the hard head-to-head:

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Haiku 4.5 has half of Sonnet's list price. But at the CLI default it wrote about five times the output tokens, and it passed 11 of 24 calls. Its cost per call was about $0.0308 (our calculation: $0.0672 per pass × 11 passes ÷ 24 calls). Sonnet's was $0.0143, and it passed all 24. Per strict pass: $0.0143 on Sonnet against $0.0672 on Haiku.

Step 5: multiply by volume

Now scale. Some examples at list price (our calculations):

  • 100 agent tasks a month on Mix A: about $264 on Sonnet 5.5 and $436 on Opus 5.5. At our success rate, about 76 of them would be resolved.
  • 10,000 hard one-turn calls on Sonnet in Claude Code: about $143.
  • 1,000 routing decisions: $5.00 on Sonnet 5.5, $8.92 on Haiku 4.5 (it thinks longer under the CLI default) and $0.0337 on Jev 1.13, a dedicated router.
Calculation
Largest value is 260x the smallest; Log shows the small bars.
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5

Hover or focus a bar for its ratio to Jev 1.13 (TypeSafe) (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest Claude Haiku 4.5 $8.92 (n 82). Lowest Jev 1.13 (TypeSafe) $0.034 (n 246).

Notesn 82–246 per row

List price × reported tokens per decision

List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions

Step 6: add the overhead you do not see

Four costs often stay out of the first estimate:

  • CLI context. For the 19,551-token Codex CLI one-liner, the context alone costs $0.002 to $0.039 per call at GPT-6.1 Sol prices, depending on the cache.
  • Pipeline steps. Our agent's notional figure was $2.81 per attempt, not $2.64, because it also counts context-compaction calls. Planning, verification and review are real work and use real tokens.
  • Gateway fees and provider choice. A gateway can add a fee on top of list price. On 2026-10-06, OpenRouter listed the vendor's own per-token price for 10 of 10 models we checked and charged 5.5% when credits are bought by card (third-party-reported): OpenRouter vs going direct. For open-weight models, the provider matters more: prices for the same model differed by up to 12.6x (cheapest open-model providers).
  • Retries and caps. In our coding calibration, the latest build cost $11.06 for three real pull requests, and one earlier capped run stopped at its cap before delivery. A cap protects the budget, but a stopped attempt still costs.

Step 7: check the caching assumption

Calculation
  • With caching (as recorded)
  • Without caching (square)
In chart order.
Claude Haiku 4.5
Claude Sonnet 5.5
Claude Opus 5.5
Claude Fable 5.1

Gap labels, Without caching vs With caching (as recorded): Without caching is x% higher (+) or lower (−) than With caching (as recorded), calculated from the two values shown (the change counted from With caching (as recorded)’s value).

List-price calculation, not a run. 4 rows, 2 series: With caching (as recorded), Without caching. With caching (as recorded): highest Claude Fable 5.1 $321. Lowest Claude Haiku 4.5 $43.61. Without caching: highest Claude Fable 5.1 $1,717. Lowest Claude Haiku 4.5 $172.

Notes

The same recorded tokens with and without cache pricing

94.0% of recorded input tokens were cache reads. Calculation, not a run.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)

At Sonnet prices, caching cut the bill 3.9x, more than the Sonnet-to-Opus gap (1.6x). Both are our calculations: Mix A over 33 attempts costs $87.23 with caching and $343.33 without at Sonnet prices ($343.33 / $87.23 ≈ 3.9), and $143.83 with caching at Opus prices ($143.83 / $87.23 ≈ 1.6). Without caching, Sonnet would cost more than Opus with caching ($143.83).

Before you trust any estimate, confirm that your tool actually hits the cache. Agent loops depend on it.

A worked estimate

Say a team plans 200 agent tasks a month, like Mix A, on Sonnet 5.5, with a 75% success rate. Our calculation:

  1. Cost per attempt: $2.64.
  2. Monthly attempts: 200 × $2.643 = $529.
  3. Cost per success: $2.64 ÷ 0.75 = $3.52.
  4. With caching lost: about 3.9 × $529 ≈ $2,080.

The gap between line 2 and line 4 is the main risk in the budget. The model choice comes second.

How the numbers were made

  • Token mixes are sums or medians of recorded tokens from our studies. Per-attempt and per-call figures are our arithmetic on those values.
  • Prices are list prices from our price table, effective 2026-09-21 (OpenAI 2026-10-03).
  • All costs are calculations. The underlying calls ran on subscriptions.

Caveats

  • Your mix is not our mix. Repository size, task length and tool use change the token counts a lot. Record your own mix when you can.
  • Another model uses other tokens. Repricing one model's tokens at another's price shows sensitivity, not that model's real cost.
  • Prices change. Check the cache-read and cache-write prices first; they move agent bills the most.
  • Small samples behind the short-call and hard-call mixes (15 to 24 calls each).

Get your real token mix

Agent shows recorded usage and cost for task work. A provider may omit token buckets or a reliable cost; treat those values as unknown. Try Agent, inspect the receipts, and use a complete recorded mix when you have one.

The data behind this post

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.