AI coding cost per developer: a formula built on recorded work
AI coding cost per developer: a monthly budget formula using recorded tokens and list-price calculations. Replace our task volume and pass rate with yours.
TL;DR
- One formula turns a unit cost into a monthly cost per developer, with your own volume.
- Unit costs from recorded work (list-price calculations): a full agent attempt costs $2.81, or $3.71 per resolved task (25 of 33 resolved, 95% interval 59% to 87%). A hard Sonnet call costs $0.01435 (n = 24). A routing decision costs $0.0050 on Sonnet (n = 82) or $0.0000337 on Jev (246 calls, 82 cases).
- Example (inputs you should replace): 5 developers, 10 tasks each per workday, 21 workdays: about $3,891 a month before routing, or $778 per developer (calculation).
- What moves it: caching (3.9 times as much without it), model price (0.50 to 3.7 times) and pass-rate sensitivity (1.5 times, using the interval endpoints).
How to estimate your AI coding bill uses a token mix. This post starts from the unit of work. The AI cost calculator covers the token view.
What if every call ran on Opus? Repricing real agent tokens
A calculation, not a run: Agent's recorded SWE-bench tokens cost $87.23 at Sonnet 5.5 prices, $143.83 at Opus 5.5 and $343.33 without caching.
Transcript
- Thought experiment · recorded tokens × list prices. What if every call ran on Opus? Agent recorded every token on 33 SWE-bench attempts. We repriced them.
- 162.9M input tokens, 94.0% of them read from the prompt cache. Input tokens recorded: 162.9M (n = 33). Output tokens recorded: 1.8M (n = 33). Input served from cache: 94.0% (n = 33). Caveat: Recorded costs are list-price estimates for subscription calls; no invoice backs them.
- Same tokens on Opus 5.5: $143.83 instead of $87.23, 1.65× the bill. On Haiku 4.5: $43.61. At Haiku 4.5 prices: $43.61 (n = 33). At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
- Per resolved instance: $3.49 on Sonnet 5.5, $5.75 on Opus 5.5, $12.85 on Fable 5.1. Chart: Thought experiment: the same tokens at other list prices. Calculation, not a run. Caveat: Different models use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.
- In this calculation caching matters more than the model: without it, Sonnet would cost $343.33, 3.9× the recorded $87.23. Chart: Thought experiment: what prompt caching saved. Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
- Price sensitivity, not predictions. Every repricing labelled as a calculation.
The formula
monthly cost = sum over unit types of (units per developer per workday × workdays × developers × cost per unit ÷ pass rate) + routing decisions × cost per decision
- Units are tasks you need finished, not tries. Dividing by the pass rate turns them into attempts. If you count tries, skip it.
- Cost per unit is the cost of one attempt. Pass rate is the share of attempts that pass your check.
- Routing decisions are optional: one for each model call that a router picks.
Unit costs from recorded work
| Unit (n) | Cost per unit | Per 100 | Per 1,000 | Check passed (95% interval) | Cost per pass |
|---|---|---|---|---|---|
| Agent attempt, Sonnet 5.5 pipeline (33) | $2.81 | $280.72 | $2,807.20 | 25/33 resolved, 76% (59% to 87%) | $3.71 |
| Claude Code session, mean of 200 | $0.0875 | $8.75 | $87.46 | Sonnet full pass 9/15 (36% to 80%) to 15/15 (80% to 100%), by memory file | Sonnet $0.0818 to $0.1386 |
| One hard call, Sonnet 5.5 in Claude Code (24) | $0.01435 | $1.435 | $14.35 | 24/24 strict passes (86% to 100%) | $0.01435 |
| One routing decision, Sonnet 5.5 via CLI (82) | $0.0050 | $0.50 | $5.00 | 77/82 exact (87% to 97%) | not applicable |
| One routing decision, Jev 1.13 (246 calls, 82 cases) | $0.0000337 | $0.0034 | $0.0337 | 221/246 exact, 89.8% (case-level 82% to 95%) | not applicable |
All costs are list-price calculations, not invoices. Jev reports tokens, not cost (routing study). Its price calculation uses those tokens. Jev repeated the same 82 cases three times; its interval uses 82 cases, not 246 independent calls. Router accuracy intervals overlap. The cases were tuned against Jev answers, so Jev has a home advantage.
The Sonnet routing row rounds $4.996 per 1,000. Scaled columns use unrounded totals, so rounded unit costs may not multiply exactly. The session row uses $17.49234 ÷ 200 sessions (120 Sonnet 5.5, 80 Haiku 4.5). It mixes models and memory conditions; it is not a typical developer session. The $2.81 includes onboarding and failed attempts; see what a resolved SWE-bench task costs.
Worked example
Example inputs you should replace (assumptions, not data): 5 developers, 21 workdays, 10 tasks per developer per workday, one routing decision per model call.
Inputs from recorded work: $92.63751 total notional cost and 1,632 model calls across 33 attempts; 25 resolved (95% interval 59% to 87%). Use the unrounded means: $92.63751 ÷ 33 per attempt and 1,632 ÷ 33 calls per attempt. Calculation:
- Tasks to finish: 5 × 10 × 21 = 1,050.
- Attempts needed: 1,050 ÷ (25/33) = 1,386.
- Agent cost: 1,386 × ($92.63751 ÷ 33) = $3,890.78, or $778 per developer.
- Routing decisions: 1,386 × (1,632 ÷ 33) = 68,544.
- Routing cost: 68,544 × $0.004996 = $342.45 on Sonnet. On Jev, 68,544 × $0.0000337 = $2.31.
- Monthly total: $4,233.22 with a Sonnet router, $3,893.09 with Jev (round only after adding).
Calculation: a Sonnet router adds 8.8% to the agent cost. Jev adds 0.06%. At the unrounded pass-rate interval endpoints, step 3 gives about $3,381 to $4,998. This is sensitivity to the pass rate, not a 95% interval for monthly cost. It holds cost per attempt fixed.
This example assumes the new workload has the same unit costs and pass rate. We did not run these developers or retries. Router choices could change the model, tokens and outcomes; this calculation holds them fixed.
What moves the result most
Calculation: caching changes the same-token cost by 3.9 times. The listed model-price scenarios span 0.50 to 3.7 times Sonnet. The pass-rate interval endpoints change projected cost by 1.5 times. These calculations use different bases, so they do not rank savings on your workload.
Pass rate: a failed attempt still costs
On the hard calls, Haiku 4.5 lists at half of Sonnet's price. It passed 11 of 24 (46%, 95% interval 28% to 65%). Its cost per strict pass was $0.0672, 4.7 times Sonnet's $0.01435 (calculation). Sonnet, Opus (default and high effort) and Fable each passed 24 of 24 (95% interval 86% to 100%). The task set hits a ceiling: pass rate cannot separate them. Their calculated costs differ.
Post: hard tasks, Haiku to Fable. Study: hard model head-to-head.
A memory file may move it too. Sonnet passed fully 9 of 15 times without one (95% interval 35.7% to 80.2%). With an 11-line file, it passed 15 of 15 (79.6% to 100%). The intervals overlap slightly. Cost per full pass was $0.1386 without memory and $0.0818 with the curated file (calculation). The author knew the tasks, so treat it as an upper bound. See does CLAUDE.md help? and the memory study.
Caching: 3.9 times
94.0% of the agent's input tokens were cache reads. Without caching, the same 33 attempts would cost $343.33 on Sonnet 5.5 against $87.23: 3.9 times as much (calculation). Across 15 turns in three Claude Code sessions, Sonnet's calculated cost was $0.1350 with cache pricing versus $0.2698 without it: 50% less. In the Claude sessions, 0 of 4 later sessions reused an earlier session's ledger cache (95% Wilson interval 0% to 49%). We did not test why. Measure your own cache reads first.
More: how much does prompt caching save? and the caching study.
Model choice: price sensitivity, not a forecast
| Priced as | Per attempt | Versus Sonnet 5.5 |
|---|---|---|
| Claude Haiku 4.5 | $1.322 | 0.50 times |
| GPT-6.1 Sol | $1.589 | 0.60 times |
| Claude Sonnet 5.5 | $2.643 | 1.00 |
| Claude Opus 5.5 | $4.359 | 1.6 times |
| Claude Fable 5.1 | $9.737 | 3.7 times |
Each row prices the same recorded Sonnet tokens from 33 attempts at that list price (calculation). Another model would use other tokens and pass a different share. On hard calls, Sonnet and default-effort Opus 5.5 both passed 24 of 24 (95% interval 86% to 100%). Opus cost $0.02824 per strict pass, about 2.0 times Sonnet's (calculation).
Read what if every call ran on Opus?. Compare Sonnet vs Opus and Haiku vs Sonnet.
The unit you pick: a pipeline or a bare agent
The 11 public bash-only runs each attempted the same 33 instances. Total API cost divided by resolved tasks ranged from $0.08 to $1.18; the mean was $0.57 (calculation). Agent's cost was $3.71, about 6.5 times the panel mean (calculation: $3.706 ÷ $0.569). Agent resolved 25/33, 76% (95% interval 59% to 87%). The panel mean was 74.1%; individual rates ranged from 21/33 to 28/33 (64% to 85%). This is a range across 11 models, not an interval. Every model's 95% interval overlaps Agent's. This sample does not show that the extra steps raised the resolve rate.
Study: SWE-bench Verified.
How we measured
- Prices are list prices effective 2026-09-21 (OpenAI rows 2026-10-03; Jev 2026-09-23). Each cost is a calculation, not an invoice: recorded tokens times list price. Method: cost thought experiments.
- The $2.81 is the platform's notional figure and counts compaction calls. The $2.643 does not. Compare repriced rows with each other.
- Plans: the calls ran on subscriptions and the dataset holds no plan prices, so there is no plan comparison.
Caveats
- Small samples: 33 attempts, 24 hard calls, 200 sessions in one small repository, 82 routing cases.
- Different checks: resolved, strict pass and exact decision differ. Do not compare pass rates across rows.
- Retries: the division assumes independent retries with a constant success probability and mean attempt cost. It estimates expected attempts, not guaranteed completion. Real retries differ.
- Public tasks: SWE-bench issues are public and may be in model training data. Contamination is uncontrolled for Agent and the panel.
- Your code is not SWE-bench. Task and repository size change the tokens. We did not price developer time, review time or plans.
What to read next
Put your own volume into the formula
Agent records the model, tokens, cost and validation result for every task. Replace our unit costs with yours. Try Agent.