• Thought experiment
  • LLM pricing
  • Calculation
  • Cost Calculator

AI coding cost per developer: a formula built on recorded work

AI coding cost per developer: a monthly budget formula using recorded tokens and list-price calculations. Replace our task volume and pass rate with yours.

TL;DR

  • One formula turns a unit cost into a monthly cost per developer, with your own volume.
  • Unit costs from recorded work (list-price calculations): a full agent attempt costs $2.81, or $3.71 per resolved task (25 of 33 resolved, 95% interval 59% to 87%). A hard Sonnet call costs $0.01435 (n = 24). A routing decision costs $0.0050 on Sonnet (n = 82) or $0.0000337 on Jev (246 calls, 82 cases).
  • Example (inputs you should replace): 5 developers, 10 tasks each per workday, 21 workdays: about $3,891 a month before routing, or $778 per developer (calculation).
  • What moves it: caching (3.9 times as much without it), model price (0.50 to 3.7 times) and pass-rate sensitivity (1.5 times, using the interval endpoints).

How to estimate your AI coding bill uses a token mix. This post starts from the unit of work. The AI cost calculator covers the token view.

Live story · 33 sWhat if every call ran on Opus? Repricing real agent tokens

What if every call ran on Opus? Repricing real agent tokens

A calculation, not a run: Agent's recorded SWE-bench tokens cost $87.23 at Sonnet 5.5 prices, $143.83 at Opus 5.5 and $343.33 without caching.

Transcript
  1. Thought experiment · recorded tokens × list prices. What if every call ran on Opus? Agent recorded every token on 33 SWE-bench attempts. We repriced them.
  2. 162.9M input tokens, 94.0% of them read from the prompt cache. Input tokens recorded: 162.9M (n = 33). Output tokens recorded: 1.8M (n = 33). Input served from cache: 94.0% (n = 33). Caveat: Recorded costs are list-price estimates for subscription calls; no invoice backs them.
  3. Same tokens on Opus 5.5: $143.83 instead of $87.23, 1.65× the bill. On Haiku 4.5: $43.61. At Haiku 4.5 prices: $43.61 (n = 33). At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  4. Per resolved instance: $3.49 on Sonnet 5.5, $5.75 on Opus 5.5, $12.85 on Fable 5.1. Chart: Thought experiment: the same tokens at other list prices. Calculation, not a run. Caveat: Different models use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.
  5. In this calculation caching matters more than the model: without it, Sonnet would cost $343.33, 3.9× the recorded $87.23. Chart: Thought experiment: what prompt caching saved. Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  6. Price sensitivity, not predictions. Every repricing labelled as a calculation.

The formula

monthly cost = sum over unit types of (units per developer per workday × workdays × developers × cost per unit ÷ pass rate) + routing decisions × cost per decision

  • Units are tasks you need finished, not tries. Dividing by the pass rate turns them into attempts. If you count tries, skip it.
  • Cost per unit is the cost of one attempt. Pass rate is the share of attempts that pass your check.
  • Routing decisions are optional: one for each model call that a router picks.

Unit costs from recorded work

Unit (n)Cost per unitPer 100Per 1,000Check passed (95% interval)Cost per pass
Agent attempt, Sonnet 5.5 pipeline (33)$2.81$280.72$2,807.2025/33 resolved, 76% (59% to 87%)$3.71
Claude Code session, mean of 200$0.0875$8.75$87.46Sonnet full pass 9/15 (36% to 80%) to 15/15 (80% to 100%), by memory fileSonnet $0.0818 to $0.1386
One hard call, Sonnet 5.5 in Claude Code (24)$0.01435$1.435$14.3524/24 strict passes (86% to 100%)$0.01435
One routing decision, Sonnet 5.5 via CLI (82)$0.0050$0.50$5.0077/82 exact (87% to 97%)not applicable
One routing decision, Jev 1.13 (246 calls, 82 cases)$0.0000337$0.0034$0.0337221/246 exact, 89.8% (case-level 82% to 95%)not applicable

All costs are list-price calculations, not invoices. Jev reports tokens, not cost (routing study). Its price calculation uses those tokens. Jev repeated the same 82 cases three times; its interval uses 82 cases, not 246 independent calls. Router accuracy intervals overlap. The cases were tuned against Jev answers, so Jev has a home advantage.

The Sonnet routing row rounds $4.996 per 1,000. Scaled columns use unrounded totals, so rounded unit costs may not multiply exactly. The session row uses $17.49234 ÷ 200 sessions (120 Sonnet 5.5, 80 Haiku 4.5). It mixes models and memory conditions; it is not a typical developer session. The $2.81 includes onboarding and failed attempts; see what a resolved SWE-bench task costs.

Worked example

Example inputs you should replace (assumptions, not data): 5 developers, 21 workdays, 10 tasks per developer per workday, one routing decision per model call.

Inputs from recorded work: $92.63751 total notional cost and 1,632 model calls across 33 attempts; 25 resolved (95% interval 59% to 87%). Use the unrounded means: $92.63751 ÷ 33 per attempt and 1,632 ÷ 33 calls per attempt. Calculation:

  1. Tasks to finish: 5 × 10 × 21 = 1,050.
  2. Attempts needed: 1,050 ÷ (25/33) = 1,386.
  3. Agent cost: 1,386 × ($92.63751 ÷ 33) = $3,890.78, or $778 per developer.
  4. Routing decisions: 1,386 × (1,632 ÷ 33) = 68,544.
  5. Routing cost: 68,544 × $0.004996 = $342.45 on Sonnet. On Jev, 68,544 × $0.0000337 = $2.31.
  6. Monthly total: $4,233.22 with a Sonnet router, $3,893.09 with Jev (round only after adding).

Calculation: a Sonnet router adds 8.8% to the agent cost. Jev adds 0.06%. At the unrounded pass-rate interval endpoints, step 3 gives about $3,381 to $4,998. This is sensitivity to the pass rate, not a 95% interval for monthly cost. It holds cost per attempt fixed.

This example assumes the new workload has the same unit costs and pass rate. We did not run these developers or retries. Router choices could change the model, tokens and outcomes; this calculation holds them fixed.

What moves the result most

Calculation: caching changes the same-token cost by 3.9 times. The listed model-price scenarios span 0.50 to 3.7 times Sonnet. The pass-rate interval endpoints change projected cost by 1.5 times. These calculations use different bases, so they do not rank savings on your workload.

Pass rate: a failed attempt still costs

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

On the hard calls, Haiku 4.5 lists at half of Sonnet's price. It passed 11 of 24 (46%, 95% interval 28% to 65%). Its cost per strict pass was $0.0672, 4.7 times Sonnet's $0.01435 (calculation). Sonnet, Opus (default and high effort) and Fable each passed 24 of 24 (95% interval 86% to 100%). The task set hits a ceiling: pass rate cannot separate them. Their calculated costs differ.

Post: hard tasks, Haiku to Fable. Study: hard model head-to-head.

A memory file may move it too. Sonnet passed fully 9 of 15 times without one (95% interval 35.7% to 80.2%). With an 11-line file, it passed 15 of 15 (79.6% to 100%). The intervals overlap slightly. Cost per full pass was $0.1386 without memory and $0.0818 with the curated file (calculation). The author knew the tasks, so treat it as an upper bound. See does CLAUDE.md help? and the memory study.

Caching: 3.9 times

Calculation
  • With caching (as recorded)
  • Without caching (square)
In chart order.
Claude Haiku 4.5
Claude Sonnet 5.5
Claude Opus 5.5
Claude Fable 5.1

Gap labels, Without caching vs With caching (as recorded): Without caching is x% higher (+) or lower (−) than With caching (as recorded), calculated from the two values shown (the change counted from With caching (as recorded)’s value).

List-price calculation, not a run. 4 rows, 2 series: With caching (as recorded), Without caching. With caching (as recorded): highest Claude Fable 5.1 $321. Lowest Claude Haiku 4.5 $43.61. Without caching: highest Claude Fable 5.1 $1,717. Lowest Claude Haiku 4.5 $172.

Notes

The same recorded tokens with and without cache pricing

94.0% of recorded input tokens were cache reads. Calculation, not a run.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)

94.0% of the agent's input tokens were cache reads. Without caching, the same 33 attempts would cost $343.33 on Sonnet 5.5 against $87.23: 3.9 times as much (calculation). Across 15 turns in three Claude Code sessions, Sonnet's calculated cost was $0.1350 with cache pricing versus $0.2698 without it: 50% less. In the Claude sessions, 0 of 4 later sessions reused an earlier session's ledger cache (95% Wilson interval 0% to 49%). We did not test why. Measure your own cache reads first.

More: how much does prompt caching save? and the caching study.

Model choice: price sensitivity, not a forecast

Calculation
Largest value is 310x the smallest; Log shows the small bars.
Claude Fable 5.1
Claude Opus 5
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol
Claude Haiku 4.5
Gemini 3.x Flash
Jev 1.13 (router)

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 8 rows. Highest Claude Fable 5.1 $12.85. Lowest Jev 1.13 (router) $0.042.

Notes

Cost per resolved SWE-bench instance if 162.9M input and 1.8M output tokens had been billed at each model's list price

Calculation, not a run: tokens recorded by Agent on claude-sonnet-5-5 (33 attempts, 25 resolved) times list prices effective 2026-09-21. Another model would use a different number of tokens and resolve a different set. Jev is a routing model and cannot do this work; its bar is a price floor only.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models), Google Gemini list prices, OpenAI list prices, Jev 1.13 list price

Priced asPer attemptVersus Sonnet 5.5
Claude Haiku 4.5$1.3220.50 times
GPT-6.1 Sol$1.5890.60 times
Claude Sonnet 5.5$2.6431.00
Claude Opus 5.5$4.3591.6 times
Claude Fable 5.1$9.7373.7 times

Each row prices the same recorded Sonnet tokens from 33 attempts at that list price (calculation). Another model would use other tokens and pass a different share. On hard calls, Sonnet and default-effort Opus 5.5 both passed 24 of 24 (95% interval 86% to 100%). Opus cost $0.02824 per strict pass, about 2.0 times Sonnet's (calculation).

Read what if every call ran on Opus?. Compare Sonnet vs Opus and Haiku vs Sonnet.

The unit you pick: a pipeline or a bare agent

Largest value is 46x the smallest; Log shows the small bars.
Agent (notional)
Claude 4.5 Opus (high)
Claude 4.5 Sonnet (high)
Claude 4.6 Opus
GLM 5 (high)
DeepSeek V3.2 (high)
GPT 5.2 (high)
Claude 4.5 Haiku (high)
Gemini 3 Flash (high)
Kimi K2.5 (high)
MiniMax M2.5 (high)
GPT 5 mini

Hover or focus a bar for its ratio to Agent (notional) (the highlighted row): a ratio of the two values shown, not a measurement.

12 rows. Highest Agent (notional) $3.71 (n 25). Lowest GPT 5 mini $0.08 (n 21).

Notesn 21–28 per row

Same 33 SWE-bench Verified instances; all attempts in the numerator

Recorded figures, not repricing. Panel costs are published API costs for a bash-only agent. Agent's figure is a list-price estimate of subscription calls and includes onboarding, planning, verification and review.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

The 11 public bash-only runs each attempted the same 33 instances. Total API cost divided by resolved tasks ranged from $0.08 to $1.18; the mean was $0.57 (calculation). Agent's cost was $3.71, about 6.5 times the panel mean (calculation: $3.706 ÷ $0.569). Agent resolved 25/33, 76% (95% interval 59% to 87%). The panel mean was 74.1%; individual rates ranged from 21/33 to 28/33 (64% to 85%). This is a range across 11 models, not an interval. Every model's 95% interval overlaps Agent's. This sample does not show that the extra steps raised the resolve rate.

Study: SWE-bench Verified.

How we measured

  • Prices are list prices effective 2026-09-21 (OpenAI rows 2026-10-03; Jev 2026-09-23). Each cost is a calculation, not an invoice: recorded tokens times list price. Method: cost thought experiments.
  • The $2.81 is the platform's notional figure and counts compaction calls. The $2.643 does not. Compare repriced rows with each other.
  • Plans: the calls ran on subscriptions and the dataset holds no plan prices, so there is no plan comparison.

Caveats

  • Small samples: 33 attempts, 24 hard calls, 200 sessions in one small repository, 82 routing cases.
  • Different checks: resolved, strict pass and exact decision differ. Do not compare pass rates across rows.
  • Retries: the division assumes independent retries with a constant success probability and mean attempt cost. It estimates expected attempts, not guaranteed completion. Real retries differ.
  • Public tasks: SWE-bench issues are public and may be in model training data. Contamination is uncontrolled for Agent and the panel.
  • Your code is not SWE-bench. Task and repository size change the tokens. We did not price developer time, review time or plans.

Put your own volume into the formula

Agent records the model, tokens, cost and validation result for every task. Replace our unit costs with yours. Try Agent.

The data behind this post

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.