Claude Code cost per task: a price ladder from one decision to one agent run
$0.087 per Claude Code coding session at API list prices, $0.0036 per short call, $2.81 per full agent attempt. Calculations on recorded tokens, not bills.
TL;DR
- At API list prices, a short Claude Code call costs a fraction of a cent. A coding session costs about nine cents. A full agent run costs dollars. The point estimates span $0.0036 per short call to $2.81 per full agent attempt.
- Every cost figure is a list-price calculation on recorded tokens. Our calls ran on subscriptions. Nothing here is a bill. These costs describe the recorded tasks, not every Claude Code task.
- Failed calls still cost. Haiku 4.5 has the lower price per token. On the hard set it still cost 4.7 times as much as Sonnet 5.5 per strict pass (calculation).
- The $2.81 attempt is not Claude Code. It is our Agent pipeline on Sonnet 5.5.
| Rung | One unit of work (Sonnet 5.5 unless stated) | List-price cost (calculation) | n |
|---|---|---|---|
| 1 | A routing decision, effort low | $0.0073 mean | 82 decisions |
| 2 | A short validated call | $0.0036 median | 15 calls |
| 3 | A hard call, per strict pass | $0.01435 | 24 calls |
| 4 | A cached 5-question session | $0.045 mean | 3 sessions |
| 5 | A Claude Code coding session (Sonnet and Haiku) | $0.087 average | 200 sessions |
| 6 | A full agent attempt on SWE-bench Verified | $2.81 ($3.71 per resolved) | 33 attempts |
Costs in the table are point calculations, with no cost confidence intervals. The ranges below show recorded variation, not 95% intervals.
Rung 1: a routing decision costs $0.0073
A routing decision with Sonnet 5.5 at low effort cost $7.324 per 1,000 decisions, or $0.0073 each (calculation, n = 82). That is a mean; individual calculated costs ranged from $0.00598 to $0.012472. It includes the CLI's own tool-schema and thinking tokens.
All 82 raw receipts report one-hour cache writes. We price those writes at $4 per million tokens, twice the input price. The study dataset uses the five-minute write price for this row. This post corrects that calculation from the receipts.
Rung 2: a short call costs $0.0036
A short validated call in Claude Code cost a median $0.0036 on Sonnet 5.5 (calculation; range $0.0034 to $0.0102, n = 15). Sonnet passed 12 of 15 short tasks (80%, 95% interval 55% to 93%). All 3 runs of the arithmetic task gave the right final answer but added working lines. The strict validator rejected their format. A failed call still costs, so the cost per passing answer was $0.0062 (calculation).
The short set also hits a ceiling: most configurations passed every call.
Rung 3: a hard call costs $0.014 per strict pass
The hard set has eight tasks with strict validators. Sonnet 5.5 passed 24 of 24 (100%, 95% interval 86% to 100%). Its cost per strict pass equals its mean cost per call: $0.01435 (calculation). Individual calls ranged from $0.0051 to $0.0424.
Haiku 4.5 has a lower price per token. It passed 11 of 24 (46%, 95% interval 28% to 65%). Five other calls had correct answers in the wrong format. They count as strict non-passes. Failed calls still cost, so Haiku came to $0.0672 per strict pass, 4.7 times Sonnet (calculation; cost has no interval).
Sonnet had the lowest point estimate of seven configurations. GPT-6.1 Sol (high) in the Codex CLI was close at $0.0151 (calculation, n = 16). We call neither one ahead. Six of seven configurations passed every call, so the set has a ceiling. Their cells have n = 16 or 24; their 95% intervals run from 81% or 86% to 100%.
Rung 4: a cached 5-question session costs $0.045
One Claude Code session sent a fixed ledger and asked 5 questions. Three Sonnet 5.5 sessions (15 turns) cost $0.1350 with the cache, as recorded. That is a mean $0.045 per session (calculation: $0.1350 ÷ 3; session range $0.0389 to $0.0491). Priced as plain input, the same tokens cost $0.2698, or about $0.090 per session (calculation).
Turn 1 averaged $0.0258 per session, or 57% of the $0.0450 mean (calculation, n = 3 sessions). It writes the ledger to a 1-hour cache at twice the input price. Turns 2 to 5 read 97% of their input from the cache on average. A new session did not reuse the ledger cache (0 of 4 later Claude sessions across both models; 95% interval 0% to 49%). We did not test why.
Rung 5: a Claude Code coding session costs about $0.087
The memory study ran 200 graded Claude Code sessions: 120 on Sonnet 5.5 and 80 on Haiku 4.5. Each session did one coding task in a small Node.js repository. At list price they cost $17.49 in total, or $0.087 per session (calculation).
For Sonnet 5.5, mean session cost ranged from $0.0815 to $0.1022 across eight memory conditions. These are condition means, not individual-session limits (calculation: total cost ÷ 15). What moved was the number of fully correct sessions.
The curated 11-line file cost $0.0818 per fully correct result: 15 of 15 correct, with a 95% interval of 80% to 100%. No memory cost $0.1386: 9 of 15 correct, with a 95% interval of 36% to 80%. These are cost calculations. With n = 15 per condition, most full-pass intervals overlap.
Haiku 4.5 had a higher point estimate of cost per fully correct result in all eight conditions. Its range was $0.0960 to $0.4317, against Sonnet’s $0.0818 to $0.1386 (calculation). Haiku ran 10 sessions per condition and had 2 to 9 full passes. The endpoint 95% intervals are 6% to 51% and 60% to 98%. No cost interval supports a general cost ranking.
Median input per Sonnet session ran from 65,045 to 125,674 tokens across conditions (n = 15 each), mostly cache reads. The top figure is the Stop hook, which sends the agent back to work.
Rung 6: a full agent attempt costs $2.81, and it is not Claude Code
Agent runs a full pipeline: onboarding, research, plan, act, verify and review. We ran it on 33 SWE-bench Verified instances with Sonnet 5.5 through a subscription CLI. The notional list-price cost was $92.64, or $2.81 per attempt (calculation).
Agent resolved 25 of 33 (76%, 95% interval 59% to 87%), so the cost per resolved task was $3.71 (calculation). Median attempt time was 9.6 minutes (range 1.6 to 54.1 minutes, n = 33). The mean was about 49 model calls per attempt. The act stage cost $44.05 (48% of the total) and research $20.56 (22%) (calculations).
We did not run Claude Code on SWE-bench.
What drives the cost
Cache writes and reads. Across the 33 agent attempts, 94.0% of 162.9M input tokens were cache reads, and output was 1.8M tokens. At Sonnet 5.5 list price, the total was $87.23 (calculation). Cache writes cost $38.99 (45%), cache reads $30.62 (35%) and output $17.61 (20%). Output is about 1% of the tokens and 20% of the dollars. Without caching, the same tokens cost $343.33, or 3.9 times as much (calculation; cost study).
Failures. A failed attempt still costs. Both figures count all 33 attempts (calculation). Dividing their total cost by 25 resolved tasks gives $3.71, against $2.81 per attempt.
How much does Claude Code cost? The subscription question
Our calls ran on flat-rate subscriptions. The dataset holds Anthropic's API list prices (recorded 2026-09-21) and no plan prices. So we make no plan comparison and cannot say if a plan costs less for you.
How to estimate yours
- Pick the rung that matches your unit of work.
- Read your token counts from your usage output: uncached input, cache reads, cache writes and output.
- Multiply each count by its list price. The AI cost calculator does this for a token mix.
- Add the cost of every failed attempt. Divide by the attempts that succeeded.
- Set the total against your plan's price.
How we measured
- Prices: Anthropic list prices of 2026-09-21. Sonnet 5.5 costs $2 per million input tokens, $0.20 per million cache-read tokens and $10 per million output tokens. Cache writes cost $2.50 per million for five minutes and $4 for one hour.
- Rung 1: Claude routers ran through the Claude Code CLI, one call per decision. We repriced the recorded tokens at list price.
- Rungs 2 to 5: Claude Code on Claude subscription accounts. Each study publishes a protocol and keeps every counted attempt without retries. The current protocol file timestamps follow the first counted calls. They do not establish that the protocols were written before those calls. Controls and probes sit outside the counted cells.
Rungs 2 and 3 ran one turn with tools off in an empty folder. Rung 4 kept five turns in one CLI session. Rung 5 ran Claude Code 2.1.286 headless in a sandbox.
- Rung 6: Agent's own notional cost, which also counts compaction calls. The token split above reprices every recorded token at Sonnet 5.5 list price: $87.23, or $2.64 per attempt.
Caveats
- Different tasks, different harnesses. The rungs are not one task at six sizes.
- Small samples: 3 sessions in rung 4, 15 calls in rung 2 and 15 sessions per condition in rung 5.
- Rung 5 mixes two models. Sonnet and Haiku sessions share the $0.087 average.
- Notional costs. No invoice backs any figure.
- SWE-bench covers 33 of 500 Verified instances. The pipeline, not the model alone, set the $2.81.
What to read next
- How to estimate your AI coding bill
- What one resolved SWE-bench task costs
- How much does prompt caching save?
- Does CLAUDE.md help an agent?
- Claude Code model page
Price your own agent work
Agent keeps receipts for your tasks: model, route, tokens, time, cost and validation result. Try Agent and see your own numbers.