• Claude Code
  • LLM pricing
  • Calculation
  • Cost Per Task

Claude Code cost per task: a price ladder from one decision to one agent run

$0.087 per Claude Code coding session at API list prices, $0.0036 per short call, $2.81 per full agent attempt. Calculations on recorded tokens, not bills.

TL;DR

  • At API list prices, a short Claude Code call costs a fraction of a cent. A coding session costs about nine cents. A full agent run costs dollars. The point estimates span $0.0036 per short call to $2.81 per full agent attempt.
  • Every cost figure is a list-price calculation on recorded tokens. Our calls ran on subscriptions. Nothing here is a bill. These costs describe the recorded tasks, not every Claude Code task.
  • Failed calls still cost. Haiku 4.5 has the lower price per token. On the hard set it still cost 4.7 times as much as Sonnet 5.5 per strict pass (calculation).
  • The $2.81 attempt is not Claude Code. It is our Agent pipeline on Sonnet 5.5.
RungOne unit of work (Sonnet 5.5 unless stated)List-price cost (calculation)n
1A routing decision, effort low$0.0073 mean82 decisions
2A short validated call$0.0036 median15 calls
3A hard call, per strict pass$0.0143524 calls
4A cached 5-question session$0.045 mean3 sessions
5A Claude Code coding session (Sonnet and Haiku)$0.087 average200 sessions
6A full agent attempt on SWE-bench Verified$2.81 ($3.71 per resolved)33 attempts

Costs in the table are point calculations, with no cost confidence intervals. The ranges below show recorded variation, not 95% intervals.

Rung 1: a routing decision costs $0.0073

A routing decision with Sonnet 5.5 at low effort cost $7.324 per 1,000 decisions, or $0.0073 each (calculation, n = 82). That is a mean; individual calculated costs ranged from $0.00598 to $0.012472. It includes the CLI's own tool-schema and thinking tokens.

All 82 raw receipts report one-hour cache writes. We price those writes at $4 per million tokens, twice the input price. The study dataset uses the five-minute write price for this row. This post corrects that calculation from the receipts.

Rung 2: a short call costs $0.0036

A short validated call in Claude Code cost a median $0.0036 on Sonnet 5.5 (calculation; range $0.0034 to $0.0102, n = 15). Sonnet passed 12 of 15 short tasks (80%, 95% interval 55% to 93%). All 3 runs of the arithmetic task gave the right final answer but added working lines. The strict validator rejected their format. A failed call still costs, so the cost per passing answer was $0.0062 (calculation).

The short set also hits a ceiling: most configurations passed every call.

Calculation
Claude Fable 5.1
Claude Sonnet 5.5
Claude Opus 5.5 (high)
Claude Opus 5.5
Claude Opus 5.5 (low)
Claude Haiku 4.5
GPT-6.1 Sol (high)
GPT-6.1 Sol (medium)
GPT-6.1 Sol (low)

Every interval overlaps every other: this chart does not order these rows.

List-price calculation, not a run. 9 rows. Highest GPT-6.1 Sol (high) · Codex CLI $0.01 (range $0.0066–$0.028, n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.0036 (range $0.0034–$0.01, n 15). All run ranges overlap.

NotesLines: lowest–highest run (not an interval)n 10–15 per row

Reported tokens × list price; the calls ran on subscriptions

Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.

Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Rung 3: a hard call costs $0.014 per strict pass

The hard set has eight tasks with strict validators. Sonnet 5.5 passed 24 of 24 (100%, 95% interval 86% to 100%). Its cost per strict pass equals its mean cost per call: $0.01435 (calculation). Individual calls ranged from $0.0051 to $0.0424.

Haiku 4.5 has a lower price per token. It passed 11 of 24 (46%, 95% interval 28% to 65%). Five other calls had correct answers in the wrong format. They count as strict non-passes. Failed calls still cost, so Haiku came to $0.0672 per strict pass, 4.7 times Sonnet (calculation; cost has no interval).

Sonnet had the lowest point estimate of seven configurations. GPT-6.1 Sol (high) in the Codex CLI was close at $0.0151 (calculation, n = 16). We call neither one ahead. Six of seven configurations passed every call, so the set has a ceiling. Their cells have n = 16 or 24; their 95% intervals run from 81% or 86% to 100%.

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Haiku 4.5 · Claude Code
Claude Fable 5.1 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

All calls in a configuration, failures and format misses included, divided by its strict passes

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Rung 4: a cached 5-question session costs $0.045

One Claude Code session sent a fixed ledger and asked 5 questions. Three Sonnet 5.5 sessions (15 turns) cost $0.1350 with the cache, as recorded. That is a mean $0.045 per session (calculation: $0.1350 ÷ 3; session range $0.0389 to $0.0491). Priced as plain input, the same tokens cost $0.2698, or about $0.090 per session (calculation).

Turn 1 averaged $0.0258 per session, or 57% of the $0.0450 mean (calculation, n = 3 sessions). It writes the ledger to a 1-hour cache at twice the input price. Turns 2 to 5 read 97% of their input from the cache on average. A new session did not reuse the ledger cache (0 of 4 later Claude sessions across both models; 95% interval 0% to 49%). We did not test why.

Calculation
  • With the cache, as recorded
  • Without a cache: every input token at the input price (square)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code

Gap labels, Without a cache: every input token at the input price vs With the cache, as recorded: Without a cache: every input token at the input price is x% higher (+) or lower (−) than With the cache, as recorded, calculated from the two values shown (the change counted from With the cache, as recorded’s value).

List-price calculation, not a run. 2 rows, 2 series: With the cache, as recorded, Without a cache: every input token at the input price. With the cache, as recorded: highest Claude Opus 5.5 · Claude Code $0.26 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.14 (n 15). Without a cache: every input token at the input price: highest Claude Opus 5.5 · Claude Code $0.54 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.27 (n 15).

Notesn = 15 per row

All recorded turns per model; the same reported tokens priced two ways

Calculation, not a bill: the calls ran on a subscription. Cache reads at the cache-read price, 1-hour cache writes at 2× the input price (every write in this run was a 1-hour write). Codex CLI is not priced here: it reports no cache-write count.

Sources: Caching sessions and repeated prompts (Claude Code and Codex CLI), Cost with and without the prompt cache (calculation), Anthropic list prices (Claude models)

Rung 5: a Claude Code coding session costs about $0.087

The memory study ran 200 graded Claude Code sessions: 120 on Sonnet 5.5 and 80 on Haiku 4.5. Each session did one coding task in a small Node.js repository. At list price they cost $17.49 in total, or $0.087 per session (calculation).

For Sonnet 5.5, mean session cost ranged from $0.0815 to $0.1022 across eight memory conditions. These are condition means, not individual-session limits (calculation: total cost ÷ 15). What moved was the number of fully correct sessions.

The curated 11-line file cost $0.0818 per fully correct result: 15 of 15 correct, with a 95% interval of 80% to 100%. No memory cost $0.1386: 9 of 15 correct, with a 95% interval of 36% to 80%. These are cost calculations. With n = 15 per condition, most full-pass intervals overlap.

Haiku 4.5 had a higher point estimate of cost per fully correct result in all eight conditions. Its range was $0.0960 to $0.4317, against Sonnet’s $0.0818 to $0.1386 (calculation). Haiku ran 10 sessions per condition and had 2 to 9 full passes. The endpoint 95% intervals are 6% to 51% and 60% to 98%. No cost interval supports a general cost ranking.

Median input per Sonnet session ran from 65,045 to 125,674 tokens across conditions (n = 15 each), mostly cache reads. The top figure is the Stop hook, which sends the agent back to work.

Calculation
  • Claude Sonnet 5.5
  • Claude Haiku 4.5 (square)
Sorted by gap, largest first.
/init CLAUDE.md
No memory
Handbook, 210 lines
Curated, 11 lines
Dreamed notes
Raw notes, 60 lines
Curated + hook
Stop hook only

Gap labels, Claude Haiku 4.5 vs Claude Sonnet 5.5: Claude Haiku 4.5 is x% higher (+) or lower (−) than Claude Sonnet 5.5, calculated from the two values shown (the change counted from Claude Sonnet 5.5’s value).

List-price calculation, not a run. 8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest No memory $0.14 (n 9). Lowest Curated, 11 lines $0.082 (n 15). Claude Haiku 4.5: highest /init CLAUDE.md $0.43 (n 2). Lowest Curated + hook $0.096 (n 9).

Notesn 2–15 per row

Sum of the CLI's cost estimates for a condition, divided by its full passes

Sessions ran on a subscription; these are the CLI's list-price estimates, not bills. A failed session still costs money, so cost per correct result falls when fewer sessions fail.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Rung 6: a full agent attempt costs $2.81, and it is not Claude Code

Agent runs a full pipeline: onboarding, research, plan, act, verify and review. We ran it on 33 SWE-bench Verified instances with Sonnet 5.5 through a subscription CLI. The notional list-price cost was $92.64, or $2.81 per attempt (calculation).

Agent resolved 25 of 33 (76%, 95% interval 59% to 87%), so the cost per resolved task was $3.71 (calculation). Median attempt time was 9.6 minutes (range 1.6 to 54.1 minutes, n = 33). The mean was about 49 model calls per attempt. The act stage cost $44.05 (48% of the total) and research $20.56 (22%) (calculations).

$92.64Total over 33 attempts (from the note)

Parts sorted by value, largest first

  1. Act (edit and run)
  2. Research
  3. Verify
  4. Other
  5. Review
  6. Context compaction
  7. Onboarding notes

Shares are calculated from the values shown; rounding can make the sum of the parts differ from the stated total by a cent.

7 rows. Highest Act (edit and run) $44.05. Lowest Onboarding notes $2.09.

Notes

Share of notional model cost by stage, all 33 attempts

Total $92.64 over 33 attempts.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

We did not run Claude Code on SWE-bench.

What drives the cost

Cache writes and reads. Across the 33 agent attempts, 94.0% of 162.9M input tokens were cache reads, and output was 1.8M tokens. At Sonnet 5.5 list price, the total was $87.23 (calculation). Cache writes cost $38.99 (45%), cache reads $30.62 (35%) and output $17.61 (20%). Output is about 1% of the tokens and 20% of the dollars. Without caching, the same tokens cost $343.33, or 3.9 times as much (calculation; cost study).

Failures. A failed attempt still costs. Both figures count all 33 attempts (calculation). Dividing their total cost by 25 resolved tasks gives $3.71, against $2.81 per attempt.

How much does Claude Code cost? The subscription question

Our calls ran on flat-rate subscriptions. The dataset holds Anthropic's API list prices (recorded 2026-09-21) and no plan prices. So we make no plan comparison and cannot say if a plan costs less for you.

How to estimate yours

  1. Pick the rung that matches your unit of work.
  2. Read your token counts from your usage output: uncached input, cache reads, cache writes and output.
  3. Multiply each count by its list price. The AI cost calculator does this for a token mix.
  4. Add the cost of every failed attempt. Divide by the attempts that succeeded.
  5. Set the total against your plan's price.

How we measured

  • Prices: Anthropic list prices of 2026-09-21. Sonnet 5.5 costs $2 per million input tokens, $0.20 per million cache-read tokens and $10 per million output tokens. Cache writes cost $2.50 per million for five minutes and $4 for one hour.
  • Rung 1: Claude routers ran through the Claude Code CLI, one call per decision. We repriced the recorded tokens at list price.
  • Rungs 2 to 5: Claude Code on Claude subscription accounts. Each study publishes a protocol and keeps every counted attempt without retries. The current protocol file timestamps follow the first counted calls. They do not establish that the protocols were written before those calls. Controls and probes sit outside the counted cells.

    Rungs 2 and 3 ran one turn with tools off in an empty folder. Rung 4 kept five turns in one CLI session. Rung 5 ran Claude Code 2.1.286 headless in a sandbox.

  • Rung 6: Agent's own notional cost, which also counts compaction calls. The token split above reprices every recorded token at Sonnet 5.5 list price: $87.23, or $2.64 per attempt.

Caveats

  • Different tasks, different harnesses. The rungs are not one task at six sizes.
  • Small samples: 3 sessions in rung 4, 15 calls in rung 2 and 15 sessions per condition in rung 5.
  • Rung 5 mixes two models. Sonnet and Haiku sessions share the $0.087 average.
  • Notional costs. No invoice backs any figure.
  • SWE-bench covers 33 of 500 Verified instances. The pipeline, not the model alone, set the $2.81.

Price your own agent work

Agent keeps receipts for your tasks: model, route, tokens, time, cost and validation result. Try Agent and see your own numbers.

The data behind this post

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.