• Cost
  • LLM pricing
  • SWE-bench
  • Prompt Caching

What one resolved SWE-bench task really costs an AI coding agent

A full agent pipeline spent a notional $3.71 per resolved SWE-bench instance. Where the money went, what caching saved, and the public panel range.

TL;DR

  • Agent spent a notional $2.81 per attempt and $3.71 per resolved instance on 33 SWE-bench Verified instances. The resolved count is 25.
  • The 11 public bash-only runs on the same instances spent between $0.08 and $1.18 per resolved instance. Their mean is $0.57.
  • Of Agent's $92.64 total, $44.05 went to the act stage (editing and running code) and $20.56 to research.
  • By token kind, cache writes cost the most ($38.99 at Sonnet 5.5 list price), then cache reads ($30.62), then output ($17.61).
  • Without prompt caching, the same tokens would cost $343.33 instead of $87.23 at Sonnet 5.5 prices. That is a calculation, not a run.

Data and every attempt: /benchmarks/swe-bench-verified and /benchmarks/cost-thought-experiments. One attempt as a receipt: /why-agent/receipt.

The number most benchmarks leave out

Resolve rate gets the headline. Cost decides whether you can use the thing every day. So we count cost the honest way: every attempt's spend goes in the numerator, and only resolved instances go in the denominator. A failed attempt still costs money, so it raises the cost of each success.

Here is how that looks next to the public panel, on the same 33 instances.

Largest value is 46x the smallest; Log shows the small bars.
Agent (notional)
Claude 4.5 Opus (high)
Claude 4.5 Sonnet (high)
Claude 4.6 Opus
GLM 5 (high)
DeepSeek V3.2 (high)
GPT 5.2 (high)
Claude 4.5 Haiku (high)
Gemini 3 Flash (high)
Kimi K2.5 (high)
MiniMax M2.5 (high)
GPT 5 mini

Hover or focus a bar for its ratio to Agent (notional) (the highlighted row): a ratio of the two values shown, not a measurement.

12 rows. Highest Agent (notional) $3.71 (n 25). Lowest GPT 5 mini $0.08 (n 21).

Notesn 21–28 per row

Same 33 SWE-bench Verified instances; all attempts in the numerator

Recorded figures, not repricing. Panel costs are published API costs for a bash-only agent. Agent's figure is a list-price estimate of subscription calls and includes onboarding, planning, verification and review.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

Agent sits at $3.71 per resolved instance. The most expensive panel run is Claude 4.5 Opus (high) at $1.18. The cheapest is GPT 5 mini at $0.08, and MiniMax M2.5 (high) is close at $0.11. The resolved counts behind each bar differ, from 21 to 28.

So yes: on a per-success basis, a full pipeline costs several times what a bare bash agent costs. We think you should know that before you read any resolve rate, ours included.

Two things make the comparison less direct than the chart suggests:

  1. Different price sources. Panel costs are published API costs. Agent's cost is a list-price estimate of calls that actually ran on a flat subscription. No invoice backs it.
  2. Different work. Agent's cost includes onboarding the repository, planning, verification and review. A bare agent skips all of that.

Where an agent's money goes

The pipeline has stages, and each stage calls the model. Here is the notional spend by stage across all 33 attempts.

$92.64Total over 33 attempts (from the note)

Parts sorted by value, largest first

  1. Act (edit and run)
  2. Research
  3. Verify
  4. Other
  5. Review
  6. Context compaction
  7. Onboarding notes

Shares are calculated from the values shown; rounding can make the sum of the parts differ from the stated total by a cent.

7 rows. Highest Act (edit and run) $44.05. Lowest Onboarding notes $2.09.

Notes

Share of notional model cost by stage, all 33 attempts

Total $92.64 over 33 attempts.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

  • Act (edit and run): $44.05. This is the core loop: change code, run it, read the output.
  • Research: $20.56. Reading the code and the issue before acting.
  • Verify: $8.72. Checking the change against tests and the plan.
  • Other: $6.24.
  • Review: $5.58. A self-review before delivery.
  • Context compaction: $5.41. Summarising a long context so the loop can continue.
  • Onboarding notes: $2.09. Repository notes, written once per repository.

Act and research are the two largest stages by far. Verification and review, the stages that make a pipeline different from a bare agent, are a smaller slice than you might think. If you want to cut cost, the first place to look is the edit-and-run loop, not the quality gates.

Where the token dollars go

Stages tell you why the model was called. Token kinds tell you what you paid for. We repriced every recorded token at Sonnet 5.5 list prices.

Calculation
$87.23Sum of the 4 parts

Parts sorted by value, largest first

  1. Cache writes (1 h)
  2. Cache reads
  3. Output
  4. Uncached input

Shares are calculated from the values shown.

List-price calculation, not a run. 4 rows. Highest Cache writes (1 h) $38.99. Lowest Uncached input $0.01.

Notes

Recorded tokens at Sonnet 5.5 list price, by token kind

List-price calculation on recorded tokens: 153.1M cache reads, 9.7M cache writes, 1.8M output, 3.2k uncached input. An agent loop re-reads its context on every call, so cache reads dominate the token count; by price the largest part is cache writes (1 h).

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)

Agent recorded 162.9M input tokens and 1.8M output tokens. Of the input, 153.1M were cache reads, 9.7M were one-hour cache writes, and only about 3.2k were uncached input. That is 94.0% of input served from cache.

By price, the picture flips:

  • Cache writes (1 h): $38.99. Few tokens, but Anthropic prices one-hour writes at twice the input rate.
  • Cache reads: $30.62. Many tokens at a low rate.
  • Output: $17.61. Only 1.8M tokens, but output is the most expensive rate per token.
  • Uncached input: $0.01.

The sum is $87.23. The platform's own notional figure is $92.64, because it also counts compaction calls.

Why so many cache reads? An agent loop re-reads its context on every call. Each step sends the conversation so far, plus the new tool output. With caching, most of that is cheap. Without caching, it is not.

What caching saved

Calculation
  • With caching (as recorded)
  • Without caching (square)
In chart order.
Claude Haiku 4.5
Claude Sonnet 5.5
Claude Opus 5.5
Claude Fable 5.1

Gap labels, Without caching vs With caching (as recorded): Without caching is x% higher (+) or lower (−) than With caching (as recorded), calculated from the two values shown (the change counted from With caching (as recorded)’s value).

List-price calculation, not a run. 4 rows, 2 series: With caching (as recorded), Without caching. With caching (as recorded): highest Claude Fable 5.1 $321. Lowest Claude Haiku 4.5 $43.61. Without caching: highest Claude Fable 5.1 $1,717. Lowest Claude Haiku 4.5 $172.

Notes

The same recorded tokens with and without cache pricing

94.0% of recorded input tokens were cache reads. Calculation, not a run.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)

At Sonnet 5.5 prices, the same tokens without cache pricing would cost $343.33 instead of $87.23. The pattern holds for every Claude model we priced: Haiku 4.5 goes from $43.61 to $171.66, Opus 5.5 from $143.83 to $686.66 and Fable 5.1 from $321.31 to $1,716.65.

This is a calculation on recorded tokens, not a run. But the lesson is practical. For a long agent loop, prompt caching is the single biggest cost lever you control. It matters more than the choice between two neighbouring models. We dig into model choice in what if every call ran on Opus?

The spread inside one run

Averages hide a lot. The per-instance table on the study page shows the range:

  • The cheapest attempt was sympy__sympy-18189 at $0.89. It was an empty patch after 1.6 minutes and 13 calls, a platform hold.
  • The most expensive was sphinx-doc__sphinx-10435 at $4.83, which resolved after 67 calls.
  • sympy__sympy-20428 cost $4.81 and did not resolve. Expensive failures exist, and they count.

Hard instances cost more, and they do not always pay off. That is why cost per resolved is the number to watch, not cost per attempt.

How we measured

  • Tokens. We summed every model call in the run telemetry of all 33 attempts: input, cache reads, one-hour cache writes and output.
  • Prices. List prices from our product price table, effective 2026-09-21. Anthropic one-hour cache writes are priced at twice the input price.
  • Formula. Uncached input × input price + cache reads × cache-read price + cache writes × write price + output × output price.
  • Cost per resolved. Every attempt's cost in the numerator, divided by the 25 resolved instances.
  • Panel. Published per-instance API costs for the same 33 instances.

Caveats

  • Notional costs. Agent ran on a subscription. Its costs are list-price estimates, not invoices.
  • Different scopes. The panel cost covers a bash agent. Agent's cost covers onboarding, planning, verification and review too.
  • n = 33. Small samples make per-instance averages noisy.
  • Repricing is a calculation. The no-cache and other-model figures reprice the same tokens. A different setup would use a different number of tokens.

See the receipts for your own work

Every Agent run records its calls, tokens and cost per stage, so you can see what each change cost and where the money went. Try Agent and look at the receipt for your first task.

The data behind this post

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.