What one resolved SWE-bench task really costs an AI coding agent
A full agent pipeline spent a notional $3.71 per resolved SWE-bench instance. Where the money went, what caching saved, and the public panel range.
TL;DR
- Agent spent a notional $2.81 per attempt and $3.71 per resolved instance on 33 SWE-bench Verified instances. The resolved count is 25.
- The 11 public bash-only runs on the same instances spent between $0.08 and $1.18 per resolved instance. Their mean is $0.57.
- Of Agent's $92.64 total, $44.05 went to the act stage (editing and running code) and $20.56 to research.
- By token kind, cache writes cost the most ($38.99 at Sonnet 5.5 list price), then cache reads ($30.62), then output ($17.61).
- Without prompt caching, the same tokens would cost $343.33 instead of $87.23 at Sonnet 5.5 prices. That is a calculation, not a run.
Data and every attempt: /benchmarks/swe-bench-verified and /benchmarks/cost-thought-experiments. One attempt as a receipt: /why-agent/receipt.
The number most benchmarks leave out
Resolve rate gets the headline. Cost decides whether you can use the thing every day. So we count cost the honest way: every attempt's spend goes in the numerator, and only resolved instances go in the denominator. A failed attempt still costs money, so it raises the cost of each success.
Here is how that looks next to the public panel, on the same 33 instances.
Agent sits at $3.71 per resolved instance. The most expensive panel run is Claude 4.5 Opus (high) at $1.18. The cheapest is GPT 5 mini at $0.08, and MiniMax M2.5 (high) is close at $0.11. The resolved counts behind each bar differ, from 21 to 28.
So yes: on a per-success basis, a full pipeline costs several times what a bare bash agent costs. We think you should know that before you read any resolve rate, ours included.
Two things make the comparison less direct than the chart suggests:
- Different price sources. Panel costs are published API costs. Agent's cost is a list-price estimate of calls that actually ran on a flat subscription. No invoice backs it.
- Different work. Agent's cost includes onboarding the repository, planning, verification and review. A bare agent skips all of that.
Where an agent's money goes
The pipeline has stages, and each stage calls the model. Here is the notional spend by stage across all 33 attempts.
- Act (edit and run): $44.05. This is the core loop: change code, run it, read the output.
- Research: $20.56. Reading the code and the issue before acting.
- Verify: $8.72. Checking the change against tests and the plan.
- Other: $6.24.
- Review: $5.58. A self-review before delivery.
- Context compaction: $5.41. Summarising a long context so the loop can continue.
- Onboarding notes: $2.09. Repository notes, written once per repository.
Act and research are the two largest stages by far. Verification and review, the stages that make a pipeline different from a bare agent, are a smaller slice than you might think. If you want to cut cost, the first place to look is the edit-and-run loop, not the quality gates.
Where the token dollars go
Stages tell you why the model was called. Token kinds tell you what you paid for. We repriced every recorded token at Sonnet 5.5 list prices.
Agent recorded 162.9M input tokens and 1.8M output tokens. Of the input, 153.1M were cache reads, 9.7M were one-hour cache writes, and only about 3.2k were uncached input. That is 94.0% of input served from cache.
By price, the picture flips:
- Cache writes (1 h): $38.99. Few tokens, but Anthropic prices one-hour writes at twice the input rate.
- Cache reads: $30.62. Many tokens at a low rate.
- Output: $17.61. Only 1.8M tokens, but output is the most expensive rate per token.
- Uncached input: $0.01.
The sum is $87.23. The platform's own notional figure is $92.64, because it also counts compaction calls.
Why so many cache reads? An agent loop re-reads its context on every call. Each step sends the conversation so far, plus the new tool output. With caching, most of that is cheap. Without caching, it is not.
What caching saved
At Sonnet 5.5 prices, the same tokens without cache pricing would cost $343.33 instead of $87.23. The pattern holds for every Claude model we priced: Haiku 4.5 goes from $43.61 to $171.66, Opus 5.5 from $143.83 to $686.66 and Fable 5.1 from $321.31 to $1,716.65.
This is a calculation on recorded tokens, not a run. But the lesson is practical. For a long agent loop, prompt caching is the single biggest cost lever you control. It matters more than the choice between two neighbouring models. We dig into model choice in what if every call ran on Opus?
The spread inside one run
Averages hide a lot. The per-instance table on the study page shows the range:
- The cheapest attempt was sympy__sympy-18189 at $0.89. It was an empty patch after 1.6 minutes and 13 calls, a platform hold.
- The most expensive was sphinx-doc__sphinx-10435 at $4.83, which resolved after 67 calls.
- sympy__sympy-20428 cost $4.81 and did not resolve. Expensive failures exist, and they count.
Hard instances cost more, and they do not always pay off. That is why cost per resolved is the number to watch, not cost per attempt.
How we measured
- Tokens. We summed every model call in the run telemetry of all 33 attempts: input, cache reads, one-hour cache writes and output.
- Prices. List prices from our product price table, effective 2026-09-21. Anthropic one-hour cache writes are priced at twice the input price.
- Formula. Uncached input × input price + cache reads × cache-read price + cache writes × write price + output × output price.
- Cost per resolved. Every attempt's cost in the numerator, divided by the 25 resolved instances.
- Panel. Published per-instance API costs for the same 33 instances.
Caveats
- Notional costs. Agent ran on a subscription. Its costs are list-price estimates, not invoices.
- Different scopes. The panel cost covers a bash agent. Agent's cost covers onboarding, planning, verification and review too.
- n = 33. Small samples make per-instance averages noisy.
- Repricing is a calculation. The no-cache and other-model figures reprice the same tokens. A different setup would use a different number of tokens.
What to read next
- SWE-bench Verified: Agent vs 11 public models on the same tasks
- What if every call ran on Opus? A thought experiment
- Harness vs model: where the gains come from
See the receipts for your own work
Every Agent run records its calls, tokens and cost per stage, so you can see what each change cost and where the money went. Try Agent and look at the receipt for your first task.