What if every call ran on Opus? Repricing real agent tokens
Agent recorded every token it used on 33 SWE-bench instances. What would the same tokens cost at other models’ list prices, and what did caching save?
Published · 5 charts · Download the data or a carousel
162.9M
The answer
Calculation, not a run. Agent used 162.9M input tokens (94.0% cache reads) and 1.8M output tokens across 33 attempts. At Sonnet 5.5 list prices that is $87.23 ($3.49 per resolved instance); the platform's own notional figure, which also counts compaction calls, is $92.64. The same tokens at Opus 5.5 prices cost $143.83, at Fable 5.1 $321.31 and at Haiku 4.5 $43.61. Without prompt caching the Sonnet bill would be $343.33. A different model would have used different tokens and resolved a different set, so these figures bound price sensitivity; they do not predict outcomes.
Live story
Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.
What if every call ran on Opus? Repricing real agent tokens
A calculation, not a run: Agent's recorded SWE-bench tokens cost $87.23 at Sonnet 5.5 prices, $143.83 at Opus 5.5 and $343.33 without caching.
Transcript
- Thought experiment · recorded tokens × list prices. What if every call ran on Opus? Agent recorded every token on 33 SWE-bench attempts. We repriced them.
- 162.9M input tokens, 94.0% of them read from the prompt cache. Input tokens recorded: 162.9M (n = 33). Output tokens recorded: 1.8M (n = 33). Input served from cache: 94.0% (n = 33). Caveat: Recorded costs are list-price estimates for subscription calls; no invoice backs them.
- Same tokens on Opus 5.5: $143.83 instead of $87.23, 1.65× the bill. On Haiku 4.5: $43.61. At Haiku 4.5 prices: $43.61 (n = 33). At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
- Per resolved instance: $3.49 on Sonnet 5.5, $5.75 on Opus 5.5, $12.85 on Fable 5.1. Chart: Thought experiment: the same tokens at other list prices. Calculation, not a run. Caveat: Different models use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.
- In this calculation caching matters more than the model: without it, Sonnet would cost $343.33, 3.9× the recorded $87.23. Chart: Thought experiment: what prompt caching saved. Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
- Price sensitivity, not predictions. Every repricing labelled as a calculation.
Key numbers
1.8M
Output tokens recorded
n = 33
94.0%
Share of input served from cache
n = 33
$87.23
Recorded tokens at Sonnet 5.5 list price
n = 33
$143.83
Same tokens at Opus 5.5 list price (calculation)
n = 33
$43.61
Same tokens at Haiku 4.5 list price (calculation)
n = 33
$343.33
Sonnet 5.5 without caching (calculation)
n = 33
$0.57
Public panel mean cost per resolved instance (recorded)
n = 11
The charts
Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.
Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.
| Item | Repriced cost per resolved instance |
|---|---|
| Claude Fable 5.1 | $12.85 |
| Claude Opus 5 | $8.72 |
| Claude Opus 5.5 | $5.75 |
| Claude Sonnet 5.5 | $3.49 |
| GPT-6.1 Sol | $2.10 |
| Claude Haiku 4.5 | $1.75 |
| Gemini 3.x Flash | $1.02 |
| Jev 1.13 (router) | $0.042 |
List-price calculation, not a run. 8 rows. Highest Claude Fable 5.1 $12.85. Lowest Jev 1.13 (router) $0.042.
Notes
Cost per resolved SWE-bench instance if 162.9M input and 1.8M output tokens had been billed at each model's list price
Calculation, not a run: tokens recorded by Agent on claude-sonnet-5-5 (33 attempts, 25 resolved) times list prices effective 2026-09-21. Another model would use a different number of tokens and resolve a different set. Jev is a routing model and cannot do this work; its bar is a price floor only.
Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models), Google Gemini list prices, OpenAI list prices, Jev 1.13 list price
Hover or focus a bar for its ratio to Agent (notional) (the highlighted row): a ratio of the two values shown, not a measurement.
| Item | Cost per resolved instance | n |
|---|---|---|
| Agent (notional) | $3.71 | 25 |
| Claude 4.5 Opus (high) | $1.18 | 24 |
| Claude 4.5 Sonnet (high) | $0.91 | 25 |
| Claude 4.6 Opus | $0.88 | 23 |
| GLM 5 (high) | $0.67 | 26 |
| DeepSeek V3.2 (high) | $0.64 | 24 |
| GPT 5.2 (high) | $0.63 | 28 |
| Claude 4.5 Haiku (high) | $0.48 | 25 |
| Gemini 3 Flash (high) | $0.44 | 27 |
| Kimi K2.5 (high) | $0.26 | 23 |
| MiniMax M2.5 (high) | $0.11 | 23 |
| GPT 5 mini | $0.08 | 21 |
12 rows. Highest Agent (notional) $3.71 (n 25). Lowest GPT 5 mini $0.08 (n 21).
Notesn 21–28 per row
Same 33 SWE-bench Verified instances; all attempts in the numerator
Recorded figures, not repricing. Panel costs are published API costs for a bash-only agent. Agent's figure is a list-price estimate of subscription calls and includes onboarding, planning, verification and review.
Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs
- With caching (as recorded)
- Without caching (square)
Gap labels, Without caching vs With caching (as recorded): Without caching is x% higher (+) or lower (−) than With caching (as recorded), calculated from the two values shown (the change counted from With caching (as recorded)’s value).
| Item | With caching (as recorded) | Without caching |
|---|---|---|
| Claude Haiku 4.5 | $43.61 | $172 |
| Claude Sonnet 5.5 | $87.23 | $343 |
| Claude Opus 5.5 | $144 | $687 |
| Claude Fable 5.1 | $321 | $1,717 |
List-price calculation, not a run. 4 rows, 2 series: With caching (as recorded), Without caching. With caching (as recorded): highest Claude Fable 5.1 $321. Lowest Claude Haiku 4.5 $43.61. Without caching: highest Claude Fable 5.1 $1,717. Lowest Claude Haiku 4.5 $172.
Notes
The same recorded tokens with and without cache pricing
94.0% of recorded input tokens were cache reads. Calculation, not a run.
Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)
Parts sorted by value, largest first
- Cache writes (1 h)
- Cache reads
- Output
- Uncached input
Shares are calculated from the values shown.
| Item | Cost |
|---|---|
| Cache writes (1 h) | $38.99 |
| Cache reads | $30.62 |
| Output | $17.61 |
| Uncached input | $0.01 |
List-price calculation, not a run. 4 rows. Highest Cache writes (1 h) $38.99. Lowest Uncached input $0.01.
Notes
Recorded tokens at Sonnet 5.5 list price, by token kind
List-price calculation on recorded tokens: 153.1M cache reads, 9.7M cache writes, 1.8M output, 3.2k uncached input. An agent loop re-reads its context on every call, so cache reads dominate the token count; by price the largest part is cache writes (1 h).
Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)
Hover or focus a bar for its ratio to 1k-token decision prompt (the lowest value): a ratio of list-price calculations, not a measurement.
| Item | Jev decision cost |
|---|---|
| 1k-token decision prompt | $0.069 |
| 2k-token decision prompt | $0.14 |
| 5k-token decision prompt | $0.34 |
List-price calculation, not a run. 3 rows. Highest 5k-token decision prompt $0.34. Lowest 1k-token decision prompt $0.069.
Notes
1,632 model calls; one Jev decision per call at an assumed prompt size
Assumption-based calculation: the decision prompt sizes are assumptions, not measurements. For scale, the recorded work cost $92.64 (notional). A router only pays off if its choices save more than this.
Sources: Repricing calculation, Jev 1.13 list price, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)
Tables
Repricing table (calculation)
| Priced as | Vendor | Input $/M | Cache read $/M | Output $/M | All 33 attempts | Per attempt | Per resolved |
|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | Anthropic | $1.00 | $0.1 | $5.00 | $43.61 | $1.32 | $1.75 |
| Claude Sonnet 5.5 | Anthropic | $2.00 | $0.2 | $10.00 | $87.23 | $2.64 | $3.49 |
| Claude Opus 5.5 | Anthropic | $4.00 | $0.2 | $20.00 | $144 | $4.36 | $5.75 |
| Claude Opus 5 | Anthropic | $5.00 | $0.5 | $25.00 | $218 | $6.61 | $8.72 |
| Claude Fable 5.1 | Anthropic | $10.00 | $0.25 | $50.00 | $321 | $9.74 | $12.85 |
| Gemini 3.x Flash | $0.75 | $0.075 | $3.75 | $25.40 | $0.77 | $1.02 | |
| GPT-6.1 Sol | OpenAI | $2.00 | $0.1 | $10.00 | $52.42 | $1.59 | $2.10 |
| Jev 1.13 (router) | TypeSafe | $0.042 | $0.0042 | $0 | $1.05 | $0.032 | $0.042 |
Method
- Tokens: the sum over every model call in the run telemetry of all 33 SWE-bench attempts (input, cache reads, one-hour cache writes, output).
- Prices: list prices recorded in the product price table, effective 2026-09-21 (Jev 2026-09-23, OpenAI rows 2026-10-03).
- Formula: uncached input × input price + cache reads × cache-read price + cache writes × write price + output × output price. Anthropic one-hour writes cost twice the input price; other vendors’ writes are priced as plain input.
- Cost per resolved keeps every attempt’s cost in the numerator and divides by the 25 instances Agent resolved.
Caveats
- Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
- Different models use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.
- Jev is a routing model. Pricing coding tokens at Jev rates shows a floor, not a feasible configuration.
- The router-overhead chart uses assumed decision prompt sizes.
- Recorded costs are list-price estimates for subscription calls; no invoice backs them.
Sources
Repricing calculation
Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.
Agent on SWE-bench Verified, campaign 1 (25 instances)
Stratified sample of 25 Verified instances (seed 20261004), one attempt each, official grading harness. Fixed model claude-sonnet-5-5, platform build f0ac3a8a.
Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)
The 6 compiled-extension instances that campaign 1 could not run, plus 2 replacement candidates. One attempt each, platform build 236c0d3f.
SWE-bench Verified leaderboard, mini-SWE-agent v2 runs
Public per-instance results of 11 models under mini-SWE-agent 2.0.0 (bash only, one attempt). Costs are API list prices as published.
Anthropic list prices (Claude models)
Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.
Gemini 3.x Flash prices as listed by the vendor on 2026-09-21. The vendor announced a doubling from 2027-01-01.
Token prices as listed by the vendor on 2026-10-03.
Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free.
Download the data
The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.
Share it as a carousel
Square slides made in your browser from the charts on this page, with the same numbers, intervals and notes, and a captions file for alt text.
Cite as: Agent public benchmarks, “What if every call ran on Opus? Repricing real agent tokens”, updated October 5, 2026, https://agent.sasid.ai/benchmarks/cost-thought-experiments.
Explainers that cite this study
Read the methods and terms in the context of these recorded results.
More comparisons based on this study (58)
These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.
- Agent vs Claude Haiku 4.5
- Agent vs Claude Opus 4.5
- Agent vs Claude Opus 4.6
- Agent vs Claude Sonnet 4.5
- Agent vs DeepSeek V3.2
- Agent vs Gemini 3 Flash
- Agent vs GLM 5
- Agent vs GPT 5.2
- Agent vs GPT 5 mini
- Agent vs Kimi K2.5
- Agent vs MiniMax M2.5
- Claude Haiku 4.5 vs Kimi K2.5
- Claude Haiku 4.5 vs MiniMax M2.5
- Claude Opus 4.5 vs Claude Opus 4.6
- Claude Opus 4.5 vs DeepSeek V3.2
- Claude Opus 4.5 vs GPT 5 mini
- Claude Opus 4.5 vs Kimi K2.5
- Claude Opus 4.5 vs MiniMax M2.5
- Claude Opus 4.6 vs DeepSeek V3.2
- Claude Opus 4.6 vs GPT 5 mini
- Claude Opus 4.6 vs Kimi K2.5
- Claude Opus 4.6 vs MiniMax M2.5
- Claude Sonnet 4.5 vs Claude Opus 4.5
- Claude Sonnet 4.5 vs Claude Opus 4.6
- Claude Sonnet 4.5 vs DeepSeek V3.2
- Claude Sonnet 4.5 vs GPT 5 mini
- Claude Sonnet 4.5 vs Kimi K2.5
- Claude Sonnet 4.5 vs MiniMax M2.5
- DeepSeek V3.2 vs GPT 5 mini
- DeepSeek V3.2 vs Kimi K2.5
- DeepSeek V3.2 vs MiniMax M2.5
- Gemini 3 Flash vs Claude Opus 4.5
- Gemini 3 Flash vs Claude Opus 4.6
- Gemini 3 Flash vs Claude Sonnet 4.5
- Gemini 3 Flash vs DeepSeek V3.2
- Gemini 3 Flash vs GLM 5
- Gemini 3 Flash vs GPT 5 mini
- Gemini 3 Flash vs Kimi K2.5
- Gemini 3 Flash vs MiniMax M2.5
- GLM 5 vs Claude Opus 4.5
- GLM 5 vs Claude Opus 4.6
- GLM 5 vs Claude Sonnet 4.5
- GLM 5 vs DeepSeek V3.2
- GLM 5 vs GPT 5 mini
- GLM 5 vs Kimi K2.5
- GLM 5 vs MiniMax M2.5
- GPT 5.2 vs Claude Opus 4.5
- GPT 5.2 vs Claude Opus 4.6
- GPT 5.2 vs Claude Sonnet 4.5
- GPT 5.2 vs DeepSeek V3.2
- GPT 5.2 vs Gemini 3 Flash
- GPT 5.2 vs GLM 5
- GPT 5.2 vs GPT 5 mini
- GPT 5.2 vs Kimi K2.5
- GPT 5.2 vs MiniMax M2.5
- Kimi K2.5 vs GPT 5 mini
- MiniMax M2.5 vs GPT 5 mini
- MiniMax M2.5 vs Kimi K2.5
More write-ups that cite this study (2)
Models and comparisons in this study
Write-ups on this study
AI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
AI coding cost per developer: a formula built on recorded work
AI coding cost per developer: a monthly budget formula using recorded tokens and list-price calculations. Replace our task volume and pass rate with yours.
Claude Code cost per task: a price ladder from one decision to one agent run
$0.087 per Claude Code coding session at API list prices, $0.0036 per short call, $2.81 per full agent attempt. Calculations on recorded tokens, not bills.
Claude Fable 5.1 vs Opus 5.5 vs Sonnet 5.5: speed, tokens and price tested
24 of 24: Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 each passed every hard task. Fable cost 3.3x Opus and 6.5x Sonnet per pass (list-price calculation).
Claude Haiku 4.5 vs Sonnet 5.5: all 80 comparison rows, and where the small model loses
Haiku 4.5 vs Sonnet 5.5 on 80 rows: Sonnet ahead on 14, Haiku on none, 31 ties. Hard tasks 11/24 vs 24/24, plus speed, memory and price.
Devin's $0.60 per task and our $3.71 per resolved task are different numbers
Devin reports $0.60 per task. Our list-price calculation gives $3.71 per resolved SWE-bench task. Compare the units with a table and buyer checklist.
Does LLM routing save money? The saving, the router and the net
Routing would save 2.8% ($3.01) on 2,362 recorded calls (a calculation). A Sonnet router on every call costs about $11.80, so the net is a loss.
GPT-6.1 Sol vs Claude Sonnet 5.5 vs Opus 5.5: every row we measured
Of 70 comparison rows for GPT-6.1 Sol, Claude Sonnet 5.5 and Opus 5.5, only 2 have a winner (speed). All 15 pass-rate rows tie. Tokens, price and route differ.
Is Claude Haiku cheaper than Sonnet? Cost per correct answer, with retries
Haiku 4.5 lists at half the price of Sonnet 5.5, yet calculated cost per correct hard answer was 4.7x higher. Retry maths, two exceptions and every source.
LLM API pricing comparison, October 2026: Claude vs GPT vs Gemini per million tokens
21 LLM API prices per million tokens, October 2026: Claude, GPT, Gemini and more. Blended prices span 100x. Price per token is not price per task.
The cheapest way to run an AI coding agent: 7 levers from measured runs
7 levers that may cut an AI coding agent's bill, sized from our data: prompt cache 3.9x, Fable/Sonnet cost per pass 6.5x, and 5 more. List-price calculations.
What is the best AI model for coding? Our data says four models tie
4 models tied at the top of our hard coding set: Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol. Only Haiku 4.5 separated. A tier list built from intervals.
Which Claude model should you use? A task-by-task guide from our measurements
Which Claude model fits your job? Sonnet passed 24/24 hard calls (95% interval 86%–100%) at $0.01435 per pass (calculation). The top models hit a ceiling.
AI coding benchmarks roundup, October 2026: sixteen studies, every number in one place
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
Claude Sonnet 5.5 vs Opus 5.5: when is Opus worth the price?
Sonnet 5.5 and Opus 5.5 tied on every quality test we ran, easy, hard and agentic. Opus cost 1.6x to 2.6x per unit of work. Where the gap comes from.
How much does prompt caching actually save? Measured in Claude Code and Codex
Claude Code read 97% of later-turn input from the cache. At list price that halved a 5-turn session and cut an agent bill about 3.9x. Turn 1 costs more.
How to estimate your AI coding bill from real token mixes
Estimate AI coding costs from recorded token mixes: an agent task at $2.64, a hard call at $0.014, a routing decision at $0.005. Calculations, limits stated.
What if every call ran on Opus? Repricing real agent tokens across models
We repriced 162.9M recorded agent tokens at Haiku, Sonnet, Opus, Fable, Gemini Flash and GPT prices. A calculation, not a run, with clear limits.
What one resolved SWE-bench task really costs an AI coding agent
A full agent pipeline spent a notional $3.71 per resolved SWE-bench instance. Where the money went, what caching saved, and the public panel range.
More studies
All benchmarksPrompt cache break-even: after how many reuses does a cached prefix cost less?
A calculation on list prices: reuses before a cached prompt prefix costs less, per Claude and GPT model, and cost per 1,000 sessions of 1 to 20 turns.
Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts
Haiku 4.5 first, then retry or escalate to Sonnet 5.5? A calculation on 64 recorded calls across 8 hard tasks: cost and time per correct answer.
How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.