What if every call ran on Opus? Repricing real agent tokens across models
We repriced 162.9M recorded agent tokens at Haiku, Sonnet, Opus, Fable, Gemini Flash and GPT prices. A calculation, not a run, with clear limits.
TL;DR
- This is a thought experiment. We took the tokens Agent actually recorded and multiplied them by other models' list prices. No other model ran.
- The 33 SWE-bench attempts used 162.9M input tokens (94.0% cache reads) and 1.8M output tokens.
- At list prices, those tokens cost $43.61 on Haiku 4.5, $87.23 on Sonnet 5.5, $143.83 on Opus 5.5 and $321.31 on Fable 5.1.
- Without prompt caching, the Sonnet figure would be $343.33.
- On a second dataset of 2,362 calls, routing side jobs to Haiku saved only 2.8%. A policy that sends strong stages to Opus cost 1.49x the all-Sonnet figure.
Full tables and formulas: /benchmarks/cost-thought-experiments.
Why a thought experiment?
Running the same 33 SWE-bench instances on every model is the right experiment. It is also slow and expensive, and we have not done it yet. Meanwhile, there is a cheaper question we can answer exactly: if the same work had been billed at another model's prices, what would it cost?
That tells you how sensitive the bill is to the price list. It does not tell you what another model would achieve. A different model would make different calls, use a different number of tokens and resolve a different set of issues. Every chart in this post carries that label, and so should every quote of it.
The same tokens at eight price lists
What if every call ran on Opus? Repricing real agent tokens
A calculation, not a run: Agent's recorded SWE-bench tokens cost $87.23 at Sonnet 5.5 prices, $143.83 at Opus 5.5 and $343.33 without caching.
Transcript
- Thought experiment · recorded tokens × list prices. What if every call ran on Opus? Agent recorded every token on 33 SWE-bench attempts. We repriced them.
- 162.9M input tokens, 94.0% of them read from the prompt cache. Input tokens recorded: 162.9M (n = 33). Output tokens recorded: 1.8M (n = 33). Input served from cache: 94.0% (n = 33). Caveat: Recorded costs are list-price estimates for subscription calls; no invoice backs them.
- Same tokens on Opus 5.5: $143.83 instead of $87.23, 1.65× the bill. On Haiku 4.5: $43.61. At Haiku 4.5 prices: $43.61 (n = 33). At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
- Per resolved instance: $3.49 on Sonnet 5.5, $5.75 on Opus 5.5, $12.85 on Fable 5.1. Chart: Thought experiment: the same tokens at other list prices. Calculation, not a run. Caveat: Different models use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.
- In this calculation caching matters more than the model: without it, Sonnet would cost $343.33, 3.9× the recorded $87.23. Chart: Thought experiment: what prompt caching saved. Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
- Price sensitivity, not predictions. Every repricing labelled as a calculation.
Cost per resolved instance, with all 33 attempts in the numerator and 25 resolved in the denominator:
- Claude Fable 5.1: $12.85
- Claude Opus 5: $8.72
- Claude Opus 5.5: $5.75
- Claude Sonnet 5.5: $3.49, the model that actually produced these tokens
- GPT-6.1 Sol: $2.10
- Claude Haiku 4.5: $1.75
- Gemini 3.x Flash: $1.02
- Jev 1.13: $0.04, a price floor only, because Jev is a router and cannot do this work
Why Opus 5.5 does not cost twice as much as Sonnet
On paper, Opus 5.5 lists input at $4 per million tokens and output at $20, twice Sonnet 5.5's $2 and $10. So why does the repriced total go from $87.23 to $143.83, not to about double?
Look at the cache-read price. In our price table, Opus 5.5 cache reads cost $0.20 per million, the same as Sonnet 5.5. An agent loop is mostly cache reads: 153.1M of the 162.9M input tokens. So the biggest token bucket costs the same on both models, and the gap narrows.
Compare Opus 5, the previous version. Its cache reads list at $0.50 per million. The same tokens cost $218.07 on Opus 5 against $143.83 on Opus 5.5. For agent workloads, the cache-read price can matter more than the headline input price. Check it before you choose a model.
Fable 5.1 shows the other side. Its input lists at $10 per million and output at $50. Even with a $0.25 cache-read price, the cache writes and output push it to $321.31.
What caching is worth
Turn caching off in the calculation and every bill multiplies:
- Haiku 4.5: $43.61 → $171.66
- Sonnet 5.5: $87.23 → $343.33
- Opus 5.5: $143.83 → $686.66
- Fable 5.1: $321.31 → $1,716.65
Sonnet with no caching would cost more than Opus 5.5 with caching. If you run long agent loops, make sure caching works before you argue about models.
Does routing save money?
Routing means using different models for different steps: a strong model for coding, a cheap one for side jobs. It sounds like an obvious saving. We tested it as a calculation on a second dataset: 50 benchmark runs and 2,362 model calls, all of which ran on Sonnet 5.5 with routing off. 93.0% of that input was cache reads.
- All Fable 5.1: $369.58
- All Opus 5.5: $170.92
- Policy (Opus on strong stages, Haiku on side jobs): $161.62, which is 1.49x all-Sonnet
- All Sonnet 5.5: $108.54
- Split (Sonnet on the main line, Haiku on side jobs): $105.53, which is 2.8% less than all-Sonnet
- All Haiku 4.5: $54.27
The split saves almost nothing. The reason is in the next chart.
At Sonnet prices, the act stage alone costs $57.31 and research $28.62. Memory and onboarding, the "economy" side job, cost $6.22. Moving a $6 slice to a cheaper model cannot save much on a $108 bill.
So the economics of routing are not "send easy work to small models". They are: find where the money goes, and route that. On our pipeline, the money is in the main coding loop. Any real saving has to come from there, and it has to keep quality, which this calculation cannot show.
There is a second cost too: the router itself. If a router like Jev made one decision before each of the 1,632 calls in the SWE-bench run, the decisions would cost about $0.07 to $0.34 in total, depending on an assumed prompt size of 1k to 5k tokens. That is small against the $92.64 recorded spend. A general LLM router would cost far more per decision, as we show in Jev vs Claude as a router. With measured per-decision costs instead of assumed prompt sizes, routing every call adds $1.67 per 1,000 tasks with Jev and $247.30 with Sonnet 5.5 through its CLI (a calculation): what does a router cost you?
How we measured
- Tokens. The sum over every model call in the run telemetry: input, cache reads, one-hour cache writes and output.
- Prices. List prices from our product price table, effective 2026-09-21 (Jev 2026-09-23, OpenAI 2026-10-03).
- Formula. Uncached input × input price + cache reads × cache-read price + cache writes × write price + output × output price. Anthropic one-hour writes cost twice the input price. Other vendors' writes are priced as plain input.
- Cost per resolved. Every attempt's cost in the numerator, divided by the 25 resolved instances.
- Routing scenarios. The same per-stage tokens repriced under each model mix. The tier for each stage comes from the platform routing policy.
Caveats
- A calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
- Different models behave differently. They use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.
- Jev is a router. Pricing coding tokens at Jev rates shows a floor, not a feasible configuration.
- Assumed router prompts. The router-overhead figures use assumed decision prompt sizes.
- No invoices. Recorded costs are list-price estimates for subscription calls.
What to read next
- Claude Sonnet vs Opus: when is Opus worth the price?
- How to estimate your AI coding bill
- What one resolved SWE-bench task really costs
- Jev vs Claude Haiku and Sonnet as a router
- What does a router cost you? Rules vs Jev vs an LLM router
- Claude Haiku vs Sonnet vs Opus vs Fable vs Codex, head to head
Know where your tokens go
Agent records tokens, cache hits and cost for every stage of every task. That is how we could run this thought experiment at all. Try Agent and get the same receipts for your own work.