• Thought experiment
  • LLM pricing
  • Opus
  • Sonnet

What if every call ran on Opus? Repricing real agent tokens across models

We repriced 162.9M recorded agent tokens at Haiku, Sonnet, Opus, Fable, Gemini Flash and GPT prices. A calculation, not a run, with clear limits.

TL;DR

  • This is a thought experiment. We took the tokens Agent actually recorded and multiplied them by other models' list prices. No other model ran.
  • The 33 SWE-bench attempts used 162.9M input tokens (94.0% cache reads) and 1.8M output tokens.
  • At list prices, those tokens cost $43.61 on Haiku 4.5, $87.23 on Sonnet 5.5, $143.83 on Opus 5.5 and $321.31 on Fable 5.1.
  • Without prompt caching, the Sonnet figure would be $343.33.
  • On a second dataset of 2,362 calls, routing side jobs to Haiku saved only 2.8%. A policy that sends strong stages to Opus cost 1.49x the all-Sonnet figure.

Full tables and formulas: /benchmarks/cost-thought-experiments.

Why a thought experiment?

Running the same 33 SWE-bench instances on every model is the right experiment. It is also slow and expensive, and we have not done it yet. Meanwhile, there is a cheaper question we can answer exactly: if the same work had been billed at another model's prices, what would it cost?

That tells you how sensitive the bill is to the price list. It does not tell you what another model would achieve. A different model would make different calls, use a different number of tokens and resolve a different set of issues. Every chart in this post carries that label, and so should every quote of it.

The same tokens at eight price lists

Calculation
Largest value is 310x the smallest; Log shows the small bars.
Claude Fable 5.1
Claude Opus 5
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol
Claude Haiku 4.5
Gemini 3.x Flash
Jev 1.13 (router)

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 8 rows. Highest Claude Fable 5.1 $12.85. Lowest Jev 1.13 (router) $0.042.

Notes

Cost per resolved SWE-bench instance if 162.9M input and 1.8M output tokens had been billed at each model's list price

Calculation, not a run: tokens recorded by Agent on claude-sonnet-5-5 (33 attempts, 25 resolved) times list prices effective 2026-09-21. Another model would use a different number of tokens and resolve a different set. Jev is a routing model and cannot do this work; its bar is a price floor only.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models), Google Gemini list prices, OpenAI list prices, Jev 1.13 list price

Live story · 33 sWhat if every call ran on Opus? Repricing real agent tokens

What if every call ran on Opus? Repricing real agent tokens

A calculation, not a run: Agent's recorded SWE-bench tokens cost $87.23 at Sonnet 5.5 prices, $143.83 at Opus 5.5 and $343.33 without caching.

Transcript
  1. Thought experiment · recorded tokens × list prices. What if every call ran on Opus? Agent recorded every token on 33 SWE-bench attempts. We repriced them.
  2. 162.9M input tokens, 94.0% of them read from the prompt cache. Input tokens recorded: 162.9M (n = 33). Output tokens recorded: 1.8M (n = 33). Input served from cache: 94.0% (n = 33). Caveat: Recorded costs are list-price estimates for subscription calls; no invoice backs them.
  3. Same tokens on Opus 5.5: $143.83 instead of $87.23, 1.65× the bill. On Haiku 4.5: $43.61. At Haiku 4.5 prices: $43.61 (n = 33). At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  4. Per resolved instance: $3.49 on Sonnet 5.5, $5.75 on Opus 5.5, $12.85 on Fable 5.1. Chart: Thought experiment: the same tokens at other list prices. Calculation, not a run. Caveat: Different models use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.
  5. In this calculation caching matters more than the model: without it, Sonnet would cost $343.33, 3.9× the recorded $87.23. Chart: Thought experiment: what prompt caching saved. Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  6. Price sensitivity, not predictions. Every repricing labelled as a calculation.

Cost per resolved instance, with all 33 attempts in the numerator and 25 resolved in the denominator:

  • Claude Fable 5.1: $12.85
  • Claude Opus 5: $8.72
  • Claude Opus 5.5: $5.75
  • Claude Sonnet 5.5: $3.49, the model that actually produced these tokens
  • GPT-6.1 Sol: $2.10
  • Claude Haiku 4.5: $1.75
  • Gemini 3.x Flash: $1.02
  • Jev 1.13: $0.04, a price floor only, because Jev is a router and cannot do this work

Why Opus 5.5 does not cost twice as much as Sonnet

On paper, Opus 5.5 lists input at $4 per million tokens and output at $20, twice Sonnet 5.5's $2 and $10. So why does the repriced total go from $87.23 to $143.83, not to about double?

Look at the cache-read price. In our price table, Opus 5.5 cache reads cost $0.20 per million, the same as Sonnet 5.5. An agent loop is mostly cache reads: 153.1M of the 162.9M input tokens. So the biggest token bucket costs the same on both models, and the gap narrows.

Compare Opus 5, the previous version. Its cache reads list at $0.50 per million. The same tokens cost $218.07 on Opus 5 against $143.83 on Opus 5.5. For agent workloads, the cache-read price can matter more than the headline input price. Check it before you choose a model.

Fable 5.1 shows the other side. Its input lists at $10 per million and output at $50. Even with a $0.25 cache-read price, the cache writes and output push it to $321.31.

What caching is worth

Calculation
  • With caching (as recorded)
  • Without caching (square)
In chart order.
Claude Haiku 4.5
Claude Sonnet 5.5
Claude Opus 5.5
Claude Fable 5.1

Gap labels, Without caching vs With caching (as recorded): Without caching is x% higher (+) or lower (−) than With caching (as recorded), calculated from the two values shown (the change counted from With caching (as recorded)’s value).

List-price calculation, not a run. 4 rows, 2 series: With caching (as recorded), Without caching. With caching (as recorded): highest Claude Fable 5.1 $321. Lowest Claude Haiku 4.5 $43.61. Without caching: highest Claude Fable 5.1 $1,717. Lowest Claude Haiku 4.5 $172.

Notes

The same recorded tokens with and without cache pricing

94.0% of recorded input tokens were cache reads. Calculation, not a run.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)

Turn caching off in the calculation and every bill multiplies:

  • Haiku 4.5: $43.61 → $171.66
  • Sonnet 5.5: $87.23 → $343.33
  • Opus 5.5: $143.83 → $686.66
  • Fable 5.1: $321.31 → $1,716.65

Sonnet with no caching would cost more than Opus 5.5 with caching. If you run long agent loops, make sure caching works before you argue about models.

Does routing save money?

Routing means using different models for different steps: a strong model for coding, a cheap one for side jobs. It sounds like an obvious saving. We tested it as a calculation on a second dataset: 50 benchmark runs and 2,362 model calls, all of which ran on Sonnet 5.5 with routing off. 93.0% of that input was cache reads.

Calculation
all Fable 5.1
all Opus 5.5
policy (Opus strong, Haiku ancillary)
all Sonnet 5.5
split (Sonnet main line, Haiku ancillary)
all Haiku 4.5

Hover or focus a bar for its ratio to all Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 6 rows. Highest all Fable 5.1 $370. Lowest all Haiku 4.5 $54.27.

Notes

50 benchmark runs, 2,362 model calls, repriced

Calculation, not a run: every recorded call ran on Sonnet 5.5 with routing off. Same tokens on every model; a different model or mix would take a different path.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Anthropic list prices (Claude models)

  • All Fable 5.1: $369.58
  • All Opus 5.5: $170.92
  • Policy (Opus on strong stages, Haiku on side jobs): $161.62, which is 1.49x all-Sonnet
  • All Sonnet 5.5: $108.54
  • Split (Sonnet on the main line, Haiku on side jobs): $105.53, which is 2.8% less than all-Sonnet
  • All Haiku 4.5: $54.27

The split saves almost nothing. The reason is in the next chart.

Calculation

Haiku 4.5

act (strong)
research (strong)
verify (strong)
review (strong)
memory and onboarding (economy)
other (standard)

Sonnet 5.5

act (strong)
research (strong)
verify (strong)
review (strong)
memory and onboarding (economy)
other (standard)

Opus 5.5

act (strong)
research (strong)
verify (strong)
review (strong)
memory and onboarding (economy)
other (standard)

Fable 5.1

act (strong)
research (strong)
verify (strong)
review (strong)
memory and onboarding (economy)
other (standard)

One panel per series, all on the same axis.

List-price calculation, not a run. 6 rows, 4 series: Haiku 4.5, Sonnet 5.5, Opus 5.5, Fable 5.1. Haiku 4.5: highest act (strong) $28.66. Lowest other (standard) $0.1. Sonnet 5.5: highest act (strong) $57.31. Lowest other (standard) $0.19.

Notes

Calculation, not a run. The tier in brackets is the routing policy tier for that stage.

Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Anthropic list prices (Claude models)

At Sonnet prices, the act stage alone costs $57.31 and research $28.62. Memory and onboarding, the "economy" side job, cost $6.22. Moving a $6 slice to a cheaper model cannot save much on a $108 bill.

So the economics of routing are not "send easy work to small models". They are: find where the money goes, and route that. On our pipeline, the money is in the main coding loop. Any real saving has to come from there, and it has to keep quality, which this calculation cannot show.

There is a second cost too: the router itself. If a router like Jev made one decision before each of the 1,632 calls in the SWE-bench run, the decisions would cost about $0.07 to $0.34 in total, depending on an assumed prompt size of 1k to 5k tokens. That is small against the $92.64 recorded spend. A general LLM router would cost far more per decision, as we show in Jev vs Claude as a router. With measured per-decision costs instead of assumed prompt sizes, routing every call adds $1.67 per 1,000 tasks with Jev and $247.30 with Sonnet 5.5 through its CLI (a calculation): what does a router cost you?

How we measured

  • Tokens. The sum over every model call in the run telemetry: input, cache reads, one-hour cache writes and output.
  • Prices. List prices from our product price table, effective 2026-09-21 (Jev 2026-09-23, OpenAI 2026-10-03).
  • Formula. Uncached input × input price + cache reads × cache-read price + cache writes × write price + output × output price. Anthropic one-hour writes cost twice the input price. Other vendors' writes are priced as plain input.
  • Cost per resolved. Every attempt's cost in the numerator, divided by the 25 resolved instances.
  • Routing scenarios. The same per-stage tokens repriced under each model mix. The tier for each stage comes from the platform routing policy.

Caveats

  • A calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  • Different models behave differently. They use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.
  • Jev is a router. Pricing coding tokens at Jev rates shows a floor, not a feasible configuration.
  • Assumed router prompts. The router-overhead figures use assumed decision prompt sizes.
  • No invoices. Recorded costs are list-price estimates for subscription calls.

Know where your tokens go

Agent records tokens, cache hits and cost for every stage of every task. That is how we could run this thought experiment at all. Try Agent and get the same receipts for your own work.

The data behind this post

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.