• Thought experiment
  • LLM pricing
  • Prompt Caching
  • Opus
  • Sonnet
  • Haiku
  • Jev

What if every call ran on Opus? Repricing real agent tokens

Agent recorded every token it used on 33 SWE-bench instances. What would the same tokens cost at other models’ list prices, and what did caching save?

Published · 5 charts · Download the data or a carousel

162.9M

n = 33

Input tokens recorded

The answer

Calculation, not a run. Agent used 162.9M input tokens (94.0% cache reads) and 1.8M output tokens across 33 attempts. At Sonnet 5.5 list prices that is $87.23 ($3.49 per resolved instance); the platform's own notional figure, which also counts compaction calls, is $92.64. The same tokens at Opus 5.5 prices cost $143.83, at Fable 5.1 $321.31 and at Haiku 4.5 $43.61. Without prompt caching the Sonnet bill would be $343.33. A different model would have used different tokens and resolved a different set, so these figures bound price sensitivity; they do not predict outcomes.

Live story

Drawn live in the page from the same data as the charts below. Play it, or download it as a video from the player.

Live story · 33 sWhat if every call ran on Opus? Repricing real agent tokens

What if every call ran on Opus? Repricing real agent tokens

A calculation, not a run: Agent's recorded SWE-bench tokens cost $87.23 at Sonnet 5.5 prices, $143.83 at Opus 5.5 and $343.33 without caching.

Transcript
  1. Thought experiment · recorded tokens × list prices. What if every call ran on Opus? Agent recorded every token on 33 SWE-bench attempts. We repriced them.
  2. 162.9M input tokens, 94.0% of them read from the prompt cache. Input tokens recorded: 162.9M (n = 33). Output tokens recorded: 1.8M (n = 33). Input served from cache: 94.0% (n = 33). Caveat: Recorded costs are list-price estimates for subscription calls; no invoice backs them.
  3. Same tokens on Opus 5.5: $143.83 instead of $87.23, 1.65× the bill. On Haiku 4.5: $43.61. At Haiku 4.5 prices: $43.61 (n = 33). At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  4. Per resolved instance: $3.49 on Sonnet 5.5, $5.75 on Opus 5.5, $12.85 on Fable 5.1. Chart: Thought experiment: the same tokens at other list prices. Calculation, not a run. Caveat: Different models use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.
  5. In this calculation caching matters more than the model: without it, Sonnet would cost $343.33, 3.9× the recorded $87.23. Chart: Thought experiment: what prompt caching saved. Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  6. Price sensitivity, not predictions. Every repricing labelled as a calculation.

Key numbers

1.8M

Output tokens recorded

n = 33

94.0%

Share of input served from cache

n = 33

$87.23

Recorded tokens at Sonnet 5.5 list price

n = 33

$143.83

Same tokens at Opus 5.5 list price (calculation)

n = 33

$43.61

Same tokens at Haiku 4.5 list price (calculation)

n = 33

$343.33

Sonnet 5.5 without caching (calculation)

n = 33

$0.57

Public panel mean cost per resolved instance (recorded)

n = 11

The charts

Hover or focus a row for its exact value, interval and sample. Each chart has a Table view.

Thought experiment: not a run. These values reprice recorded tokens at list prices. No model was called again.

Calculation
Largest value is 310x the smallest; Log shows the small bars.
Claude Fable 5.1
Claude Opus 5
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol
Claude Haiku 4.5
Gemini 3.x Flash
Jev 1.13 (router)

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the highlighted row): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 8 rows. Highest Claude Fable 5.1 $12.85. Lowest Jev 1.13 (router) $0.042.

Notes

Cost per resolved SWE-bench instance if 162.9M input and 1.8M output tokens had been billed at each model's list price

Calculation, not a run: tokens recorded by Agent on claude-sonnet-5-5 (33 attempts, 25 resolved) times list prices effective 2026-09-21. Another model would use a different number of tokens and resolve a different set. Jev is a routing model and cannot do this work; its bar is a price floor only.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models), Google Gemini list prices, OpenAI list prices, Jev 1.13 list price

Share card (PNG)
Largest value is 46x the smallest; Log shows the small bars.
Agent (notional)
Claude 4.5 Opus (high)
Claude 4.5 Sonnet (high)
Claude 4.6 Opus
GLM 5 (high)
DeepSeek V3.2 (high)
GPT 5.2 (high)
Claude 4.5 Haiku (high)
Gemini 3 Flash (high)
Kimi K2.5 (high)
MiniMax M2.5 (high)
GPT 5 mini

Hover or focus a bar for its ratio to Agent (notional) (the highlighted row): a ratio of the two values shown, not a measurement.

12 rows. Highest Agent (notional) $3.71 (n 25). Lowest GPT 5 mini $0.08 (n 21).

Notesn 21–28 per row

Same 33 SWE-bench Verified instances; all attempts in the numerator

Recorded figures, not repricing. Panel costs are published API costs for a bash-only agent. Agent's figure is a list-price estimate of subscription calls and includes onboarding, planning, verification and review.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

Share card (PNG)
Calculation
  • With caching (as recorded)
  • Without caching (square)
In chart order.
Claude Haiku 4.5
Claude Sonnet 5.5
Claude Opus 5.5
Claude Fable 5.1

Gap labels, Without caching vs With caching (as recorded): Without caching is x% higher (+) or lower (−) than With caching (as recorded), calculated from the two values shown (the change counted from With caching (as recorded)’s value).

List-price calculation, not a run. 4 rows, 2 series: With caching (as recorded), Without caching. With caching (as recorded): highest Claude Fable 5.1 $321. Lowest Claude Haiku 4.5 $43.61. Without caching: highest Claude Fable 5.1 $1,717. Lowest Claude Haiku 4.5 $172.

Notes

The same recorded tokens with and without cache pricing

94.0% of recorded input tokens were cache reads. Calculation, not a run.

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)

Share card (PNG)
Calculation
$87.23Sum of the 4 parts

Parts sorted by value, largest first

  1. Cache writes (1 h)
  2. Cache reads
  3. Output
  4. Uncached input

Shares are calculated from the values shown.

List-price calculation, not a run. 4 rows. Highest Cache writes (1 h) $38.99. Lowest Uncached input $0.01.

Notes

Recorded tokens at Sonnet 5.5 list price, by token kind

List-price calculation on recorded tokens: 153.1M cache reads, 9.7M cache writes, 1.8M output, 3.2k uncached input. An agent loop re-reads its context on every call, so cache reads dominate the token count; by price the largest part is cache writes (1 h).

Sources: Repricing calculation, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), Anthropic list prices (Claude models)

Share card (PNG)
Calculation
1k-token decision prompt
2k-token decision prompt
5k-token decision prompt

Hover or focus a bar for its ratio to 1k-token decision prompt (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 3 rows. Highest 5k-token decision prompt $0.34. Lowest 1k-token decision prompt $0.069.

Notes

1,632 model calls; one Jev decision per call at an assumed prompt size

Assumption-based calculation: the decision prompt sizes are assumptions, not measurements. For scale, the recorded work cost $92.64 (notional). A router only pays off if its choices save more than this.

Sources: Repricing calculation, Jev 1.13 list price, Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

Share card (PNG)

Tables

Repricing table (calculation)

Priced asVendorInput $/MCache read $/MOutput $/MAll 33 attemptsPer attemptPer resolved
Claude Haiku 4.5Anthropic$1.00$0.1$5.00$43.61$1.32$1.75
Claude Sonnet 5.5Anthropic$2.00$0.2$10.00$87.23$2.64$3.49
Claude Opus 5.5Anthropic$4.00$0.2$20.00$144$4.36$5.75
Claude Opus 5Anthropic$5.00$0.5$25.00$218$6.61$8.72
Claude Fable 5.1Anthropic$10.00$0.25$50.00$321$9.74$12.85
Gemini 3.x FlashGoogle$0.75$0.075$3.75$25.40$0.77$1.02
GPT-6.1 SolOpenAI$2.00$0.1$10.00$52.42$1.59$2.10
Jev 1.13 (router)TypeSafe$0.042$0.0042$0$1.05$0.032$0.042

Method

  1. Tokens: the sum over every model call in the run telemetry of all 33 SWE-bench attempts (input, cache reads, one-hour cache writes, output).
  2. Prices: list prices recorded in the product price table, effective 2026-09-21 (Jev 2026-09-23, OpenAI rows 2026-10-03).
  3. Formula: uncached input × input price + cache reads × cache-read price + cache writes × write price + output × output price. Anthropic one-hour writes cost twice the input price; other vendors’ writes are priced as plain input.
  4. Cost per resolved keeps every attempt’s cost in the numerator and divides by the 25 instances Agent resolved.

Caveats

  • Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  • Different models use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.
  • Jev is a routing model. Pricing coding tokens at Jev rates shows a floor, not a feasible configuration.
  • The router-overhead chart uses assumed decision prompt sizes.
  • Recorded costs are list-price estimates for subscription calls; no invoice backs them.

Sources

  • Repricing calculation

    Calculation ·

    Recorded token counts multiplied by the list prices in the price-list sources above. A calculation, not a run: a different model would have used a different number of tokens and reached different outcomes.

  • Agent on SWE-bench Verified, campaign 1 (25 instances)

    Our recorded runs ·

    Stratified sample of 25 Verified instances (seed 20261004), one attempt each, official grading harness. Fixed model claude-sonnet-5-5, platform build f0ac3a8a.

    Raw data: swebench/attempts.json, swebench/exclusions.json

  • Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

    Our recorded runs ·

    The 6 compiled-extension instances that campaign 1 could not run, plus 2 replacement candidates. One attempt each, platform build 236c0d3f.

    Raw data: swebench/attempts.json

  • SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

    Public leaderboard ·

    Public per-instance results of 11 models under mini-SWE-agent 2.0.0 (bash only, one attempt). Costs are API list prices as published.

    Raw data: swebench/panel.json

  • Anthropic list prices (Claude models)

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

  • Google Gemini list prices

    Vendor price list ·

    Gemini 3.x Flash prices as listed by the vendor on 2026-09-21. The vendor announced a doubling from 2027-01-01.

  • OpenAI list prices

    Vendor price list ·

    Token prices as listed by the vendor on 2026-10-03.

  • Jev 1.13 list price

    Vendor price list ·

    Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free.

Download the data

The study as JSON (with its sources), every chart point as CSV, or the whole public dataset. Free to reuse under CC BY 4.0: credit Agent and link the study.

Cite as: Agent public benchmarks, “What if every call ran on Opus? Repricing real agent tokens”, updated October 5, 2026, https://agent.sasid.ai/benchmarks/cost-thought-experiments.

Explainers that cite this study

Read the methods and terms in the context of these recorded results.

More comparisons based on this study (58)

These pages reuse this study’s recorded rows. Read each page’s original sample, comparability and ceiling limits.

More write-ups that cite this study (2)

Models and comparisons in this study

More studies

All benchmarks
Includes calculations
  • Prompt Caching
  • Break Even

Prompt cache break-even: after how many reuses does a cached prefix cost less?

A calculation on list prices: reuses before a cached prompt prefix costs less, per Claude and GPT model, and cost per 1,000 sessions of 1 to 20 turns.

2reuses (the 3rd request) · Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation)

3 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • Claude Haiku

Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts

Haiku 4.5 first, then retry or escalate to Sonnet 5.5? A calculation on 64 recorded calls across 8 hard tasks: cost and time per correct answer.

46% (11/24)Strict passes, Claude Haiku 4.5 · Claude Code, eight hard tasks · n = 24

6 chartsUpdated October 7, 2026

Includes calculations
  • Thought experiment
  • Calculation

How much of an AI bill is thinking? Reasoning tokens by model and effort

Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.

92%(Claude Haiku 4.5 · Claude Code; range 76% to 99%) · Highest median reasoning share of output tokens, hard tasks (calculation) · n = 24

5 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.