Model · Anthropic

Claude Opus 5

Anthropic model, measured as a blind code-review critic and in a list-price calculation.

2 values from 2 studies (1 are list-price calculations) · Updated

At a glance

The best-supported value per category: a 95% interval first, then a run range, then the larger n. Three separate values from separate studies, never one score.

Qualitypass rates, accuracy and scores

70% (28/40)

95% CI 55%–82% · n = 40

Does the judge’s model family matter?

blind review panel · AI pull requests vs merged human pull requests, judged blind

Speedtime per call or decision

Not measured in any study yet.

CostUS dollars per call, pass or decision

Not measured in any study yet.

Claude Opus 5 price per 1M tokens

Claude Opus 5 costs $5.00 per 1M input tokens and $25.00 per 1M output tokens at standard Claude API list prices (observed 2026-10-08). Cache reads cost $0.50 per 1M tokens.

Vendor list price · observed 2026-10-08Third-party-reported · snapshot 2026-10-06

Vendor list price

$5.00

Input, per 1M tokens

Vendor list price

$25.00

Output, per 1M tokens

Vendor list price

$0.50

Cache read, per 1M tokens

Vendor specification · 2026-10-08

1M

Context window, tokens

Cache writes are separate: $6.25 per 1M tokens for five minutes; $10.00 per 1M tokens for one hour. Standard global Claude API prices; other tiers, tools and taxes can add charges.

Source: Official Claude API prices · observed .

Context: Official model specification · observed . Route and provider limits can be lower than the model specification.

Recorded study price source: Anthropic list prices (Claude models) (). Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price.

Providers in the dated snapshot

5 standard-tier providers · snapshot 2026-10-06 · USD per 1M tokens

Same snapshot price at all 5 providers
  • Anthropicfirst-party
  • Amazon Bedrock
  • Azure
  • Claude Platform on AWS
  • Google Vertex

Third-party-reported, not measured by Agent · Standard tier only: regional, flex, fast and priority tiers are left out · One row per provider: its cheapest standard endpoint

Claude Opus 5: 5 standard-tier providers, all at the same price ($5.00 input, $25.00 output per 1M tokens). Third-party-reported, snapshot 2026-10-06.

Source: OpenRouter public API: models and provider endpoints (snapshot) (). The diamond marks the first-party provider in this snapshot. Current availability is not checked.

Where it sits

Every measured value, grouped by study. Each row puts the value on its own track, with the other configurations of the same chart as muted dots. A range is the fastest to slowest recorded run and p50–p95 is the median to the 95th percentile; neither is a confidence interval. Use Table for the plain values.

Does the judge’s model family matter?n = 40 · 95% CI 55%–82% · blind review panel
70% (28/40)

Whiskers: 95% Wilson intervaln beside each valueMuted dots: the other configurations on the same chart

Claude Opus 5 in AI pull requests vs merged human pull requests, judged blind: 1 value, first Does the judge’s model family matter? 70% (28/40).

Thought experiment: the same tokens at other list pricescalculation: Agent’s recorded tokens at this model’s list price
$8.72Calculation

n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation

Claude Opus 5 in What if every call ran on Opus? Repricing real agent tokens: 1 value, first Thought experiment: the same tokens at other list prices $8.72.

Watch

Live story · 33 sWhat if every call ran on Opus? Repricing real agent tokens

What if every call ran on Opus? Repricing real agent tokens

A calculation, not a run: Agent's recorded SWE-bench tokens cost $87.23 at Sonnet 5.5 prices, $143.83 at Opus 5.5 and $343.33 without caching.

Transcript
  1. Thought experiment · recorded tokens × list prices. What if every call ran on Opus? Agent recorded every token on 33 SWE-bench attempts. We repriced them.
  2. 162.9M input tokens, 94.0% of them read from the prompt cache. Input tokens recorded: 162.9M (n = 33). Output tokens recorded: 1.8M (n = 33). Input served from cache: 94.0% (n = 33). Caveat: Recorded costs are list-price estimates for subscription calls; no invoice backs them.
  3. Same tokens on Opus 5.5: $143.83 instead of $87.23, 1.65× the bill. On Haiku 4.5: $43.61. At Haiku 4.5 prices: $43.61 (n = 33). At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  4. Per resolved instance: $3.49 on Sonnet 5.5, $5.75 on Opus 5.5, $12.85 on Fable 5.1. Chart: Thought experiment: the same tokens at other list prices. Calculation, not a run. Caveat: Different models use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.
  5. In this calculation caching matters more than the model: without it, Sonnet would cost $343.33, 3.9× the recorded $87.23. Chart: Thought experiment: what prompt caching saved. Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  6. Price sensitivity, not predictions. Every repricing labelled as a calculation.
Live story · 34 sAI pull requests vs merged human pull requests, judged blind

AI pull requests vs merged human pull requests, judged blind

Blind critics preferred the AI change on 9 of 12 tasks at the latest attempt and 6 of 12 at the first. Same-family bias disclosed.

Transcript
  1. Blind review · 12 real merged pull requests. AI change vs the merged human change. Critic models review both, unlabelled, once in each order.
  2. Latest attempt: the panel preferred the AI change on 9 of 12 tasks. First attempt: 6 of 12. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Latest attempts are not independent first tries: they had lessons, review replays and some operator answers from earlier attempts on the same task.
  3. The first attempt is the cleaner estimate: 50%, with a 95% interval from 25% to 75%. Chart: First attempt vs latest attempt (n = 5–12 each). Caveat: n = 12 tasks: intervals are wide.
  4. Vote by vote: 9 of 12 tasks went unanimously to the AI change, 2 unanimously to the human one. Chart: Blind panel votes per task (latest attempt) (n = 6–8 each). Caveat: 7 of the 12 tasks come from one private Go service. They carry neutral labels (private-go-a and so on), and only their language and kind are published.
  5. Cross-family check: OpenAI critics preferred the AI change in 10 of 12 verdicts. Small n, wide intervals. Chart: Does the judge’s model family matter? (n = 2–40 each). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
  6. Open benchmarks: intervals, sources and every failure kept.

Write-ups that use this data

All posts

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.