Model · Anthropic
Claude Opus 4.5
Anthropic model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).
3 values from 2 studies · Updated
At a glance
The best-supported value per category: a 95% interval first, then a run range, then the larger n. Three separate values from separate studies, never one score.
Qualitypass rates, accuracy and scores
73% (24/33)
95% CI 56%–85% · n = 33
Resolved rate on the same 33 SWE-bench Verified instances
effort high · public mini-SWE-agent v2 run, same instances · Agent on SWE-bench Verified vs 11 public models
Speedtime per call or decision
Not measured in any study yet.
CostUS dollars per call, pass or decision
$1.18
n = 24
Recorded cost per resolved instance: Agent vs the public panel
effort high · public mini-SWE-agent v2 run, same instances · What if every call ran on Opus? Repricing real agent tokens
Where it sits
Every measured value, grouped by study. Each row puts the value on its own track, with the other configurations of the same chart as muted dots. A range is the fastest to slowest recorded run and p50–p95 is the median to the 95th percentile; neither is a confidence interval. Use Table for the plain values.
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 73% (24/33) | 33 | 56%–85% (95% CI) | effort high · public mini-SWE-agent v2 run, same instances |
| Model calls per instance | 35.9 | 33 | — | effort high · public mini-SWE-agent v2 run, same instances |
Whiskers: 95% Wilson intervaln beside each valueMuted dots: the other configurations on the same chart
Claude Opus 4.5 in Agent on SWE-bench Verified vs 11 public models: 2 values, first Resolved rate on the same 33 SWE-bench Verified instances 73% (24/33).
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Recorded cost per resolved instance: Agent vs the public panel | $1.18 | 24 | — | effort high · public mini-SWE-agent v2 run, same instances |
n beside each valueMuted dots: the other configurations on the same chart
Claude Opus 4.5 in What if every call ran on Opus? Repricing real agent tokens: 1 value, first Recorded cost per resolved instance: Agent vs the public panel $1.18.
Compare Claude Opus 4.5
Each bar counts the rows of one comparison: a side ahead only where its interval or range is apart, otherwise a tie or unclear.
vs model
Claude Haiku 4.5 vs Claude Opus 4.5
1 tie · 2 unclear
vs model
Claude Opus 4.5 vs Claude Opus 4.6
1 tie · 2 unclear
vs model
Claude Opus 4.5 vs DeepSeek V3.2
1 tie · 2 unclear
vs model
Claude Opus 4.5 vs GPT 5 mini
1 tie · 2 unclear
vs model
Claude Opus 4.5 vs Kimi K2.5
1 tie · 2 unclear
vs model
Claude Opus 4.5 vs MiniMax M2.5
1 tie · 2 unclear
vs model
Claude Sonnet 4.5 vs Claude Opus 4.5
1 tie · 2 unclear
vs model
Gemini 3 Flash vs Claude Opus 4.5
1 tie · 2 unclear
vs model
GLM 5 vs Claude Opus 4.5
1 tie · 2 unclear
vs model
GPT 5.2 vs Claude Opus 4.5
1 tie · 2 unclear
vs agent harness
Agent vs Claude Opus 4.5
1 tie · 2 unclear
Watch
SWE-bench Verified: Agent vs 11 public model runs
Agent resolved 25 of 33 SWE-bench Verified instances; every 95% interval overlaps the public panel, so the data does not rank it.
Transcript
- SWE-bench Verified · same 33 instances. Agent vs 11 public model runs. One attempt each. No retries. 95% intervals on individual run rates. The panel mean has no interval.
- Agent resolved 25 of 33. The public panel averaged 74.1% on the same instances. Agent resolved (full pipeline, Sonnet 5.5): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Agent model cost per attempt (calculation, notional): $2.81 (n = 33). Caveat: Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
- Panel runs resolved 21 to 28 of 33. Every interval overlaps, so this sample cannot rank Agent. Chart: Resolved rate on the same 33 SWE-bench Verified instances (n = 33 each). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
- Hardest band: on the 4 instances no panel model solved, Agent solved 1. Too few to call a rate. Chart: Resolved rate by difficulty band (n = 4–14 each). Caveat: Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
- Open benchmarks: intervals, sources and every failure kept.
Write-ups that use this data
All postsAI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
AI coding cost per developer: a formula built on recorded work
AI coding cost per developer: a monthly budget formula using recorded tokens and list-price calculations. Replace our task volume and pass rate with yours.
Claude Code cost per task: a price ladder from one decision to one agent run
$0.087 per Claude Code coding session at API list prices, $0.0036 per short call, $2.81 per full agent attempt. Calculations on recorded tokens, not bills.