Model · Anthropic

Claude Opus 4.5

Anthropic model in the public SWE-bench Verified panel (mini-SWE-agent v2, high effort).

3 values from 2 studies · Updated

At a glance

The best-supported value per category: a 95% interval first, then a run range, then the larger n. Three separate values from separate studies, never one score.

Qualitypass rates, accuracy and scores

73% (24/33)

95% CI 56%–85% · n = 33

Resolved rate on the same 33 SWE-bench Verified instances

effort high · public mini-SWE-agent v2 run, same instances · Agent on SWE-bench Verified vs 11 public models

Speedtime per call or decision

Not measured in any study yet.

CostUS dollars per call, pass or decision

$1.18

n = 24

Recorded cost per resolved instance: Agent vs the public panel

effort high · public mini-SWE-agent v2 run, same instances · What if every call ran on Opus? Repricing real agent tokens

Where it sits

Every measured value, grouped by study. Each row puts the value on its own track, with the other configurations of the same chart as muted dots. A range is the fastest to slowest recorded run and p50–p95 is the median to the 95th percentile; neither is a confidence interval. Use Table for the plain values.

Resolved rate on the same 33 SWE-bench Verified instancesn = 33 · 95% CI 56%–85% · effort high · public mini-SWE-agent v2 run, same instances
73% (24/33)
Model calls per instancen = 33 · effort high · public mini-SWE-agent v2 run, same instances
35.9

Whiskers: 95% Wilson intervaln beside each valueMuted dots: the other configurations on the same chart

Claude Opus 4.5 in Agent on SWE-bench Verified vs 11 public models: 2 values, first Resolved rate on the same 33 SWE-bench Verified instances 73% (24/33).

Recorded cost per resolved instance: Agent vs the public paneln = 24 · effort high · public mini-SWE-agent v2 run, same instances
$1.18

n beside each valueMuted dots: the other configurations on the same chart

Claude Opus 4.5 in What if every call ran on Opus? Repricing real agent tokens: 1 value, first Recorded cost per resolved instance: Agent vs the public panel $1.18.

Compare Claude Opus 4.5

Each bar counts the rows of one comparison: a side ahead only where its interval or range is apart, otherwise a tie or unclear.

Watch

Live story · 28 sSWE-bench Verified: Agent vs 11 public model runs

SWE-bench Verified: Agent vs 11 public model runs

Agent resolved 25 of 33 SWE-bench Verified instances; every 95% interval overlaps the public panel, so the data does not rank it.

Transcript
  1. SWE-bench Verified · same 33 instances. Agent vs 11 public model runs. One attempt each. No retries. 95% intervals on individual run rates. The panel mean has no interval.
  2. Agent resolved 25 of 33. The public panel averaged 74.1% on the same instances. Agent resolved (full pipeline, Sonnet 5.5): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Agent model cost per attempt (calculation, notional): $2.81 (n = 33). Caveat: Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
  3. Panel runs resolved 21 to 28 of 33. Every interval overlaps, so this sample cannot rank Agent. Chart: Resolved rate on the same 33 SWE-bench Verified instances (n = 33 each). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
  4. Hardest band: on the 4 instances no panel model solved, Agent solved 1. Too few to call a rate. Chart: Resolved rate by difficulty band (n = 4–14 each). Caveat: Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
  5. Open benchmarks: intervals, sources and every failure kept.

Write-ups that use this data

All posts

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.