2 measured metrics · 1 calculated · 2 studies

AgentvsClaude Haiku 4.5

No row separates them: 1 tie, 2 unclear.

The verdict

Agent and Claude Haiku 4.5 share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; Claude Haiku 4.5 ran under a different, simpler harness, so this compares systems, not models.

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates)n is shown per side on every rowHollow marks: list-price calculations, not runs

3 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Recorded cost per resolved instance: Agent vs the public panel, 7.7x (Agent larger).

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Agent
  • Claude Haiku 4.5
  • 95% interval
  • where the two overlap
  • hollow: list-price calculation

Agent on SWE-bench Verified vs 11 public models

1 tie · 1 unclear
  • Resolved rate on the same 33 SWE-bench Verified instances: Agent 76% (25/33) (n 33, 95% interval 59%–87%); Claude Haiku 4.5 76% (25/33) (n 33, 95% interval 59%–87%). Tie.
  • Model calls per instance: Agent 49.5 (n 33); Claude Haiku 4.5 68.5 (n 33). Unclear.

What if every call ran on Opus? Repricing real agent tokens

1 unclear
  • Recorded cost per resolved instance: Agent vs the public panel, calculation: Agent $3.71 (n 25); Claude Haiku 4.5 $0.48 (n 25). Unclear.

Marks: 95% intervals (Wilson for rates)n is shown per side on every rowHollow marks: list-price calculations, not runs

3 rows from 2 studies. No row separates them: 1 tie, 2 unclear.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Agent

No row in this data puts Agent ahead of Claude Haiku 4.5. Pick on other grounds (price, access, the tasks you run), or measure your own workload.

When to pick Claude Haiku 4.5

  • Recorded cost per resolved instance: Agent vs the public panel: $0.48 vs $3.71. A list-price calculation, not a measured difference. Calculation

Side by side

The study charts, showing only these two. Open a study for every configuration.

Agent (Sonnet 5.5, full pipeline)
Claude 4.5 Haiku (high)

2 rows. All at 76%.

NotesWhiskers: 95% Wilson intervaln = 33 per row

Agent vs 11 public mini-SWE-agent v2 runs, one attempt each

Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

Model calls per instance

Mean over the same 33 instances

Claude 4.5 Haiku (high)
Agent

2 rows. Highest Claude 4.5 Haiku (high) 68.5 (n 33). Lowest Agent 49.5 (n 33).

Notesn = 33 per row

A panel call is one bash-agent step. An Agent call is one model request of any stage (research, plan, act, verify, review).

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

Agent (notional)
Claude 4.5 Haiku (high)

2 rows. Highest Agent (notional) $3.71 (n 25). Lowest Claude 4.5 Haiku (high) $0.48 (n 25).

Notesn = 25 per row

Same 33 SWE-bench Verified instances; all attempts in the numerator

Recorded figures, not repricing. Panel costs are published API costs for a bash-only agent. Agent's figure is a list-price estimate of subscription calls and includes onboarding, planning, verification and review.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Agent or Claude Haiku 4.5?
Agent and Claude Haiku 4.5 share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; Claude Haiku 4.5 ran under a different, simpler harness, so this compares systems, not models.
How were Agent and Claude Haiku 4.5 measured?
They share 2 measured metrics and 1 list-price calculation from 2 public studies: Agent on SWE-bench Verified vs 11 public models; What if every call ran on Opus? Repricing real agent tokens. Every row names its configuration, its sample size and its interval or range.
How do Agent and Claude Haiku 4.5 compare on resolved rate on the same 33 SWE-bench Verified instances?
Agent: 76% (25/33) (full pipeline on Claude Sonnet 5.5; n = 33; 95% interval 59% to 87%). Claude Haiku 4.5: 76% (25/33) (effort high · public mini-SWE-agent v2 run, same instances; n = 33; 95% interval 59% to 87%). The 95% intervals overlap (Agent 59% to 87%; Claude Haiku 4.5 59% to 87%), so this sample cannot separate them.

The studies behind this page

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.