2 measured metrics · 1 calculated · 2 studies

AgentvsGPT 5 mini

No row separates them: 1 tie, 2 unclear.

The verdict

Agent and GPT 5 mini share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; GPT 5 mini ran under a different, simpler harness, so this compares systems, not models.

Headline metrics

How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.

Includes calculations

Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.

Marks: 95% intervals (Wilson for rates)n is shown per side on every rowHollow marks: list-price calculations, not runs

3 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Recorded cost per resolved instance: Agent vs the public panel, 46x (Agent larger).

Metric by metric

Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.

  • Agent
  • GPT 5 mini
  • 95% interval
  • where the two overlap
  • hollow: list-price calculation

These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.

Agent on SWE-bench Verified vs 11 public models

1 tie · 1 unclear
  • Resolved rate on the same 33 SWE-bench Verified instances: Agent 76% (25/33) (n 33, 95% interval 59%–87%); GPT 5 mini 64% (21/33) (n 33, 95% interval 47%–78%). Tie.

    Calculation: at these rates, about 208 runs per side would separate them.

  • Model calls per instance: Agent 49.5 (n 33); GPT 5 mini 20.8 (n 33). Unclear.

What if every call ran on Opus? Repricing real agent tokens

1 unclear
  • Recorded cost per resolved instance: Agent vs the public panel, calculation: Agent $3.71 (n 25); GPT 5 mini $0.080 (n 21). Unclear.

Marks: 95% intervals (Wilson for rates)n is shown per side on every rowHollow marks: list-price calculations, not runs

3 rows from 2 studies. No row separates them: 1 tie, 2 unclear.

When to pick which

Only from the rows above. A tie is not a reason to pick either side.

When to pick Agent

No row in this data puts Agent ahead of GPT 5 mini. Pick on other grounds (price, access, the tasks you run), or measure your own workload.

When to pick GPT 5 mini

  • Recorded cost per resolved instance: Agent vs the public panel: $0.080 vs $3.71. A list-price calculation, not a measured difference. Calculation

Side by side

The study charts, showing only these two. Open a study for every configuration.

Agent (Sonnet 5.5, full pipeline)
GPT 5 mini

2 rows. Highest Agent (Sonnet 5.5, full pipeline) 76% (95% interval 59%–87%, n 33). Lowest GPT 5 mini 64% (95% interval 47%–78%, n 33). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 33 per row

Agent vs 11 public mini-SWE-agent v2 runs, one attempt each

Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

Model calls per instance

Mean over the same 33 instances

Agent
GPT 5 mini

2 rows. Highest Agent 49.5 (n 33). Lowest GPT 5 mini 20.8 (n 33).

Notesn = 33 per row

A panel call is one bash-agent step. An Agent call is one model request of any stage (research, plan, act, verify, review).

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

Largest value is 46x the smallest; Log shows the small bars.
Agent (notional)
GPT 5 mini

2 rows. Highest Agent (notional) $3.71 (n 25). Lowest GPT 5 mini $0.08 (n 21).

Notesn 21–25 per row

Same 33 SWE-bench Verified instances; all attempts in the numerator

Recorded figures, not repricing. Panel costs are published API costs for a bash-only agent. Agent's figure is a list-price estimate of subscription calls and includes onboarding, planning, verification and review.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

How a row is called

  • AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
  • TieThe values match, or both sit at the same ceiling.
  • UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
  • CalculationDerived from list prices and recorded counts. Not a bill and not a run.

The dataset compiler makes every call; this page only draws it. No composite score, no rank.

Questions

Which is better, Agent or GPT 5 mini?
Agent and GPT 5 mini share 2 measured metrics and 1 list-price calculation from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Agent is a full pipeline on one model; GPT 5 mini ran under a different, simpler harness, so this compares systems, not models.
How were Agent and GPT 5 mini measured?
They share 2 measured metrics and 1 list-price calculation from 2 public studies: Agent on SWE-bench Verified vs 11 public models; What if every call ran on Opus? Repricing real agent tokens. Every row names its configuration, its sample size and its interval or range.
How do Agent and GPT 5 mini compare on resolved rate on the same 33 SWE-bench Verified instances?
Agent: 76% (25/33) (full pipeline on Claude Sonnet 5.5; n = 33; 95% interval 59% to 87%). GPT 5 mini: 64% (21/33) (public mini-SWE-agent v2 run, same instances; n = 33; 95% interval 47% to 78%). The 95% intervals overlap (Agent 59% to 87%; GPT 5 mini 47% to 78%), so this sample cannot separate them.

The studies behind this page

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

All comparisons

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.