Agent harness · Agent

Agent

The Agent coding pipeline: onboarding, research, plan, act, verify and review, on Claude Sonnet 5.5 through a Claude subscription.

23 values from 4 studies (5 are list-price calculations) · Updated

At a glance

The best-supported value per category: a 95% interval first, then a run range, then the larger n. Three separate values from separate studies, never one score.

Qualitypass rates, accuracy and scores

69% (91/132)

95% CI 61%–76% · n = 132

Single critic verdicts that preferred the AI change

blind panel: Agent change vs merged human change · blind review panel · AI pull requests vs merged human pull requests, judged blind

Speedtime per call or decision

9.6min

n = 33

Median worker time per attempt

full pipeline on Claude Sonnet 5.5 · Agent on SWE-bench Verified vs 11 public models

CostUS dollars per call, pass or decision

$2.81

n = 33 · list-price calculation

Agent model cost per attempt (notional)

full pipeline on Claude Sonnet 5.5 · Agent on SWE-bench Verified vs 11 public models

Where it sits

Every measured value, grouped by study. Each row puts the value on its own track, with the other configurations of the same chart as muted dots. A range is the fastest to slowest recorded run and p50–p95 is the median to the 95th percentile; neither is a confidence interval. Use Table for the plain values.

Resolved rate on the same 33 SWE-bench Verified instancesn = 33 · 95% CI 59%–87% · full pipeline on Claude Sonnet 5.5
76% (25/33)
Resolved rate by difficulty band: No panel model solved itn = 4 · 95% CI 4.6%–70% · full pipeline on Claude Sonnet 5.5
25% (1/4)
Resolved rate by difficulty band: Under half solved itn = 4 · 95% CI 30%–95% · full pipeline on Claude Sonnet 5.5
75% (3/4)
Resolved rate by difficulty band: Half or more solved itn = 11 · 95% CI 52%–95% · full pipeline on Claude Sonnet 5.5
82% (9/11)
Resolved rate by difficulty band: Every panel model solved itn = 14 · 95% CI 60%–96% · full pipeline on Claude Sonnet 5.5
86% (12/14)
Model calls per instancen = 33 · full pipeline on Claude Sonnet 5.5
49.5

Whiskers: 95% Wilson intervaln beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation

Agent in Agent on SWE-bench Verified vs 11 public models: 13 values, first Resolved rate on the same 33 SWE-bench Verified instances 76% (25/33).

Tasks where the panel preferred the AI change (latest attempt)n = 12 · 95% CI 47%–91% · blind panel: Agent change vs merged human change · blind review panel
75% (9/12)
Tasks where the panel preferred the AI change (first scored attempt)n = 12 · 95% CI 25%–75% · blind panel: Agent change vs merged human change · blind review panel
50% (6/12)
All scored pairs where the panel preferred the AI changen = 20 · 95% CI 39%–78% · blind panel: Agent change vs merged human change · blind review panel
60% (12/20)
Public OSS tasks, latest attemptn = 5 · 95% CI 38%–96% · blind panel: Agent change vs merged human change · blind review panel
80% (4/5)
Single critic verdicts that preferred the AI changen = 132 · 95% CI 61%–76% · blind panel: Agent change vs merged human change · blind review panel
69% (91/132)

Whiskers: 95% Wilson intervaln beside each valueMuted dots: the other configurations on the same chart

Agent in AI pull requests vs merged human pull requests, judged blind: 5 values, first Tasks where the panel preferred the AI change (latest attempt) 75% (9/12).

Recorded cost per resolved instance: Agent vs the public paneln = 25 · full pipeline on Claude Sonnet 5.5 · notional
$3.71Calculation

n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation

Agent in What if every call ran on Opus? Repricing real agent tokens: 1 value, first Recorded cost per resolved instance: Agent vs the public panel $3.71.

Verified deliveries, latest buildn = 3 · three real tasks, platform builds compared
1 of 3
Notional cost, latest build, all 3 tasksn = 3 · three real tasks, platform builds compared
$11.06Calculation
Guardrail refusals, first vs latest slicen = 3 · three real tasks, platform builds compared
26 → 19
First calibration run (capped, fastify/session)n = 1 · three real tasks, platform builds compared
$4.89, stopped at capCalculation

n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation

Agent in Coding calibration: what broke on three real pull requests: 4 values, first Verified deliveries, latest build 1 of 3.

Compare Agent

Each bar counts the rows of one comparison: a side ahead only where its interval or range is apart, otherwise a tie or unclear.

Watch

Live story · 28 sSWE-bench Verified: Agent vs 11 public model runs

SWE-bench Verified: Agent vs 11 public model runs

Agent resolved 25 of 33 SWE-bench Verified instances; every 95% interval overlaps the public panel, so the data does not rank it.

Transcript
  1. SWE-bench Verified · same 33 instances. Agent vs 11 public model runs. One attempt each. No retries. 95% intervals on individual run rates. The panel mean has no interval.
  2. Agent resolved 25 of 33. The public panel averaged 74.1% on the same instances. Agent resolved (full pipeline, Sonnet 5.5): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Agent model cost per attempt (calculation, notional): $2.81 (n = 33). Caveat: Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
  3. Panel runs resolved 21 to 28 of 33. Every interval overlaps, so this sample cannot rank Agent. Chart: Resolved rate on the same 33 SWE-bench Verified instances (n = 33 each). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
  4. Hardest band: on the 4 instances no panel model solved, Agent solved 1. Too few to call a rate. Chart: Resolved rate by difficulty band (n = 4–14 each). Caveat: Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
  5. Open benchmarks: intervals, sources and every failure kept.

Write-ups that use this data

All posts

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.