Agent harness · Agent
Agent
The Agent coding pipeline: onboarding, research, plan, act, verify and review, on Claude Sonnet 5.5 through a Claude subscription.
23 values from 4 studies (5 are list-price calculations) · Updated
At a glance
The best-supported value per category: a 95% interval first, then a run range, then the larger n. Three separate values from separate studies, never one score.
Qualitypass rates, accuracy and scores
69% (91/132)
95% CI 61%–76% · n = 132
Single critic verdicts that preferred the AI change
blind panel: Agent change vs merged human change · blind review panel · AI pull requests vs merged human pull requests, judged blind
Speedtime per call or decision
9.6min
n = 33
Median worker time per attempt
full pipeline on Claude Sonnet 5.5 · Agent on SWE-bench Verified vs 11 public models
CostUS dollars per call, pass or decision
$2.81
n = 33 · list-price calculation
Agent model cost per attempt (notional)
full pipeline on Claude Sonnet 5.5 · Agent on SWE-bench Verified vs 11 public models
Where it sits
Every measured value, grouped by study. Each row puts the value on its own track, with the other configurations of the same chart as muted dots. A range is the fastest to slowest recorded run and p50–p95 is the median to the 95th percentile; neither is a confidence interval. Use Table for the plain values.
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Resolved rate on the same 33 SWE-bench Verified instances | 76% (25/33) | 33 | 59%–87% (95% CI) | full pipeline on Claude Sonnet 5.5 |
| Resolved rate by difficulty band: No panel model solved it | 25% (1/4) | 4 | 4.6%–70% (95% CI) | full pipeline on Claude Sonnet 5.5 |
| Resolved rate by difficulty band: Under half solved it | 75% (3/4) | 4 | 30%–95% (95% CI) | full pipeline on Claude Sonnet 5.5 |
| Resolved rate by difficulty band: Half or more solved it | 82% (9/11) | 11 | 52%–95% (95% CI) | full pipeline on Claude Sonnet 5.5 |
| Resolved rate by difficulty band: Every panel model solved it | 86% (12/14) | 14 | 60%–96% (95% CI) | full pipeline on Claude Sonnet 5.5 |
| Model calls per instance | 49.5 | 33 | — | full pipeline on Claude Sonnet 5.5 |
| Every way to slice the run, with intervals: Campaign 1: 25-instance sample | 72% (18/25) | 25 | 52%–86% (95% CI) | full pipeline on Claude Sonnet 5.5 |
| Every way to slice the run, with intervals: Campaign 2: 8 compiled-extension instances | 88% (7/8) | 8 | 53%–98% (95% CI) | full pipeline on Claude Sonnet 5.5 |
| Every way to slice the run, with intervals: Original seed draw of 25 | 76% (19/25) | 25 | 57%–89% (95% CI) | full pipeline on Claude Sonnet 5.5 |
| Every way to slice the run, with intervals: All 33 attempted | 76% (25/33) | 33 | 59%–87% (95% CI) | full pipeline on Claude Sonnet 5.5 |
| Agent model cost per attempt (notional) Calculation | $2.81 | 33 | — | full pipeline on Claude Sonnet 5.5 |
| Agent model cost per resolved instance (notional) Calculation | $3.71 | 25 | — | full pipeline on Claude Sonnet 5.5 |
| Median worker time per attempt | 9.6 min | 33 | — | full pipeline on Claude Sonnet 5.5 |
Whiskers: 95% Wilson intervaln beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Agent in Agent on SWE-bench Verified vs 11 public models: 13 values, first Resolved rate on the same 33 SWE-bench Verified instances 76% (25/33).
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Tasks where the panel preferred the AI change (latest attempt) | 75% (9/12) | 12 | 47%–91% (95% CI) | blind panel: Agent change vs merged human change · blind review panel |
| Tasks where the panel preferred the AI change (first scored attempt) | 50% (6/12) | 12 | 25%–75% (95% CI) | blind panel: Agent change vs merged human change · blind review panel |
| All scored pairs where the panel preferred the AI change | 60% (12/20) | 20 | 39%–78% (95% CI) | blind panel: Agent change vs merged human change · blind review panel |
| Public OSS tasks, latest attempt | 80% (4/5) | 5 | 38%–96% (95% CI) | blind panel: Agent change vs merged human change · blind review panel |
| Single critic verdicts that preferred the AI change | 69% (91/132) | 132 | 61%–76% (95% CI) | blind panel: Agent change vs merged human change · blind review panel |
Whiskers: 95% Wilson intervaln beside each valueMuted dots: the other configurations on the same chart
Agent in AI pull requests vs merged human pull requests, judged blind: 5 values, first Tasks where the panel preferred the AI change (latest attempt) 75% (9/12).
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Recorded cost per resolved instance: Agent vs the public panel Calculation | $3.71 | 25 | — | full pipeline on Claude Sonnet 5.5 · notional |
n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Agent in What if every call ran on Opus? Repricing real agent tokens: 1 value, first Recorded cost per resolved instance: Agent vs the public panel $3.71.
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Verified deliveries, latest build | 1 of 3 | 3 | — | three real tasks, platform builds compared |
| Notional cost, latest build, all 3 tasks Calculation | $11.06 | 3 | — | three real tasks, platform builds compared |
| Guardrail refusals, first vs latest slice | 26 → 19 | 3 | — | three real tasks, platform builds compared |
| First calibration run (capped, fastify/session) Calculation | $4.89, stopped at cap | 1 | — | three real tasks, platform builds compared |
n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Agent in Coding calibration: what broke on three real pull requests: 4 values, first Verified deliveries, latest build 1 of 3.
Compare Agent
Each bar counts the rows of one comparison: a side ahead only where its interval or range is apart, otherwise a tie or unclear.
vs model
Agent vs Claude Haiku 4.5
1 tie · 2 unclear
vs model
Agent vs Claude Opus 4.5
1 tie · 2 unclear
vs model
Agent vs Claude Opus 4.6
1 tie · 2 unclear
vs model
Agent vs Claude Sonnet 4.5
1 tie · 2 unclear
vs model
Agent vs DeepSeek V3.2
1 tie · 2 unclear
vs model
Agent vs Gemini 3 Flash
1 tie · 2 unclear
vs model
Agent vs GLM 5
1 tie · 2 unclear
vs model
Agent vs GPT 5.2
1 tie · 2 unclear
vs model
Agent vs GPT 5 mini
1 tie · 2 unclear
vs model
Agent vs Kimi K2.5
1 tie · 2 unclear
vs model
Agent vs MiniMax M2.5
1 tie · 2 unclear
Watch
SWE-bench Verified: Agent vs 11 public model runs
Agent resolved 25 of 33 SWE-bench Verified instances; every 95% interval overlaps the public panel, so the data does not rank it.
Transcript
- SWE-bench Verified · same 33 instances. Agent vs 11 public model runs. One attempt each. No retries. 95% intervals on individual run rates. The panel mean has no interval.
- Agent resolved 25 of 33. The public panel averaged 74.1% on the same instances. Agent resolved (full pipeline, Sonnet 5.5): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Agent model cost per attempt (calculation, notional): $2.81 (n = 33). Caveat: Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
- Panel runs resolved 21 to 28 of 33. Every interval overlaps, so this sample cannot rank Agent. Chart: Resolved rate on the same 33 SWE-bench Verified instances (n = 33 each). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
- Hardest band: on the 4 instances no panel model solved, Agent solved 1. Too few to call a rate. Chart: Resolved rate by difficulty band (n = 4–14 each). Caveat: Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
- Open benchmarks: intervals, sources and every failure kept.
Write-ups that use this data
All postsAI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
AI coding cost per developer: a formula built on recorded work
AI coding cost per developer: a monthly budget formula using recorded tokens and list-price calculations. Replace our task volume and pass rate with yours.
Claude Code cost per task: a price ladder from one decision to one agent run
$0.087 per Claude Code coding session at API list prices, $0.0036 per short call, $2.81 per full agent attempt. Calculations on recorded tokens, not bills.