Explainer · Agent harness
What is an agent harness? The code around the model, measured
Definition
Agent harness is the code that wraps a language model and turns it into an agent. The model answers; the harness runs the loop, offers tools, builds prompts, keeps memory, manages context and checks the result. Two systems can share one model and still differ in time and tokens, so a benchmark scores a system, not a model.
Agent team · · 5 min read · Every number is from the public studies
How it works
Where a full pipeline spends
Agent's notional model cost by stage, summed over 33 SWE-bench Verified attempts.
Onboarding
$2.09
The pipeline reads the repository and writes notes for later calls.
- Notional list-price estimate
Research
$20.56
The model reads code and docs before it changes anything.
- Second-largest share of spend
Act
$44.05
The model edits files and runs commands.
- The largest share of spend
Verify
$8.72
The pipeline checks the change before it delivers.
- Checks cost model calls too
Review
$5.58
A review pass reads the change last.
- Smaller than act and research
| Stage | What it does | Checks |
|---|---|---|
| 1. Onboarding | The pipeline reads the repository and writes notes for later calls. ($2.09) | Notional list-price estimate |
| 2. Research | The model reads code and docs before it changes anything. ($20.56) | Second-largest share of spend |
| 3. Act | The model edits files and runs commands. ($44.05) | The largest share of spend |
| 4. Verify | The pipeline checks the change before it delivers. ($8.72) | Checks cost model calls too |
| 5. Review | A review pass reads the change last. ($5.58) | Smaller than act and research |
Five stages of one full pipeline, with Agent's notional model cost for each. The act stage, where the model edits and runs code, costs the most: $44.05, about 48% of the total spend (calculation). List-price estimates for subscription calls, not invoices.
Context compaction ($5.41) and other calls ($6.24) are not drawn as stages, and the chart has no separate line for planning. Source: chart swebench-cost-by-stage; n 33 attempts.
Source: SWE-bench Verified study
What an agent harness does
The model reads text and writes text. It cannot open a file or run a test. The harness does those jobs, and some teams call it a scaffold:
- The loop. It sends a request, runs the tool the model asks for and returns the result.
- Tools and prompts. It decides which actions exist and writes the system prompt.
- Memory and context. It keeps notes between calls and chooses what stays in context.
- Checks. It runs tests, reviews the patch and decides when to stop.
Three kinds of agent harness
- A bare API call. Your code sends a prompt and reads the answer.
- A CLI agent. Claude Code and the Codex CLI add a system prompt, tools and a loop.
- A full pipeline. Agent, one system in our data, runs onboarding, research, a plan, edits, verification and review on one model.
The 11 public runs in our panel use mini-SWE-agent v2, a bash-only harness: one model acting only through a shell.
What the harness changed in our data
Same model, different route: time and tokens
GPT-6.1 Sol repaired the same scheduler task through the OpenAI API and the Codex CLI. Only the route changed:
- OpenAI API: median 17.3 s, range 16.3 to 18.6 s, n 3.
- Codex CLI: median 61.2 s, range 59.9 to 69.5 s, n 3.
All 9 runs passed all 296 checks, so this task has a ceiling for quality. A range is not a confidence interval. The ranges do not overlap, but 3 runs per side is too few for a tested ranking, so read this as directional. The chart also shows Claude Code with Sonnet 5.5, a different model, so it does not isolate the route.
The harness also adds tokens, a context tax. For a one-line request, the Codex CLI sent a median 19,551 input tokens (n 15) and the API sent 17. Part of the CLI input came from cache. Over 15 runs each, the CLI took a median 3.9 s and the API 1.1 s. The ranges, a calculation over three cells, are 2.9 to 4.7 s and 0.7 to 2.2 s.
Full pipeline, bash-only harness: cost, not resolved rate
On 33 SWE-bench Verified instances, Agent (a full pipeline on Sonnet 5.5) resolved 25 of 33: 75.8%, 95% interval 59.0% to 87.2%. Eleven public runs, one model each in a bash-only harness, resolved 21 to 28 of the same 33 (n 33 each). The lowest, 21 of 33, has a 95% interval of 46.6% to 77.8%; the highest, 28 of 33, 69.1% to 93.3%.
Every interval overlaps, so this sample cannot rank any system. With n 33, one interval spans about 28 points (calculation: 87.2 minus 59.0), so the set cannot see a small gap. The two sides use different models, so this compares systems, not harnesses alone. Verified issues are public, so contamination is uncontrolled for every system.
Cost and effort:
- Cost per attempt: Agent $2.81 notional (n 33). Panel runs average $0.051 to $0.861 in published API cost.
- Cost per resolved instance: Agent $3.71 notional (n 25). The panel mean is $0.57 (recorded, 11 runs). See the cost experiments.
- Calls per attempt: Agent averages 49.5, inside the panel's 20.8 to 88.2, so calls alone do not explain the cost gap. A panel call is one bash step; an Agent call is any model request.
- Time: Agent's median is 9.6 minutes per attempt (range 1.6 to 54.1, not an interval).
The cost bases differ: Agent's numbers are list-price estimates for subscription calls, not invoices.
Where a pipeline spends
Agent's notional spend was $92.64 over 33 attempts. These shares are a calculation on the chart values:
- Act (edit and run): $44.05, 48%.
- Research: $20.56, 22%.
- Verify and review: $14.30 ($8.72 and $5.58), 15%.
- Compaction, onboarding notes, other calls: $13.74, 15%.
The checks are not where most money goes. Agent recorded 162.9M input tokens and 1.8M output tokens. Cache served 94.0% of the input, because the loop re-reads its context on every call. At list price (a calculation), cache writes ($38.99) and cache reads ($30.62) are about 80% of $87.23, the cost of the recorded tokens. The $92.64 total also counts compaction calls. We have no panel token counts, so we cannot compare context sizes.
How to compare harnesses fairly
- Change one thing. Keep the model, tasks and validators fixed. This study has no bash-only run of Sonnet 5.5, the model Agent used.
- Count every attempt. Our 33 include 3 empty patches from platform holds. They count as failures.
- Use one price basis. Report cost per resolved task and time as a median with its range.
- Show n and the 95% interval. Say tie when intervals overlap. Say when a task set hits a ceiling.
In our coding calibration, the model stayed fixed across four platform builds, and verified deliveries went from 0 of 3 to 1 of 3. That is one attempt per task per build, and the builds differ in caps and fixes too. It finds defects. It measures no rate and does not show a cause.
Frequently asked questions
Is the harness or the model more important?
We did not measure that directly. With the model fixed, the route changed time: GPT-6.1 Sol took a median 17.3 s through the API and 61.2 s through the Codex CLI (ranges 16.3 to 18.6 s and 59.9 to 69.5 s; n 3 each, directional). On resolved rate, the pipeline (25 of 33, 59.0% to 87.2%) and the bash-only runs (21 to 28 of 33) overlap, and the models differ. The data cannot split model from harness.
Why does a pipeline cost more per task?
A pipeline adds stages, but we did not isolate the cause. Agent's notional cost was $2.81 per attempt (n 33), against $0.051 to $0.861 in API cost for the bash-only runs. Verify and review took 15% of the spend (calculation), so the added checks are not the main cost. The cost bases differ, so the gap is not exact.
Is Claude Code an agent harness?
Yes. Claude Code is a CLI agent: a loop, tools and a system prompt around a Claude model. In our scheduler repair it took a median 15.0 s with Sonnet 5.5 (range 13.9 to 15.9 s, n 3). That model differs from the GPT-6.1 Sol runs, so it does not compare harnesses.
How do I test a harness?
Fix the model, tasks and validators. Run the harness and a bare baseline on the same tasks. Count every attempt. Report cost per resolved task, the median time with its range, n and the 95% interval.
Watch the data
SWE-bench Verified: Agent vs 11 public model runs
Agent resolved 25 of 33 SWE-bench Verified instances; every 95% interval overlaps the public panel, so the data does not rank it.
Transcript
- SWE-bench Verified · same 33 instances. Agent vs 11 public model runs. One attempt each. No retries. 95% intervals on individual run rates. The panel mean has no interval.
- Agent resolved 25 of 33. The public panel averaged 74.1% on the same instances. Agent resolved (full pipeline, Sonnet 5.5): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Agent model cost per attempt (calculation, notional): $2.81 (n = 33). Caveat: Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
- Panel runs resolved 21 to 28 of 33. Every interval overlaps, so this sample cannot rank Agent. Chart: Resolved rate on the same 33 SWE-bench Verified instances (n = 33 each). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
- Hardest band: on the 4 instances no panel model solved, Agent solved 1. Too few to call a rate. Chart: Resolved rate by difficulty band (n = 4–14 each). Caveat: Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
- Open benchmarks: intervals, sources and every failure kept.