Harness vs model: where do AI coding agent gains really come from?
Is it the model or the harness? Our data shows the harness clearly moves speed, tokens and cost. Whether it moves accuracy, our samples cannot yet say.
TL;DR
- People argue about whether coding-agent progress comes from better models or better harnesses. We looked at what our own data can and cannot say.
- The harness clearly moves speed and tokens. The same GPT-6.1 Sol model took 17.32 s through the OpenAI API and 61.16 s through the Codex CLI on the same repair task. Both passed all 296 checks.
- The harness clearly moves cost. Agent's full pipeline spent a notional $2.81 per SWE-bench attempt. Bash-only panel runs on the same instances spent $0.051 to $0.861.
- Whether the harness moves accuracy, our samples cannot say yet. Agent resolved 25 of 33, and so did the bash-only Claude 4.5 Sonnet run. The intervals overlap.
- We say what experiment would settle it, and we have not run it yet.
Studies used: /benchmarks/swe-bench-verified, /benchmarks/cli-model-latency-tokens, /benchmarks/coding-calibration.
Two parts of every agent
Every coding agent has two parts:
- The model, which reads text and writes text.
- The harness, which decides what the model sees, which tools it can use, when to stop, and what counts as done.
A bare harness, like mini-SWE-agent, gives the model a shell and the issue and gets out of the way. A full pipeline, like Agent, adds stages: onboarding, research, a plan, edit-and-run, verification and review. A coding CLI, like Claude Code or Codex, sits in between, with its own system prompt and tool context.
When a benchmark result improves, which part did it? The honest answer is usually "both, and we did not separate them". Here is what our data separates and what it does not.
What the harness clearly changes: cost
On the same 33 SWE-bench Verified instances, the 11 public bash-only runs and Agent land in the same band of resolve rates: 63.6% to 84.8% for the panel, 75.8% for Agent. On cost, they do not overlap at all. The panel's published mean cost per instance runs from $0.051 (GPT 5 mini) to $0.861 (Claude 4.5 Opus high). Agent's notional cost is $2.81.
Where does the pipeline spend it? The act stage, $44.05 of the $92.64 total, and research, $20.56. Verification ($8.72) and review ($5.58) are smaller than you might expect. The pipeline's extra cost is not mostly in its quality gates. It is in carrying more context through every call. A bare agent call is one bash step. An Agent call carries notes, a plan and verification output, so each call is heavier, even though the call counts are similar: 49.5 per instance for Agent and 51 for the bash-only Claude 4.5 Sonnet run.
What the harness clearly changes: speed
The cleanest harness test we have holds the model fixed and changes only the route.
Same model, same effort, same small coding task:
- GPT-6.1 Sol low: 6.0 s through the API, 14.15 s through the Codex CLI.
- GPT-6.1 Sol high: 9.56 s through the API, 17.85 s through the Codex CLI.
- GPT-6 Luna: 4.01 s through the API, 9.23 s through the Codex CLI.
Every run passed its validator. The model did not change. The wrapper changed the time by a large factor. On the scheduler repair, the same pattern held: 17.32 s through the API and 61.16 s through the Codex CLI. For a one-line answer, the CLI sent about 19,551 input tokens where the API sent 17.
More in Claude Code vs Codex CLI vs the API.
What we cannot yet separate: accuracy
Here is the tempting comparison. Agent, a full pipeline on Sonnet 5.5, resolved 25 of 33. Claude 4.5 Sonnet (high) in the bash-only harness also resolved 25 of 33. So does the pipeline add nothing?
We cannot conclude that, for three reasons:
- Different models. Agent ran Sonnet 5.5. The panel ran Claude 4.5 Sonnet. There is no public bash-only run of Sonnet 5.5 on these instances.
- Small n. The interval for 25 of 33 runs from 59% to 87%. A real difference of several points would hide inside it.
- Different failure modes. The rates match, but the instances do not.
That third point is the interesting one.
Agent resolved 1 of the 4 instances that no panel model solved, and 3 of 4 that under half solved. It missed 2 of the 14 that every panel model solved. One of those misses was an empty patch after 1.6 minutes, a platform hold. A pipeline adds ways to succeed on hard problems, through research and verification. It also adds ways to fail on easy ones, through gates and holds that a bare agent does not have. On this sample, those effects roughly cancel. With four instances per hard band, that is a hypothesis, not a finding.
Same model, different harness builds
Our coding calibration holds the model fixed, claude-sonnet-5-5, and changes the platform build. One attempt per task per build.
Across four builds, verified delivery went from 0 of 3 to 1 of 3. Costs per task stayed in a narrow band, from $2.44 to $4.88. The changes between builds were all harness changes: caps, gates, context handling. One build fixed fastify/session delivery. The same build broke h3, when a formatting failure was lost across a context fold.
So the harness clearly changes outcomes on the same model. It can change them in both directions, and with one attempt per cell we cannot call it a rate. We write about why we publish these anyway in why we count every failed attempt.
The blind review study shows a related effect. With lessons and review replays from earlier attempts, the share of tasks where critics preferred the AI change rose from 6 of 12 to 9 of 12. That is harness memory at work, but it is mixed with operator answers on some tasks, so it is not a clean measurement either. See do blind AI critics prefer AI pull requests?
The experiment that would settle it
To separate harness from model on accuracy, you need:
- The same model in a bare harness and in the full pipeline.
- The same instances, with one attempt each.
- Enough instances for the intervals to separate, which means far more than 33.
- Paired analysis, such as McNemar's test, on the instances where the two disagree.
That is on our list. Until it runs, we will not claim the pipeline makes the model more accurate.
What builders can take from this today
- Measure the wrapper. Same model, different route, very different latency.
- Budget for context, not just calls. A pipeline's calls are heavier, so cost grows faster than call count.
- Expect new failure modes. Every gate you add can catch a real problem or hold a good change. Count both.
- Do not credit the harness for a model upgrade, or the model for a harness fix. Change one thing at a time when you can.
How we measured
- SWE-bench. 33 Verified instances, one attempt each, official harness. Panel figures are public mini-SWE-agent v2 per-instance results on the same instances.
- Route latency. Matched cohorts of the provider explorer: same model, same effort, same prompt, back to back on one host.
- Calibration. fastify/session, h3 and uvicorn, fixed model claude-sonnet-5-5, four platform builds, offline gates against the merged reference.
Caveats
- Small samples in every study used here.
- Different model versions between Agent and the panel.
- Notional costs for subscription calls, published API costs for the panel.
- Calibration builds differ in more than one way, and the last two remove the caps.
What to read next
- SWE-bench Verified: Agent vs 11 public models on the same tasks
- Claude Code vs Codex CLI vs the API
- What one resolved SWE-bench task really costs
See the harness at work
Open an Agent task to inspect its saved plan, actions, checks and decision records. A check that did not run stays unverified. Try Agent and review the recorded output before accepting the work.