• Coding Agents
  • Harness
  • SWE-bench
  • Latency

Harness vs model: where do AI coding agent gains really come from?

Is it the model or the harness? Our data shows the harness clearly moves speed, tokens and cost. Whether it moves accuracy, our samples cannot yet say.

TL;DR

  • People argue about whether coding-agent progress comes from better models or better harnesses. We looked at what our own data can and cannot say.
  • The harness clearly moves speed and tokens. The same GPT-6.1 Sol model took 17.32 s through the OpenAI API and 61.16 s through the Codex CLI on the same repair task. Both passed all 296 checks.
  • The harness clearly moves cost. Agent's full pipeline spent a notional $2.81 per SWE-bench attempt. Bash-only panel runs on the same instances spent $0.051 to $0.861.
  • Whether the harness moves accuracy, our samples cannot say yet. Agent resolved 25 of 33, and so did the bash-only Claude 4.5 Sonnet run. The intervals overlap.
  • We say what experiment would settle it, and we have not run it yet.

Studies used: /benchmarks/swe-bench-verified, /benchmarks/cli-model-latency-tokens, /benchmarks/coding-calibration.

Two parts of every agent

Every coding agent has two parts:

  1. The model, which reads text and writes text.
  2. The harness, which decides what the model sees, which tools it can use, when to stop, and what counts as done.

A bare harness, like mini-SWE-agent, gives the model a shell and the issue and gets out of the way. A full pipeline, like Agent, adds stages: onboarding, research, a plan, edit-and-run, verification and review. A coding CLI, like Claude Code or Codex, sits in between, with its own system prompt and tool context.

When a benchmark result improves, which part did it? The honest answer is usually "both, and we did not separate them". Here is what our data separates and what it does not.

What the harness clearly changes: cost

  • Public panel (mini-SWE-agent v2)
  • Agent
Better: upper left

12 points: Resolved rate against Mean model cost per instance (USD). Mean model cost per instance (USD) runs from $0.051 to $2.81; Resolved rate from 64% to 85%. Highlighted: Agent.

Notesn = 33 per point

Same 33 instances. Panel = API list price; Agent = notional subscription estimate

Agent's cost includes repository onboarding, planning, verification and review; it is a list-price estimate for subscription calls, not an invoice. Panel costs are published API costs.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, Anthropic list prices (Claude models)

On the same 33 SWE-bench Verified instances, the 11 public bash-only runs and Agent land in the same band of resolve rates: 63.6% to 84.8% for the panel, 75.8% for Agent. On cost, they do not overlap at all. The panel's published mean cost per instance runs from $0.051 (GPT 5 mini) to $0.861 (Claude 4.5 Opus high). Agent's notional cost is $2.81.

Where does the pipeline spend it? The act stage, $44.05 of the $92.64 total, and research, $20.56. Verification ($8.72) and review ($5.58) are smaller than you might expect. The pipeline's extra cost is not mostly in its quality gates. It is in carrying more context through every call. A bare agent call is one bash step. An Agent call carries notes, a plan and verification output, so each call is heavier, even though the call counts are similar: 49.5 per instance for Agent and 51 for the bash-only Claude 4.5 Sonnet run.

What the harness clearly changes: speed

The cleanest harness test we have holds the model fixed and changes only the route.

  • Total time
  • First useful output
OpenAI API · GPT-6 Luna · none
OpenAI API · GPT-6.1 Sol · low
Codex CLI · GPT-6 Luna · none
OpenAI API · GPT-6.1 Sol · high
Codex CLI · GPT-6.1 Sol · low
Codex CLI · GPT-6.1 Sol · high

Seconds · log scale: each gridline is 10 times the one before

6 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · high 17.9 s (range 17.7 s–22.4 s, n 3). Fastest OpenAI API · GPT-6 Luna · none 4 s (range 3.8 s–4.4 s, n 3). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · high 17.3 s (range 17.1 s–21.9 s, n 3). Fastest OpenAI API · GPT-6 Luna · none 0.7 s (range 0.6 s–0.8 s, n 3). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Matched cohort, small coding task, 3 runs per configuration

Dot = median; whiskers = fastest and slowest run (a range, not a confidence interval). Every run in these cohorts passed its validator.

Source: Provider explorer receipts: CLI vs API

Same model, same effort, same small coding task:

  • GPT-6.1 Sol low: 6.0 s through the API, 14.15 s through the Codex CLI.
  • GPT-6.1 Sol high: 9.56 s through the API, 17.85 s through the Codex CLI.
  • GPT-6 Luna: 4.01 s through the API, 9.23 s through the Codex CLI.

Every run passed its validator. The model did not change. The wrapper changed the time by a large factor. On the scheduler repair, the same pattern held: 17.32 s through the API and 61.16 s through the Codex CLI. For a one-line answer, the CLI sent about 19,551 input tokens where the API sent 17.

More in Claude Code vs Codex CLI vs the API.

What we cannot yet separate: accuracy

Here is the tempting comparison. Agent, a full pipeline on Sonnet 5.5, resolved 25 of 33. Claude 4.5 Sonnet (high) in the bash-only harness also resolved 25 of 33. So does the pipeline add nothing?

We cannot conclude that, for three reasons:

  1. Different models. Agent ran Sonnet 5.5. The panel ran Claude 4.5 Sonnet. There is no public bash-only run of Sonnet 5.5 on these instances.
  2. Small n. The interval for 25 of 33 runs from 59% to 87%. A real difference of several points would hide inside it.
  3. Different failure modes. The rates match, but the instances do not.

That third point is the interesting one.

  • Agent
  • Public panel mean (square)
In chart order.
No panel model solved it
Under half solved it
Half or more solved it
Every panel model solved it

Gap labels, Public panel mean vs Agent: Public panel mean is x percentage points higher (+) or lower (−) than Agent, calculated from the two values shown; lines are the 95% Wilson interval.

4 difficulty bands, 2 series: Agent, Public panel mean. Agent: highest Every panel model solved it 86% (95% interval 60%–96%, n 14). Lowest No panel model solved it 25% (95% interval 4.6%–70%, n 4). All intervals overlap. Public panel mean: highest Every panel model solved it 100% (n 14). Lowest No panel model solved it 0% (n 4). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 4–14 per row

Band = how many of the 11 public panel models solved the instance

Agent whiskers are 95% Wilson intervals. The "no panel model solved it" band has 4 instances; Agent resolved 1 (matplotlib__matplotlib-21568).

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, SWE-bench campaign rules and sample design

Agent resolved 1 of the 4 instances that no panel model solved, and 3 of 4 that under half solved. It missed 2 of the 14 that every panel model solved. One of those misses was an empty patch after 1.6 minutes, a platform hold. A pipeline adds ways to succeed on hard problems, through research and verification. It also adds ways to fail on easy ones, through gates and holds that a bare agent does not have. On this sample, those effects roughly cancel. With four instances per hard band, that is a hypothesis, not a finding.

Same model, different harness builds

Our coding calibration holds the model fixed, claude-sonnet-5-5, and changes the platform build. One attempt per task per build.

Baseline (capped)

fastify/session
h3js/h3
Kludex/uvicorn

Fix wave 1 (capped)

fastify/session
h3js/h3
Kludex/uvicorn

Uncapped, build f0ac3a8a

fastify/session
h3js/h3
Kludex/uvicorn

Uncapped, build 236c0d3f

fastify/session
h3js/h3
Kludex/uvicorn

One panel per series, all on the same axis.

3 rows, 4 series: Baseline (capped), Fix wave 1 (capped), Uncapped, build f0ac3a8a, Uncapped, build 236c0d3f. Baseline (capped): highest Kludex/uvicorn $4.88. Lowest fastify/session $3.43. Fix wave 1 (capped): highest fastify/session $4.85. Lowest h3js/h3 $3.29.

Notes

List-price estimates of subscription calls. Failed and capped attempts count.

Source: Coding calibration: fastify/session, h3, uvicorn

Across four builds, verified delivery went from 0 of 3 to 1 of 3. Costs per task stayed in a narrow band, from $2.44 to $4.88. The changes between builds were all harness changes: caps, gates, context handling. One build fixed fastify/session delivery. The same build broke h3, when a formatting failure was lost across a context fold.

So the harness clearly changes outcomes on the same model. It can change them in both directions, and with one attempt per cell we cannot call it a rate. We write about why we publish these anyway in why we count every failed attempt.

The blind review study shows a related effect. With lessons and review replays from earlier attempts, the share of tasks where critics preferred the AI change rose from 6 of 12 to 9 of 12. That is harness memory at work, but it is mixed with operator answers on some tasks, so it is not a clean measurement either. See do blind AI critics prefer AI pull requests?

The experiment that would settle it

To separate harness from model on accuracy, you need:

  1. The same model in a bare harness and in the full pipeline.
  2. The same instances, with one attempt each.
  3. Enough instances for the intervals to separate, which means far more than 33.
  4. Paired analysis, such as McNemar's test, on the instances where the two disagree.

That is on our list. Until it runs, we will not claim the pipeline makes the model more accurate.

What builders can take from this today

  • Measure the wrapper. Same model, different route, very different latency.
  • Budget for context, not just calls. A pipeline's calls are heavier, so cost grows faster than call count.
  • Expect new failure modes. Every gate you add can catch a real problem or hold a good change. Count both.
  • Do not credit the harness for a model upgrade, or the model for a harness fix. Change one thing at a time when you can.

How we measured

  • SWE-bench. 33 Verified instances, one attempt each, official harness. Panel figures are public mini-SWE-agent v2 per-instance results on the same instances.
  • Route latency. Matched cohorts of the provider explorer: same model, same effort, same prompt, back to back on one host.
  • Calibration. fastify/session, h3 and uvicorn, fixed model claude-sonnet-5-5, four platform builds, offline gates against the merged reference.

Caveats

  • Small samples in every study used here.
  • Different model versions between Agent and the panel.
  • Notional costs for subscription calls, published API costs for the panel.
  • Calibration builds differ in more than one way, and the last two remove the caps.

See the harness at work

Open an Agent task to inspect its saved plan, actions, checks and decision records. A check that did not run stays unverified. Try Agent and review the recorded output before accepting the work.

The data behind this post

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

  • Calibration
  • Coding Agents

Coding calibration: what broke on three real pull requests

One attempt per task, four platform builds, failures kept: how an AI worker did on real fastify/session, h3 and uvicorn issues, and what broke.

1 of 3Verified deliveries, latest build · n = 3

4 chartsUpdated October 5, 2026

  • Code Review
  • AI vs human

AI pull requests vs merged human pull requests, judged blind

A blind panel of Claude and GPT critics preferred Agent's change over the merged human change on 9 of 12 real tasks. Votes, scores, caveats.

75% (9/12)Tasks where the panel preferred the AI change (latest attempt) · n = 12

4 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.