Explainer · Agent harness

What is an agent harness? The code around the model, measured

Definition

Agent harness is the code that wraps a language model and turns it into an agent. The model answers; the harness runs the loop, offers tools, builds prompts, keeps memory, manages context and checks the result. Two systems can share one model and still differ in time and tokens, so a benchmark scores a system, not a model.

Agent team · · 5 min read · Every number is from the public studies

How it works

Where a full pipeline spends

Agent's notional model cost by stage, summed over 33 SWE-bench Verified attempts.

MeasuredNotional cost
  1. Onboarding

    $2.09

    The pipeline reads the repository and writes notes for later calls.

    • Notional list-price estimate
  2. Research

    $20.56

    The model reads code and docs before it changes anything.

    • Second-largest share of spend
  3. Act

    $44.05

    The model edits files and runs commands.

    • The largest share of spend
  4. Verify

    $8.72

    The pipeline checks the change before it delivers.

    • Checks cost model calls too
  5. Review

    $5.58

    A review pass reads the change last.

    • Smaller than act and research

Five stages of one full pipeline, with Agent's notional model cost for each. The act stage, where the model edits and runs code, costs the most: $44.05, about 48% of the total spend (calculation). List-price estimates for subscription calls, not invoices.

Context compaction ($5.41) and other calls ($6.24) are not drawn as stages, and the chart has no separate line for planning. Source: chart swebench-cost-by-stage; n 33 attempts.

Source: SWE-bench Verified study

What an agent harness does

The model reads text and writes text. It cannot open a file or run a test. The harness does those jobs, and some teams call it a scaffold:

  • The loop. It sends a request, runs the tool the model asks for and returns the result.
  • Tools and prompts. It decides which actions exist and writes the system prompt.
  • Memory and context. It keeps notes between calls and chooses what stays in context.
  • Checks. It runs tests, reviews the patch and decides when to stop.

Three kinds of agent harness

  • A bare API call. Your code sends a prompt and reads the answer.
  • A CLI agent. Claude Code and the Codex CLI add a system prompt, tools and a loop.
  • A full pipeline. Agent, one system in our data, runs onboarding, research, a plan, edits, verification and review on one model.

The 11 public runs in our panel use mini-SWE-agent v2, a bash-only harness: one model acting only through a shell.

What the harness changed in our data

Same model, different route: time and tokens

GPT-6.1 Sol repaired the same scheduler task through the OpenAI API and the Codex CLI. Only the route changed:

  • Total time
  • First useful output
Entrance: medians race at 44× real time
Claude Code CLI · Sonnet 5.5 · medium
Codex CLI · GPT-6.1 Sol · medium
OpenAI API · GPT-6.1 Sol · medium

3 rows, 2 series: Total time, First useful output. Total time: slowest Codex CLI · GPT-6.1 Sol · medium 61.2 s (range 59.9 s–69.5 s, n 3). Fastest Claude Code CLI · Sonnet 5.5 · medium 15 s (range 13.9 s–15.9 s, n 3). Not all run ranges overlap. First useful output: slowest Codex CLI · GPT-6.1 Sol · medium 15.6 s (range 13.7 s–23 s, n 3). Fastest OpenAI API · GPT-6.1 Sol · medium 7.5 s (range 6.7 s–9.1 s, n 3). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 3 per row

Same prompt, medium effort, 296 behavioral checks, 3 runs each

All 9 runs passed all 296 checks. Dot = median, whiskers = range. Different models (Sonnet 5.5 vs GPT-6.1 Sol), so this compares route + model pairs, not routes alone.

Source: Provider explorer receipts: CLI vs API

  • OpenAI API: median 17.3 s, range 16.3 to 18.6 s, n 3.
  • Codex CLI: median 61.2 s, range 59.9 to 69.5 s, n 3.

All 9 runs passed all 296 checks, so this task has a ceiling for quality. A range is not a confidence interval. The ranges do not overlap, but 3 runs per side is too few for a tested ranking, so read this as directional. The chart also shows Claude Code with Sonnet 5.5, a different model, so it does not isolate the route.

The harness also adds tokens, a context tax. For a one-line request, the Codex CLI sent a median 19,551 input tokens (n 15) and the API sent 17. Part of the CLI input came from cache. Over 15 runs each, the CLI took a median 3.9 s and the API 1.1 s. The ranges, a calculation over three cells, are 2.9 to 4.7 s and 0.7 to 2.2 s.

Full pipeline, bash-only harness: cost, not resolved rate

On 33 SWE-bench Verified instances, Agent (a full pipeline on Sonnet 5.5) resolved 25 of 33: 75.8%, 95% interval 59.0% to 87.2%. Eleven public runs, one model each in a bash-only harness, resolved 21 to 28 of the same 33 (n 33 each). The lowest, 21 of 33, has a 95% interval of 46.6% to 77.8%; the highest, 28 of 33, 69.1% to 93.3%.

Every interval overlaps, so this sample cannot rank any system. With n 33, one interval spans about 28 points (calculation: 87.2 minus 59.0), so the set cannot see a small gap. The two sides use different models, so this compares systems, not harnesses alone. Verified issues are public, so contamination is uncontrolled for every system.

  • Public panel (mini-SWE-agent v2)
  • Agent
Better: upper left

12 points: Resolved rate against Mean model cost per instance (USD). Mean model cost per instance (USD) runs from $0.051 to $2.81; Resolved rate from 64% to 85%. Highlighted: Agent.

Notesn = 33 per point

Same 33 instances. Panel = API list price; Agent = notional subscription estimate

Agent's cost includes repository onboarding, planning, verification and review; it is a list-price estimate for subscription calls, not an invoice. Panel costs are published API costs.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, Anthropic list prices (Claude models)

Cost and effort:

  • Cost per attempt: Agent $2.81 notional (n 33). Panel runs average $0.051 to $0.861 in published API cost.
  • Cost per resolved instance: Agent $3.71 notional (n 25). The panel mean is $0.57 (recorded, 11 runs). See the cost experiments.
  • Calls per attempt: Agent averages 49.5, inside the panel's 20.8 to 88.2, so calls alone do not explain the cost gap. A panel call is one bash step; an Agent call is any model request.
  • Time: Agent's median is 9.6 minutes per attempt (range 1.6 to 54.1, not an interval).

The cost bases differ: Agent's numbers are list-price estimates for subscription calls, not invoices.

Where a pipeline spends

$92.64Total over 33 attempts (from the note)

Parts sorted by value, largest first

  1. Act (edit and run)
  2. Research
  3. Verify
  4. Other
  5. Review
  6. Context compaction
  7. Onboarding notes

Shares are calculated from the values shown; rounding can make the sum of the parts differ from the stated total by a cent.

7 rows. Highest Act (edit and run) $44.05. Lowest Onboarding notes $2.09.

Notes

Share of notional model cost by stage, all 33 attempts

Total $92.64 over 33 attempts.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

Agent's notional spend was $92.64 over 33 attempts. These shares are a calculation on the chart values:

  • Act (edit and run): $44.05, 48%.
  • Research: $20.56, 22%.
  • Verify and review: $14.30 ($8.72 and $5.58), 15%.
  • Compaction, onboarding notes, other calls: $13.74, 15%.

The checks are not where most money goes. Agent recorded 162.9M input tokens and 1.8M output tokens. Cache served 94.0% of the input, because the loop re-reads its context on every call. At list price (a calculation), cache writes ($38.99) and cache reads ($30.62) are about 80% of $87.23, the cost of the recorded tokens. The $92.64 total also counts compaction calls. We have no panel token counts, so we cannot compare context sizes.

How to compare harnesses fairly

  1. Change one thing. Keep the model, tasks and validators fixed. This study has no bash-only run of Sonnet 5.5, the model Agent used.
  2. Count every attempt. Our 33 include 3 empty patches from platform holds. They count as failures.
  3. Use one price basis. Report cost per resolved task and time as a median with its range.
  4. Show n and the 95% interval. Say tie when intervals overlap. Say when a task set hits a ceiling.

In our coding calibration, the model stayed fixed across four platform builds, and verified deliveries went from 0 of 3 to 1 of 3. That is one attempt per task per build, and the builds differ in caps and fixes too. It finds defects. It measures no rate and does not show a cause.

Frequently asked questions

Is the harness or the model more important?

We did not measure that directly. With the model fixed, the route changed time: GPT-6.1 Sol took a median 17.3 s through the API and 61.2 s through the Codex CLI (ranges 16.3 to 18.6 s and 59.9 to 69.5 s; n 3 each, directional). On resolved rate, the pipeline (25 of 33, 59.0% to 87.2%) and the bash-only runs (21 to 28 of 33) overlap, and the models differ. The data cannot split model from harness.

Why does a pipeline cost more per task?

A pipeline adds stages, but we did not isolate the cause. Agent's notional cost was $2.81 per attempt (n 33), against $0.051 to $0.861 in API cost for the bash-only runs. Verify and review took 15% of the spend (calculation), so the added checks are not the main cost. The cost bases differ, so the gap is not exact.

Is Claude Code an agent harness?

Yes. Claude Code is a CLI agent: a loop, tools and a system prompt around a Claude model. In our scheduler repair it took a median 15.0 s with Sonnet 5.5 (range 13.9 to 15.9 s, n 3). That model differs from the GPT-6.1 Sol runs, so it does not compare harnesses.

How do I test a harness?

Fix the model, tasks and validators. Run the harness and a bare baseline on the same tasks. Count every attempt. Report cost per resolved task, the median time with its range, n and the 95% interval.

Watch the data

Live story · 28 sSWE-bench Verified: Agent vs 11 public model runs

SWE-bench Verified: Agent vs 11 public model runs

Agent resolved 25 of 33 SWE-bench Verified instances; every 95% interval overlaps the public panel, so the data does not rank it.

Transcript
  1. SWE-bench Verified · same 33 instances. Agent vs 11 public model runs. One attempt each. No retries. 95% intervals on individual run rates. The panel mean has no interval.
  2. Agent resolved 25 of 33. The public panel averaged 74.1% on the same instances. Agent resolved (full pipeline, Sonnet 5.5): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Agent model cost per attempt (calculation, notional): $2.81 (n = 33). Caveat: Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
  3. Panel runs resolved 21 to 28 of 33. Every interval overlaps, so this sample cannot rank Agent. Chart: Resolved rate on the same 33 SWE-bench Verified instances (n = 33 each). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
  4. Hardest band: on the 4 instances no panel model solved, Agent solved 1. Too few to call a rate. Chart: Resolved rate by difficulty band (n = 4–14 each). Caveat: Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
  5. Open benchmarks: intervals, sources and every failure kept.

The data behind this explainer

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.