One real attempt, annotated

A receipt for one real AI agent task

Agent tried one public coding task from SWE-bench Verified. This record shows the result, the minutes, the model calls and the notional cost. It also shows where the attempt sits among all 33.

Receipt

django__django-14559

SWE-bench Verified · one attempt · claude-sonnet-5-5

Resolved

$2.71

n = 1

Notional model cost of this attempt

  • 7.2min

    Worker time

    n = 1

  • 47

    Model calls

    n = 1

  • 11/11

    Public panel models that resolved it

Position among all 33 attemptsCalculation

  • 17th lowest cost of 33
  • 9th shortest time of 33
  • 14th fewest calls of 33 (2 others tied)

The official SWE-bench harness graded this patch as resolved.

Source: Agent on SWE-bench Verified, campaign 1 (25 instances), recorded 2026-10-04 · Raw attempts file

How to read this receipt

What the receipt shows

  • The outcome. The official SWE-bench harness grades the patch that Agent delivered.
  • The minutes. This is the worker time of the attempt.
  • The model calls. One call is one model request, of any stage.
  • The notional cost. This is the list-price estimate for the attempt, onboarding included.

What it does not show

  • An invoice. The calls ran on a subscription. We price them at list price, so no bill backs this number.
  • A rate. This is one attempt. Across all 33 attempts, Agent resolved 76% (25/33), with a 95% interval of 59% to 87%.
  • The difficulty. The study measures difficulty by how many of the 11 public runs solved an instance. All 11 solved this one, so it sits at the easy end. This receipt says nothing about a hard task.
  • Customer work. This is a public benchmark task, not work for a client.
  • The steps. This page shows four values for the attempt. It does not list each action.

Grading uses the official SWE-bench harness and instance images (under amd64 emulation). The gold patch resolved on the host for every instance.

This attempt among all 33

Each row sorts the 33 attempts by one value. The large dot is this attempt. The muted dots are the other 32.

All 33 attempts: cost, time and model calls

One row per value. Lower values sit on the left.

Cost (notional)n = 1 · list-price estimate
$2.71
Timen = 1 · worker time of the attempt
7.2 min
Model callsn = 1 · model requests, any stage
47

n = 1 attempt per dotDots are table values. Positions are counts of table values (calculation).Cost is notional.

This attempt: cost $2.71, 7.2 min, 47 model calls. Position among 33 attempts: 17th lowest cost, 9th shortest time, 14th fewest calls (2 others tied).

Calculation Of the 33 attempts, 16 cost less and 16 cost more.

Across all 33 attempts: median time 9.6 min, mean cost $2.81, mean 49 model calls. These are study stats (n = 33).

A panel call is one bash-agent step. An Agent call is one model request of any stage (research, plan, act, verify, review). Model calls per instance, next to the public panel

A failed attempt gets a receipt too

8 of 33 attempts did not resolve (5 unresolved, 3 empty patch). The totals on the study page include them.

Failure receipt

astropy__astropy-7336

SWE-bench Verified · one attempt · claude-sonnet-5-5

Not resolved

$2.03

Notional model cost of this attempt

n = 1

  • 23min

    Worker time

    n = 1

  • 48

    Model calls

    n = 1

  • 11/11

    Public panel models that resolved it

Position among all 33 attemptsCalculation

  • 7th lowest cost of 33
  • 28th shortest time of 33
  • 17th fewest calls of 33

The official SWE-bench harness graded this patch as not resolved.

Source: Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), recorded 2026-10-05 · Raw attempts file

All 11 public panel models resolved this instance. This attempt did not.

Calculation The 8 attempts that did not resolve cost $20.93 together (calculation: the sum of their table values).

Cost per resolved instance divides all spend, failed attempts included, by the 25 resolved. That gives $3.71 (n = 25). Cost per attempt is $2.81 (n = 33).

For scale, the public bash-only runs spent a mean of $0.57 per resolved instance (published API cost, n = 11 runs). Agent runs a full pipeline, so it spends more. See the comparison

Where the money goes, for all 33 attempts together

The dataset keeps the stage split for all 33 attempts together. It has no split for this attempt, so this chart is not this receipt.

All 33 attempts, not this one

Largest value is 21x the smallest; Log shows the small bars.
Act (edit and run)
Research
Verify
Other
Review
Context compaction
Onboarding notes

Hover or focus a bar for its ratio to Onboarding notes (the lowest value): a ratio of the two values shown, not a measurement.

7 rows. Highest Act (edit and run) $44.05. Lowest Onboarding notes $2.09.

Notes

Share of notional model cost by stage, all 33 attempts

Total $92.64 over 33 attempts.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

94.0% of the input came from the cache, at the lower cache-read price. This holds for all 33 attempts together, not for this attempt.

What could make this receipt misleading

These caveats come from the study, in its own words.
  • n = 33: intervals are wide. This is a defect-finding run, not a ranking.
  • Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
  • Verified issues are public (2015 to 2023) and likely in every model’s training data; contamination is uncontrolled for all systems.
  • Three campaign-1 empty patches were platform holds before delivery (missing lint tools, a too-literal plan gate, an unanswered question), not wrong fixes. They count as failures here.
All caveats and the method

Read next