• Cost
  • LLM pricing
  • SWE-bench
  • Benchmarks

Devin's $0.60 per task and our $3.71 per resolved task are different numbers

Devin reports $0.60 per task. Our list-price calculation gives $3.71 per resolved SWE-bench task. Compare the units with a table and buyer checklist.

TL;DR

  • The two numbers count different things, so you cannot rank them. Cognition reports an average of $0.60 per task for Devin Fusion on FrontierCode 1.1 Extended, with a weighted score of 68.8 (vendor-reported, read 2026-10-07). Our calculation gives $3.71 per resolved task for Agent on SWE-bench Verified: 25 of 33 resolved (76%, 95% interval 59% to 87%). It uses notional list-price costs, not an invoice.
  • A score is not a pass rate. $0.60 ÷ 0.688 = $0.87 is a calculation, not a cost per solved task. Cognition does not report it in the post text we read, and we do not use it.
  • One run, five costs. Our own 33 attempts give five different costs, from $2.81 to $13.73, depending on what you count (calculations on dataset values and raw rows).
  • A vendor number is a snapshot. Cognition's Devin Fusion post (dated June 29, chart data updated August 7) lists Devin Fusion at $1.35 per task and a score of 63.1. Its September 28 post lists $0.60 and 68.8 (first-party, read 2026-10-07).
  • Ask eight questions before you trust any cost-per-task claim, ours included. The checklist is below.
  • Disclosure: I build Agent. We did not run Devin. This post does not say that one system is cheaper or better.

The short answer

You see a cost per task in vendor posts, benchmark charts and sales calls. It sounds like one thing. It is at least five.

Cognition's September 28, 2026 post says Devin Fusion reached a score of 68.8 on FrontierCode 1.1 Extended, "at $0.60 per task on average". Our SWE-bench Verified cost calculation gives $3.71 per resolved task and $2.81 per attempt, at notional list prices.

A reader wants to divide one by the other. Do not. The benchmarks differ. The units differ. The quality measures differ. The price bases differ. The sections below show where.

The same post says Fusion leads FrontierCode 1.1 on score. That is Cognition's claim. This post does not test it.

This post sets a claim next to receipts. It does not compare Devin with Agent, because we did not run Devin.

What each number counts

Read this table row by row. Two rows cause most of the confusion. One is the unit: per task or per solved task. The other is the quality number: a share of tasks solved, or not.

QuestionCognition: $0.60 (Devin Fusion)Agent: $2.81 and $3.71
Who reports itCognition, first party. We found no independent reproduction. Read 2026-10-07The author, who builds Agent. Every attempt is in the open dataset
BenchmarkFrontierCode 1.1 Extended: 150 tasks, kept private, graded against a rubricSWE-bench Verified: 33 of 500 public issues, graded by the official tests
UnitAverage per task. A second chart in the June 29 Fusion post says per rollout, which we read as one run of one taskCalculations: $2.81 per attempt. $3.71 per resolved task
Quality number beside itWeighted score 68.8. A failed blocker zeroes or caps a task's score (the two method posts word it differently). This is not a pass rateResolved 25 of 33 (76%, 95% interval 59% to 87%)
Failed attempts in the costNot stated for this pointYes. All 33 attempts, including 5 unresolved and 3 empty patches
Retries and repeatsThe method post says 5 runs per model at each available effort. The count for this point is not publishedOne attempt per instance. No retries
Prompt cacheNot reported. The cache diagrams in the post are labelled illustrative94.0% of input tokens read from the cache, priced at list cache rates
What the cost coversNot stated: grading, test and VM compute, setupModel tokens only, including onboarding, planning, verification and review. No test compute, no grading, no human time
Invoice or estimateNot shown to be an invoice. Devin bills in credits or Agent Compute Units (billing docs read 2026-10-07)A list-price estimate of subscription calls. No invoice
n and intervalTask set: 150. The point's run count is not published. No interval in the post text we readn = 33, 25 resolved. Interval on the rate only

Where the table says "not stated", Cognition may know the answer. The public post does not give it. We do not guess.

Why you cannot divide a cost by a score

The division that works

On our run, the calculated cost per attempt is $2.807 and the resolved rate is 25/33 = 0.7576. Calculation: $2.807 ÷ 0.7576 = $3.71. That equals our cost per resolved task.

It works for two reasons. First, 25/33 counts tasks: 25 solved out of 33. Second, the $2.807 holds the spend of every attempt, failed ones too.

The division that does not

Try the same step on Cognition's figures: $0.60 ÷ 0.688 = $0.87 (a calculation). It looks the same. It is not, for two reasons.

  1. The 68.8 is a score, not a pass rate. The FrontierCode method post defines the score as a weighted aggregate of rubric items. It says a solution that fails a blocker scores zero. The 1.1 post says a failed blocker caps the score. Either way, 68.8 does not mean that Fusion solved 68.8% of tasks. The announcement does not give the pass rate for the Fusion point.
  2. The $0.60 may or may not hold every attempt. The post does not say how failed, aborted or repeated runs count.

A cost divided by something that is not a share of tasks does not mean dollars per solved task. Report a cost per resolved task only when you know both parts. You need the total spend over all attempts. You also need the number of tasks that passed.

A sensitivity calculation, not a cost interval

The rate 25/33 has a 95% interval of 59% to 87%. Hold the cost per attempt fixed (a simplification) and the same division gives $3.22 to $4.76 per resolved task. That is a sensitivity calculation: $2.807 ÷ 0.8717 and $2.807 ÷ 0.5898. The range is $1.54 wide (calculation). It is not a 95% cost interval.

It ignores cost variation and the link between cost and success. The observed run still gives $3.71 per resolved task.

Our receipts: one run, five costs

Here is the same data as a live story. Then the numbers.

Live story · 28 sSWE-bench Verified: Agent vs 11 public model runs

SWE-bench Verified: Agent vs 11 public model runs

Agent resolved 25 of 33 SWE-bench Verified instances; every 95% interval overlaps the public panel, so the data does not rank it.

Transcript
  1. SWE-bench Verified · same 33 instances. Agent vs 11 public model runs. One attempt each. No retries. 95% intervals on individual run rates. The panel mean has no interval.
  2. Agent resolved 25 of 33. The public panel averaged 74.1% on the same instances. Agent resolved (full pipeline, Sonnet 5.5): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Agent model cost per attempt (calculation, notional): $2.81 (n = 33). Caveat: Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.
  3. Panel runs resolved 21 to 28 of 33. Every interval overlaps, so this sample cannot rank Agent. Chart: Resolved rate on the same 33 SWE-bench Verified instances (n = 33 each). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
  4. Hardest band: on the 4 instances no panel model solved, Agent solved 1. Too few to call a rate. Chart: Resolved rate by difficulty band (n = 4–14 each). Caveat: Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.
  5. Open benchmarks: intervals, sources and every failure kept.

We ran 33 SWE-bench Verified instances through Agent's full pipeline, one attempt each. The platform logged $92.64 of notional model cost. We counted it five ways.

What we countCalculationResult
Per attempt, every attempt$92.64 ÷ 33$2.81
Per resolved task, every attempt's cost in the numerator$92.64 ÷ 25$3.71
Per resolved task, resolved attempts' cost only$71.72 ÷ 25$2.87
Per resolved task, model tokens only at Sonnet 5.5 list price (no compaction calls)$87.23 ÷ 25$3.49
Per resolved task, if no prompt were ever cached (a counterfactual)$343.33 ÷ 25$13.73

These are calculations on dataset values and the public raw attempt rows. Sum raw costs before rounding to get the $92.64 total. Summing the dataset's already-rounded per-instance costs gives $92.63. That rounding also changes the resolved and failed subtotals.

Each choice moves the answer:

  • Failed attempts. 8 of the 33 attempts did not resolve: 5 unresolved and 3 empty patches. Their notional cost totals $20.92, or 22.6% of the spend (calculations on raw rows). Leave them out and the cost per resolved task falls from $3.71 to $2.87. Both numbers are calculations. Only $3.71 shows what the whole run spent for each solved task.
  • Cache. 94.0% of our input tokens came from the cache. At Sonnet 5.5 list prices, the same tokens cost $343.33 without caching, against $87.23 with it. That is 3.9 times as much (calculation). A claim that assumes a cache hit rate you do not get will not match your bill.
  • Scope. Onboarding notes ($2.09), verification ($8.72) and review ($5.58) total $16.39, or 17.7% of the $92.64 (calculations). These stages form part of our accounting. We did not isolate their effect in a matched experiment.
$92.64Total over 33 attempts (from the note)

Parts sorted by value, largest first

  1. Act (edit and run)
  2. Research
  3. Verify
  4. Other
  5. Review
  6. Context compaction
  7. Onboarding notes

Shares are calculated from the values shown; rounding can make the sum of the parts differ from the stated total by a cent.

7 rows. Highest Act (edit and run) $44.05. Lowest Onboarding notes $2.09.

Notes

Share of notional model cost by stage, all 33 attempts

Total $92.64 over 33 attempts.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

Averages hide spread. Across n = 33, a single attempt's notional list-price cost calculation ranged from $0.89 (an empty patch after 13 calls) to $4.83 (resolved after 67 calls). The median attempt took 9.6 minutes, with a range of 1.6 to 54.1 minutes. This run range is not a 95% interval. Attempts averaged 49.5 model calls (calculation: 1,632 ÷ 33).

Where this sits next to public runs

On the same 33 instances, 11 public mini-SWE-agent v2 runs resolved between 21 and 28. Their published API costs per attempt ran from $0.05 to $0.86.

  • Public panel (mini-SWE-agent v2)
  • Agent
Better: upper left

12 points: Resolved rate against Mean model cost per instance (USD). Mean model cost per instance (USD) runs from $0.051 to $2.81; Resolved rate from 64% to 85%. Highlighted: Agent.

Notesn = 33 per point

Same 33 instances. Panel = API list price; Agent = notional subscription estimate

Agent's cost includes repository onboarding, planning, verification and review; it is a list-price estimate for subscription calls, not an invoice. Panel costs are published API costs.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs, Anthropic list prices (Claude models)

Agent resolved 25 of 33 (76%, 95% interval 59% to 87%). Every panel interval overlaps Agent's, so this sample cannot rank Agent above or below any panel model. On cost, our point sits furthest to the right.

Largest value is 46x the smallest; Log shows the small bars.
Agent (notional)
Claude 4.5 Opus (high)
Claude 4.5 Sonnet (high)
Claude 4.6 Opus
GLM 5 (high)
DeepSeek V3.2 (high)
GPT 5.2 (high)
Claude 4.5 Haiku (high)
Gemini 3 Flash (high)
Kimi K2.5 (high)
MiniMax M2.5 (high)
GPT 5 mini

Hover or focus a bar for its ratio to Agent (notional) (the highlighted row): a ratio of the two values shown, not a measurement.

12 rows. Highest Agent (notional) $3.71 (n 25). Lowest GPT 5 mini $0.08 (n 21).

Notesn 21–28 per row

Same 33 SWE-bench Verified instances; all attempts in the numerator

Recorded figures, not repricing. Panel costs are published API costs for a bash-only agent. Agent's figure is a list-price estimate of subscription calls and includes onboarding, planning, verification and review.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

Per resolved task, the panel's published costs run from $0.08 to $1.18, with a mean of $0.57 (11 runs). Our list-price calculation gives $3.71. That is the highest plotted point, 3.1 times the next one (a calculation: $3.706 ÷ $1.184). No cost interval is recorded, so the gap is descriptive. I would rather you read that from us. It does not say anything about Devin.

Two limits apply to these bars. Our full pipeline accounts for onboarding, planning, verification and review; the panel uses a bash-only harness. Our cost is a list-price estimate for subscription calls, with no invoice behind it. These data do not isolate why the costs differ.

For a row-by-row view against one public run, see Agent vs Claude Sonnet 4.5. It marks each row a tie or unclear, and it says where no interval was recorded.

A vendor number is a snapshot

Cognition's June 29 Devin Fusion post charts the average cost per task on FrontierCode 1.1 Extended. A second chart in the same post gives the cost per rollout, which we read as one run of one task. The chart data, last updated 2026-08-07, lists Devin Fusion at $1.35 and a score of 63.1. The September 28 post lists $0.60 and 68.8. I read both on 2026-10-07. Both are first-party. Neither post says if the two points share one configuration.

Calculation: $0.60 is 55.6% below $1.35, and 52 days separate August 7 and September 28. Cognition says its September cost cuts come from newer models, including SWE-2, and from changes to its harness. We did not test that explanation.

It shows what a cost per task is. It belongs to one configuration on one date. Ask for both. Our $3.71 is a snapshot too: it mixes two platform builds, one before and one after a sandbox fix.

A buyer checklist for any cost-per-task claim

Use these questions on Cognition, on us and on any other vendor.

  1. Cost per resolved task, not per task. Ask for the total spend over all attempts, divided by the tasks that passed. A per-task average hides how many tasks failed.
  2. n and the interval. Ask how many tasks and how many runs per task. Our rate on 33 tasks has a 95% interval from 59% to 87%, about 28 percentage points wide.
  3. Failed, aborted and retried attempts. Ask if they are in the cost. Our calculations give $2.87 without failed attempts and $3.71 with them. No retries were run.
  4. The type of the quality number. Ask for the pass rate next to any score. Ask what a failed blocker does to the score.
  5. Cache share. Ask what share of input tokens the cache served. Ours is 94.0%. Without the cache, the same tokens cost 3.9 times as much (a calculation).
  6. List price or invoice. Ask which one the number is. Then ask how the chart's dollars map to your bill. Devin bills in credits and Agent Compute Units. Our run used a subscription, so ours is an estimate.
  7. What is inside the cost. Ask about setup, onboarding, verification, review, grading, test compute and human time. Ours covers model tokens only.
  8. The task set, the configuration and the date. Ask if the tasks look like yours. Public tasks can sit in training data; SWE-bench Verified issues date from 2015 to 2023. You cannot check private tasks. Ask which models, which effort and which build ran, and when.

Here is a receipt line you can ask any vendor to fill in:

Cost per resolved task: $X (total $S over A attempts; R resolved; resolved rate L% to U% at 95%)
Counted: failed attempts, aborted attempts, retries, cache share, setup, review, grading, human time (yes or no each)
Price basis: invoice or list price, and the date
Task set: name, n, public or private
Configuration: models, effort, harness build, run date

And here is ours, filled in:

Cost per resolved task (calculation): $3.71 (notional total $92.64 over 33 attempts; 25 resolved; resolved rate 59% to 87% at 95%)
Counted: failed attempts yes; retries none were run; cache share 94.0%; onboarding, planning, verification, review yes; test compute, grading, human time no
Price basis: list-price estimate of subscription calls, no invoice
Task set: SWE-bench Verified, 33 of 500 public issues; initial stratified draw of 25 (seed 20261004), then replacements and a compiled-instance campaign
Configuration: Claude Sonnet 5.5 through a subscription CLI; platform builds f0ac3a8a and 236c0d3f; run dates 2026-10-04 and 2026-10-05

The only cost per task that matters to you is the one on your own tasks. Take a small sample of your tickets. Run every attempt. Keep every receipt. Divide total spend by solved tasks, and show the interval. How to estimate your AI coding bill and the AI cost calculator show the steps.

A hypothesis to measure: lead and sidekick

Cognition describes Devin Fusion as two agents: a capable lead model and a cheaper sidekick. Each keeps its own cached context. The lead plans, settles ambiguity and reviews. The sidekick takes mechanical work. A classifier can switch the model when the agent compacts its context (June 29 post, read 2026-10-07).

That is a cost hypothesis, and it is worth measuring in any agent. Does handing mechanical work to a cheaper model lower the cost per resolved task? Does the resolved count stay the same? The vendor post shows five selected sample tasks (n = 5, vendor-reported). Reported cost fell on all five, by 25% to 62%. The score change ranged from -27 to +12 points. These are example ranges, not 95% intervals. In the one case where the score fell by 27 points, the post says the lead handed off work where judgment was the deliverable. Five chosen examples illustrate a design. They do not measure a rate.

Our data shows only price sensitivity, not a ceiling on cost or a test of the idea. All of our SWE-bench work ran on Sonnet 5.5. Repricing the same recorded tokens at Haiku 4.5 list prices gives $43.61 against $87.23 (a calculation). A cheaper model would use different tokens and resolve different tasks, so this bounds price sensitivity. It does not predict a result. We have not run a lead and sidekick split.

To test the idea, you need the same tasks, every attempt counted, n with an interval, and a review of the merged result.

Limits and disclosure

  • I build Agent. This post sits on a site that sells Agent. Treat that as a conflict of interest.
  • Home advantages. We chose the sample: 25 of the 500 Verified instances, stratified by public difficulty. Six compiled-extension instances could not import in the worker checkout, so a rule we declared replaced them. A second campaign ran the compiled ones after a sandbox fix. That is why the run has 33 attempts and two platform builds. We wrote the cost accounting. All three empty patches were platform holds before delivery, not wrong fixes, and they count as failures. Counting every failed attempt was our choice. It makes our number higher, not lower.
  • We did not run Devin. No Devin account or CLI was used, and nobody paid for a run. We read public pages on 2026-10-07. The public leaderboard data has 42 FrontierCode 1.1 entries and none for Devin Fusion. The public data cannot rebuild the 68.8 and $0.60 point.
  • Different benchmarks. SWE-bench Verified checks issue resolution with tests. FrontierCode grades against a maintainer-style rubric on private tasks. A score on one cannot become a rate on the other.
  • n = 33. This is a defect-finding run, not a ranking. Verified issues are public. We do not control contamination for any system, ours included.
  • No ranking. Nothing here says that one system is cheaper or better than another.

Sources read on 2026-10-07: Cognition's September 28 efficiency post, its FrontierCode method post, its FrontierCode 1.1 post, its Devin Fusion post, its public leaderboard data and the Devin billing documentation. Our numbers come from /benchmarks/swe-bench-verified and /benchmarks/cost-thought-experiments. Agent calculations use those dataset values and the linked raw rows. Calculations on vendor claims use the attributed figures and dates. The vendor source reads are recorded in saved notes; we did not run Devin.

See the receipt for your own task

Every Agent run records its calls, tokens and cost per stage. You see what each change cost and where the money went, failed attempts included. Try Agent and read the receipt for your first task.

The data behind this post

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.