{"i":0,"study":{"slug":"swe-bench-verified","title":"Agent on SWE-bench Verified vs 11 public models","seoTitle":"SWE-bench Verified: Agent vs GPT, Claude and Gemini","description":"Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.","question":"How does Agent, a full worker pipeline on one model, do on SWE-bench Verified next to public single-model runs on the very same instances?","answer":"Agent resolved 25 of 33 attempted instances (75.8%, 95% interval 59% to 87%). On the same instances the 11 public mini-SWE-agent v2 runs resolved between 21 and 28 (panel mean 74.1%). Every interval overlaps, so this sample cannot rank Agent above or below any panel model. Agent spent a notional $2.81 and 49 model calls per attempt, with a median of 9.6 minutes; it is slower and more expensive per instance than a bare bash agent because it onboards, plans, verifies and reviews. It resolved 1 of 4 instances that no panel model solved.","date":"2026-10-05","updated":"2026-10-05","tags":["swe-bench","coding-agents","leaderboard","cost","claude-sonnet"],"caveats":["n = 33: intervals are wide. This is a defect-finding run, not a ranking.","Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.","Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.","Campaign 1 ended up 19 django, 3 sympy, 2 sphinx and 1 xarray after replacements. Campaign 2 covers the compiled repositories.","Verified issues are public (2015 to 2023) and likely in every model’s training data; contamination is uncontrolled for all systems.","Difficulty bands come from the panel’s own results, so a system outside the panel tends to look better than the panel on hard bands and worse on easy ones (regression to the mean). Read the band chart with that selection effect in mind.","Three campaign-1 empty patches were platform holds before delivery (missing lint tools, a too-literal plan gate, an unanswered question), not wrong fixes. They count as failures here."],"sourceIds":["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","swebench-protocol","price-anthropic"],"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["agent-rate-33","Agent resolved, all 33 attempted instances",0.7576,"rate","76% (25/33)",33,[0.5898,0.8717],"Both campaigns, one attempt each, failures and empty patches included."],["agent-rate-c1","Agent resolved, campaign 1 sample of 25",0.72,"rate","72% (18/25)",25,[0.5242,0.8572],"\u0001"],["agent-rate-original25","Agent resolved, original seed draw of 25 (no replacements)",0.76,"rate","76% (19/25)",25,[0.5657,0.885],"Mixes two platform builds."],["panel-mean-33","Public panel mean on the same 33 instances",0.741,"rate","74.1%",33,"\u0001","Mean of 11 public mini-SWE-agent v2 runs."],["cost-per-attempt","Agent model cost per attempt (notional)",2.81,"usd","$2.81",33,"\u0001","Subscription calls priced at list price; onboarding and failed attempts included."],["cost-per-resolved","Agent model cost per resolved instance (notional)",3.71,"usd","$3.71",25,"\u0001","\u0001"],["median-minutes","Median worker time per attempt",9.6,"minutes","9.6 min",33,"\u0001","Range 1.6 to 54.1 min."],["calls-per-attempt","Model calls per attempt",49.5,"calls","49",33,"\u0001","\u0001"]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds","xLabel"],"$r":[["swebench-same-instance-leaderboard","Resolved rate on the same 33 SWE-bench Verified instances","Agent vs 11 public mini-SWE-agent v2 runs, one attempt each","dot-range","rate","Resolved",[{"name":"Resolved rate","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["GPT 5.2 (high)",0.8485,0.6908,0.9335,33,"\u0001"],["Gemini 3 Flash (high)",0.8182,0.6561,0.9139,33,"\u0001"],["GLM 5 (high)",0.7879,0.6225,0.8932,33,"\u0001"],["Agent (Sonnet 5.5, full pipeline)",0.7576,0.5898,0.8717,33,true],["Claude 4.5 Sonnet (high)",0.7576,0.5898,0.8717,33,"\u0001"],["Claude 4.5 Haiku (high)",0.7576,0.5898,0.8717,33,"\u0001"],["Claude 4.5 Opus (high)",0.7273,0.5578,0.8493,33,"\u0001"],["DeepSeek V3.2 (high)",0.7273,0.5578,0.8493,33,"\u0001"],["MiniMax M2.5 (high)",0.697,0.5266,0.8262,33,"\u0001"],["Claude 4.6 Opus",0.697,0.5266,0.8262,33,"\u0001"],["Kimi K2.5 (high)",0.697,0.5266,0.8262,33,"\u0001"],["GPT 5 mini",0.6364,0.4662,0.7781,33,"\u0001"]]}}],"Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard"],"\u0001"],["swebench-by-difficulty-band","Resolved rate by difficulty band","Band = how many of the 11 public panel models solved the instance","grouped-bar","rate","Resolved",[{"name":"Agent","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["No panel model solved it",0.25,0.0456,0.6994,4,true],["Under half solved it",0.75,0.3006,0.9544,4,true],["Half or more solved it",0.8182,0.523,0.9486,11,true],["Every panel model solved it",0.8571,0.6006,0.9599,14,true]]}},{"name":"Public panel mean","points":{"$k":["label","value","n"],"$r":[["No panel model solved it",0,4],["Under half solved it",0.2727,4],["Half or more solved it",0.8512,11],["Every panel model solved it",1,14]]}}],"Agent whiskers are 95% Wilson intervals. The \"no panel model solved it\" band has 4 instances; Agent resolved 1 (matplotlib__matplotlib-21568).",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","swebench-protocol"],"Difficulty band"],["swebench-cost-vs-resolved","Cost per instance vs resolved rate","Same 33 instances. Panel = API list price; Agent = notional subscription estimate","scatter","rate","Resolved rate",[{"name":"Public panel (mini-SWE-agent v2)","points":{"$k":["label","x","value","n"],"$r":[["Claude 4.5 Opus (high)",0.861,0.7273,33],["Gemini 3 Flash (high)",0.357,0.8182,33],["MiniMax M2.5 (high)",0.075,0.697,33],["Claude 4.6 Opus",0.61,0.697,33],["GLM 5 (high)",0.525,0.7879,33],["GPT 5.2 (high)",0.533,0.8485,33],["Claude 4.5 Sonnet (high)",0.692,0.7576,33],["Kimi K2.5 (high)",0.179,0.697,33],["DeepSeek V3.2 (high)",0.463,0.7273,33],["Claude 4.5 Haiku (high)",0.363,0.7576,33],["GPT 5 mini",0.051,0.6364,33]]}},{"name":"Agent","points":[{"label":"Agent","x":2.807,"value":0.7576,"n":33,"highlight":true}]}],"Agent's cost includes repository onboarding, planning, verification and review; it is a list-price estimate for subscription calls, not an invoice. Panel costs are published API costs.",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","price-anthropic"],"Mean model cost per instance (USD)"],["swebench-model-calls","Model calls per instance","Mean over the same 33 instances","bar","calls","Calls per instance",[{"name":"Mean calls","points":{"$k":["label","value","n","highlight"],"$r":[["DeepSeek V3.2 (high)",88.2,33,"\u0001"],["GLM 5 (high)",77.5,33,"\u0001"],["Claude 4.5 Haiku (high)",68.5,33,"\u0001"],["MiniMax M2.5 (high)",58.4,33,"\u0001"],["Kimi K2.5 (high)",56.7,33,"\u0001"],["Gemini 3 Flash (high)",54.2,33,"\u0001"],["Claude 4.5 Sonnet (high)",51,33,"\u0001"],["Agent",49.5,33,true],["Claude 4.5 Opus (high)",35.9,33,"\u0001"],["GPT 5.2 (high)",35.6,33,"\u0001"],["Claude 4.6 Opus",28.9,33,"\u0001"],["GPT 5 mini",20.8,33,"\u0001"]]}}],"A panel call is one bash-agent step. An Agent call is one model request of any stage (research, plan, act, verify, review).",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard"],"\u0001"],["swebench-cost-by-stage","Where Agent's model spend goes","Share of notional model cost by stage, all 33 attempts","bar","usd","USD (notional)",[{"name":"Cost","points":{"$k":["label","value"],"$r":[["Act (edit and run)",44.05],["Research",20.56],["Verify",8.72],["Other",6.24],["Review",5.58],["Context compaction",5.41],["Onboarding notes",2.09]]}}],"Total $92.64 over 33 attempts.",["agent-swebench-c1","agent-swebench-c2"],"\u0001"],["swebench-views","Every way to slice the run, with intervals","Agent resolved rate and 95% Wilson interval per declared view","dot-range","rate","Resolved",[{"name":"Agent","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Campaign 1: 25-instance sample",0.72,0.5242,0.8572,25,"\u0001"],["Campaign 2: 8 compiled-extension instances",0.875,0.5291,0.9776,8,"\u0001"],["Original seed draw of 25",0.76,0.5657,0.885,25,"\u0001"],["All 33 attempted",0.7576,0.5898,0.8717,33,true]]}},{"name":"Public panel mean, same instances","points":{"$k":["label","value","n"],"$r":[["Campaign 1: 25-instance sample",0.7091,25],["Campaign 2: 8 compiled-extension instances",0.8409,8],["Original seed draw of 25",0.7091,25],["All 33 attempted",0.741,33]]}}],"The original draw and \"all 33\" mix two platform builds.",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","swebench-protocol"],"\u0001"]]}}}