{"i":13,"slug":"cost-thought-experiments","chart":{"id":"cost-per-resolved-agent-vs-panel","title":"Recorded cost per resolved instance: Agent vs the public panel","subtitle":"Same 33 SWE-bench Verified instances; all attempts in the numerator","kind":"bar","unit":"usd","yLabel":"USD per resolved instance","series":[{"name":"Cost per resolved instance","points":{"$k":["label","value","n","highlight"],"$r":[["Agent (notional)",3.706,25,true],["Claude 4.5 Opus (high)",1.184,24,"\u0001"],["Claude 4.5 Sonnet (high)",0.913,25,"\u0001"],["Claude 4.6 Opus",0.875,23,"\u0001"],["GLM 5 (high)",0.667,26,"\u0001"],["DeepSeek V3.2 (high)",0.637,24,"\u0001"],["GPT 5.2 (high)",0.628,28,"\u0001"],["Claude 4.5 Haiku (high)",0.479,25,"\u0001"],["Gemini 3 Flash (high)",0.436,27,"\u0001"],["Kimi K2.5 (high)",0.256,23,"\u0001"],["MiniMax M2.5 (high)",0.107,23,"\u0001"],["GPT 5 mini",0.08,21,"\u0001"]]}}],"note":"Recorded figures, not repricing. Panel costs are published API costs for a bash-only agent. Agent's figure is a list-price estimate of subscription calls and includes onboarding, planning, verification and review.","sourceIds":["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard"]}}