{"i":1,"study":{"slug":"swe-bench-opus-vs-sonnet","title":"Opus 5.5 vs Sonnet 5.5 as the Agent brain on SWE-bench Verified (interim)","seoTitle":"Opus 5.5 vs Sonnet 5.5 on SWE-bench Verified (interim)","description":"Interim: 3 of 8 paired SWE-bench Verified instances graded. Opus 5.5 resolved 2, Sonnet 5.5 1 (McNemar p = 1.0). Cost 2.6×, a list-price calculation.","question":"Does Agent resolve more SWE-bench Verified instances with Claude Opus 5.5 as its brain than with Claude Sonnet 5.5, and at what cost and time?","answer":"Interim, not a final result: 3 of 8 declared pairs are graded. With Claude Opus 5.5 as its brain, Agent resolved 2 of 3 (95% Wilson 21–94%); with Claude Sonnet 5.5 it resolved 1 of 3 (6–79%) on the same instances. The one discordant instance (django__django-10554) went to Opus, and the exact McNemar p is 1.0, so these pairs show no difference. On the same instances Opus cost 2.6× as much, a list-price calculation ($22.76 vs $8.64), and used 1.9× the worker minutes. The arms ran on different platform builds, and the other 5 pairs remain not started in this dataset; the recorded usage-reset date was 2026-10-09.","date":"2026-10-06","updated":"2026-10-06","tags":["swe-bench","claude-opus","claude-sonnet","agent-harness","interim"],"caveats":["Interim: 3 of 8 declared pairs. n = 3 supports no ranking: the 95% Wilson intervals overlap almost completely.","Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.","2 of the Sonnet misses in the declared set (django__django-14631, sympy__sympy-18189) were empty patches from platform holds, not wrong fixes. They are not graded for Opus yet.","The instances were chosen to hold Sonnet misses and resolves in equal numbers, so neither rate estimates SWE-bench Verified as a whole.","Costs are list-price calculations on subscription runs, not invoices. Grading ran under amd64 emulation; contamination is not controlled."],"sourceIds":["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2","calc-repricing","price-anthropic"],"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["swebench-opus-resolved","Claude Opus 5.5 in Agent: resolved, interim (3 of 8 pairs graded)",0.6667,"rate","67% (2/3)",3,[0.2077,0.9385],"\u0001"],["swebench-sonnet-resolved","Claude Sonnet 5.5 in Agent: resolved on the same 3 instances",0.3333,"rate","33% (1/3)",3,[0.0615,0.7923],"\u0001"],["swebench-opus-sonnet-mcnemar","Exact McNemar p on the graded pairs",1,"score","p = 1.0",3,"\u0001","1 pair where only Opus resolved, 0 pairs where only Sonnet did. The test reads only these; p = 1.0 is no evidence of a difference."],["swebench-opus-sonnet-graded","Declared pairs graded so far",3,"count","3 of 8",8,"\u0001","5 not started: the usage gate stopped the campaign before django__django-11885. The recorded usage-reset date was 2026-10-09; no resumed attempts are included here. Every started attempt counts."],["swebench-opus-sonnet-cost-ratio","Opus vs Sonnet list-price cost on the same instances (calculation)",2.63,"ratio","2.6×",3,"\u0001","$22.76 vs $8.64 over the 3 graded instances: notional cost from the platform price table, not an invoice."],["swebench-opus-sonnet-minutes-ratio","Opus vs Sonnet worker minutes on the same instances (ratio of totals)",1.9,"ratio","1.9×",3,"\u0001","55.3 vs 29.1 worker minutes."]]},"charts":{"$k":["id","title","subtitle","kind","unit","whisker","yLabel","viz","series","note","sourceIds","xLabel"],"$r":[["swebench-opus-sonnet-resolved","Resolved on the same 3 SWE-bench Verified instances (interim)","One attempt per arm per instance, official grader · 95% Wilson intervals","dot-range","rate","ci95","Resolved","IntervalDotPlot",[{"name":"Resolved","points":[{"label":"Claude Opus 5.5 (Agent, new build)","value":0.6667,"lo":0.2077,"hi":0.9385,"n":3},{"label":"Claude Sonnet 5.5 (Agent, older builds)","value":0.3333,"lo":0.0615,"hi":0.7923,"n":3}]}],"Interim: 3 of 8 declared pairs are graded; 5 were never started; no resumed attempts are included here. With n = 3 the intervals span most of the axis, so this chart supports no ranking. The instances were chosen to hold Sonnet misses and resolves in equal numbers, so neither rate estimates SWE-bench Verified as a whole.",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2"],"\u0001"],["swebench-opus-sonnet-cost-per-attempt","List-price cost per attempt (calculation)","Mean over the same 3 instances; notional cost from the platform price table","bar","usd","\u0001","USD per attempt","CostBars",[{"name":"List-price cost per attempt","points":[{"label":"Claude Opus 5.5 (Agent, new build)","value":7.59,"n":3},{"label":"Claude Sonnet 5.5 (Agent, older builds)","value":2.88,"n":3}]}],"Calculation, not an invoice: the platform price table at each run commit applied to the recorded tokens of subscription runs (Opus 5.5: $4 input, $20 output, $0.20 cache read per million tokens). Totals $22.76 vs $8.64: 2.6×, a ratio of two calculations.",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2","calc-repricing","price-anthropic"],"\u0001"],["swebench-opus-sonnet-cost-by-instance","List-price cost per instance (calculation)","One attempt per arm; notional cost from the platform price table","grouped-bar","usd","\u0001","USD","DumbbellPairs",[{"name":"Claude Sonnet 5.5 (Agent, older builds)","points":{"$k":["label","value","n"],"$r":[["django-10554",1.52,1],["matplotlib-21568",3.02,1],["django-15022",4.1,1]]}},{"name":"Claude Opus 5.5 (Agent, new build)","points":{"$k":["label","value","n"],"$r":[["django-10554",9.99,1],["matplotlib-21568",4.9,1],["django-15022",7.87,1]]}}],"Calculation, not an invoice. Outcomes: django-10554 Opus resolved, Sonnet unresolved; matplotlib-21568 Opus resolved, Sonnet resolved; django-15022 Opus unresolved, Sonnet unresolved.",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2","calc-repricing","price-anthropic"],"Instance"],["swebench-opus-sonnet-minutes","Worker time per attempt","Median minutes; whiskers = fastest and slowest of 3 attempts (not an interval)","dot-range","minutes","minmax","Minutes","LatencyLanes",[{"name":"Worker minutes per attempt","points":[{"label":"Claude Opus 5.5 (Agent, new build)","value":20.26,"lo":9.76,"hi":25.29,"n":3},{"label":"Claude Sonnet 5.5 (Agent, older builds)","value":9.37,"lo":4.74,"hi":15,"n":3}]}],"Worker minutes from the run reports. Opus attempts ran one at a time on the newer build; the Sonnet attempts ran in earlier campaigns. 3 attempts per arm is too few to call a difference; a range is not a confidence interval.",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2"],"\u0001"],["swebench-opus-sonnet-stage-cost","Where the cost goes: pipeline stages (calculation)","List-price cost per stage, summed over the same 3 instances","grouped-bar","usd","\u0001","USD","DumbbellPairs",[{"name":"Claude Sonnet 5.5 (Agent, older builds)","points":{"$k":["label","value","n"],"$r":[["Research",2.34,3],["Plan",1.16,3],["Implement",1.32,3],["Validate",1.42,3],["Review",0.51,3],["Deliver",1.56,3]]}},{"name":"Claude Opus 5.5 (Agent, new build)","points":{"$k":["label","value","n"],"$r":[["Research",6.28,3],["Plan",1.11,3],["Implement",4.94,3],["Validate",2.97,3],["Review",2.01,3],["Deliver",4.34,3]]}}],"Calculation, not an invoice: stage costs from the run reports (run-mode stages). Onboarding and calls outside a stage are not in these bars, so the stages sum to less than the attempt totals.",["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2","calc-repricing","price-anthropic"],"Pipeline stage"]]},"related":["swe-bench-verified","coding-agents-head-to-head","effort-ladder"],"hero":{"statIds":["swebench-opus-resolved","swebench-sonnet-resolved"],"testStatId":"swebench-opus-sonnet-mcnemar"}}}