{"i":136,"comparison":{"slug":"gpt-5-2-vs-claude-opus-4-5","a":"gpt-5-2","b":"claude-opus-4-5","title":"GPT 5.2 vs Claude Opus 4.5","seoTitle":"GPT 5.2 vs Claude Opus 4.5: measured benchmarks","description":"GPT 5.2 vs Claude Opus 4.5: 3 measured metrics from 2 studies, with sample sizes, intervals and every failure counted.","verdict":"GPT 5.2 and Claude Opus 4.5 share 3 measured metrics from 2 studies. No row separates them: every interval or run range overlaps, too few runs were recorded, no interval was recorded, or more is not better for that metric. The rows are 1 tie and 2 unclear; each row says why.","rows":{"$k":["metric","aValue","bValue","unit","aDisplay","bDisplay","winner","basis","studySlug","n","chartId","aN","bN","aContext","bContext","rangeKind","spanKind","aRange","bRange"],"$r":[["Resolved rate on the same 33 SWE-bench Verified instances",0.8485,0.7273,"rate","85% (28/33)","73% (24/33)","tie","The 95% intervals overlap (GPT 5.2 69% to 93%; Claude Opus 4.5 56% to 85%), so this sample cannot separate them.","swe-bench-verified",33,"swebench-same-instance-leaderboard",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","ci95","ci95",[0.6908,0.9335],[0.5578,0.8493]],["Model calls per instance",35.6,35.9,"calls","35.6","35.9","unclear","More or fewer calls is not better or worse by itself; this row describes behaviour, not a winner.","swe-bench-verified",33,"swebench-model-calls",33,33,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"],["Recorded cost per resolved instance: Agent vs the public panel",0.628,1.184,"usd","$0.63","$1.18","unclear","No interval or range was recorded for either side, so the gap ($0.63 vs $1.18) is not tested against run-to-run variation.","cost-thought-experiments","\u0001","cost-per-resolved-agent-vs-panel",28,24,"effort high · public mini-SWE-agent v2 run, same instances","effort high · public mini-SWE-agent v2 run, same instances","\u0001","\u0001","\u0001","\u0001"]]}}}