{"i":1,"slug":"swe-bench-opus-vs-sonnet","chart":{"id":"swebench-opus-sonnet-minutes","title":"Worker time per attempt","subtitle":"Median minutes; whiskers = fastest and slowest of 3 attempts (not an interval)","kind":"dot-range","unit":"minutes","whisker":"minmax","yLabel":"Minutes","viz":"LatencyLanes","series":[{"name":"Worker minutes per attempt","points":[{"label":"Claude Opus 5.5 (Agent, new build)","value":20.26,"lo":9.76,"hi":25.29,"n":3},{"label":"Claude Sonnet 5.5 (Agent, older builds)","value":9.37,"lo":4.74,"hi":15,"n":3}]}],"note":"Worker minutes from the run reports. Opus attempts ran one at a time on the newer build; the Sonnet attempts ran in earlier campaigns. 3 attempts per arm is too few to call a difference; a range is not a confidence interval.","sourceIds":["agent-swebench-opus","agent-swebench-c1","agent-swebench-c2"]}}