{"i":25,"slug":"harder-tasks-head-to-head","chart":{"id":"harder-h2h-pass-rate","title":"Pass rate on 4 harder tasks","subtitle":"Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Passed","whisker":"ci95","series":[{"name":"Strict pass","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.4167,0.1933,0.6805,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}},{"name":"Lenient (format misses counted)","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0.6875,0.444,0.8584,16],["Claude Opus 5.5 · Claude Code",0.5,0.2538,0.7462,12],["Claude Sonnet 5.5 · Claude Code",0.375,0.1848,0.6136,16],["Claude Haiku 4.5 · Claude Code",0,0,0.2425,12]]}}],"note":"Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row.\n\nCounted calls are new calls.","sourceIds":["agent-harder-tasks"]}}