{"i":25,"slug":"harder-tasks-head-to-head","chart":{"id":"harder-h2h-tool-attempts","title":"Calls that tried a tool although tools were off","subtitle":"Share of calls whose reply held tool-call markup or whose tool call the CLI could not parse","kind":"dot-range","unit":"rate","yLabel":"Calls with a tool attempt","polarity":"none","whisker":"ci95","series":[{"name":"Tool attempt","points":{"$k":["label","value","lo","hi","n"],"$r":[["GPT-6.1 Sol (medium) · Codex CLI",0,0,0.1936,16],["Claude Opus 5.5 · Claude Code",0.4167,0.1933,0.6805,12],["Claude Sonnet 5.5 · Claude Code",0.3125,0.1416,0.556,16],["Claude Haiku 4.5 · Claude Code",0.0833,0.0149,0.3539,12]]}}],"note":"Every prompt asked for the answer only. Both CLIs disabled tools. The Codex runner adds a developer instruction not to call tools. The Claude Code runner adds no such instruction. The validator decides the score. Tool-call markup is a separate flag; a flagged call can still pass if its whole reply meets the validator.\n\nIt is a behaviour, not a quality score: more is not better. No call with a tool attempt passed. Whiskers are 95% Wilson intervals. The flag comes from the reply text, which is not published.","sourceIds":["agent-harder-tasks"]}}