{"i":4,"slug":"hard-model-head-to-head","chart":{"id":"hard-h2h-outcomes","title":"What happened on every call","subtitle":"Counts per configuration: strict passes, format misses and wrong answers","kind":"stacked-bar","unit":"count","yLabel":"Calls","series":{"$k":["name","points"],"$r":[["Strict pass",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",24,24],["Claude Opus 5.5 · Claude Code",24,24],["Claude Opus 5.5 (high) · Claude Code",24,24],["GPT-6.1 Sol (medium) · Codex CLI",16,16],["Claude Fable 5.1 · Claude Code",24,24],["GPT-6.1 Sol (high) · Codex CLI",16,16],["Claude Haiku 4.5 · Claude Code",11,24]]}],["Format miss (correct answer, wrong format)",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",0,24],["Claude Opus 5.5 · Claude Code",0,24],["Claude Opus 5.5 (high) · Claude Code",0,24],["GPT-6.1 Sol (medium) · Codex CLI",0,16],["Claude Fable 5.1 · Claude Code",0,24],["GPT-6.1 Sol (high) · Codex CLI",0,16],["Claude Haiku 4.5 · Claude Code",5,24]]}],["Wrong answer",{"$k":["label","value","n"],"$r":[["Claude Sonnet 5.5 · Claude Code",0,24],["Claude Opus 5.5 · Claude Code",0,24],["Claude Opus 5.5 (high) · Claude Code",0,24],["GPT-6.1 Sol (medium) · Codex CLI",0,16],["Claude Fable 5.1 · Claude Code",0,24],["GPT-6.1 Sol (high) · Codex CLI",0,16],["Claude Haiku 4.5 · Claude Code",8,24]]}]]},"note":"A format miss is a reply that the strict validator rejected (for example, wrapped in a code fence or in prose although the prompt said not to) whose extracted answer passes the same validator. It is not a pass.","sourceIds":["agent-provider-h2h-hard"]}}