{"i":18,"slug":"json-schema-vs-instructions","chart":{"id":"structured-output-pass-rate","title":"Does a JSON schema raise the pass rate? Instructions vs schema mode","subtitle":"Three extraction prompts pooled; whiskers are 95% Wilson intervals","kind":"dot-range","unit":"rate","polarity":"higher","yLabel":"Passed","series":[{"name":"Strict pass: the whole reply is the right JSON","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0,0,0.138,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0.75,0.551,0.88,24],["Claude Sonnet 5.5 (instructions) · Claude Code",1,0.7575,1,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",1,0.7575,1,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",1,0.7575,1,12]]}},{"name":"Right answer in any format (strict pass or format miss)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0.7083,0.5083,0.8509,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0.75,0.551,0.88,24],["Claude Sonnet 5.5 (instructions) · Claude Code",1,0.7575,1,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",1,0.7575,1,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",1,0.7575,1,12]]}}],"note":"Whiskers are 95% Wilson intervals (a calculation) over 24 calls and 12 calls per configuration; every error counts as a fail. Strict: the whole reply parses as JSON and matches the expected answer exactly. A format miss is a right answer inside a code fence or prose, so it is never a strict pass.","whisker":"ci95","sourceIds":["agent-structured-output"]}}