{"i":18,"study":{"slug":"json-schema-vs-instructions","title":"Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI","seoTitle":"LLM structured output: JSON schema vs instructions","description":"96 calls. Strict passes, schema vs instructions: Haiku 18/24 vs 0/24, Sonnet 12/12 vs 12/12, GPT-6.1 Sol 12/12 vs 12/12. With intervals.","question":"For the same extraction tasks, does enforcing a JSON schema through the CLI change the strict pass rate, the format misses and the wrong values, compared with asking for JSON in the prompt?","answer":"96 counted calls on three JSON extraction prompts, each asked with instructions only and with the CLI’s JSON schema mode. Format misses (a right answer in the wrong format): 17 of 48 calls with instructions (95% interval 23% to 50%), 0 of 48 with a schema (95% interval 0% to 7%). Code fences: 24/48 instruction replies (95% interval 36% to 64%). Schema replies: 0/48 (95% interval 0% to 7%). The instructions said not to use a fence. Strict passes: Haiku: 0/24 (95% interval 0% to 14%; 17 format misses; 7 wrong-value replies) with instructions and 18/24 (95% interval 55% to 88%; 6 wrong-value replies) with a schema; Sonnet: 12/12 (95% interval 76% to 100%) with instructions and 12/12 (95% interval 76% to 100%) with a schema; GPT-6.1 Sol: 12/12 (95% interval 76% to 100%) with instructions and 12/12 (95% interval 76% to 100%) with a schema. Wrong values: 7 of 48 with instructions (95% interval 7% to 27%), 6 of 48 with a schema (95% interval 6% to 25%). These schemas constrain the reply’s shape. They do not verify totals or due dates. Where the 95% intervals do not overlap: Haiku, schema ahead (18/24 vs 0/24; paired McNemar p = 0.00000762939453125, calculation); every other pair overlaps, so the data does not rank it. Sonnet and GPT-6.1 Sol passed every call in both modes, a ceiling on this task set that cannot show a schema effect. Recommendation: test the CLI’s schema mode on your own tasks. On this task set, use it for Haiku, where strict passes were ahead (the 95% intervals do not overlap); for Sonnet and GPT-6.1 Sol the intervals overlap, so this sample does not establish a difference. Keep a validator for totals and dates; these schemas do not check them.","date":"2026-10-07","updated":"2026-10-07","tags":["structured-output","json-schema","format-miss","claude-code","codex-cli","claude-haiku","claude-sonnet","gpt-6-1-sol"],"caveats":["Small samples: Haiku 24 calls per mode, Sonnet 12 calls per mode and GPT-6.1 Sol 12 calls per mode. A 12/12 result has a 95% interval of 76% to 100%. This is not a minimum detectable difference.","The calls repeat only three fixed prompts. Wilson intervals describe call outcomes under a binomial assumption; they do not measure accuracy across unseen tasks. The paired p-values also assume independent pairs and do not remove this limit.","The prompts are hand-made. The order prompt reuses a case with known Haiku format misses from an earlier study; this is not a blind holdout.","Both routes used one shared Mac. The files do not establish host isolation from other work. Timings include CLI start-up and network time.","The current protocol file was born after the Codex batch. An earlier version may have existed, but these files cannot verify pre-call registration. Amendment 1 says 00:44 UTC and before counted calls; the Codex batch started at 00:43 UTC. Amendment 4 says 06:10 UTC and before the Claude probe; the probe receipt was saved at 06:04 UTC and counted Claude calls began at 06:04 UTC. These timing conflicts limit the protocol claims. The calls remain descriptive evidence.","The first Claude lane stopped after a 90-minute wait for another study to release the route. It made no inference call. The lane restarted later; no counted call was retried.","The comparison changes both the schema flag and the format sentence. It cannot isolate the effect of the flag alone.","These schemas constrain keys, types and currency labels. They do not verify totals or due dates. Both modes use the same task text. A change in wrong values is a measured result, not a design aim.","Sonnet and GPT-6.1 Sol passed every call in both modes. These three tasks are within reach of those models, so they cannot show a schema effect for them; harder extraction may differ.","The instruction wording is one sentence; another wording, or a schema with the format sentence kept, was not tested. The format-miss counts here are not comparable with the caching and consistency study, which used a different wording on one of these prompts.","In schema mode Claude Code returns the answer through a structured-output tool call. Its median was 2 model turns (range 2 to 2, n = 36). Instructions had a median of 1 (range 1 to 1, n = 36). Its time and token counts include that extra step.","The Codex CLI’s exec mode does not echo the model or the reasoning effort. The request named GPT-6.1 Sol at low effort, and the CLI’s own catalog lists that model with that effort, but a reroute would not have been visible. Codex timings include its start-up and its larger system prompt, so Codex rows compare route and model pairs, not models alone."],"sourceIds":["agent-structured-output"],"stats":{"$k":["id","label","value","unit","display","n","ci","note"],"$r":[["structured-output-calls","Counted calls in this study (every one counted)",96,"calls","96 (72 Claude Code, 24 Codex CLI)","\u0001","\u0001","\u0001"],["structured-output-format-miss-instructions","Format-miss rate with instructions only, all models",0.3542,"rate","35% (17/48)",48,[0.2343,0.4956],"A right answer in a code fence or prose. Every error counted as a call."],["structured-output-format-miss-schema","Format-miss rate with a JSON schema, all models",0,"rate","0% (0/48)",48,[0,0.0741],"A right answer in a code fence or prose. Every error counted as a call."],["structured-output-fenced-instructions","Replies in a code fence with instructions only, all models",0.5,"rate","50% (24/48)",48,[0.3639,0.6361],"The prompt said: no code fence. Completed replies only; a fence can hold a right or a wrong answer."],["structured-output-fenced-schema","Replies in a code fence with a JSON schema, all models",0,"rate","0% (0/48)",48,[0,0.0741],"Completed replies only. Claude Code uses its structured_output field. Codex CLI uses its final message."],["structured-output-wrong-values-instructions","Wrong-values rate with instructions only, all models",0.1458,"rate","15% (7/48)",48,[0.0725,0.2717],"A completed reply that is neither a strict pass nor a format miss. Denominator includes all attempts; errors are a separate outcome."],["structured-output-wrong-values-schema","Wrong-values rate with a JSON schema, all models",0.125,"rate","13% (6/48)",48,[0.0586,0.247],"A completed reply that is neither a strict pass nor a format miss. Denominator includes all attempts; errors are a separate outcome."],["structured-output-strict-instructions","Strict pass rate with instructions only, all models",0.5,"rate","50% (24/48)",48,[0.3639,0.6361],"Pooled over three models and three prompts; the models differ in n, so read the per-model chart."],["structured-output-strict-schema","Strict pass rate with a JSON schema, all models",0.875,"rate","88% (42/48)",48,[0.753,0.9414],"Pooled over three models and three prompts; the models differ in n, so read the per-model chart."]]},"charts":{"$k":["id","title","subtitle","kind","unit","polarity","yLabel","series","note","whisker","sourceIds"],"$r":[["structured-output-pass-rate","Does a JSON schema raise the pass rate? Instructions vs schema mode","Three extraction prompts pooled; whiskers are 95% Wilson intervals","dot-range","rate","higher","Passed",[{"name":"Strict pass: the whole reply is the right JSON","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0,0,0.138,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0.75,0.551,0.88,24],["Claude Sonnet 5.5 (instructions) · Claude Code",1,0.7575,1,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",1,0.7575,1,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",1,0.7575,1,12]]}},{"name":"Right answer in any format (strict pass or format miss)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0.7083,0.5083,0.8509,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0.75,0.551,0.88,24],["Claude Sonnet 5.5 (instructions) · Claude Code",1,0.7575,1,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",1,0.7575,1,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",1,0.7575,1,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",1,0.7575,1,12]]}}],"Whiskers are 95% Wilson intervals (a calculation) over 24 calls and 12 calls per configuration; every error counts as a fail. Strict: the whole reply parses as JSON and matches the expected answer exactly. A format miss is a right answer inside a code fence or prose, so it is never a strict pass.","ci95",["agent-structured-output"]],["structured-output-outcomes","What each call produced: strict pass, format miss, wrong values or error","Counts of calls per configuration; the three prompts pooled","stacked-bar","count","\u0001","Calls",{"$k":["name","points"],"$r":[["Strict pass",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",18,24],["Claude Sonnet 5.5 (instructions) · Claude Code",12,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",12,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",12,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",12,12]]}],["Format miss",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",17,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0,24],["Claude Sonnet 5.5 (instructions) · Claude Code",0,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",0,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",0,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",0,12]]}],["Wrong values",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",7,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",6,24],["Claude Sonnet 5.5 (instructions) · Claude Code",0,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",0,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",0,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",0,12]]}],["Error",{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",0,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",0,24],["Claude Sonnet 5.5 (instructions) · Claude Code",0,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",0,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",0,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",0,12]]}]]},"Counts of calls, not rates; the pass-rate chart carries the same results with 95% intervals. Format miss: the right answer inside a code fence or prose. Wrong values: any other completed reply, with a wrong value, key or type (a reply that sits in a code fence and also has a wrong value is counted here). Error: the call did not complete.","\u0001",["agent-structured-output"]],["structured-output-time","Time per call, instructions vs schema mode","Median; whiskers = fastest and slowest completed call","dot-range","seconds","\u0001","Seconds",[{"name":"Median time per call (the three prompts pooled)","points":{"$k":["label","value","lo","hi","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",9.52,5.67,17,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",8.46,5.9,12.23,24],["Claude Sonnet 5.5 (instructions) · Claude Code",3.52,2.67,4.12,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",4.2,2.95,6.14,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",6.21,4.2,12.27,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",5.96,4.62,20.97,12]]}}],"Whiskers are a range (fastest and slowest call), not a confidence interval. Wall time from process start to exit, so it includes CLI start-up; Codex CLI timings include its larger system prompt. The three prompts differ in length, which widens every range.","minmax",["agent-structured-output"]],["structured-output-tokens","Output and reasoning tokens per call, instructions vs schema mode","Median per call; ranges and sample sizes are in the note","grouped-bar","tokens","\u0001","Tokens",[{"name":"Median output tokens per call","points":{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",1128,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",1036,24],["Claude Sonnet 5.5 (instructions) · Claude Code",368,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",424,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",117,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",123,12]]}},{"name":"Median reasoning tokens per call (thinking)","points":{"$k":["label","value","n"],"$r":[["Claude Haiku 4.5 (instructions) · Claude Code",934,24],["Claude Haiku 4.5 (JSON schema) · Claude Code",727,24],["Claude Sonnet 5.5 (instructions) · Claude Code",182,12],["Claude Sonnet 5.5 (JSON schema) · Claude Code",109,12],["GPT-6.1 Sol (low, instructions) · Codex CLI",21,12],["GPT-6.1 Sol (low, JSON schema) · Codex CLI",23,12]]}}],"Output tokens include reasoning tokens. Claude schema-mode output includes the CLI’s structured-output tool call. Input counts include CLI context and are not compared. Ranges (not intervals): Claude Haiku 4.5 (instructions) · Claude Code: output 698 to 2059 (n = 24); reasoning 559 to 1830 (n = 24); Claude Haiku 4.5 (JSON schema) · Claude Code: output 716 to 1462 (n = 24); reasoning 522 to 1187 (n = 24); Claude Sonnet 5.5 (instructions) · Claude Code: output 154 to 456 (n = 12); reasoning 54 to 278 (n = 12); Claude Sonnet 5.5 (JSON schema) · Claude Code: output 273 to 541 (n = 12); reasoning 0 to 260 (n = 12); GPT-6.1 Sol (low, instructions) · Codex CLI: output 69 to 259 (n = 12); reasoning 0 to 60 (n = 12); GPT-6.1 Sol (low, JSON schema) · Codex CLI: output 98 to 181 (n = 12); reasoning 0 to 49 (n = 12).","\u0001",["agent-structured-output"]]]},"related":["caching-consistency","hard-model-head-to-head"]}}