{
  "schema": "agent-public-bench@1",
  "generatedAt": "2026-10-07T00:00:00.000Z",
  "url": "https://agent.sasid.ai/benchmarks/json-schema-vs-instructions",
  "study": {
    "slug": "json-schema-vs-instructions",
    "title": "Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI",
    "seoTitle": "LLM structured output: JSON schema vs instructions",
    "description": "96 calls. Strict passes, schema vs instructions: Haiku 18/24 vs 0/24, Sonnet 12/12 vs 12/12, GPT-6.1 Sol 12/12 vs 12/12. With intervals.",
    "question": "For the same extraction tasks, does enforcing a JSON schema through the CLI change the strict pass rate, the format misses and the wrong values, compared with asking for JSON in the prompt?",
    "answer": "96 counted calls on three JSON extraction prompts, each asked with instructions only and with the CLI’s JSON schema mode. Format misses (a right answer in the wrong format): 17 of 48 calls with instructions (95% interval 23% to 50%), 0 of 48 with a schema (95% interval 0% to 7%). Code fences: 24/48 instruction replies (95% interval 36% to 64%). Schema replies: 0/48 (95% interval 0% to 7%). The instructions said not to use a fence. Strict passes: Haiku: 0/24 (95% interval 0% to 14%; 17 format misses; 7 wrong-value replies) with instructions and 18/24 (95% interval 55% to 88%; 6 wrong-value replies) with a schema; Sonnet: 12/12 (95% interval 76% to 100%) with instructions and 12/12 (95% interval 76% to 100%) with a schema; GPT-6.1 Sol: 12/12 (95% interval 76% to 100%) with instructions and 12/12 (95% interval 76% to 100%) with a schema. Wrong values: 7 of 48 with instructions (95% interval 7% to 27%), 6 of 48 with a schema (95% interval 6% to 25%). These schemas constrain the reply’s shape. They do not verify totals or due dates. Where the 95% intervals do not overlap: Haiku, schema ahead (18/24 vs 0/24; paired McNemar p = 0.00000762939453125, calculation); every other pair overlaps, so the data does not rank it. Sonnet and GPT-6.1 Sol passed every call in both modes, a ceiling on this task set that cannot show a schema effect. Recommendation: test the CLI’s schema mode on your own tasks. On this task set, use it for Haiku, where strict passes were ahead (the 95% intervals do not overlap); for Sonnet and GPT-6.1 Sol the intervals overlap, so this sample does not establish a difference. Keep a validator for totals and dates; these schemas do not check them.",
    "date": "2026-10-07",
    "updated": "2026-10-07",
    "tags": [
      "structured-output",
      "json-schema",
      "format-miss",
      "claude-code",
      "codex-cli",
      "claude-haiku",
      "claude-sonnet",
      "gpt-6-1-sol"
    ],
    "method": [
      "The protocol states a pre-call declaration. Its current file birth time is 2026-10-07T01:47:06.955Z. The first counted call began at 2026-10-07T00:43:42.104Z. These files cannot verify a protocol declaration before the Codex calls. The extract keeps every counted call.",
      "Three prompts each have one exact expected JSON answer and a sandboxed validator. The order task merges lines, drops a cancelled line and applies a discount. The invoice task applies a line discount and then tax, rounded half up. It also includes an untaxed fee and shipping. The meeting task extracts tickets, owners and due dates, resolves relative dates, and skips a decision and a note. We do not publish prompts or replies.",
      "Two modes use the same task text. Instructions mode adds “Reply with the JSON object only, no code fence, no other text.” Schema mode removes that sentence and uses the CLI’s schema flag. Claude Code uses --json-schema; the harness reads its structured_output field. Codex exec uses --output-schema; the harness reads its final message.",
      "Repetitions per prompt, in each mode: Haiku: 8; Sonnet: 4; GPT-6.1 Sol: 4. The order alternates by repetition, so time drift falls on both modes. Claude rows use the CLI’s default effort. GPT-6.1 Sol rows use low effort.",
      "Scoring: a strict pass needs the whole reply to be the right JSON. A format miss is a right answer inside a code fence or prose. Wrong values means any other completed reply, including a fenced reply with a wrong value. An error means the call did not complete; it counts as a fail.",
      "Rates carry Wilson 95% intervals (a calculation). A side is ahead only when the intervals do not overlap and the paired exact McNemar p is below 0.05. The paired table gives this p-value calculation. Time is a median with the fastest and slowest call: a range, not an interval.",
      "Controls ran before inference. Each reference answer passes, including through the schema path. A reference inside a code fence counts as a format miss. Four planted wrong answers per prompt fail. One uncounted probe per route (2 here) checked how the harness reads the CLI output. Probes stay outside every cell.",
      "Each call uses a fresh empty working folder, with tools off, no MCP servers and no session persistence. The harness reads the answer from the CLI output. Claude Code uses its JSON output format in both modes; earlier studies used its stream format. Schema mode adds the schema flag and removes the format sentence.",
      "Every counted batch completed all planned cells."
    ],
    "caveats": [
      "Small samples: Haiku 24 calls per mode, Sonnet 12 calls per mode and GPT-6.1 Sol 12 calls per mode. A 12/12 result has a 95% interval of 76% to 100%. This is not a minimum detectable difference.",
      "The calls repeat only three fixed prompts. Wilson intervals describe call outcomes under a binomial assumption; they do not measure accuracy across unseen tasks. The paired p-values also assume independent pairs and do not remove this limit.",
      "The prompts are hand-made. The order prompt reuses a case with known Haiku format misses from an earlier study; this is not a blind holdout.",
      "Both routes used one shared Mac. The files do not establish host isolation from other work. Timings include CLI start-up and network time.",
      "The current protocol file was born after the Codex batch. An earlier version may have existed, but these files cannot verify pre-call registration. Amendment 1 says 00:44 UTC and before counted calls; the Codex batch started at 00:43 UTC. Amendment 4 says 06:10 UTC and before the Claude probe; the probe receipt was saved at 06:04 UTC and counted Claude calls began at 06:04 UTC. These timing conflicts limit the protocol claims. The calls remain descriptive evidence.",
      "The first Claude lane stopped after a 90-minute wait for another study to release the route. It made no inference call. The lane restarted later; no counted call was retried.",
      "The comparison changes both the schema flag and the format sentence. It cannot isolate the effect of the flag alone.",
      "These schemas constrain keys, types and currency labels. They do not verify totals or due dates. Both modes use the same task text. A change in wrong values is a measured result, not a design aim.",
      "Sonnet and GPT-6.1 Sol passed every call in both modes. These three tasks are within reach of those models, so they cannot show a schema effect for them; harder extraction may differ.",
      "The instruction wording is one sentence; another wording, or a schema with the format sentence kept, was not tested. The format-miss counts here are not comparable with the caching and consistency study, which used a different wording on one of these prompts.",
      "In schema mode Claude Code returns the answer through a structured-output tool call. Its median was 2 model turns (range 2 to 2, n = 36). Instructions had a median of 1 (range 1 to 1, n = 36). Its time and token counts include that extra step.",
      "The Codex CLI’s exec mode does not echo the model or the reasoning effort. The request named GPT-6.1 Sol at low effort, and the CLI’s own catalog lists that model with that effort, but a reroute would not have been visible. Codex timings include its start-up and its larger system prompt, so Codex rows compare route and model pairs, not models alone."
    ],
    "sourceIds": [
      "agent-structured-output"
    ],
    "stats": [
      {
        "id": "structured-output-calls",
        "label": "Counted calls in this study (every one counted)",
        "value": 96,
        "unit": "calls",
        "display": "96 (72 Claude Code, 24 Codex CLI)"
      },
      {
        "id": "structured-output-format-miss-instructions",
        "label": "Format-miss rate with instructions only, all models",
        "value": 0.3542,
        "unit": "rate",
        "display": "35% (17/48)",
        "n": 48,
        "ci": [
          0.2343,
          0.4956
        ],
        "note": "A right answer in a code fence or prose. Every error counted as a call."
      },
      {
        "id": "structured-output-format-miss-schema",
        "label": "Format-miss rate with a JSON schema, all models",
        "value": 0,
        "unit": "rate",
        "display": "0% (0/48)",
        "n": 48,
        "ci": [
          0,
          0.0741
        ],
        "note": "A right answer in a code fence or prose. Every error counted as a call."
      },
      {
        "id": "structured-output-fenced-instructions",
        "label": "Replies in a code fence with instructions only, all models",
        "value": 0.5,
        "unit": "rate",
        "display": "50% (24/48)",
        "n": 48,
        "ci": [
          0.3639,
          0.6361
        ],
        "note": "The prompt said: no code fence. Completed replies only; a fence can hold a right or a wrong answer."
      },
      {
        "id": "structured-output-fenced-schema",
        "label": "Replies in a code fence with a JSON schema, all models",
        "value": 0,
        "unit": "rate",
        "display": "0% (0/48)",
        "n": 48,
        "ci": [
          0,
          0.0741
        ],
        "note": "Completed replies only. Claude Code uses its structured_output field. Codex CLI uses its final message."
      },
      {
        "id": "structured-output-wrong-values-instructions",
        "label": "Wrong-values rate with instructions only, all models",
        "value": 0.1458,
        "unit": "rate",
        "display": "15% (7/48)",
        "n": 48,
        "ci": [
          0.0725,
          0.2717
        ],
        "note": "A completed reply that is neither a strict pass nor a format miss. Denominator includes all attempts; errors are a separate outcome."
      },
      {
        "id": "structured-output-wrong-values-schema",
        "label": "Wrong-values rate with a JSON schema, all models",
        "value": 0.125,
        "unit": "rate",
        "display": "13% (6/48)",
        "n": 48,
        "ci": [
          0.0586,
          0.247
        ],
        "note": "A completed reply that is neither a strict pass nor a format miss. Denominator includes all attempts; errors are a separate outcome."
      },
      {
        "id": "structured-output-strict-instructions",
        "label": "Strict pass rate with instructions only, all models",
        "value": 0.5,
        "unit": "rate",
        "display": "50% (24/48)",
        "n": 48,
        "ci": [
          0.3639,
          0.6361
        ],
        "note": "Pooled over three models and three prompts; the models differ in n, so read the per-model chart."
      },
      {
        "id": "structured-output-strict-schema",
        "label": "Strict pass rate with a JSON schema, all models",
        "value": 0.875,
        "unit": "rate",
        "display": "88% (42/48)",
        "n": 48,
        "ci": [
          0.753,
          0.9414
        ],
        "note": "Pooled over three models and three prompts; the models differ in n, so read the per-model chart."
      }
    ],
    "charts": [
      {
        "id": "structured-output-pass-rate",
        "title": "Does a JSON schema raise the pass rate? Instructions vs schema mode",
        "subtitle": "Three extraction prompts pooled; whiskers are 95% Wilson intervals",
        "kind": "dot-range",
        "unit": "rate",
        "polarity": "higher",
        "yLabel": "Passed",
        "series": [
          {
            "name": "Strict pass: the whole reply is the right JSON",
            "points": [
              {
                "label": "Claude Haiku 4.5 (instructions) · Claude Code",
                "value": 0,
                "lo": 0,
                "hi": 0.138,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (JSON schema) · Claude Code",
                "value": 0.75,
                "lo": 0.551,
                "hi": 0.88,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (instructions) · Claude Code",
                "value": 1,
                "lo": 0.7575,
                "hi": 1,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 (JSON schema) · Claude Code",
                "value": 1,
                "lo": 0.7575,
                "hi": 1,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, instructions) · Codex CLI",
                "value": 1,
                "lo": 0.7575,
                "hi": 1,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, JSON schema) · Codex CLI",
                "value": 1,
                "lo": 0.7575,
                "hi": 1,
                "n": 12
              }
            ]
          },
          {
            "name": "Right answer in any format (strict pass or format miss)",
            "points": [
              {
                "label": "Claude Haiku 4.5 (instructions) · Claude Code",
                "value": 0.7083,
                "lo": 0.5083,
                "hi": 0.8509,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (JSON schema) · Claude Code",
                "value": 0.75,
                "lo": 0.551,
                "hi": 0.88,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (instructions) · Claude Code",
                "value": 1,
                "lo": 0.7575,
                "hi": 1,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 (JSON schema) · Claude Code",
                "value": 1,
                "lo": 0.7575,
                "hi": 1,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, instructions) · Codex CLI",
                "value": 1,
                "lo": 0.7575,
                "hi": 1,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, JSON schema) · Codex CLI",
                "value": 1,
                "lo": 0.7575,
                "hi": 1,
                "n": 12
              }
            ]
          }
        ],
        "note": "Whiskers are 95% Wilson intervals (a calculation) over 24 calls and 12 calls per configuration; every error counts as a fail. Strict: the whole reply parses as JSON and matches the expected answer exactly. A format miss is a right answer inside a code fence or prose, so it is never a strict pass.",
        "whisker": "ci95",
        "sourceIds": [
          "agent-structured-output"
        ]
      },
      {
        "id": "structured-output-outcomes",
        "title": "What each call produced: strict pass, format miss, wrong values or error",
        "subtitle": "Counts of calls per configuration; the three prompts pooled",
        "kind": "stacked-bar",
        "unit": "count",
        "yLabel": "Calls",
        "series": [
          {
            "name": "Strict pass",
            "points": [
              {
                "label": "Claude Haiku 4.5 (instructions) · Claude Code",
                "value": 0,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (JSON schema) · Claude Code",
                "value": 18,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (instructions) · Claude Code",
                "value": 12,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 (JSON schema) · Claude Code",
                "value": 12,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, instructions) · Codex CLI",
                "value": 12,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, JSON schema) · Codex CLI",
                "value": 12,
                "n": 12
              }
            ]
          },
          {
            "name": "Format miss",
            "points": [
              {
                "label": "Claude Haiku 4.5 (instructions) · Claude Code",
                "value": 17,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (JSON schema) · Claude Code",
                "value": 0,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (instructions) · Claude Code",
                "value": 0,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 (JSON schema) · Claude Code",
                "value": 0,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, instructions) · Codex CLI",
                "value": 0,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, JSON schema) · Codex CLI",
                "value": 0,
                "n": 12
              }
            ]
          },
          {
            "name": "Wrong values",
            "points": [
              {
                "label": "Claude Haiku 4.5 (instructions) · Claude Code",
                "value": 7,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (JSON schema) · Claude Code",
                "value": 6,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (instructions) · Claude Code",
                "value": 0,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 (JSON schema) · Claude Code",
                "value": 0,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, instructions) · Codex CLI",
                "value": 0,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, JSON schema) · Codex CLI",
                "value": 0,
                "n": 12
              }
            ]
          },
          {
            "name": "Error",
            "points": [
              {
                "label": "Claude Haiku 4.5 (instructions) · Claude Code",
                "value": 0,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (JSON schema) · Claude Code",
                "value": 0,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (instructions) · Claude Code",
                "value": 0,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 (JSON schema) · Claude Code",
                "value": 0,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, instructions) · Codex CLI",
                "value": 0,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, JSON schema) · Codex CLI",
                "value": 0,
                "n": 12
              }
            ]
          }
        ],
        "note": "Counts of calls, not rates; the pass-rate chart carries the same results with 95% intervals. Format miss: the right answer inside a code fence or prose. Wrong values: any other completed reply, with a wrong value, key or type (a reply that sits in a code fence and also has a wrong value is counted here). Error: the call did not complete.",
        "sourceIds": [
          "agent-structured-output"
        ]
      },
      {
        "id": "structured-output-time",
        "title": "Time per call, instructions vs schema mode",
        "subtitle": "Median; whiskers = fastest and slowest completed call",
        "kind": "dot-range",
        "unit": "seconds",
        "yLabel": "Seconds",
        "series": [
          {
            "name": "Median time per call (the three prompts pooled)",
            "points": [
              {
                "label": "Claude Haiku 4.5 (instructions) · Claude Code",
                "value": 9.52,
                "lo": 5.67,
                "hi": 17,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (JSON schema) · Claude Code",
                "value": 8.46,
                "lo": 5.9,
                "hi": 12.23,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (instructions) · Claude Code",
                "value": 3.52,
                "lo": 2.67,
                "hi": 4.12,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 (JSON schema) · Claude Code",
                "value": 4.2,
                "lo": 2.95,
                "hi": 6.14,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, instructions) · Codex CLI",
                "value": 6.21,
                "lo": 4.2,
                "hi": 12.27,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, JSON schema) · Codex CLI",
                "value": 5.96,
                "lo": 4.62,
                "hi": 20.97,
                "n": 12
              }
            ]
          }
        ],
        "note": "Whiskers are a range (fastest and slowest call), not a confidence interval. Wall time from process start to exit, so it includes CLI start-up; Codex CLI timings include its larger system prompt. The three prompts differ in length, which widens every range.",
        "whisker": "minmax",
        "sourceIds": [
          "agent-structured-output"
        ]
      },
      {
        "id": "structured-output-tokens",
        "title": "Output and reasoning tokens per call, instructions vs schema mode",
        "subtitle": "Median per call; ranges and sample sizes are in the note",
        "kind": "grouped-bar",
        "unit": "tokens",
        "yLabel": "Tokens",
        "series": [
          {
            "name": "Median output tokens per call",
            "points": [
              {
                "label": "Claude Haiku 4.5 (instructions) · Claude Code",
                "value": 1128,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (JSON schema) · Claude Code",
                "value": 1036,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (instructions) · Claude Code",
                "value": 368,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 (JSON schema) · Claude Code",
                "value": 424,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, instructions) · Codex CLI",
                "value": 117,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, JSON schema) · Codex CLI",
                "value": 123,
                "n": 12
              }
            ]
          },
          {
            "name": "Median reasoning tokens per call (thinking)",
            "points": [
              {
                "label": "Claude Haiku 4.5 (instructions) · Claude Code",
                "value": 934,
                "n": 24
              },
              {
                "label": "Claude Haiku 4.5 (JSON schema) · Claude Code",
                "value": 727,
                "n": 24
              },
              {
                "label": "Claude Sonnet 5.5 (instructions) · Claude Code",
                "value": 182,
                "n": 12
              },
              {
                "label": "Claude Sonnet 5.5 (JSON schema) · Claude Code",
                "value": 109,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, instructions) · Codex CLI",
                "value": 21,
                "n": 12
              },
              {
                "label": "GPT-6.1 Sol (low, JSON schema) · Codex CLI",
                "value": 23,
                "n": 12
              }
            ]
          }
        ],
        "note": "Output tokens include reasoning tokens. Claude schema-mode output includes the CLI’s structured-output tool call. Input counts include CLI context and are not compared. Ranges (not intervals): Claude Haiku 4.5 (instructions) · Claude Code: output 698 to 2059 (n = 24); reasoning 559 to 1830 (n = 24); Claude Haiku 4.5 (JSON schema) · Claude Code: output 716 to 1462 (n = 24); reasoning 522 to 1187 (n = 24); Claude Sonnet 5.5 (instructions) · Claude Code: output 154 to 456 (n = 12); reasoning 54 to 278 (n = 12); Claude Sonnet 5.5 (JSON schema) · Claude Code: output 273 to 541 (n = 12); reasoning 0 to 260 (n = 12); GPT-6.1 Sol (low, instructions) · Codex CLI: output 69 to 259 (n = 12); reasoning 0 to 60 (n = 12); GPT-6.1 Sol (low, JSON schema) · Codex CLI: output 98 to 181 (n = 12); reasoning 0 to 49 (n = 12).",
        "sourceIds": [
          "agent-structured-output"
        ]
      }
    ],
    "tables": [
      {
        "id": "structured-output-cells",
        "title": "Every cell: configuration by prompt",
        "columns": [
          {
            "key": "config",
            "label": "Configuration",
            "unit": "text"
          },
          {
            "key": "prompt",
            "label": "Prompt",
            "unit": "text"
          },
          {
            "key": "strict",
            "label": "Strict passes",
            "unit": "text"
          },
          {
            "key": "ci",
            "label": "95% interval",
            "unit": "text"
          },
          {
            "key": "formatMisses",
            "label": "Format misses",
            "unit": "count"
          },
          {
            "key": "wrong",
            "label": "Wrong values",
            "unit": "count"
          },
          {
            "key": "errors",
            "label": "Errors",
            "unit": "count"
          },
          {
            "key": "fenced",
            "label": "Replies in a code fence",
            "unit": "count"
          },
          {
            "key": "completedN",
            "label": "Completed calls (n for time and tokens)",
            "unit": "count"
          },
          {
            "key": "medianTotal",
            "label": "Median time (s)",
            "unit": "seconds"
          },
          {
            "key": "timeRange",
            "label": "Time range (s, not an interval)",
            "unit": "text"
          },
          {
            "key": "medianOutput",
            "label": "Median output tokens",
            "unit": "tokens"
          },
          {
            "key": "outputRange",
            "label": "Output token range",
            "unit": "text"
          },
          {
            "key": "medianReasoning",
            "label": "Median reasoning tokens",
            "unit": "tokens"
          },
          {
            "key": "reasoningRange",
            "label": "Reasoning token range",
            "unit": "text"
          }
        ],
        "rows": [
          {
            "config": "Claude Haiku 4.5 (instructions) · Claude Code",
            "prompt": "Order to JSON",
            "strict": "0/8",
            "ci": "0% to 32%",
            "formatMisses": 8,
            "wrong": 0,
            "errors": 0,
            "fenced": 8,
            "medianTotal": 7.01,
            "timeRange": "5.67 to 9.8",
            "medianOutput": 897,
            "outputRange": "698 to 1134",
            "medianReasoning": 758.5,
            "reasoningRange": "559 to 996",
            "completedN": 8
          },
          {
            "config": "Claude Haiku 4.5 (instructions) · Claude Code",
            "prompt": "Invoice to JSON",
            "strict": "0/8",
            "ci": "0% to 32%",
            "formatMisses": 8,
            "wrong": 0,
            "errors": 0,
            "fenced": 8,
            "medianTotal": 9.92,
            "timeRange": "7.19 to 13.03",
            "medianOutput": 1280,
            "outputRange": "1088 to 1636",
            "medianReasoning": 1047,
            "reasoningRange": "856 to 1403",
            "completedN": 8
          },
          {
            "config": "Claude Haiku 4.5 (instructions) · Claude Code",
            "prompt": "Meeting notes to action items",
            "strict": "0/8",
            "ci": "0% to 32%",
            "formatMisses": 1,
            "wrong": 7,
            "errors": 0,
            "fenced": 8,
            "medianTotal": 10.34,
            "timeRange": "8.2 to 17",
            "medianOutput": 1235,
            "outputRange": "991 to 2059",
            "medianReasoning": 1030,
            "reasoningRange": "761 to 1830",
            "completedN": 8
          },
          {
            "config": "Claude Haiku 4.5 (JSON schema) · Claude Code",
            "prompt": "Order to JSON",
            "strict": "8/8",
            "ci": "68% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 7.19,
            "timeRange": "6.45 to 10.5",
            "medianOutput": 791.5,
            "outputRange": "716 to 1222",
            "medianReasoning": 586,
            "reasoningRange": "522 to 1010",
            "completedN": 8
          },
          {
            "config": "Claude Haiku 4.5 (JSON schema) · Claude Code",
            "prompt": "Invoice to JSON",
            "strict": "8/8",
            "ci": "68% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 9.39,
            "timeRange": "7.7 to 9.91",
            "medianOutput": 1120.5,
            "outputRange": "909 to 1226",
            "medianReasoning": 788,
            "reasoningRange": "576 to 893",
            "completedN": 8
          },
          {
            "config": "Claude Haiku 4.5 (JSON schema) · Claude Code",
            "prompt": "Meeting notes to action items",
            "strict": "2/8",
            "ci": "7% to 59%",
            "formatMisses": 0,
            "wrong": 6,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 9.27,
            "timeRange": "5.9 to 12.23",
            "medianOutput": 1164.5,
            "outputRange": "806 to 1462",
            "medianReasoning": 890.5,
            "reasoningRange": "532 to 1187",
            "completedN": 8
          },
          {
            "config": "Claude Sonnet 5.5 (instructions) · Claude Code",
            "prompt": "Order to JSON",
            "strict": "4/4",
            "ci": "51% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 3.27,
            "timeRange": "2.67 to 3.8",
            "medianOutput": 166.5,
            "outputRange": "154 to 175",
            "medianReasoning": 66.5,
            "reasoningRange": "54 to 75",
            "completedN": 4
          },
          {
            "config": "Claude Sonnet 5.5 (instructions) · Claude Code",
            "prompt": "Invoice to JSON",
            "strict": "4/4",
            "ci": "51% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 3.65,
            "timeRange": "3.32 to 4.12",
            "medianOutput": 367.5,
            "outputRange": "362 to 373",
            "medianReasoning": 181.5,
            "reasoningRange": "176 to 187",
            "completedN": 4
          },
          {
            "config": "Claude Sonnet 5.5 (instructions) · Claude Code",
            "prompt": "Meeting notes to action items",
            "strict": "4/4",
            "ci": "51% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 3.7,
            "timeRange": "3.49 to 3.93",
            "medianOutput": 417.5,
            "outputRange": "392 to 456",
            "medianReasoning": 239.5,
            "reasoningRange": "214 to 278",
            "completedN": 4
          },
          {
            "config": "Claude Sonnet 5.5 (JSON schema) · Claude Code",
            "prompt": "Order to JSON",
            "strict": "4/4",
            "ci": "51% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 3.22,
            "timeRange": "2.95 to 6.14",
            "medianOutput": 273.5,
            "outputRange": "273 to 275",
            "medianReasoning": 64.5,
            "reasoningRange": "64 to 66",
            "completedN": 4
          },
          {
            "config": "Claude Sonnet 5.5 (JSON schema) · Claude Code",
            "prompt": "Invoice to JSON",
            "strict": "4/4",
            "ci": "51% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 4.37,
            "timeRange": "4.2 to 4.54",
            "medianOutput": 503.5,
            "outputRange": "495 to 541",
            "medianReasoning": 76,
            "reasoningRange": "0 to 154",
            "completedN": 4
          },
          {
            "config": "Claude Sonnet 5.5 (JSON schema) · Claude Code",
            "prompt": "Meeting notes to action items",
            "strict": "4/4",
            "ci": "51% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 4.25,
            "timeRange": "3.9 to 5.35",
            "medianOutput": 423.5,
            "outputRange": "408 to 498",
            "medianReasoning": 185.5,
            "reasoningRange": "170 to 260",
            "completedN": 4
          },
          {
            "config": "GPT-6.1 Sol (low, instructions) · Codex CLI",
            "prompt": "Order to JSON",
            "strict": "4/4",
            "ci": "51% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 6.21,
            "timeRange": "4.2 to 8.24",
            "medianOutput": 92,
            "outputRange": "69 to 93",
            "medianReasoning": 21,
            "reasoningRange": "0 to 22",
            "completedN": 4
          },
          {
            "config": "GPT-6.1 Sol (low, instructions) · Codex CLI",
            "prompt": "Invoice to JSON",
            "strict": "4/4",
            "ci": "51% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 10.19,
            "timeRange": "6.75 to 12.27",
            "medianOutput": 248,
            "outputRange": "235 to 259",
            "medianReasoning": 49,
            "reasoningRange": "36 to 60",
            "completedN": 4
          },
          {
            "config": "GPT-6.1 Sol (low, instructions) · Codex CLI",
            "prompt": "Meeting notes to action items",
            "strict": "4/4",
            "ci": "51% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 5.01,
            "timeRange": "4.71 to 5.81",
            "medianOutput": 117,
            "outputRange": "117 to 117",
            "medianReasoning": 0,
            "reasoningRange": "0 to 0",
            "completedN": 4
          },
          {
            "config": "GPT-6.1 Sol (low, JSON schema) · Codex CLI",
            "prompt": "Order to JSON",
            "strict": "4/4",
            "ci": "51% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 5.73,
            "timeRange": "4.62 to 20.97",
            "medianOutput": 99.5,
            "outputRange": "98 to 102",
            "medianReasoning": 22.5,
            "reasoningRange": "21 to 25",
            "completedN": 4
          },
          {
            "config": "GPT-6.1 Sol (low, JSON schema) · Codex CLI",
            "prompt": "Invoice to JSON",
            "strict": "4/4",
            "ci": "51% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 6.13,
            "timeRange": "5.77 to 6.37",
            "medianOutput": 172.5,
            "outputRange": "166 to 181",
            "medianReasoning": 40.5,
            "reasoningRange": "34 to 49",
            "completedN": 4
          },
          {
            "config": "GPT-6.1 Sol (low, JSON schema) · Codex CLI",
            "prompt": "Meeting notes to action items",
            "strict": "4/4",
            "ci": "51% to 100%",
            "formatMisses": 0,
            "wrong": 0,
            "errors": 0,
            "fenced": 0,
            "medianTotal": 5.76,
            "timeRange": "4.87 to 6.51",
            "medianOutput": 123,
            "outputRange": "123 to 123",
            "medianReasoning": 0,
            "reasoningRange": "0 to 0",
            "completedN": 4
          }
        ]
      },
      {
        "id": "structured-output-pairs",
        "title": "Paired view: the same prompt and repetition, asked both ways",
        "columns": [
          {
            "key": "model",
            "label": "Model and route",
            "unit": "text"
          },
          {
            "key": "pairs",
            "label": "Pairs",
            "unit": "count"
          },
          {
            "key": "both",
            "label": "Both passed",
            "unit": "count"
          },
          {
            "key": "onlyInstructions",
            "label": "Only instructions passed",
            "unit": "count"
          },
          {
            "key": "onlySchema",
            "label": "Only schema passed",
            "unit": "count"
          },
          {
            "key": "neither",
            "label": "Neither passed",
            "unit": "count"
          },
          {
            "key": "mcnemarP",
            "label": "Exact two-sided McNemar p (calculation)",
            "unit": "score"
          }
        ],
        "rows": [
          {
            "model": "Claude Haiku 4.5 · Claude Code",
            "pairs": 24,
            "both": 0,
            "onlyInstructions": 0,
            "onlySchema": 18,
            "neither": 6,
            "mcnemarP": 0.00000762939453125
          },
          {
            "model": "Claude Sonnet 5.5 · Claude Code",
            "pairs": 12,
            "both": 12,
            "onlyInstructions": 0,
            "onlySchema": 0,
            "neither": 0,
            "mcnemarP": 1
          },
          {
            "model": "GPT-6.1 Sol · Codex CLI",
            "pairs": 12,
            "both": 12,
            "onlyInstructions": 0,
            "onlySchema": 0,
            "neither": 0,
            "mcnemarP": 1
          }
        ]
      }
    ],
    "related": [
      "caching-consistency",
      "hard-model-head-to-head"
    ]
  },
  "sources": [
    {
      "id": "agent-structured-output",
      "title": "JSON schema vs instructions",
      "kind": "run",
      "date": "2026-10-07",
      "data": [
        "/benchmarks/raw/structured-output/receipts.json"
      ],
      "note": "Receipts copied from a run of the same three JSON extraction prompts, each asked with instructions only (mode I) and with the CLI’s JSON schema mode (mode S). Prompts, model output and failure reasons are not published; a failed check is named, never quoted. Every attempt is kept. Both routes ran one call at a time. The current protocol file does not verify pre-call registration; see protocolAudit. Calls repeat three fixed hand-made prompts, including a known format-miss case. Both routes used a shared Mac."
    }
  ]
}
