• Structured Output
  • Json Schema
  • Format Miss
  • Claude Haiku

Does a JSON schema stop format misses? Haiku went from 0/24 to 18/24

96 calls on 3 JSON tasks. Haiku 4.5 passed 0/24 with instructions and 18/24 with a JSON schema. Sonnet 5.5 and GPT-6.1 Sol passed 12/12 either way.

TL;DR

  • With a JSON schema, we observed no format misses. We ran 96 counted calls on 3 JSON extraction prompts. Each prompt ran two ways: with instructions only, and with the CLI's schema mode. Format misses: 17 of 48 calls with instructions (95% interval 23% to 50%), 0 of 48 with a schema (0% to 7%). Models are pooled with different sample sizes.
  • Haiku 4.5 gained the most. Strict passes: 0/24 with instructions (95% interval 0% to 14%) and 18/24 with a schema (55% to 88%). The intervals do not overlap, so the schema is ahead.
  • Haiku ignored "no code fence" 24 times out of 24. The prompt said "no code fence". All 24 replies came in a fence anyway (95% interval 86% to 100%). 17 held the right answer. 7 held a wrong one.
  • A schema constrains the shape; it does not verify values. 6 of 24 Haiku schema replies were still wrong (95% interval 12% to 45%). All 6 gave Saturday 10 October for "by Friday".
  • Sonnet 5.5 and GPT-6.1 Sol (low) passed 12/12 in both modes (76% to 100%). These tasks are at a ceiling for them: the data shows no schema effect.
  • Time and tokens did not separate the modes. In every pair, the fastest-to-slowest ranges overlap.
  • Recommendation: use the schema mode for Haiku, where the intervals do not overlap. For Sonnet and GPT-6.1 Sol the intervals overlap, so this sample shows no difference. Keep a validator on the values in every case.

Full study, charts and data: the JSON schema vs instructions study.

The question

Format misses were a large share of our failures. In the consistency study, Claude Haiku 4.5 passed a JSON prompt 1 time in 10 with a strict parser (95% interval 2% to 40%). The other 9 replies had the right content in the wrong format (60% to 98%).

  • Exact number
  • JSON object
  • Code fix
Claude Haiku 4.5
Claude Sonnet 5.5
GPT-6.1 Sol (medium)

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 10 per row

One series per prompt; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

On our hard task set, 5 of 13 non-passes were format misses too (95% interval 18% to 64%). The strict validators rejected those right answers.

Claude Code and the Codex CLI both have a schema mode. You give the CLI a JSON schema. It returns the answer in that shape. So we asked one question: does the schema mode stop format misses, and does it change the values?

What we ran

  • 3 extraction prompts. Each has one exact expected JSON and a sandboxed validator.
    • Order note to JSON: merge two lines for one SKU, drop a cancelled line, apply a discount to one SKU, give the total in cents. 9 checks.
    • Invoice to JSON: a line discount, 8% tax after the discount rounded half up, an untaxed fee, shipping. Per-line and total cents. 13 checks.
    • Meeting notes to action items: ticket, owner and due date. Resolve "by Friday", "two weeks" and "tomorrow" against the meeting date. Skip a decision and a note. Use null for no deadline. 9 checks.
  • 2 modes, same task text.
    • Instructions only: the task plus one sentence, "Reply with the JSON object only, no code fence, no other text."
    • JSON schema: no format sentence. The CLI's own schema mode: claude --json-schema for Claude Code, and codex exec --output-schema for Codex.
  • 3 configurations. Claude Haiku 4.5 (8 repetitions per prompt and mode) and Claude Sonnet 5.5 (4), both through Claude Code at default effort. GPT-6.1 Sol at low effort through the Codex CLI (4).
  • Controls before the first call. Each reference answer passes. Each planted wrong answer fails. A reference inside a code fence counts as a format miss.
  • One call at a time per account. A fresh empty folder per call, tools off, no session. The order of the two modes alternates by repetition.

We counted all 96 inference calls. None failed or was retried. Both counted batches completed. Two harness probes stay outside the counts. An earlier Claude lane stopped after a 90-minute wait; it made no inference call.

ResultMeaning
Strict passThe whole reply parses as JSON and matches the expected answer exactly.
Format missThe answer is right, but it sits inside a code fence or other text. A strict parser rejects it.
Wrong valuesAny other completed reply: a wrong number, key, type or date.
ErrorThe call did not finish. It counts as a fail.

All intervals in this post are 95% Wilson intervals (calculation).

Result: schema mode raised strict passes for Haiku in this sample

Calculation
  • Strict pass: the whole reply is the right JSON
  • Right answer in any format (strict pass or format miss)
Claude Haiku 4.5 (instruction…
Claude Haiku 4.5 (JSON schema)
Claude Sonnet 5.5 (instructio…
Claude Sonnet 5.5 (JSON schem…
GPT-6.1 Sol (low, instruction…
GPT-6.1 Sol (low, JSON schema)

List-price calculation, not a run. 6 rows, 2 series: Strict pass: the whole reply is the right JSON, Right answer in any format (strict pass or format miss). Strict pass: the whole reply is the right JSON: highest Claude Sonnet 5.5 (instructions) · Claude Code 100% (95% interval 76%–100%, n 12). Lowest Claude Haiku 4.5 (instructions) · Claude Code 0% (95% interval 0%–14%, n 24). Not all intervals overlap. Right answer in any format (strict pass or format miss): highest Claude Sonnet 5.5 (instructions) · Claude Code 100% (95% interval 76%–100%, n 12). Lowest Claude Haiku 4.5 (instructions) · Claude Code 71% (95% interval 51%–85%, n 24). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 12–24 per row4 of 6 (Strict pass: the whole reply is the right JSON) at 100%: this task set cannot separate them.

Three extraction prompts pooled; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals (a calculation) over 24 calls and 12 calls per configuration; every error counts as a fail. Strict: the whole reply parses as JSON and matches the expected answer exactly. A format miss is a right answer inside a code fence or prose, so it is never a strict pass.

Source: JSON schema vs instructions

ConfigurationWith instructionsWith a JSON schema
Claude Haiku 4.5 · Claude Code (n = 24 per mode)0/24 (0% to 14%)18/24 (55% to 88%)
Claude Sonnet 5.5 · Claude Code (n = 12 per mode)12/12 (76% to 100%)12/12 (76% to 100%)
GPT-6.1 Sol (low) · Codex CLI (n = 12 per mode)12/12 (76% to 100%)12/12 (76% to 100%)

Haiku's two intervals do not overlap. That is the only pair where the data ranks the two modes. In the paired view, the same prompt and repetition asked both ways, 18 of 24 pairs passed only with the schema. 6 passed in neither mode. 0 passed only with instructions. The exact two-sided McNemar p is 0.00000763 (calculation).

The Haiku fence habit

All 24 Haiku replies with instructions began with a code fence. The prompt said not to use one. With the schema, none of the 24 scored answers had a fence (95% interval 0% to 14%). Claude Code reads those answers from a structured field rather than the reply text.

17 of the 24 fenced replies held the right answer. Those are 17 format misses (17/24, 95% interval 51% to 85%). A lenient parser that strips the fence would have passed them. A strict parser rejects them.

  • Strict pass
  • Format miss
  • Wrong values
  • Error
Claude Haiku 4.5 (instructions) · Claude Code
Claude Haiku 4.5 (JSON schema) · Claude Code
Claude Sonnet 5.5 (instructions) · Claude Code
Claude Sonnet 5.5 (JSON schema) · Claude Code
GPT-6.1 Sol (low, instructions) · Codex CLI
GPT-6.1 Sol (low, JSON schema) · Codex CLI

One square per call; counts at the right are exact and in legend order.

6 rows, 4 series: Strict pass, Format miss, Wrong values, Error. Strict pass: highest Claude Haiku 4.5 (JSON schema) · Claude Code 18 (n 24). Lowest Claude Haiku 4.5 (instructions) · Claude Code 0 (n 24). Format miss: highest Claude Haiku 4.5 (instructions) · Claude Code 17 (n 24). Lowest GPT-6.1 Sol (low, JSON schema) · Codex CLI 0 (n 12).

Notesn 12–24 per row

Counts of calls per configuration; the three prompts pooled

Counts of calls, not rates; the pass-rate chart carries the same results with 95% intervals. Format miss: the right answer inside a code fence or prose. Wrong values: any other completed reply, with a wrong value, key or type (a reply that sits in a code fence and also has a wrong value is counted here). Error: the call did not complete.

Source: JSON schema vs instructions

Haiku 4.5 promptInstructions: strict passesWhat the other calls wereJSON schema: strict passes
Order to JSON0/8 (0% to 32%)8 format misses (68% to 100%)8/8 (68% to 100%)
Invoice to JSON0/8 (0% to 32%)8 format misses (68% to 100%)8/8 (68% to 100%)
Meeting notes to action items0/8 (0% to 32%)1 format miss (2% to 47%), 7 wrong values (53% to 98%)2/8 (7% to 59%), 6 wrong values (41% to 93%)

On the order and invoice prompts, instructions produced 16 format misses; schema mode produced 16 strict passes. The meeting prompt shows the limit.

What a schema does not fix

The meeting prompt asks for the due date of "by Friday". The meeting was on Monday 2026-10-05. The right answer is 2026-10-09. In 13 of 16 Haiku meeting replies, Haiku wrote 2026-10-10, a Saturday (95% interval 57% to 93%). That was true in both modes: 7 instructions replies and 6 schema replies. The public data names the failed check, the first action item. We read the date from the kept replies.

Every schema-mode answer had a valid shape, but some dates were wrong. So the number of right answers in any format barely moved: 17/24 with instructions (51% to 85%) and 18/24 with a schema (55% to 88%). Those intervals overlap. The clear gain was in strict passes. These results do not establish a gain in value accuracy. Both the schema flag and the format sentence changed, so the test cannot isolate the flag alone.

Sonnet 5.5 and GPT-6.1 Sol each gave 2026-10-09 in 8 of 8 meeting calls (95% interval 68% to 100%).

Sonnet and GPT-6.1 Sol: a ceiling

Neither model wrapped a reply in a fence. Neither made a wrong value. Both passed all 12 calls in both modes.

A 12 out of 12 has a 95% interval of 76% to 100%. That interval is not a minimum detectable difference. This sample does not establish equal reliability. These three tasks are within reach of both models, so they cannot show a schema effect for them. A harder extraction may differ. We make no claim either way.

Time and tokens

Entrance: medians race at 6.8× real timeMotion reduced: press Replay to animateThe slowest median is 9.5 s. The clock runs at the recorded speed.
Claude Haiku 4.5 (instructions) · Claude Code
Claude Haiku 4.5 (JSON schema) · Claude Code
Claude Sonnet 5.5 (instructions) · Claude Code
Claude Sonnet 5.5 (JSON schema) · Claude Code
GPT-6.1 Sol (low, instructions) · Codex CLI
GPT-6.1 Sol (low, JSON schema) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

6 rows. Slowest Claude Haiku 4.5 (instructions) · Claude Code 9.5 s (range 5.7 s–17 s, n 24). Fastest Claude Sonnet 5.5 (instructions) · Claude Code 3.5 s (range 2.7 s–4.1 s, n 12). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n 12–24 per row

Median; whiskers = fastest and slowest completed call

Whiskers are a range (fastest and slowest call), not a confidence interval. Wall time from process start to exit, so it includes CLI start-up; Codex CLI timings include its larger system prompt. The three prompts differ in length, which widens every range.

Source: JSON schema vs instructions

Median time per call, with the fastest and slowest call. These are ranges, not confidence intervals. Each mode has n = 24 for Haiku and n = 12 for each other model:

  • Haiku: 9.52 s (5.67 to 17.00) with instructions, 8.46 s (5.90 to 12.23) with a schema.
  • Sonnet: 3.52 s (2.67 to 4.12) with instructions, 4.20 s (2.95 to 6.14) with a schema.
  • GPT-6.1 Sol: 6.21 s (4.20 to 12.27) with instructions, 5.96 s (4.62 to 20.97) with a schema.

In every pair the ranges overlap, so the data does not rank the modes on time. The schema path in Claude Code takes 2 model turns (a median) where instructions take 1 (n = 36 per mode; ranges 2 to 2 and 1 to 1), because the CLI returns the answer through a structured-output tool call. That extra step did not show up as a clear time cost here. The Codex timings include the CLI's start-up and its larger system prompt, so compare modes within a row, not models across routes.

  • Median output tokens per call
  • of which median reasoning tokens per call (thinking) (inner bar)
Claude Haiku 4.5 (instructions) · Claude Code
Claude Haiku 4.5 (JSON schema) · Claude Code
Claude Sonnet 5.5 (instructions) · Claude Code
Claude Sonnet 5.5 (JSON schema) · Claude Code
GPT-6.1 Sol (low, instructions) · Codex CLI
GPT-6.1 Sol (low, JSON schema) · Codex CLI

6 rows, 2 series: Median output tokens per call, Median reasoning tokens per call (thinking). Median output tokens per call: highest Claude Haiku 4.5 (instructions) · Claude Code 1,128 (n 24). Lowest GPT-6.1 Sol (low, instructions) · Codex CLI 117 (n 12). Median reasoning tokens per call (thinking): highest Claude Haiku 4.5 (instructions) · Claude Code 934 (n 24). Lowest GPT-6.1 Sol (low, instructions) · Codex CLI 21 (n 12).

Notesn 12–24 per row

Median per call; ranges and sample sizes are in the note

Output tokens include reasoning tokens. Claude schema-mode output includes the CLI’s structured-output tool call. Input counts include CLI context and are not compared. Ranges (not intervals): Claude Haiku 4.5 (instructions) · Claude Code: output 698 to 2059 (n = 24); reasoning 559 to 1830 (n = 24); Claude Haiku 4.5 (JSON schema) · Claude Code: output 716 to 1462 (n = 24); reasoning 522 to 1187 (n = 24); Claude Sonnet 5.5 (instructions) · Claude Code: output 154 to 456 (n = 12); reasoning 54 to 278 (n = 12); Claude Sonnet 5.5 (JSON schema) · Claude Code: output 273 to 541 (n = 12); reasoning 0 to 260 (n = 12); GPT-6.1 Sol (low, instructions) · Codex CLI: output 69 to 259 (n = 12); reasoning 0 to 60 (n = 12); GPT-6.1 Sol (low, JSON schema) · Codex CLI: output 98 to 181 (n = 12); reasoning 0 to 49 (n = 12).

Source: JSON schema vs instructions

Median output tokens per call, with observed ranges (not confidence intervals):

ConfigurationInstructions: median (range)Schema: median (range)
Haiku (n = 24 per mode)1,128 (698 to 2,059)1,036 (716 to 1,462)
Sonnet (n = 12 per mode)368 (154 to 456)424 (273 to 541)
GPT-6.1 Sol (n = 12 per mode)117 (69 to 259)123 (98 to 181)

Most of Haiku's output tokens were thinking tokens. Counts are provider-reported and include reasoning. Claude schema-mode counts also include the structured-output tool call. Token ranges overlap within every model; tokenizers differ across vendors.

What to do

  1. If you run Haiku 4.5 through Claude Code, test the schema mode on your tasks (--json-schema). In this sample it removed all 17 format misses, and the intervals do not overlap.
  2. Do not trust "no code fence" in a prompt. Haiku ignored it 24 times out of 24. If you cannot use a schema, parse leniently and strip a fence before you validate. In this sample that recovers 17/24 right answers (95% interval 51% to 85%), against 18/24 (55% to 88%) with the schema. The intervals overlap.
  3. Keep a validator on the values. A schema checks keys and types. It does not check that Friday is the 9th. 6 of 24 Haiku schema replies were well-formed and wrong (95% interval 12% to 45%).
  4. For Sonnet and GPT-6.1 Sol, this sample shows no difference. Both passed everything. Test your own prompt in both modes before you add a schema for them.
  5. Repeat your prompts and keep the uncertainty visible. The Haiku meeting prompt passed 2 of 8 times with a schema (95% interval 7% to 59%). One run could have passed or failed. Eight runs still leave a wide interval; this study does not set a minimum sample size.

Limits

  • Protocol timing conflicts. The current protocol file was born at 01:47 UTC, after the first counted call at 00:43 UTC on 2026-10-07. Amendment 1 says 00:44 UTC and before counted calls. Amendment 4 says 06:10 UTC and before the Claude probe, but its receipt and counted calls began at 06:04 UTC. An earlier version may have existed; these files cannot verify pre-call registration. Treat the calls as descriptive evidence.
  • Repeated fixed prompts. Wilson intervals and the paired p-value are calculations that assume independent outcomes or pairs. Repeating three prompts does not measure accuracy across unseen tasks.
  • Known failure case. The hand-made order prompt reuses a case with known Haiku format misses. It is not a blind holdout.
  • Shared host. Both routes used one Mac. The files do not establish isolation from other work. Timings include CLI start-up and network time.
  • Two changes together. Schema mode adds the schema flag and removes the format sentence. This test cannot isolate either change.
  • Small samples. n is 24 calls per mode for Haiku and 12 for each other model. The intervals are not minimum detectable differences.
  • One wording. The instruction wording is one sentence. Another wording, or a schema with the format sentence kept, was not tested. The format-miss counts here are not comparable with the consistency study, which used a different wording on one of these prompts.
  • CLI flags, not the API. We tested claude --json-schema and codex exec --output-schema. We did not test a provider's API structured-output feature.
  • Codex does not echo the model or effort. The call named GPT-6.1 Sol at low effort, and the CLI's catalog lists it, but a reroute would not have been visible.
  • Three moderate prompts. Harder or longer extraction may behave differently.

Prompts and replies are not published. The public data holds counts, times, tokens and the name of each failed check.

Disclosure

I build Agent, the product behind this site. No Agent worker pipeline ran in this study. Every row is a model through its vendor's CLI, so the numbers say nothing about Agent's own pipeline. Every counted inference call is in the data. The files cannot verify that the protocol was written before the first counted call.

Run it on your own work

Agent keeps the same receipts for your tasks: the model, the route, the tokens, the time, the cost and the validation result. Try Agent and see your own numbers.

The data behind this post

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.