Does a JSON schema stop format misses? Haiku went from 0/24 to 18/24
96 calls on 3 JSON tasks. Haiku 4.5 passed 0/24 with instructions and 18/24 with a JSON schema. Sonnet 5.5 and GPT-6.1 Sol passed 12/12 either way.
TL;DR
- With a JSON schema, we observed no format misses. We ran 96 counted calls on 3 JSON extraction prompts. Each prompt ran two ways: with instructions only, and with the CLI's schema mode. Format misses: 17 of 48 calls with instructions (95% interval 23% to 50%), 0 of 48 with a schema (0% to 7%). Models are pooled with different sample sizes.
- Haiku 4.5 gained the most. Strict passes: 0/24 with instructions (95% interval 0% to 14%) and 18/24 with a schema (55% to 88%). The intervals do not overlap, so the schema is ahead.
- Haiku ignored "no code fence" 24 times out of 24. The prompt said "no code fence". All 24 replies came in a fence anyway (95% interval 86% to 100%). 17 held the right answer. 7 held a wrong one.
- A schema constrains the shape; it does not verify values. 6 of 24 Haiku schema replies were still wrong (95% interval 12% to 45%). All 6 gave Saturday 10 October for "by Friday".
- Sonnet 5.5 and GPT-6.1 Sol (low) passed 12/12 in both modes (76% to 100%). These tasks are at a ceiling for them: the data shows no schema effect.
- Time and tokens did not separate the modes. In every pair, the fastest-to-slowest ranges overlap.
- Recommendation: use the schema mode for Haiku, where the intervals do not overlap. For Sonnet and GPT-6.1 Sol the intervals overlap, so this sample shows no difference. Keep a validator on the values in every case.
Full study, charts and data: the JSON schema vs instructions study.
The question
Format misses were a large share of our failures. In the consistency study, Claude Haiku 4.5 passed a JSON prompt 1 time in 10 with a strict parser (95% interval 2% to 40%). The other 9 replies had the right content in the wrong format (60% to 98%).
On our hard task set, 5 of 13 non-passes were format misses too (95% interval 18% to 64%). The strict validators rejected those right answers.
Claude Code and the Codex CLI both have a schema mode. You give the CLI a JSON schema. It returns the answer in that shape. So we asked one question: does the schema mode stop format misses, and does it change the values?
What we ran
- 3 extraction prompts. Each has one exact expected JSON and a sandboxed validator.
- Order note to JSON: merge two lines for one SKU, drop a cancelled line, apply a discount to one SKU, give the total in cents. 9 checks.
- Invoice to JSON: a line discount, 8% tax after the discount rounded half up, an untaxed fee, shipping. Per-line and total cents. 13 checks.
- Meeting notes to action items: ticket, owner and due date. Resolve "by Friday", "two weeks" and "tomorrow" against the meeting date. Skip a decision and a note. Use
nullfor no deadline. 9 checks.
- 2 modes, same task text.
- Instructions only: the task plus one sentence, "Reply with the JSON object only, no code fence, no other text."
- JSON schema: no format sentence. The CLI's own schema mode:
claude --json-schemafor Claude Code, andcodex exec --output-schemafor Codex.
- 3 configurations. Claude Haiku 4.5 (8 repetitions per prompt and mode) and Claude Sonnet 5.5 (4), both through Claude Code at default effort. GPT-6.1 Sol at low effort through the Codex CLI (4).
- Controls before the first call. Each reference answer passes. Each planted wrong answer fails. A reference inside a code fence counts as a format miss.
- One call at a time per account. A fresh empty folder per call, tools off, no session. The order of the two modes alternates by repetition.
We counted all 96 inference calls. None failed or was retried. Both counted batches completed. Two harness probes stay outside the counts. An earlier Claude lane stopped after a 90-minute wait; it made no inference call.
| Result | Meaning |
|---|---|
| Strict pass | The whole reply parses as JSON and matches the expected answer exactly. |
| Format miss | The answer is right, but it sits inside a code fence or other text. A strict parser rejects it. |
| Wrong values | Any other completed reply: a wrong number, key, type or date. |
| Error | The call did not finish. It counts as a fail. |
All intervals in this post are 95% Wilson intervals (calculation).
Result: schema mode raised strict passes for Haiku in this sample
| Configuration | With instructions | With a JSON schema |
|---|---|---|
| Claude Haiku 4.5 · Claude Code (n = 24 per mode) | 0/24 (0% to 14%) | 18/24 (55% to 88%) |
| Claude Sonnet 5.5 · Claude Code (n = 12 per mode) | 12/12 (76% to 100%) | 12/12 (76% to 100%) |
| GPT-6.1 Sol (low) · Codex CLI (n = 12 per mode) | 12/12 (76% to 100%) | 12/12 (76% to 100%) |
Haiku's two intervals do not overlap. That is the only pair where the data ranks the two modes. In the paired view, the same prompt and repetition asked both ways, 18 of 24 pairs passed only with the schema. 6 passed in neither mode. 0 passed only with instructions. The exact two-sided McNemar p is 0.00000763 (calculation).
The Haiku fence habit
All 24 Haiku replies with instructions began with a code fence. The prompt said not to use one. With the schema, none of the 24 scored answers had a fence (95% interval 0% to 14%). Claude Code reads those answers from a structured field rather than the reply text.
17 of the 24 fenced replies held the right answer. Those are 17 format misses (17/24, 95% interval 51% to 85%). A lenient parser that strips the fence would have passed them. A strict parser rejects them.
| Haiku 4.5 prompt | Instructions: strict passes | What the other calls were | JSON schema: strict passes |
|---|---|---|---|
| Order to JSON | 0/8 (0% to 32%) | 8 format misses (68% to 100%) | 8/8 (68% to 100%) |
| Invoice to JSON | 0/8 (0% to 32%) | 8 format misses (68% to 100%) | 8/8 (68% to 100%) |
| Meeting notes to action items | 0/8 (0% to 32%) | 1 format miss (2% to 47%), 7 wrong values (53% to 98%) | 2/8 (7% to 59%), 6 wrong values (41% to 93%) |
On the order and invoice prompts, instructions produced 16 format misses; schema mode produced 16 strict passes. The meeting prompt shows the limit.
What a schema does not fix
The meeting prompt asks for the due date of "by Friday". The meeting was on Monday 2026-10-05. The right answer is 2026-10-09. In 13 of 16 Haiku meeting replies, Haiku wrote 2026-10-10, a Saturday (95% interval 57% to 93%). That was true in both modes: 7 instructions replies and 6 schema replies. The public data names the failed check, the first action item. We read the date from the kept replies.
Every schema-mode answer had a valid shape, but some dates were wrong. So the number of right answers in any format barely moved: 17/24 with instructions (51% to 85%) and 18/24 with a schema (55% to 88%). Those intervals overlap. The clear gain was in strict passes. These results do not establish a gain in value accuracy. Both the schema flag and the format sentence changed, so the test cannot isolate the flag alone.
Sonnet 5.5 and GPT-6.1 Sol each gave 2026-10-09 in 8 of 8 meeting calls (95% interval 68% to 100%).
Sonnet and GPT-6.1 Sol: a ceiling
Neither model wrapped a reply in a fence. Neither made a wrong value. Both passed all 12 calls in both modes.
A 12 out of 12 has a 95% interval of 76% to 100%. That interval is not a minimum detectable difference. This sample does not establish equal reliability. These three tasks are within reach of both models, so they cannot show a schema effect for them. A harder extraction may differ. We make no claim either way.
Time and tokens
Median time per call, with the fastest and slowest call. These are ranges, not confidence intervals. Each mode has n = 24 for Haiku and n = 12 for each other model:
- Haiku: 9.52 s (5.67 to 17.00) with instructions, 8.46 s (5.90 to 12.23) with a schema.
- Sonnet: 3.52 s (2.67 to 4.12) with instructions, 4.20 s (2.95 to 6.14) with a schema.
- GPT-6.1 Sol: 6.21 s (4.20 to 12.27) with instructions, 5.96 s (4.62 to 20.97) with a schema.
In every pair the ranges overlap, so the data does not rank the modes on time. The schema path in Claude Code takes 2 model turns (a median) where instructions take 1 (n = 36 per mode; ranges 2 to 2 and 1 to 1), because the CLI returns the answer through a structured-output tool call. That extra step did not show up as a clear time cost here. The Codex timings include the CLI's start-up and its larger system prompt, so compare modes within a row, not models across routes.
Median output tokens per call, with observed ranges (not confidence intervals):
| Configuration | Instructions: median (range) | Schema: median (range) |
|---|---|---|
| Haiku (n = 24 per mode) | 1,128 (698 to 2,059) | 1,036 (716 to 1,462) |
| Sonnet (n = 12 per mode) | 368 (154 to 456) | 424 (273 to 541) |
| GPT-6.1 Sol (n = 12 per mode) | 117 (69 to 259) | 123 (98 to 181) |
Most of Haiku's output tokens were thinking tokens. Counts are provider-reported and include reasoning. Claude schema-mode counts also include the structured-output tool call. Token ranges overlap within every model; tokenizers differ across vendors.
What to do
- If you run Haiku 4.5 through Claude Code, test the schema mode on your tasks (
--json-schema). In this sample it removed all 17 format misses, and the intervals do not overlap. - Do not trust "no code fence" in a prompt. Haiku ignored it 24 times out of 24. If you cannot use a schema, parse leniently and strip a fence before you validate. In this sample that recovers 17/24 right answers (95% interval 51% to 85%), against 18/24 (55% to 88%) with the schema. The intervals overlap.
- Keep a validator on the values. A schema checks keys and types. It does not check that Friday is the 9th. 6 of 24 Haiku schema replies were well-formed and wrong (95% interval 12% to 45%).
- For Sonnet and GPT-6.1 Sol, this sample shows no difference. Both passed everything. Test your own prompt in both modes before you add a schema for them.
- Repeat your prompts and keep the uncertainty visible. The Haiku meeting prompt passed 2 of 8 times with a schema (95% interval 7% to 59%). One run could have passed or failed. Eight runs still leave a wide interval; this study does not set a minimum sample size.
Limits
- Protocol timing conflicts. The current protocol file was born at 01:47 UTC, after the first counted call at 00:43 UTC on 2026-10-07. Amendment 1 says 00:44 UTC and before counted calls. Amendment 4 says 06:10 UTC and before the Claude probe, but its receipt and counted calls began at 06:04 UTC. An earlier version may have existed; these files cannot verify pre-call registration. Treat the calls as descriptive evidence.
- Repeated fixed prompts. Wilson intervals and the paired p-value are calculations that assume independent outcomes or pairs. Repeating three prompts does not measure accuracy across unseen tasks.
- Known failure case. The hand-made order prompt reuses a case with known Haiku format misses. It is not a blind holdout.
- Shared host. Both routes used one Mac. The files do not establish isolation from other work. Timings include CLI start-up and network time.
- Two changes together. Schema mode adds the schema flag and removes the format sentence. This test cannot isolate either change.
- Small samples. n is 24 calls per mode for Haiku and 12 for each other model. The intervals are not minimum detectable differences.
- One wording. The instruction wording is one sentence. Another wording, or a schema with the format sentence kept, was not tested. The format-miss counts here are not comparable with the consistency study, which used a different wording on one of these prompts.
- CLI flags, not the API. We tested
claude --json-schemaandcodex exec --output-schema. We did not test a provider's API structured-output feature. - Codex does not echo the model or effort. The call named GPT-6.1 Sol at low effort, and the CLI's catalog lists it, but a reroute would not have been visible.
- Three moderate prompts. Harder or longer extraction may behave differently.
Prompts and replies are not published. The public data holds counts, times, tokens and the name of each failed check.
Disclosure
I build Agent, the product behind this site. No Agent worker pipeline ran in this study. Every row is a model through its vendor's CLI, so the numbers say nothing about Agent's own pipeline. Every counted inference call is in the data. The files cannot verify that the protocol was written before the first counted call.
Run it on your own work
Agent keeps the same receipts for your tasks: the model, the route, the tokens, the time, the cost and the validation result. Try Agent and see your own numbers.
What to read next
- Same prompt, ten answers: how consistent are Claude and Codex?
- Caching and consistency study
- Best LLM for JSON output? Haiku, Sonnet and GPT-6.1 Sol, 10 runs each
- Format misses vs wrong answers: your eval may fail right answers
- Haiku vs Sonnet, every measured row
- Claude Code vs Codex CLI
- How a Wilson interval works