Best LLM for JSON output? Haiku, Sonnet and GPT-6.1 Sol, 10 runs each
Best LLM for JSON output? On one prompt, Sonnet and GPT-6.1 Sol passed 10/10 (95%: 72–100%); Haiku passed 1/10 (2–40%). CLI results, not JSON mode.
TL;DR
- One JSON prompt, 10 runs each, strict parser: Claude Sonnet 5.5 passed 10/10 and GPT-6.1 Sol (medium, Codex CLI) passed 10/10 (95% interval 72% to 100%). Claude Haiku 4.5 passed 1/10 (2% to 40%).
- Haiku had 9 format misses and 0 wrong answers. The JSON was right. The reply shape was not.
- Haiku passed a different JSON task. All 26 invoice calls passed (95% interval 87% to 100%; pooled calculation). Haiku passed 3/3 (44% to 100%). That study removed a wrapping code fence before it checked, so the two results are not like for like.
- Speed does not rank them. Medians and run ranges, n = 10 each: Sonnet 2.89 s (2.68–5.30), GPT-6.1 Sol 6.42 s (5.25–8.26), Haiku 7.03 s (5.28–12.27). These ranges overlap; they are not confidence intervals.
- Our advice: parse leniently, validate against a schema, and test your own prompt 10 times.
Prompt caching and consistency: what the cache saves, and how much answers vary
Calculation at list price: the cache cut a 5-question session 50% on Sonnet 5.5 and 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every repetition.
Transcript
- Caching and consistency · 135 calls. What the cache saves, and how much answers vary. Five-question sessions on a fixed context. Then the same prompt, 10 times.
- 135 calls: 45 cache turns, 90 repeated prompts. Every call counted. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- Claude Code turn 1 reads 19% from the cache, the CLI’s own prefix. Turns 2–5 read 90%–99%. Codex CLI: 98%–99%. Chart: Share of input read from the cache · mean of 3 sessions per turn (n = 3 each). Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- A calculation: Sonnet 5.5 $0.1350 with the cache vs $0.2698 without, 50% less. Opus 5.5 $0.2551 vs $0.5442, 53% less. Chart: List-price cost of the recorded sessions · calculation (n = 15 each). Calculation, not a run. Caveat: Costs are list-price calculations; the calls used flat subscriptions.
- No clear speed effect: Sonnet 5.5 1.6 s on turn 1 vs 1.6 s later; Opus 5.5 1.9 s vs 2.4 s. The ranges overlap. Sonnet 5.5: median turn 1 vs turns 2–5: 1.6 s vs 1.6 s (ranges 1.6–1.8 s and 1.4–5.6 s · n = 3 and 12). Opus 5.5: median turn 1 vs turns 2–5: 1.9 s vs 2.4 s (ranges 1.8–4.4 s and 1.6–12.7 s · n = 3 and 12). Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- 7 of 9 cells passed 10/10. Haiku 4.5: 0/10 on the exact number (10 wrong, 1 distinct answer), 1/10 on JSON (9 format misses). Chart: Same prompt, 10 times · strict passes · a 10/10 is 72%–100% at 95% (n = 10 each). Caveat: Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
- Consistent is not correct: Haiku 4.5 gave the same wrong number all 10 times. Code fix, distinct correct bodies: Haiku 4.5 6, Sonnet 5.5 3, GPT-6.1 Sol (medium) 6. Chart: Same prompt, 10 times · distinct answers (n = 10 each). Caveat: The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
- Cache the fixed context; check answers, not agreement. Every call online.
The answer: two models passed every run
The prompt asked each model to turn an order note into one JSON object with exact keys. The note had merged lines and a cancelled line. A strict pass means the whole reply, trimmed, parses as JSON and passes the validator. A format miss means the strict check failed, but a lenient extractor found an answer that passes the same validator. We never count a format miss as a pass.
| Configuration | Strict passes | 95% interval | Format misses | Wrong answers |
|---|---|---|---|---|
| Claude Sonnet 5.5 (Claude Code) | 10/10 | 72% to 100% | 0 | 0 |
| GPT-6.1 Sol, medium (Codex CLI) | 10/10 | 72% to 100% | 0 | 0 |
| Claude Haiku 4.5 (Claude Code) | 1/10 | 2% to 40% | 9 | 0 |
The Haiku interval does not overlap the other two. So the Haiku vs Sonnet and Haiku vs GPT-6.1 Sol pages name the other model ahead on this row. The Sonnet vs GPT-6.1 Sol page marks their row as a tie. Their intervals overlap; this does not prove equal reliability. This prompt has a ceiling for them: it cannot separate them.
Haiku's JSON was right in all 10 replies, and all 10 reduced to one distinct answer. Only the shape changed. Haiku gave 3 distinct raw replies in 10 runs; Sonnet and GPT-6.1 Sol gave 1 each. In all 9 misses, the run receipts show the parser stopped at a code fence that opened the reply.
The invoice task: every call passed
The five-task head-to-head had one invoice task, "Extract invoice fields to JSON". Every configuration passed it: 26 of 26 calls (87% to 100%; calculation, the column sum). Haiku passed 3 of 3 (44% to 100%).
So one model can pass one JSON task and miss another. The five-task study removed one wrapping fence before it checked; the other two studies did not. The raw receipts show a wrapping fence in all 3 Haiku invoice replies. The validator removed it before parsing. We did not test which prompt wording causes a fence.
The hard set has one JSON task, a room schedule. Haiku passed 0/3 strictly (0% to 56%, calculation): 2 format misses and 1 wrong answer. The other six configurations passed 16 of 16 calls (81% to 100%, a pooled calculation). With n = 3 for Haiku, the interval is wide. The other configurations reached a ceiling on this task; these calls cannot separate them.
Over all 8 hard tasks, Haiku had 5 format misses. Two were in the JSON room schedule, 2 in the SQL task and 1 in the CSV parser task. The other six configurations had 0.
The lenient extractor read each of the 5 from a fenced block. The room-schedule prompt asked for only one JSON object and "No other text". It did not explicitly ban a code fence. The SQL and CSV-parser prompts did ban fences. Haiku returned fences on all three tasks. See When tasks get hard.
How fast is a JSON reply?
Median time per call, with the fastest and slowest of 10 (ranges, not confidence intervals): Sonnet 2.89 s (2.68 to 5.30), GPT-6.1 Sol 6.42 s (5.25 to 8.26), Haiku 7.03 s (5.28 to 12.27).
The comparison pages mark all three JSON time rows unclear, because the run ranges overlap. Haiku and Sonnet overlap slightly at the ends of their ranges. We name no fastest model. The Codex CLI times include its start-up and a larger system prompt, so they compare a route and a model together.
A one-retry projection adds 7.03 s if the retry takes Haiku's observed median (calculation, n = 10; range 5.28–12.27 s). We did not run retries or measure parser time. A lenient parser can avoid another model call when the extracted answer passes validation.
What to do about format misses
These are our recommendations. We did not test a parsing library or a vendor feature.
- Parse leniently, then validate. Strip one wrapping code fence, parse, then check keys and types against a schema. A lenient extractor found a passing answer in all 14 Haiku format misses across the consistency and hard studies. Of these, 11 were in JSON tasks (9 here, 2 on the hard set) and 3 in SQL or JavaScript tasks (calculation).
- Never skip validation. A lenient parse fixes the shape of the reply, not its content. Haiku also gave 8 wrong answers on the hard set. Count its format misses as passes and it reaches 16/24 (67%, 95% interval 47% to 82%; calculation). Sonnet's 24/24 (86% to 100%) is still ahead, because the intervals do not overlap.
- Test the format on 10 repeats, not one run. One run shows a pass or a miss. Ten runs show a rate with an interval: Haiku's was 1 in 10 (2% to 40%).
- Pick the model by strict pass rate on your own prompt. Keep the route in the result: Claude Code and Codex CLI each add a system prompt.
- Do not rely on "no code fence" in the prompt alone. Haiku returned fences despite explicit bans in the hard set's SQL and CSV-parser prompts. We did not test other wording.
How we measured
- Design: the study ran 3 prompts, 3 configurations and 10 repeats, 90 calls. The main table uses the 30 JSON calls. A deterministic validator decides pass or fail.
- Controls: before the first call on each route, every reference answer passed, every plausible wrong answer failed, and a wrapped reference was a format miss.
- Isolation: fresh empty folder, tools off, no session persistence, one call at a time. Claude Code ran at default effort, GPT-6.1 Sol at medium. We kept every call and retried none.
- Diversity counts distinct normalized answers. The public data holds ordinal ids, never reply text.
Studies: caching and consistency, five-task head-to-head and hard head-to-head.
Caveats
- n = 10 per cell. A 10/10 has a 95% interval of 72% to 100%.
- One JSON prompt. A longer or nested object can give another result.
- Route and model pairs. Rows compare each CLI with its model, not the models alone.
- Strict validator. It rejects a fence even when the JSON is right; the five-task set removed one, so its rows are not like for like.
- Small cells. The invoice task has 2 or 3 calls per configuration; the room schedule has 3 for Haiku.
- Hard-set exclusions: its scored totals exclude 30 blocked, unscored Codex attempts. The public receipts retain them. The follow-up Codex batch supplies the scored results; this is not a first-attempt success rate.
- Timing of the written protocol: the consistency protocol files were created after each route's first call. Control receipts predate the calls, but we do not claim a protocol written before inference.
- Scope: these are repeated calls on synthetic prompts, from one host and network on 2026-10-05 and 2026-10-06. Pooled intervals summarize calls across configurations, not independent task samples.
- Not covered by this post's data: vendor structured-output or JSON-mode features, schema-constrained decoding and other models.
What to read next
- Same prompt, ten answers: how consistent are Claude and Codex?
- Claude vs Codex on hard tasks
- Why we count every failed attempt
Test your own JSON prompt
Measured figures above come from per-call receipts: model, route, tokens, time and validation result. We label pooled figures and projections as calculations. Agent leaves a receipt for each task in the same way. Try Agent on your own prompt.