• Json
  • Structured Output
  • Claude Haiku
  • Claude Sonnet

Best LLM for JSON output? Haiku, Sonnet and GPT-6.1 Sol, 10 runs each

Best LLM for JSON output? On one prompt, Sonnet and GPT-6.1 Sol passed 10/10 (95%: 72–100%); Haiku passed 1/10 (2–40%). CLI results, not JSON mode.

TL;DR

  • One JSON prompt, 10 runs each, strict parser: Claude Sonnet 5.5 passed 10/10 and GPT-6.1 Sol (medium, Codex CLI) passed 10/10 (95% interval 72% to 100%). Claude Haiku 4.5 passed 1/10 (2% to 40%).
  • Haiku had 9 format misses and 0 wrong answers. The JSON was right. The reply shape was not.
  • Haiku passed a different JSON task. All 26 invoice calls passed (95% interval 87% to 100%; pooled calculation). Haiku passed 3/3 (44% to 100%). That study removed a wrapping code fence before it checked, so the two results are not like for like.
  • Speed does not rank them. Medians and run ranges, n = 10 each: Sonnet 2.89 s (2.68–5.30), GPT-6.1 Sol 6.42 s (5.25–8.26), Haiku 7.03 s (5.28–12.27). These ranges overlap; they are not confidence intervals.
  • Our advice: parse leniently, validate against a schema, and test your own prompt 10 times.
Live story · 49 sPrompt caching and consistency: what the cache saves, and how much answers vary

Prompt caching and consistency: what the cache saves, and how much answers vary

Calculation at list price: the cache cut a 5-question session 50% on Sonnet 5.5 and 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every repetition.

Transcript
  1. Caching and consistency · 135 calls. What the cache saves, and how much answers vary. Five-question sessions on a fixed context. Then the same prompt, 10 times.
  2. 135 calls: 45 cache turns, 90 repeated prompts. Every call counted. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
  3. Claude Code turn 1 reads 19% from the cache, the CLI’s own prefix. Turns 2–5 read 90%–99%. Codex CLI: 98%–99%. Chart: Share of input read from the cache · mean of 3 sessions per turn (n = 3 each). Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  4. A calculation: Sonnet 5.5 $0.1350 with the cache vs $0.2698 without, 50% less. Opus 5.5 $0.2551 vs $0.5442, 53% less. Chart: List-price cost of the recorded sessions · calculation (n = 15 each). Calculation, not a run. Caveat: Costs are list-price calculations; the calls used flat subscriptions.
  5. No clear speed effect: Sonnet 5.5 1.6 s on turn 1 vs 1.6 s later; Opus 5.5 1.9 s vs 2.4 s. The ranges overlap. Sonnet 5.5: median turn 1 vs turns 2–5: 1.6 s vs 1.6 s (ranges 1.6–1.8 s and 1.4–5.6 s · n = 3 and 12). Opus 5.5: median turn 1 vs turns 2–5: 1.9 s vs 2.4 s (ranges 1.8–4.4 s and 1.6–12.7 s · n = 3 and 12). Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
  6. 7 of 9 cells passed 10/10. Haiku 4.5: 0/10 on the exact number (10 wrong, 1 distinct answer), 1/10 on JSON (9 format misses). Chart: Same prompt, 10 times · strict passes · a 10/10 is 72%–100% at 95% (n = 10 each). Caveat: Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
  7. Consistent is not correct: Haiku 4.5 gave the same wrong number all 10 times. Code fix, distinct correct bodies: Haiku 4.5 6, Sonnet 5.5 3, GPT-6.1 Sol (medium) 6. Chart: Same prompt, 10 times · distinct answers (n = 10 each). Caveat: The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
  8. Cache the fixed context; check answers, not agreement. Every call online.

The answer: two models passed every run

The prompt asked each model to turn an order note into one JSON object with exact keys. The note had merged lines and a cancelled line. A strict pass means the whole reply, trimmed, parses as JSON and passes the validator. A format miss means the strict check failed, but a lenient extractor found an answer that passes the same validator. We never count a format miss as a pass.

ConfigurationStrict passes95% intervalFormat missesWrong answers
Claude Sonnet 5.5 (Claude Code)10/1072% to 100%00
GPT-6.1 Sol, medium (Codex CLI)10/1072% to 100%00
Claude Haiku 4.5 (Claude Code)1/102% to 40%90
  • Exact number
  • JSON object
  • Code fix
Claude Haiku 4.5
Claude Sonnet 5.5
GPT-6.1 Sol (medium)

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 10 per row

One series per prompt; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

The Haiku interval does not overlap the other two. So the Haiku vs Sonnet and Haiku vs GPT-6.1 Sol pages name the other model ahead on this row. The Sonnet vs GPT-6.1 Sol page marks their row as a tie. Their intervals overlap; this does not prove equal reliability. This prompt has a ceiling for them: it cannot separate them.

Haiku's JSON was right in all 10 replies, and all 10 reduced to one distinct answer. Only the shape changed. Haiku gave 3 distinct raw replies in 10 runs; Sonnet and GPT-6.1 Sol gave 1 each. In all 9 misses, the run receipts show the parser stopped at a code fence that opened the reply.

Exact number

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

JSON object

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Code fix

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

One panel per series, all on the same axis.

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: all at 1. JSON object: all at 1.

Notesn = 10 per row

Distinct normalized answers over 10 repetitions (1 = the same answer every time)

Normalized: the last number for the exact-number prompt; key-sorted JSON for the JSON prompt; for the code fix, the code without fences, comments, whitespace, quote style and line-end semicolons. One distinct answer is not the same as a correct answer: an answer can be the same and wrong every time. Different code can be equally correct.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

The invoice task: every call passed

The five-task head-to-head had one invoice task, "Extract invoice fields to JSON". Every configuration passed it: 26 of 26 calls (87% to 100%; calculation, the column sum). Haiku passed 3 of 3 (44% to 100%).

So one model can pass one JSON task and miss another. The five-task study removed one wrapping fence before it checked; the other two studies did not. The raw receipts show a wrapping fence in all 3 Haiku invoice replies. The validator removed it before parsing. We did not test which prompt wording causes a fence.

The hard set has one JSON task, a room schedule. Haiku passed 0/3 strictly (0% to 56%, calculation): 2 format misses and 1 wrong answer. The other six configurations passed 16 of 16 calls (81% to 100%, a pooled calculation). With n = 3 for Haiku, the interval is wide. The other configurations reached a ceiling on this task; these calls cannot separate them.

Over all 8 hard tasks, Haiku had 5 format misses. Two were in the JSON room schedule, 2 in the SQL task and 1 in the CSV parser task. The other six configurations had 0.

The lenient extractor read each of the 5 from a fenced block. The room-schedule prompt asked for only one JSON object and "No other text". It did not explicitly ban a code fence. The SQL and CSV-parser prompts did ban fences. Haiku returned fences on all three tasks. See When tasks get hard.

How fast is a JSON reply?

  • Exact number
  • JSON object
  • Code fix
Entrance: medians race at 9.6× real timeMotion reduced: press Replay to animateThe slowest median is 13.4 s. The clock runs at the recorded speed.
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: slowest GPT-6.1 Sol (medium) · Codex CLI 13.4 s (range 12.3 s–18 s, n 10). Fastest Claude Haiku 4.5 · Claude Code 5.1 s (range 4.4 s–6.2 s, n 10). Not all run ranges overlap. JSON object: slowest Claude Haiku 4.5 · Claude Code 7 s (range 5.3 s–12.3 s, n 10). Fastest Claude Sonnet 5.5 · Claude Code 2.9 s (range 2.7 s–5.3 s, n 10). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 10 per row

Median; whiskers = fastest and slowest of 10 calls

Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Median time per call, with the fastest and slowest of 10 (ranges, not confidence intervals): Sonnet 2.89 s (2.68 to 5.30), GPT-6.1 Sol 6.42 s (5.25 to 8.26), Haiku 7.03 s (5.28 to 12.27).

The comparison pages mark all three JSON time rows unclear, because the run ranges overlap. Haiku and Sonnet overlap slightly at the ends of their ranges. We name no fastest model. The Codex CLI times include its start-up and a larger system prompt, so they compare a route and a model together.

A one-retry projection adds 7.03 s if the retry takes Haiku's observed median (calculation, n = 10; range 5.28–12.27 s). We did not run retries or measure parser time. A lenient parser can avoid another model call when the extracted answer passes validation.

What to do about format misses

These are our recommendations. We did not test a parsing library or a vendor feature.

  1. Parse leniently, then validate. Strip one wrapping code fence, parse, then check keys and types against a schema. A lenient extractor found a passing answer in all 14 Haiku format misses across the consistency and hard studies. Of these, 11 were in JSON tasks (9 here, 2 on the hard set) and 3 in SQL or JavaScript tasks (calculation).
  2. Never skip validation. A lenient parse fixes the shape of the reply, not its content. Haiku also gave 8 wrong answers on the hard set. Count its format misses as passes and it reaches 16/24 (67%, 95% interval 47% to 82%; calculation). Sonnet's 24/24 (86% to 100%) is still ahead, because the intervals do not overlap.
  3. Test the format on 10 repeats, not one run. One run shows a pass or a miss. Ten runs show a rate with an interval: Haiku's was 1 in 10 (2% to 40%).
  4. Pick the model by strict pass rate on your own prompt. Keep the route in the result: Claude Code and Codex CLI each add a system prompt.
  5. Do not rely on "no code fence" in the prompt alone. Haiku returned fences despite explicit bans in the hard set's SQL and CSV-parser prompts. We did not test other wording.

How we measured

  • Design: the study ran 3 prompts, 3 configurations and 10 repeats, 90 calls. The main table uses the 30 JSON calls. A deterministic validator decides pass or fail.
  • Controls: before the first call on each route, every reference answer passed, every plausible wrong answer failed, and a wrapped reference was a format miss.
  • Isolation: fresh empty folder, tools off, no session persistence, one call at a time. Claude Code ran at default effort, GPT-6.1 Sol at medium. We kept every call and retried none.
  • Diversity counts distinct normalized answers. The public data holds ordinal ids, never reply text.

Studies: caching and consistency, five-task head-to-head and hard head-to-head.

Caveats

  • n = 10 per cell. A 10/10 has a 95% interval of 72% to 100%.
  • One JSON prompt. A longer or nested object can give another result.
  • Route and model pairs. Rows compare each CLI with its model, not the models alone.
  • Strict validator. It rejects a fence even when the JSON is right; the five-task set removed one, so its rows are not like for like.
  • Small cells. The invoice task has 2 or 3 calls per configuration; the room schedule has 3 for Haiku.
  • Hard-set exclusions: its scored totals exclude 30 blocked, unscored Codex attempts. The public receipts retain them. The follow-up Codex batch supplies the scored results; this is not a first-attempt success rate.
  • Timing of the written protocol: the consistency protocol files were created after each route's first call. Control receipts predate the calls, but we do not claim a protocol written before inference.
  • Scope: these are repeated calls on synthetic prompts, from one host and network on 2026-10-05 and 2026-10-06. Pooled intervals summarize calls across configurations, not independent task samples.
  • Not covered by this post's data: vendor structured-output or JSON-mode features, schema-constrained decoding and other models.

Test your own JSON prompt

Measured figures above come from per-call receipts: model, route, tokens, time and validation result. We label pooled figures and projections as calculations. Agent leaves a receipt for each task in the same way. Try Agent on your own prompt.

The data behind this post

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.