• Evaluation
  • Format Misses
  • LLM Evals
  • Validators

Format misses vs wrong answers: your LLM eval may be failing right answers

5 of 13 failed calls on our hard set were right answers in the wrong format. How to grade LLM output strictly, test validators first and report both numbers.

TL;DR

  • Strict format grading can fail right answers. On our hard set, 139 of 152 calls passed strictly (91%, 95% interval 85.9% to 94.9%). Read leniently, 144 of 152 had a correct answer (95%, 90.0% to 97.3%).
  • 5 of the 13 non-passes were format misses: right answer, wrong shape. The other 8 were wrong answers. All 13 came from one configuration, Claude Haiku 4.5.
  • The reading can flip a verdict. On a repeated JSON prompt, Haiku scored 1/10 strictly (2% to 40%) and 10/10 leniently (72% to 100%, a calculation). Against Sonnet's 10/10, strict puts Sonnet ahead under our interval rule. Lenient does not separate them; it does not prove equality.
  • Test the validator before the first model call. Ours passed 8 of 8 references and rejected 26 of 26 plausible wrong answers.
  • Our rule: if format is part of the task, grade strictly and count misses apart. If not, report both numbers.

Strict vs lenient grading: one run, two numbers

One validator, two readings of the same calls:

  • Strict: the whole reply, trimmed, passes as given.
  • Lenient: also counts a format miss. The strict check failed, but an extractor finds an answer in the reply that passes the same validator.

Each scored call has one of three outcomes: strict pass, format miss or wrong answer. Here, "wrong answer" means neither the strict check nor the extractor found a valid answer.

  • Strict pass
  • Format miss (correct answer, wrong format)
  • Wrong answer
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

One square per call; counts at the right are exact and in legend order.

7 rows, 3 series: Strict pass, Format miss (correct answer, wrong format), Wrong answer. Strict pass: highest Claude Sonnet 5.5 · Claude Code 24 (n 24). Lowest Claude Haiku 4.5 · Claude Code 11 (n 24). Format miss (correct answer, wrong format): highest Claude Haiku 4.5 · Claude Code 5 (n 24). Lowest GPT-6.1 Sol (high) · Codex CLI 0 (n 16).

Notesn 16–24 per row

Counts per configuration: strict passes, format misses and wrong answers

A format miss is a reply that the strict validator rejected (for example, wrapped in a code fence or in prose although the prompt said not to) whose extracted answer passes the same validator. It is not a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

The two rates differ by 5 calls, or 3.3 points (a calculation). These rates use the same calls. Interval overlap is not a test of their paired difference; we make no significance claim.

By configuration, Haiku 4.5 passed 11 of 24 strictly (46%, 28% to 65%). It gave a correct answer in 16 of 24 (67%, 47% to 82%). Claude Sonnet 5.5 passed 24 of 24 (86% to 100%). Both Haiku intervals sit below Sonnet's. Sonnet stays ahead under our interval rule even after the extractor recovers Haiku's format misses (Haiku vs Sonnet).

What an LLM output format error looks like

A format miss is a right answer in the wrong shape. We saw two forms.

A wrapper. Every hard-set prompt said no code fence and no other text. Our extractor looks for a fenced block, an outer JSON value, a first code line or a single output line.

Extra lines. On the five-task set, Sonnet's 3 non-passes were all on one task. Each ended on the expected answer but added working lines. The exact-text validator rejects that by design. Those 3 were the only non-passes in 130 calls.

Claude Fable 5.1
Claude Sonnet 5.5
Claude Opus 5.5 (high)
Claude Opus 5.5
Claude Opus 5.5 (low)
Claude Haiku 4.5
GPT-6.1 Sol (high)
GPT-6.1 Sol (medium)
GPT-6.1 Sol (low)

Every interval overlaps every other: this chart does not order these rows.

9 rows. Highest Claude Fable 5.1 · Claude Code 100% (95% interval 80%–100%, n 15). Lowest Claude Sonnet 5.5 · Claude Code 80% (95% interval 55%–93%, n 15). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln 10–15 per row8 of 9 at 100%: this task set cannot separate them.

Every call counts; failures and timeouts are non-passes

Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.

Source: Provider head-to-head: Claude Code models vs Codex efforts

Read strictly, Sonnet has the lowest rate on that set: 12 of 15 (80%, 55% to 93%). The grading records classify all 3 as format misses. Its interval overlaps the 15/15 configurations (80% to 100%). It also overlaps the 10/10 configuration (72% to 100%). The strict rates do not separate them under our interval rule; overlap does not prove equality.

Small cells hide the same effect. Haiku's room-schedule and SQL cells each went from 0/3 strict (95% interval 0% to 56%) to 2/3 lenient (21% to 94%). These lenient rates are calculations. The hard-tasks post has the per-task detail.

Test the validator first

A grader can pass a wrong answer or fail a right one. On the hard set, we ran three controls on every task before any model call:

  1. The reference answer must pass. 8 of 8 did.
  2. Plausible wrong answers must fail. 26 of 26 did.
  3. A reference wrapped in a fence or in prose must count as a format miss. 8 of 8 did.

Control 3 checks both sides: strict fails the wrapped answer, and the extractor still finds it. Without controls, a 100% can come from a validator that accepts everything.

Decide whether format is part of the task

Ask first: does the reader of the reply care about its shape?

  • Yes. A program parses the reply (JSON, a regex), and a code fence breaks it. Keep strict as the score. Count format misses next to it.
  • No. A person reads the reply, or you strip wrappers before use. Report both numbers and say which one leads.

Our repeated-prompt test shows why. Haiku ran each prompt 10 times. On the JSON prompt it had 1 strict pass, 9 format misses and 0 wrong answers. That is 1/10 strict (2% to 40%) and 10/10 lenient (72% to 100%, a calculation). Sonnet passed 10/10 (72% to 100%). Strict puts Sonnet ahead under our interval rule. Lenient reaches a ceiling at 10/10 for both; it cannot separate them or prove equality.

On the exact-number prompt, Haiku had 0 strict passes, 0 format misses and 10 wrong answers (0/10, 0% to 28%). That cell is an answer problem, not a format problem.

  • Exact number
  • JSON object
  • Code fix
Claude Haiku 4.5
Claude Sonnet 5.5
GPT-6.1 Sol (medium)

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 10 per row

One series per prompt; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Report the zeros too. In the effort ladder, 0 of 176 calls were format misses or wrong answers (176/176 strict, 95% interval 97.9% to 100%). All 11 configurations hit the ceiling: 16/16 each (81% to 100%).

A checklist for your LLM eval

Our recommendations:

  1. Test the validator first, with the three controls above.
  2. State the format in the prompt. Say what the reply must not contain.
  3. Pick the headline before the first call. Write "strict" or "lenient" in the protocol.
  4. Report three counts: strict passes, format misses and wrong answers.
  5. Use one validator for both readings. The extractor changes where you look, not what counts as right.
  6. Show n and a 95% interval for both rates. Say the rates do not separate when intervals overlap. Do not claim equality.
  7. Split by configuration. One configuration held all 13 of our non-passes. The pooled 91% hid it.
  8. Report zeros. "0 format misses" shows that the format rule held.

How we measured

  • Hard set: 8 tasks, each with a deterministic validator in a sandbox without network. 152 calls reached a model: 24 per Claude Code configuration and 16 per Codex CLI configuration. 30 earlier Codex tries never reached a model. We do not score them. One successful Codex probe also sits outside the scored matrix.
  • Other sets: five-task head-to-head, 130 calls. Repeated prompts, 90 calls. Effort ladder, 176 calls (80 reused from the hard set). The scored matrices kept every recorded call, with no recorded call repeated or replaced. The hard-set Codex route ran a later batch after the blocked attempts. One interrupted batch resumed by skipping recorded calls.
  • Grading: strict and lenient as defined above, with 95% Wilson intervals. The five-task study removed one wrapping fence before validation; its strict score differs from the hard-set rule.
  • Timing disclosure: control receipts place the hard-set checks before the first inference call. Available protocol file creation times for the hard, consistency and effort sets fall after their first calls. These files do not verify a claim that the protocol was written before inference.
  • Calculations: the 3.3-point gap, Haiku's lenient JSON rate (1 + 9 of 10), its lenient room-schedule and SQL rates and intervals, and each comparison under the interval rule.
  • Sources: the public dataset, hard-set receipts, five-task receipts, consistency receipts and effort receipts. We quote no model output.

Caveats

  • Small counts. 13 non-passes, 5 format misses and 2 to 3 calls per task cell. They describe these runs only.
  • One extractor, same calls. A different extractor could recover other replies. The two rates use the same calls, so we make no significance claim.
  • A ceiling. 6 of 7 hard-set configurations passed every call (24/24: 86% to 100%; 16/16: 81% to 100%).
  • Different strictness. The five-task validator removed one wrapping fence before it checked. The hard-set validator did not. The tasks differ too, so we cannot say why Sonnet had misses on one set only.
  • JSON rule. Our JSON validator fails a fenced reply even when the JSON is right. The table counts such replies as format misses, not wrong answers.

Grade the shape, then the answer

Apply these checks to your own tasks. Try Agent on your own tasks.

The data behind this post

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.