Format misses vs wrong answers: your LLM eval may be failing right answers
5 of 13 failed calls on our hard set were right answers in the wrong format. How to grade LLM output strictly, test validators first and report both numbers.
TL;DR
- Strict format grading can fail right answers. On our hard set, 139 of 152 calls passed strictly (91%, 95% interval 85.9% to 94.9%). Read leniently, 144 of 152 had a correct answer (95%, 90.0% to 97.3%).
- 5 of the 13 non-passes were format misses: right answer, wrong shape. The other 8 were wrong answers. All 13 came from one configuration, Claude Haiku 4.5.
- The reading can flip a verdict. On a repeated JSON prompt, Haiku scored 1/10 strictly (2% to 40%) and 10/10 leniently (72% to 100%, a calculation). Against Sonnet's 10/10, strict puts Sonnet ahead under our interval rule. Lenient does not separate them; it does not prove equality.
- Test the validator before the first model call. Ours passed 8 of 8 references and rejected 26 of 26 plausible wrong answers.
- Our rule: if format is part of the task, grade strictly and count misses apart. If not, report both numbers.
Strict vs lenient grading: one run, two numbers
One validator, two readings of the same calls:
- Strict: the whole reply, trimmed, passes as given.
- Lenient: also counts a format miss. The strict check failed, but an extractor finds an answer in the reply that passes the same validator.
Each scored call has one of three outcomes: strict pass, format miss or wrong answer. Here, "wrong answer" means neither the strict check nor the extractor found a valid answer.
The two rates differ by 5 calls, or 3.3 points (a calculation). These rates use the same calls. Interval overlap is not a test of their paired difference; we make no significance claim.
By configuration, Haiku 4.5 passed 11 of 24 strictly (46%, 28% to 65%). It gave a correct answer in 16 of 24 (67%, 47% to 82%). Claude Sonnet 5.5 passed 24 of 24 (86% to 100%). Both Haiku intervals sit below Sonnet's. Sonnet stays ahead under our interval rule even after the extractor recovers Haiku's format misses (Haiku vs Sonnet).
What an LLM output format error looks like
A format miss is a right answer in the wrong shape. We saw two forms.
A wrapper. Every hard-set prompt said no code fence and no other text. Our extractor looks for a fenced block, an outer JSON value, a first code line or a single output line.
Extra lines. On the five-task set, Sonnet's 3 non-passes were all on one task. Each ended on the expected answer but added working lines. The exact-text validator rejects that by design. Those 3 were the only non-passes in 130 calls.
Read strictly, Sonnet has the lowest rate on that set: 12 of 15 (80%, 55% to 93%). The grading records classify all 3 as format misses. Its interval overlaps the 15/15 configurations (80% to 100%). It also overlaps the 10/10 configuration (72% to 100%). The strict rates do not separate them under our interval rule; overlap does not prove equality.
Small cells hide the same effect. Haiku's room-schedule and SQL cells each went from 0/3 strict (95% interval 0% to 56%) to 2/3 lenient (21% to 94%). These lenient rates are calculations. The hard-tasks post has the per-task detail.
Test the validator first
A grader can pass a wrong answer or fail a right one. On the hard set, we ran three controls on every task before any model call:
- The reference answer must pass. 8 of 8 did.
- Plausible wrong answers must fail. 26 of 26 did.
- A reference wrapped in a fence or in prose must count as a format miss. 8 of 8 did.
Control 3 checks both sides: strict fails the wrapped answer, and the extractor still finds it. Without controls, a 100% can come from a validator that accepts everything.
Decide whether format is part of the task
Ask first: does the reader of the reply care about its shape?
- Yes. A program parses the reply (JSON, a regex), and a code fence breaks it. Keep strict as the score. Count format misses next to it.
- No. A person reads the reply, or you strip wrappers before use. Report both numbers and say which one leads.
Our repeated-prompt test shows why. Haiku ran each prompt 10 times. On the JSON prompt it had 1 strict pass, 9 format misses and 0 wrong answers. That is 1/10 strict (2% to 40%) and 10/10 lenient (72% to 100%, a calculation). Sonnet passed 10/10 (72% to 100%). Strict puts Sonnet ahead under our interval rule. Lenient reaches a ceiling at 10/10 for both; it cannot separate them or prove equality.
On the exact-number prompt, Haiku had 0 strict passes, 0 format misses and 10 wrong answers (0/10, 0% to 28%). That cell is an answer problem, not a format problem.
Report the zeros too. In the effort ladder, 0 of 176 calls were format misses or wrong answers (176/176 strict, 95% interval 97.9% to 100%). All 11 configurations hit the ceiling: 16/16 each (81% to 100%).
A checklist for your LLM eval
Our recommendations:
- Test the validator first, with the three controls above.
- State the format in the prompt. Say what the reply must not contain.
- Pick the headline before the first call. Write "strict" or "lenient" in the protocol.
- Report three counts: strict passes, format misses and wrong answers.
- Use one validator for both readings. The extractor changes where you look, not what counts as right.
- Show n and a 95% interval for both rates. Say the rates do not separate when intervals overlap. Do not claim equality.
- Split by configuration. One configuration held all 13 of our non-passes. The pooled 91% hid it.
- Report zeros. "0 format misses" shows that the format rule held.
How we measured
- Hard set: 8 tasks, each with a deterministic validator in a sandbox without network. 152 calls reached a model: 24 per Claude Code configuration and 16 per Codex CLI configuration. 30 earlier Codex tries never reached a model. We do not score them. One successful Codex probe also sits outside the scored matrix.
- Other sets: five-task head-to-head, 130 calls. Repeated prompts, 90 calls. Effort ladder, 176 calls (80 reused from the hard set). The scored matrices kept every recorded call, with no recorded call repeated or replaced. The hard-set Codex route ran a later batch after the blocked attempts. One interrupted batch resumed by skipping recorded calls.
- Grading: strict and lenient as defined above, with 95% Wilson intervals. The five-task study removed one wrapping fence before validation; its strict score differs from the hard-set rule.
- Timing disclosure: control receipts place the hard-set checks before the first inference call. Available protocol file creation times for the hard, consistency and effort sets fall after their first calls. These files do not verify a claim that the protocol was written before inference.
- Calculations: the 3.3-point gap, Haiku's lenient JSON rate (1 + 9 of 10), its lenient room-schedule and SQL rates and intervals, and each comparison under the interval rule.
- Sources: the public dataset, hard-set receipts, five-task receipts, consistency receipts and effort receipts. We quote no model output.
Caveats
- Small counts. 13 non-passes, 5 format misses and 2 to 3 calls per task cell. They describe these runs only.
- One extractor, same calls. A different extractor could recover other replies. The two rates use the same calls, so we make no significance claim.
- A ceiling. 6 of 7 hard-set configurations passed every call (24/24: 86% to 100%; 16/16: 81% to 100%).
- Different strictness. The five-task validator removed one wrapping fence before it checked. The hard-set validator did not. The tasks differ too, so we cannot say why Sonnet had misses on one set only.
- JSON rule. Our JSON validator fails a fenced reply even when the JSON is right. The table counts such replies as format misses, not wrong answers.
What to read next
- Hard tasks: Claude Haiku vs Sonnet vs Opus vs Fable
- Same prompt, ten answers
- How to read AI benchmarks honestly
- Studies: hard, five-task, consistency, effort.
Grade the shape, then the answer
Apply these checks to your own tasks. Try Agent on your own tasks.