Explainer · Strict grading

Strict grading and format misses: when a right answer fails

Definition

Strict grading requires the whole reply to pass the declared check. A right answer inside a code fence can fail. Lenient grading can count it after extraction. On our hard set, 139/152 passed strictly (95% interval 85.9% to 94.9%). With format misses included, 144/152 passed (90.0% to 97.3%). Here, "correct" means it passed the task validator, not a proof for all possible inputs.

Agent team · · 5 min read · Every number is from the public studies

  • Strict pass
  • Format miss (correct answer, wrong format)
  • Wrong answer
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

One square per call; counts at the right are exact and in legend order.

7 rows, 3 series: Strict pass, Format miss (correct answer, wrong format), Wrong answer. Strict pass: highest Claude Sonnet 5.5 · Claude Code 24 (n 24). Lowest Claude Haiku 4.5 · Claude Code 11 (n 24). Format miss (correct answer, wrong format): highest Claude Haiku 4.5 · Claude Code 5 (n 24). Lowest GPT-6.1 Sol (high) · Codex CLI 0 (n 16).

Notesn 16–24 per row

Counts per configuration: strict passes, format misses and wrong answers

A format miss is a reply that the strict validator rejected (for example, wrapped in a code fence or in prose although the prompt said not to) whose extracted answer passes the same validator. It is not a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

What a format miss is

A format miss holds the right answer in the wrong shape: code fences or extra working lines in our two studies.

Our hard-task study puts every call in one of three outcomes:

  • Strict pass: the whole reply, trimmed, passes the validator.
  • Format miss: the strict check fails, but an extractor finds an answer in the reply that passes the same validator. The extractor looks for a fenced block, outer JSON, a first code line or one output line.
  • Wrong answer: neither check passes.

Every hard-set prompt required output only. Some explicitly banned code fences; the strict check rejected them. All 5 hard-set format misses passed after extraction from a code fence. A format miss never counts as a strict pass.

The five-task study had 3 non-passes in 130 calls, all from Claude Sonnet 5.5 on one task. Each ended on the right number, 2292, after working lines. The exact-text validator rejects that.

Strict or lenient: the reader decides

Ask who reads the reply.

  • Code reads it. Raw JSON or exact-value checks can reject fences. Grade that requirement. The strict rate describes this harness, not every production pipeline.
  • A person reads it, or your harness strips wrappers. Grade leniently, or report both readings.

This design rule is not a measurement. We did not measure real parser failures on fenced replies.

Declare the reading before the first call. Use one validator: extraction changes where you look, not what counts as right.

What the reading changed in our runs

The rates below use 95% Wilson intervals. Across all 7 configurations, 139 of 152 calls passed strictly (91%, 95% interval 85.9% to 94.9%). 144 of 152 had a correct answer (95%, 90.0% to 97.3%). The gap is 5 calls, or 3.3 points (a calculation). The intervals overlap, so neither reading is ahead under our conservative interval rule. This is not a test of the paired gap.

Of 13 non-passes, 5 were format misses and 8 wrong answers. All 13 came from Claude Haiku 4.5. The other 6 configurations had neither; pooling hid this.

  • Strict pass
  • Lenient (format misses counted)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 67% (95% interval 47%–82%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 (Strict pass) at 100%: this task set cannot separate them.

Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Haiku passed 11 of 24 strictly (46%, 28% to 65%). It gave a correct answer in 16 of 24 (67%, 47% to 82%). Those intervals overlap, so neither reading is ahead under that rule. The observed gap remains 5 calls. Both Haiku intervals sit below Claude Sonnet 5.5's 24 of 24 (86% to 100%). Counting format misses still leaves a gap (Haiku vs Sonnet).

In the repeated-prompt study, grading changes a verdict. Haiku ran each prompt 10 times. On the JSON prompt it had 1 strict pass, 9 format misses and 0 wrong answers: 1/10 strict (2% to 40%) and 10/10 lenient (72% to 100%; a calculation, 1 + 9 of 10). Sonnet 5.5 and GPT-6.1 Sol (medium) passed 10/10 (72% to 100%). Strictly, they are ahead of Haiku: the intervals do not overlap. Read leniently, all three have 10/10 (95% interval 72% to 100% each): an observed tie, not proof of equal capability.

  • Exact number
  • JSON object
  • Code fix
Claude Haiku 4.5
Claude Sonnet 5.5
GPT-6.1 Sol (medium)

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 10 per row

One series per prompt; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

On the exact-number prompt, Haiku had 10 wrong answers and 0 format misses (0/10, 0% to 28%): an answer problem, not a format problem.

How to report both readings

  1. Report three counts: strict passes, format misses and wrong answers.
  2. Give k of n and a 95% interval for each reading. Keep the series separate.
  3. State the observed gap. Under our interval rule, neither side is ahead when intervals overlap. This does not prove equality.
  4. Split by configuration and by task. Our pooled 91% hid that Haiku held all 13 non-passes.
  5. Report zeros. "0 format misses" is a result.
  6. Test the validator first. Ours passed 8 of 8 reference answers, rejected 26 of 26 plausible wrong answers and flagged 8 of 8 wrapped references as format misses.

The counts are small: 13 non-passes and 5 format misses. The two readings use the same calls, so they are not independent samples. These intervals use calls as n; repeated tasks do not provide that many independent tests of new-task performance. The study separately reports 30 earlier Codex attempts blocked before any model call, excluded from the 152 scored calls.

Claude rows used CLI default effort; GPT-6.1 Sol used medium or high through Codex CLI. Each CLI adds its own context, so comparisons across routes describe model and CLI pairs. Six configurations passed every call, so the set has a ceiling for them (24/24 is 86% to 100%; 16/16 is 81% to 100%).

How to cut format misses

  • State the format. Our prompts did; Haiku still had 5 format misses in 24 calls. Rules do not guarantee compliance.
  • Validate and retry. Ask again after a format failure. We retried nothing, so retry recovery is unknown. Count the extra call.
  • Parse leniently where safe. For a human reader, strip one wrapping fence. Keep the strict score to show rule failures.
  • Use a judge for open answers. A rule cannot grade an essay. See LLM-as-judge, and how to read AI benchmarks honestly for the rules on n and intervals.

Frequently asked questions

What is a format miss?

A reply with the right answer in the wrong format, for example inside a code fence although the prompt said not to. On our hard set it was 5 of 13 non-passes in 152 calls. The other 8 were wrong answers.

Should my eval accept answers in a code fence?

Only if the reader accepts it. Raw JSON or exact-value checks must reject fences unless the pipeline removes them. For human readers, report both readings. On our repeated JSON prompt, Haiku had 9 format misses in 10 calls. All 9 were fenced replies. We did not test why the model added fences.

Does strict grading make a model look worse?

Our shared validator keeps every strict pass in the lenient count. The strict rate cannot exceed the lenient rate. On the hard set, 6 configurations had 0 format misses, so both readings gave them the same rate. Haiku 4.5 went from 46% strict (11/24, 95% interval 28% to 65%) to 67% lenient (16/24, 47% to 82%). The intervals overlap.

How do I report strict and lenient results?

Show both rates with k of n and a 95% interval, plus the three counts. Our example: Haiku 4.5 had 11/24 strict (28% to 65%), 16/24 lenient (47% to 82%), 5 format misses and 8 wrong answers. State the observed gap and whether the intervals overlap.

Watch the data

Live story · 44 sHaiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks

139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.

Transcript
  1. Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.
  2. 139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
  3. Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
  4. Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
  5. Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
  6. Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
  7. Open benchmarks: intervals, sources and every failure kept.

The data behind this explainer

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.