Same prompt, ten answers: how consistent are Claude and Codex?
3 prompts, 10 runs each, on Haiku, Sonnet and GPT-6.1 Sol. 7 of 9 cells passed 10/10. Haiku gave the same wrong number 10 times: consistent is not correct.
TL;DR
- We sent 3 prompts 10 times each to Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI: 90 calls, each checked by a deterministic validator.
- 7 of 9 model-and-prompt cells passed 10 of 10 (95% interval 72% to 100%).
- Haiku failed the exact-number prompt 0/10, and gave the same wrong number all 10 times (289; the answer is 282). Consistent is not the same as correct.
- Haiku passed the JSON prompt 1/10; the other 9 replies had the right content in the wrong format.
- On the code-fix prompt, every model passed 10/10, but wrote different code bodies: Haiku 6, Sonnet 3, GPT-6.1 Sol 6 distinct versions.
- Under the comparison rules, Sonnet and GPT-6.1 Sol beat Haiku on the exact-number and JSON prompts. Sonnet and Haiku were faster than GPT-6.1 Sol through the Codex CLI on two prompts.
Study: /benchmarks/caching-consistency.
The question
A benchmark usually runs each task once or a few times. In production, the same prompt runs hundreds of times. So two questions matter:
- Does the pass rate hold when you ask the same thing again?
- Does the answer itself change, even when it passes?
We measured both, plus the spread in time per call.
What we ran
- 3 prompts, each with a deterministic validator:
- Exact number: how many integers from 1 to 2026 have a digit sum divisible by 7? One correct integer.
- JSON object: turn an order note into one JSON object with exact keys, types and values (merged lines, a cancelled line, a discount).
- Code fix: repair a
medianfunction (numeric sort, even length, no mutation, aRangeErrorcase), checked by 10 checks in a sandbox.
- 3 configurations: Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code (both at default effort) and GPT-6.1 Sol (medium) · Codex CLI.
- 10 repetitions per cell, one-shot calls, one at a time per account. Fresh empty folder, tools off, no session persistence.
- Controls before inference on each route: every reference answer passes, every plausible wrong answer fails, and a wrapped reference is flagged as a format miss.
- Answer diversity counts distinct normalized answers. The public extract holds ordinal answer ids, never the reply text.
Nothing was retried or trimmed.
Pass rate: mostly 10 of 10
| Configuration | Exact number | JSON object | Code fix |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 0/10 (0% to 28%) | 1/10 (2% to 40%) | 10/10 (72% to 100%) |
| Claude Sonnet 5.5 · Claude Code | 10/10 (72% to 100%) | 10/10 (72% to 100%) | 10/10 (72% to 100%) |
| GPT-6.1 Sol (medium) · Codex CLI | 10/10 (72% to 100%) | 10/10 (72% to 100%) | 10/10 (72% to 100%) |
Sonnet and GPT-6.1 Sol passed all 30 of their calls each. Haiku passed the code fix every time, and failed the other two almost every time.
Where the intervals do not overlap, the comparison pages name a winner. That gives four rows: Sonnet beats Haiku and GPT-6.1 Sol beats Haiku, each on the exact-number and the JSON prompt. See Haiku vs Sonnet and Haiku vs GPT-6.1 Sol. Between Sonnet and GPT-6.1 Sol, every pass-rate row is a tie.
A 10/10 still has a 95% interval of 72% to 100% (how a Wilson interval works). Ten repetitions show that a failure is not common; they cannot show that it never happens.
Consistent is not correct
The most useful result is Haiku on the exact-number prompt. All 10 replies were worded differently: 10 distinct raw replies. But all 10 reduced to one answer: 289. The correct answer is 282.
So Haiku was perfectly consistent, and wrong every time. If you test a prompt by running it twice and checking that the answers agree, this case passes your check. Only a validator with the right answer catches it.
The JSON case is different. Haiku's content was right, and it gave one distinct answer in all 10 replies. But 9 of the 10 wrapped the object in a code fence, which this validator rejects by design. That is 9 format misses and 0 wrong answers. A more lenient parser would have accepted them; a strict pipeline would not.
Same pass, different code
On the code-fix prompt, every call passed. The code was not the same each time:
- Claude Haiku 4.5: 6 distinct code bodies in 10 runs
- Claude Sonnet 5.5: 3
- GPT-6.1 Sol (medium): 6
All of them pass the 10 checks. For most uses that is fine. It matters when you diff outputs, cache by output, or review changes: the same request can produce a different patch on each run. The comparison pages mark these rows unclear, because fewer distinct answers is not better or worse by itself.
Time per call: how much it moves
Median time per call, with the fastest and slowest of 10:
| Configuration | Exact number | JSON object | Code fix |
|---|---|---|---|
| Haiku 4.5 · Claude Code | 5.06 s (4.42 to 6.20) | 7.03 s (5.28 to 12.27) | 5.95 s (4.89 to 7.33) |
| Sonnet 5.5 · Claude Code | 6.89 s (5.81 to 7.81) | 2.89 s (2.68 to 5.30) | 2.67 s (2.32 to 4.34) |
| GPT-6.1 Sol · Codex CLI | 13.38 s (12.29 to 17.97) | 6.42 s (5.25 to 8.26) | 11.29 s (9.08 to 14.85) |
With one prompt repeated, the spread is much narrower than across different tasks. The coefficient of variation stayed between 0.09 and 0.28 in every cell.
Because the ranges are narrow, some rows separate under the comparison rules (no range overlap, 10 runs per side):
- Sonnet faster than GPT-6.1 Sol on the exact-number and code-fix prompts.
- Haiku faster than GPT-6.1 Sol on the same two prompts.
- Sonnet faster than Haiku on the code fix.
These rows compare route + model pairs. The Codex CLI adds its own system prompt and start-up time, so part of the gap is the CLI, not the model. See Claude Code vs Codex CLI: the hidden context tax.
What this means for your prompts
- Test against the right answer, not against the last answer. Agreement between runs does not show correctness.
- Count format misses apart from wrong answers. They need a different fix: a parser or an instruction, not a bigger model.
- Expect different code for the same request. Validate behaviour with tests, not by text match.
- Run a prompt more than once before you trust it. Ten repetitions found Haiku's two weak prompts at once.
Caveats
- 10 repetitions per cell. A 10/10 has a 95% interval of 72% to 100%.
- 3 prompts. They are short and have one correct answer each. Open-ended work will vary more.
- Effort differs by route. Claude rows ran at default effort, GPT-6.1 Sol at medium.
- Strict format. The JSON validator rejects a reply in a code fence even when the JSON is right.
What to read next
- Does reasoning effort buy quality? Claude and Codex on hard tasks
- How much does prompt caching actually save?
- When tasks get hard: Haiku vs Sonnet vs Opus vs Fable
- Why we count every failed attempt
Validate every answer
Agent checks each step it delivers against a validator, and records the result next to the model, the time and the cost. Try Agent and see how consistent your own tasks are.