• Consistency
  • Variance
  • Claude Haiku
  • Claude Sonnet

Same prompt, ten answers: how consistent are Claude and Codex?

3 prompts, 10 runs each, on Haiku, Sonnet and GPT-6.1 Sol. 7 of 9 cells passed 10/10. Haiku gave the same wrong number 10 times: consistent is not correct.

TL;DR

  • We sent 3 prompts 10 times each to Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI: 90 calls, each checked by a deterministic validator.
  • 7 of 9 model-and-prompt cells passed 10 of 10 (95% interval 72% to 100%).
  • Haiku failed the exact-number prompt 0/10, and gave the same wrong number all 10 times (289; the answer is 282). Consistent is not the same as correct.
  • Haiku passed the JSON prompt 1/10; the other 9 replies had the right content in the wrong format.
  • On the code-fix prompt, every model passed 10/10, but wrote different code bodies: Haiku 6, Sonnet 3, GPT-6.1 Sol 6 distinct versions.
  • Under the comparison rules, Sonnet and GPT-6.1 Sol beat Haiku on the exact-number and JSON prompts. Sonnet and Haiku were faster than GPT-6.1 Sol through the Codex CLI on two prompts.

Study: /benchmarks/caching-consistency.

The question

A benchmark usually runs each task once or a few times. In production, the same prompt runs hundreds of times. So two questions matter:

  1. Does the pass rate hold when you ask the same thing again?
  2. Does the answer itself change, even when it passes?

We measured both, plus the spread in time per call.

What we ran

  • 3 prompts, each with a deterministic validator:
    • Exact number: how many integers from 1 to 2026 have a digit sum divisible by 7? One correct integer.
    • JSON object: turn an order note into one JSON object with exact keys, types and values (merged lines, a cancelled line, a discount).
    • Code fix: repair a median function (numeric sort, even length, no mutation, a RangeError case), checked by 10 checks in a sandbox.
  • 3 configurations: Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code (both at default effort) and GPT-6.1 Sol (medium) · Codex CLI.
  • 10 repetitions per cell, one-shot calls, one at a time per account. Fresh empty folder, tools off, no session persistence.
  • Controls before inference on each route: every reference answer passes, every plausible wrong answer fails, and a wrapped reference is flagged as a format miss.
  • Answer diversity counts distinct normalized answers. The public extract holds ordinal answer ids, never the reply text.

Nothing was retried or trimmed.

Pass rate: mostly 10 of 10

  • Exact number
  • JSON object
  • Code fix
Claude Haiku 4.5
Claude Sonnet 5.5
GPT-6.1 Sol (medium)

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 10 per row

One series per prompt; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

ConfigurationExact numberJSON objectCode fix
Claude Haiku 4.5 · Claude Code0/10 (0% to 28%)1/10 (2% to 40%)10/10 (72% to 100%)
Claude Sonnet 5.5 · Claude Code10/10 (72% to 100%)10/10 (72% to 100%)10/10 (72% to 100%)
GPT-6.1 Sol (medium) · Codex CLI10/10 (72% to 100%)10/10 (72% to 100%)10/10 (72% to 100%)

Sonnet and GPT-6.1 Sol passed all 30 of their calls each. Haiku passed the code fix every time, and failed the other two almost every time.

Where the intervals do not overlap, the comparison pages name a winner. That gives four rows: Sonnet beats Haiku and GPT-6.1 Sol beats Haiku, each on the exact-number and the JSON prompt. See Haiku vs Sonnet and Haiku vs GPT-6.1 Sol. Between Sonnet and GPT-6.1 Sol, every pass-rate row is a tie.

A 10/10 still has a 95% interval of 72% to 100% (how a Wilson interval works). Ten repetitions show that a failure is not common; they cannot show that it never happens.

Consistent is not correct

Exact number

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

JSON object

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Code fix

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

One panel per series, all on the same axis.

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: all at 1. JSON object: all at 1.

Notesn = 10 per row

Distinct normalized answers over 10 repetitions (1 = the same answer every time)

Normalized: the last number for the exact-number prompt; key-sorted JSON for the JSON prompt; for the code fix, the code without fences, comments, whitespace, quote style and line-end semicolons. One distinct answer is not the same as a correct answer: an answer can be the same and wrong every time. Different code can be equally correct.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

The most useful result is Haiku on the exact-number prompt. All 10 replies were worded differently: 10 distinct raw replies. But all 10 reduced to one answer: 289. The correct answer is 282.

So Haiku was perfectly consistent, and wrong every time. If you test a prompt by running it twice and checking that the answers agree, this case passes your check. Only a validator with the right answer catches it.

The JSON case is different. Haiku's content was right, and it gave one distinct answer in all 10 replies. But 9 of the 10 wrapped the object in a code fence, which this validator rejects by design. That is 9 format misses and 0 wrong answers. A more lenient parser would have accepted them; a strict pipeline would not.

Same pass, different code

On the code-fix prompt, every call passed. The code was not the same each time:

  • Claude Haiku 4.5: 6 distinct code bodies in 10 runs
  • Claude Sonnet 5.5: 3
  • GPT-6.1 Sol (medium): 6

All of them pass the 10 checks. For most uses that is fine. It matters when you diff outputs, cache by output, or review changes: the same request can produce a different patch on each run. The comparison pages mark these rows unclear, because fewer distinct answers is not better or worse by itself.

Time per call: how much it moves

  • Exact number
  • JSON object
  • Code fix
Entrance: medians race at 9.6× real timeMotion reduced: press Replay to animateThe slowest median is 13.4 s. The clock runs at the recorded speed.
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: slowest GPT-6.1 Sol (medium) · Codex CLI 13.4 s (range 12.3 s–18 s, n 10). Fastest Claude Haiku 4.5 · Claude Code 5.1 s (range 4.4 s–6.2 s, n 10). Not all run ranges overlap. JSON object: slowest Claude Haiku 4.5 · Claude Code 7 s (range 5.3 s–12.3 s, n 10). Fastest Claude Sonnet 5.5 · Claude Code 2.9 s (range 2.7 s–5.3 s, n 10). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 10 per row

Median; whiskers = fastest and slowest of 10 calls

Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Median time per call, with the fastest and slowest of 10:

ConfigurationExact numberJSON objectCode fix
Haiku 4.5 · Claude Code5.06 s (4.42 to 6.20)7.03 s (5.28 to 12.27)5.95 s (4.89 to 7.33)
Sonnet 5.5 · Claude Code6.89 s (5.81 to 7.81)2.89 s (2.68 to 5.30)2.67 s (2.32 to 4.34)
GPT-6.1 Sol · Codex CLI13.38 s (12.29 to 17.97)6.42 s (5.25 to 8.26)11.29 s (9.08 to 14.85)

With one prompt repeated, the spread is much narrower than across different tasks. The coefficient of variation stayed between 0.09 and 0.28 in every cell.

Because the ranges are narrow, some rows separate under the comparison rules (no range overlap, 10 runs per side):

  • Sonnet faster than GPT-6.1 Sol on the exact-number and code-fix prompts.
  • Haiku faster than GPT-6.1 Sol on the same two prompts.
  • Sonnet faster than Haiku on the code fix.

These rows compare route + model pairs. The Codex CLI adds its own system prompt and start-up time, so part of the gap is the CLI, not the model. See Claude Code vs Codex CLI: the hidden context tax.

What this means for your prompts

  1. Test against the right answer, not against the last answer. Agreement between runs does not show correctness.
  2. Count format misses apart from wrong answers. They need a different fix: a parser or an instruction, not a bigger model.
  3. Expect different code for the same request. Validate behaviour with tests, not by text match.
  4. Run a prompt more than once before you trust it. Ten repetitions found Haiku's two weak prompts at once.

Caveats

  • 10 repetitions per cell. A 10/10 has a 95% interval of 72% to 100%.
  • 3 prompts. They are short and have one correct answer each. Open-ended work will vary more.
  • Effort differs by route. Claude rows ran at default effort, GPT-6.1 Sol at medium.
  • Strict format. The JSON validator rejects a reply in a code fence even when the JSON is right.

Validate every answer

Agent checks each step it delivers against a validator, and records the result next to the model, the time and the cost. Try Agent and see how consistent your own tasks are.

The data behind this post

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.