Explainer · LLM nondeterminism

Why the same prompt gives different answers: LLM nondeterminism, measured

Definition

LLM nondeterminism means a model can answer the same prompt differently. General causes include how the model samples each next token and how batching on shared servers can change tiny numeric details. In 90 repeated calls, the wording changed in 5 of 9 prompt-and-model pairs and the answer in 3 of 9.

Agent team · · 5 min read · Every number is from the public studies

Exact number

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

JSON object

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Code fix

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

One panel per series, all on the same axis.

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: all at 1. JSON object: all at 1.

Notesn = 10 per row

Distinct normalized answers over 10 repetitions (1 = the same answer every time)

Normalized: the last number for the exact-number prompt; key-sorted JSON for the JSON prompt; for the code fix, the code without fences, comments, whitespace, quote style and line-end semicolons. One distinct answer is not the same as a correct answer: an answer can be the same and wrong every time. Different code can be equally correct.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Wording changes more than the answer

A reply has two parts. The wording is the raw text. The answer is what a script keeps after it normalizes the reply: the last number, sorted JSON, or code without fences, comments and spacing. The two change at different rates. We did not measure the causes.

We sent 3 prompts to 3 setups, 10 times each, one call at a time. That is 90 calls in the caching and consistency study. The prompts were an exact number, a JSON object with fixed keys and a small code fix. A validator in code checked every reply.

We did not set or vary temperature or seeds. The CLIs ran at their defaults. Claude ran at its default effort and GPT-6.1 Sol at medium. So this page shows no effect of any setting.

Exact number

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

JSON object

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

Code fix

Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

One panel per series, all on the same axis.

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: all at 1. JSON object: all at 1.

Notesn = 10 per row

Distinct normalized answers over 10 repetitions (1 = the same answer every time)

Normalized: the last number for the exact-number prompt; key-sorted JSON for the JSON prompt; for the code fix, the code without fences, comments, whitespace, quote style and line-end semicolons. One distinct answer is not the same as a correct answer: an answer can be the same and wrong every time. Different code can be equally correct.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

A cell is one prompt on one setup: 10 calls (n = 10). A count of 1 means the same answer every time.

  • Exact number and JSON object: all 6 cells gave 1 distinct answer.
  • Wording on those two prompts: Sonnet 5.5 and GPT-6.1 Sol (medium) wrote the same raw reply 10 times in 10. Haiku 4.5 wrote 10 different raw replies on the exact-number prompt and 3 on the JSON prompt. Each set held one answer.
  • Code fix: the code changed. We counted 6 distinct code bodies for Haiku 4.5, 3 for Sonnet 5.5 and 6 for GPT-6.1 Sol. All passed the validator, so different code was equally correct.

Consistent is not correct

  • Exact number
  • JSON object
  • Code fix
Claude Haiku 4.5
Claude Sonnet 5.5
GPT-6.1 Sol (medium)

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 10 per row

One series per prompt; whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Haiku 4.5 gave 289 on the exact-number prompt in all 10 calls. The expected answer is 282. That is 0 of 10 strict passes (95% Wilson interval 0% to 28%) and 10 wrong answers. The model was consistent and wrong every time.

A test that asks only "did the runs agree?" passes this cell. Only a validator that knows the right answer catches it. The most common answer of the 10 calls was also the wrong one.

The JSON prompt shows a third case. Haiku 4.5 passed 1 of 10 (95% interval 2% to 40%). The other 9 were format misses: the right content in the wrong format, for example a reply in a code fence. There were 0 wrong answers.

In total, 7 of 9 cells passed 10 of 10. A 10/10 has a 95% interval of 72% to 100%. Ten calls cannot show that a failure never happens.

Sonnet 5.5 and GPT-6.1 Sol passed all 3 prompts 10 of 10 (72% to 100% each). They tie on strict passes, and this task set hits a ceiling for them. On the exact-number prompt, Sonnet 5.5 (10/10, 72% to 100%) is ahead of Haiku 4.5 (0/10, 0% to 28%), because the intervals do not overlap. See Haiku 4.5 vs Sonnet 5.5.

Time also varies

  • Exact number
  • JSON object
  • Code fix
Entrance: medians race at 9.6× real timeMotion reduced: press Replay to animateThe slowest median is 13.4 s. The clock runs at the recorded speed.
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (medium) · Codex CLI

3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: slowest GPT-6.1 Sol (medium) · Codex CLI 13.4 s (range 12.3 s–18 s, n 10). Fastest Claude Haiku 4.5 · Claude Code 5.1 s (range 4.4 s–6.2 s, n 10). Not all run ranges overlap. JSON object: slowest Claude Haiku 4.5 · Claude Code 7 s (range 5.3 s–12.3 s, n 10). Fastest Claude Sonnet 5.5 · Claude Code 2.9 s (range 2.7 s–5.3 s, n 10). All run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 10 per row

Median; whiskers = fastest and slowest of 10 calls

Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.

Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)

Time per call moves even when the answer does not. The chart shows the median and the shortest-to-longest range of 10 calls. A range is not a confidence interval.

  • Sonnet 5.5, exact number: median 6.89 s, range 5.81 s to 7.81 s.
  • Haiku 4.5, JSON object: median 7.03 s, range 5.28 s to 12.27 s. The longest call took about 2.3 times as long as the shortest (calculation: 12.27 / 5.28).
  • All 9 cells: coefficient of variation from 0.09 to 0.28, and interquartile range from 0.54 s to 2.25 s.

Codex CLI times include its start-up, so do not compare times across routes.

How to test your own prompt

  1. Fix the setup. Use the same model, settings and tools, and a fresh context, for every call. We used an empty working folder, tools off and no saved sessions.
  2. Write the validator first. The reference answer must pass, and each plausible wrong answer must fail. We ran these controls before the first call.
  3. Repeat the call at least 10 times, one at a time.
  4. Normalize each reply. Count the distinct answers and the distinct raw replies.
  5. Count strict passes, format misses and wrong answers apart. Report k/n with its 95% Wilson interval.
  6. Keep every call, failures included. We retried and dropped none.

See also how to read AI benchmarks honestly.

What cuts variance, and what our data shows

Our 90 calls support some steps. They do not test others.

  • Check the format with code. A strict validator turned Haiku 4.5's JSON replies into 9 counted format misses. This makes the problem visible. We did not test a format instruction.
  • Test code by behavior, not by text. Haiku 4.5 wrote 6 different code bodies and all passed. A text match against one body rejects the rest.
  • Prefer short tasks with one exact answer. All 6 such cells gave 1 distinct answer. The 3 code-fix cells gave 3 to 6. The prompts differ in kind, so this is a pattern, not a controlled test.

We did not test temperature, seeds or longer prompts.

Frequently asked questions

Is LLM output deterministic at temperature 0?

We did not test it. We did not set or vary temperature, so our 90 calls say nothing about temperature 0. At the CLI defaults we ran, the wording changed in 5 of 9 cells (n = 10 each). Check your provider's documentation. Then measure your own prompt.

Why did the model give the same wrong answer every time?

Haiku 4.5 gave 289 on all 10 calls of the exact-number prompt (n = 10). The expected answer is 282. We did not test the cause. Repeating the call did not change the answer, so check every answer with a validator.

How many times should I repeat a prompt to test it?

Start with 10. A clean 10 of 10 still has a 95% interval of 72% to 100%. A clean 30 of 30 has 89% to 100% (calculation: Wilson interval for 30 of 30). Repeat more for a tighter bound. We tested 10 only.

Does variance mean the model is unreliable?

Not by itself. Different wording or code can both be correct: every code-fix cell passed 10 of 10 with 6, 3 and 6 distinct code bodies. The strict pass rate shows reliability. Haiku 4.5 gave one answer on the exact-number prompt and still passed 0 of 10.

Watch the data

Live story · 49 sPrompt caching and consistency: what the cache saves, and how much answers vary

Prompt caching and consistency: what the cache saves, and how much answers vary

Calculation at list price: the cache cut a 5-question session 50% on Sonnet 5.5 and 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every repetition.

Transcript
  1. Caching and consistency · 135 calls. What the cache saves, and how much answers vary. Five-question sessions on a fixed context. Then the same prompt, 10 times.
  2. 135 calls: 45 cache turns, 90 repeated prompts. Every call counted. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
  3. Claude Code turn 1 reads 19% from the cache, the CLI’s own prefix. Turns 2–5 read 90%–99%. Codex CLI: 98%–99%. Chart: Share of input read from the cache · mean of 3 sessions per turn (n = 3 each). Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  4. A calculation: Sonnet 5.5 $0.1350 with the cache vs $0.2698 without, 50% less. Opus 5.5 $0.2551 vs $0.5442, 53% less. Chart: List-price cost of the recorded sessions · calculation (n = 15 each). Calculation, not a run. Caveat: Costs are list-price calculations; the calls used flat subscriptions.
  5. No clear speed effect: Sonnet 5.5 1.6 s on turn 1 vs 1.6 s later; Opus 5.5 1.9 s vs 2.4 s. The ranges overlap. Sonnet 5.5: median turn 1 vs turns 2–5: 1.6 s vs 1.6 s (ranges 1.6–1.8 s and 1.4–5.6 s · n = 3 and 12). Opus 5.5: median turn 1 vs turns 2–5: 1.9 s vs 2.4 s (ranges 1.8–4.4 s and 1.6–12.7 s · n = 3 and 12). Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
  6. 7 of 9 cells passed 10/10. Haiku 4.5: 0/10 on the exact number (10 wrong, 1 distinct answer), 1/10 on JSON (9 format misses). Chart: Same prompt, 10 times · strict passes · a 10/10 is 72%–100% at 95% (n = 10 each). Caveat: Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
  7. Consistent is not correct: Haiku 4.5 gave the same wrong number all 10 times. Code fix, distinct correct bodies: Haiku 4.5 6, Sonnet 5.5 3, GPT-6.1 Sol (medium) 6. Chart: Same prompt, 10 times · distinct answers (n = 10 each). Caveat: The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
  8. Cache the fixed context; check answers, not agreement. Every call online.

The data behind this explainer

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

More explainers

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.