Explainer · LLM nondeterminism
Why the same prompt gives different answers: LLM nondeterminism, measured
Definition
LLM nondeterminism means a model can answer the same prompt differently. General causes include how the model samples each next token and how batching on shared servers can change tiny numeric details. In 90 repeated calls, the wording changed in 5 of 9 prompt-and-model pairs and the answer in 3 of 9.
Agent team · · 5 min read · Every number is from the public studies
Exact number
JSON object
Code fix
One panel per series, all on the same axis.
| Item | Exact number | JSON object | Code fix | n |
|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 1 | 1 | 6 | 10 |
| Claude Sonnet 5.5 · Claude Code | 1 | 1 | 3 | 10 |
| GPT-6.1 Sol (medium) · Codex CLI | 1 | 1 | 6 | 10 |
3 rows, 3 series: Exact number, JSON object, Code fix. Exact number: all at 1. JSON object: all at 1.
Notesn = 10 per row
Distinct normalized answers over 10 repetitions (1 = the same answer every time)
Normalized: the last number for the exact-number prompt; key-sorted JSON for the JSON prompt; for the code fix, the code without fences, comments, whitespace, quote style and line-end semicolons. One distinct answer is not the same as a correct answer: an answer can be the same and wrong every time. Different code can be equally correct.
Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)
Wording changes more than the answer
A reply has two parts. The wording is the raw text. The answer is what a script keeps after it normalizes the reply: the last number, sorted JSON, or code without fences, comments and spacing. The two change at different rates. We did not measure the causes.
We sent 3 prompts to 3 setups, 10 times each, one call at a time. That is 90 calls in the caching and consistency study. The prompts were an exact number, a JSON object with fixed keys and a small code fix. A validator in code checked every reply.
We did not set or vary temperature or seeds. The CLIs ran at their defaults. Claude ran at its default effort and GPT-6.1 Sol at medium. So this page shows no effect of any setting.
A cell is one prompt on one setup: 10 calls (n = 10). A count of 1 means the same answer every time.
- Exact number and JSON object: all 6 cells gave 1 distinct answer.
- Wording on those two prompts: Sonnet 5.5 and GPT-6.1 Sol (medium) wrote the same raw reply 10 times in 10. Haiku 4.5 wrote 10 different raw replies on the exact-number prompt and 3 on the JSON prompt. Each set held one answer.
- Code fix: the code changed. We counted 6 distinct code bodies for Haiku 4.5, 3 for Sonnet 5.5 and 6 for GPT-6.1 Sol. All passed the validator, so different code was equally correct.
Consistent is not correct
Haiku 4.5 gave 289 on the exact-number prompt in all 10 calls. The expected answer is 282. That is 0 of 10 strict passes (95% Wilson interval 0% to 28%) and 10 wrong answers. The model was consistent and wrong every time.
A test that asks only "did the runs agree?" passes this cell. Only a validator that knows the right answer catches it. The most common answer of the 10 calls was also the wrong one.
The JSON prompt shows a third case. Haiku 4.5 passed 1 of 10 (95% interval 2% to 40%). The other 9 were format misses: the right content in the wrong format, for example a reply in a code fence. There were 0 wrong answers.
In total, 7 of 9 cells passed 10 of 10. A 10/10 has a 95% interval of 72% to 100%. Ten calls cannot show that a failure never happens.
Sonnet 5.5 and GPT-6.1 Sol passed all 3 prompts 10 of 10 (72% to 100% each). They tie on strict passes, and this task set hits a ceiling for them. On the exact-number prompt, Sonnet 5.5 (10/10, 72% to 100%) is ahead of Haiku 4.5 (0/10, 0% to 28%), because the intervals do not overlap. See Haiku 4.5 vs Sonnet 5.5.
Time also varies
Time per call moves even when the answer does not. The chart shows the median and the shortest-to-longest range of 10 calls. A range is not a confidence interval.
- Sonnet 5.5, exact number: median 6.89 s, range 5.81 s to 7.81 s.
- Haiku 4.5, JSON object: median 7.03 s, range 5.28 s to 12.27 s. The longest call took about 2.3 times as long as the shortest (calculation: 12.27 / 5.28).
- All 9 cells: coefficient of variation from 0.09 to 0.28, and interquartile range from 0.54 s to 2.25 s.
Codex CLI times include its start-up, so do not compare times across routes.
How to test your own prompt
- Fix the setup. Use the same model, settings and tools, and a fresh context, for every call. We used an empty working folder, tools off and no saved sessions.
- Write the validator first. The reference answer must pass, and each plausible wrong answer must fail. We ran these controls before the first call.
- Repeat the call at least 10 times, one at a time.
- Normalize each reply. Count the distinct answers and the distinct raw replies.
- Count strict passes, format misses and wrong answers apart. Report k/n with its 95% Wilson interval.
- Keep every call, failures included. We retried and dropped none.
See also how to read AI benchmarks honestly.
What cuts variance, and what our data shows
Our 90 calls support some steps. They do not test others.
- Check the format with code. A strict validator turned Haiku 4.5's JSON replies into 9 counted format misses. This makes the problem visible. We did not test a format instruction.
- Test code by behavior, not by text. Haiku 4.5 wrote 6 different code bodies and all passed. A text match against one body rejects the rest.
- Prefer short tasks with one exact answer. All 6 such cells gave 1 distinct answer. The 3 code-fix cells gave 3 to 6. The prompts differ in kind, so this is a pattern, not a controlled test.
We did not test temperature, seeds or longer prompts.
Frequently asked questions
Is LLM output deterministic at temperature 0?
We did not test it. We did not set or vary temperature, so our 90 calls say nothing about temperature 0. At the CLI defaults we ran, the wording changed in 5 of 9 cells (n = 10 each). Check your provider's documentation. Then measure your own prompt.
Why did the model give the same wrong answer every time?
Haiku 4.5 gave 289 on all 10 calls of the exact-number prompt (n = 10). The expected answer is 282. We did not test the cause. Repeating the call did not change the answer, so check every answer with a validator.
How many times should I repeat a prompt to test it?
Start with 10. A clean 10 of 10 still has a 95% interval of 72% to 100%. A clean 30 of 30 has 89% to 100% (calculation: Wilson interval for 30 of 30). Repeat more for a tighter bound. We tested 10 only.
Does variance mean the model is unreliable?
Not by itself. Different wording or code can both be correct: every code-fix cell passed 10 of 10 with 6, 3 and 6 distinct code bodies. The strict pass rate shows reliability. Haiku 4.5 gave one answer on the exact-number prompt and still passed 0 of 10.
Watch the data
Prompt caching and consistency: what the cache saves, and how much answers vary
Calculation at list price: the cache cut a 5-question session 50% on Sonnet 5.5 and 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every repetition.
Transcript
- Caching and consistency · 135 calls. What the cache saves, and how much answers vary. Five-question sessions on a fixed context. Then the same prompt, 10 times.
- 135 calls: 45 cache turns, 90 repeated prompts. Every call counted. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- Claude Code turn 1 reads 19% from the cache, the CLI’s own prefix. Turns 2–5 read 90%–99%. Codex CLI: 98%–99%. Chart: Share of input read from the cache · mean of 3 sessions per turn (n = 3 each). Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- A calculation: Sonnet 5.5 $0.1350 with the cache vs $0.2698 without, 50% less. Opus 5.5 $0.2551 vs $0.5442, 53% less. Chart: List-price cost of the recorded sessions · calculation (n = 15 each). Calculation, not a run. Caveat: Costs are list-price calculations; the calls used flat subscriptions.
- No clear speed effect: Sonnet 5.5 1.6 s on turn 1 vs 1.6 s later; Opus 5.5 1.9 s vs 2.4 s. The ranges overlap. Sonnet 5.5: median turn 1 vs turns 2–5: 1.6 s vs 1.6 s (ranges 1.6–1.8 s and 1.4–5.6 s · n = 3 and 12). Opus 5.5: median turn 1 vs turns 2–5: 1.9 s vs 2.4 s (ranges 1.8–4.4 s and 1.6–12.7 s · n = 3 and 12). Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- 7 of 9 cells passed 10/10. Haiku 4.5: 0/10 on the exact number (10 wrong, 1 distinct answer), 1/10 on JSON (9 format misses). Chart: Same prompt, 10 times · strict passes · a 10/10 is 72%–100% at 95% (n = 10 each). Caveat: Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
- Consistent is not correct: Haiku 4.5 gave the same wrong number all 10 times. Code fix, distinct correct bodies: Haiku 4.5 6, Sonnet 5.5 3, GPT-6.1 Sol (medium) 6. Chart: Same prompt, 10 times · distinct answers (n = 10 each). Caveat: The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
- Cache the fixed context; check answers, not agreement. Every call online.