Would majority voting fix it? A self-consistency thought experiment on 90 real calls
Calculation on 90 recorded calls: a vote over 10 Haiku replies returns 289, not 282. Repeated errors and format misses limit self-consistency voting.
TL;DR
- A vote cannot fix a mistake that every recorded reply repeats. Claude Haiku 4.5 gave 289 in all 10 exact-number replies. It passed 0/10 (95% interval 0% to 28%). The answer is 282. A vote over those 10 numbers returns 289 (calculation).
- That vote uses 10 times the calls (calculation). The 10 Haiku calls took a median 5.06 s (range 4.42 to 6.20 s). Projected time is about 51 s in sequence or 6.2 s in parallel. Both are calculations, not measured voting times.
- JSON prompt: 9 of 10 Haiku replies had the right content in the wrong format. Strict passes: 1/10 (95% interval 2% to 40%). A raw-text vote picks a failing reply (calculation). Check the format.
- Code fix: all 30 calls passed: 10/10 per configuration (95% interval 72% to 100% each). Each had 3 to 6 distinct normalized bodies. Run the tests.
- Sonnet 5.5 and GPT-6.1 Sol passed every repeat. Each prompt had 10/10 passes (95% interval 72% to 100%). A vote adds no observed accuracy gain (calculation).
- This is a thought experiment on 90 existing calls. We ran no vote.
Our same-prompt study found that consistent is not correct. Would a majority vote fix that?
Prompt caching and consistency: what the cache saves, and how much answers vary
Calculation at list price: the cache cut a 5-question session 50% on Sonnet 5.5 and 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every repetition.
Transcript
- Caching and consistency · 135 calls. What the cache saves, and how much answers vary. Five-question sessions on a fixed context. Then the same prompt, 10 times.
- 135 calls: 45 cache turns, 90 repeated prompts. Every call counted. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- Claude Code turn 1 reads 19% from the cache, the CLI’s own prefix. Turns 2–5 read 90%–99%. Codex CLI: 98%–99%. Chart: Share of input read from the cache · mean of 3 sessions per turn (n = 3 each). Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- A calculation: Sonnet 5.5 $0.1350 with the cache vs $0.2698 without, 50% less. Opus 5.5 $0.2551 vs $0.5442, 53% less. Chart: List-price cost of the recorded sessions · calculation (n = 15 each). Calculation, not a run. Caveat: Costs are list-price calculations; the calls used flat subscriptions.
- No clear speed effect: Sonnet 5.5 1.6 s on turn 1 vs 1.6 s later; Opus 5.5 1.9 s vs 2.4 s. The ranges overlap. Sonnet 5.5: median turn 1 vs turns 2–5: 1.6 s vs 1.6 s (ranges 1.6–1.8 s and 1.4–5.6 s · n = 3 and 12). Opus 5.5: median turn 1 vs turns 2–5: 1.9 s vs 2.4 s (ranges 1.8–4.4 s and 1.6–12.7 s · n = 3 and 12). Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.
- 7 of 9 cells passed 10/10. Haiku 4.5: 0/10 on the exact number (10 wrong, 1 distinct answer), 1/10 on JSON (9 format misses). Chart: Same prompt, 10 times · strict passes · a 10/10 is 72%–100% at 95% (n = 10 each). Caveat: Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.
- Consistent is not correct: Haiku 4.5 gave the same wrong number all 10 times. Code fix, distinct correct bodies: Haiku 4.5 6, Sonnet 5.5 3, GPT-6.1 Sol (medium) 6. Chart: Same prompt, 10 times · distinct answers (n = 10 each). Caveat: The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.
- Cache the fixed context; check answers, not agreement. Every call online.
What self-consistency voting assumes
Here, voting means keeping the most common answer from repeated calls. A strict majority needs more than half the votes. The largest group alone is a plurality.
Self-consistency prompting samples different reasoning paths, then votes on their final answers. These existing calls repeated fixed prompts at each CLI's default sampling settings. We did not test that full method or tune sampling for diverse reasoning.
Voting can help when correct answers outnumber each competing wrong answer. Independent errors do not guarantee that condition, and independence is not required. These repeats do not measure error independence.
The repeated-prompt part of the caching and consistency study made 90 calls: 3 prompts, 3 configurations, 10 repeats each. We can group each cell's 10 recorded replies as a hypothetical ballot.
Case 1: the same wrong answer, 10 times
Haiku 4.5 passed the exact-number prompt 0 of 10 times (95% interval 0% to 28%). It gave 10 wrong answers, 0 format misses and 1 distinct answer: 289. The expected answer is 282.
The 10 replies used 10 distinct raw texts and one normalized number. A vote counts 10 for 289 and 0 for 282 (calculation). Any non-empty subset of these recorded numbers also returns 289. This does not predict every future call.
A vote can reduce some scattered errors. It cannot correct a number that every reply in this ballot gets wrong.
What 10 votes cost
The consistency cells have no per-call price, so we count calls and seconds.
| Exact-number prompt (n = 10 each) | One call: median (range) | 10 calls, in sequence (calculation) | 10 calls, in parallel (calculation) |
|---|---|---|---|
| Haiku 4.5 · Claude Code | 5.06 s (4.42 to 6.20) | about 51 s | about 6.2 s |
| Sonnet 5.5 · Claude Code | 6.89 s (5.81 to 7.81) | about 69 s | about 7.8 s |
| GPT-6.1 Sol (medium) · Codex CLI | 13.38 s (12.29 to 17.97) | about 134 s | about 18.0 s |
Calculation. The sequence projection is 10 times the median, not the sum of the recorded calls. The parallel projection uses the slowest recorded call. It assumes simultaneous starts and unchanged call times. Neither includes vote processing or extra queuing time.
We ran one call at a time per account. The observed ranges are not confidence intervals.
Sonnet and GPT-6.1 Sol passed 10 of 10 on every prompt (95% interval 72% to 100% each). Their exact-number vote returns 282, as each recorded call did. It uses 9 extra calls per final answer with no observed accuracy gain (calculation). These cells hit a ceiling; they cannot show that voting never helps on harder prompts.
Case 2: the JSON prompt needs a format fix
Haiku passed the JSON prompt 1 of 10 times (95% interval 2% to 40%). It gave 9 format misses, 0 wrong answers. All 10 held the same JSON content: 1 distinct answer.
There are two ways to vote:
- On the normalized answer. The vote selects the right JSON content (calculation). A pipeline still needs a valid reply. Extraction and serialization could fix the format, but voting alone does not do that.
- On the raw text. There are 3 distinct raw replies. One failing reply has 8 votes, the passing reply has 1 and a second failing reply has 1. The vote returns a reply that the validator rejects (a calculation from the published receipts).
The fix belongs in the format step. In our data, a format miss means a lenient extractor finds an answer that passes the validator. We tested no format fix.
Case 3: the code fix needs tests
All 30 code-fix calls passed (10/10 per configuration, 95% interval 72% to 100% each). These cells hit a ceiling. The normalized code differed: Haiku 6, Sonnet 3, GPT-6.1 Sol 6 distinct bodies.
The largest group of normalized bodies has 5 votes for Haiku, 8 for Sonnet and 4 for GPT-6.1 Sol (a calculation from the receipts). A strict majority needs 6, so only Sonnet has one. Every body passes the 10 checks anyway. Run the checks instead of voting on text.
The hard set: a vote cannot invent an answer
On the 8 hard tasks, Haiku passed 11 of 24 (46%, 95% interval 28% to 65%). It passed 0/3 on three tasks (95% interval 0% to 56% each). On event-loop order, no reply was a format miss, so all 3 were wrong. A vote that only selects a recorded answer cannot invent a correct one. On the room schedule and the SQLite query, 2 of 3 replies in each task were format misses. Each has a 95% interval of 21% to 94%.
Sonnet passed 24/24 (95% interval 86% to 100%). GPT-6.1 Sol passed 16/16 at each of medium and high (81% to 100% each). These configurations hit a ceiling, and their intervals overlap. A vote adds no observed accuracy gain; this set cannot rank them or predict harder work.
When voting helps, and what to do instead
Voting can help when the correct answer is the most common. For code, two passing bodies can differ and split the vote. Haiku passed 2/3 on the refactor task (95% interval 21% to 94%). That pass count does not show which code body would win a text vote.
Repeat your prompt and count distinct answers as an initial check. One distinct wrong answer means voting cannot fix that recorded ballot. More calls or different sampling could produce different answers.
- Check each answer against the right answer. A test beats agreement between runs.
- Count format misses apart from wrong answers. A vote does not fix a format miss.
- Move a failing prompt to a stronger model. On the exact-number prompt, Sonnet passed 10/10 (95% interval 72% to 100%), one call per try. Haiku passed 0/10 (0% to 28%). The intervals do not overlap on this prompt.
How we measured
- Data: the repeated-prompt part of the caching and consistency study: Haiku 4.5 and Sonnet 5.5 in Claude Code, GPT-6.1 Sol (medium) in Codex CLI. One-shot calls, tools off, nothing retried or trimmed.
- Votes: calculations group the recorded replies by normalized answer id or raw-reply id. The largest group is not always a strict majority. Normalization uses the last number, key-sorted JSON, or code without fences, comments, whitespace, quote style and line-end semicolons.
- Sources: the study's
consistency-cellstable (passes, intervals, times), published consistency receipts (group sizes), and the hard study'shard-h2h-pass-matrixtable. The receipts record 10 checks per code-fix reply.
Caveats
- n = 10 per cell, 3 prompts, one day (2026-10-06). A 0/10 has a 95% interval of 0% to 28%; a 10/10, 72% to 100%.
- One cell shows that a vote can fail. It does not show that voting fails in general.
- One setting per model: Claude at its default effort, GPT-6.1 Sol at medium. Other settings may spread the wrong answers. We did not test them.
- Parallel time is an untested projection for 10 calls at once. Concurrency limits and server load could change it.
- Routes differ: Claude Code and Codex CLI add their own context and start-up time. Cross-route rows compare a model plus CLI.
- No new voting run: this calculation does not measure voting accuracy, independence, or improvements from diverse reasoning samples.
What to read next
- Same prompt, ten answers: how consistent are Claude and Codex?
- When tasks get hard: Haiku vs Sonnet vs Opus vs Fable
- How to read AI benchmarks honestly
Check answers, not votes
Agent publishes the validator results next to the model and the time. Explore the study, then check whether your own prompts repeat a wrong answer.