Model · Anthropic
Claude Fable 5
Anthropic model, measured as a blind code-review critic.
1 values from 1 study · Updated
At a glance
The best-supported value per category: a 95% interval first, then a run range, then the larger n. Three separate values from separate studies, never one score.
Qualitypass rates, accuracy and scores
65% (26/40)
95% CI 50%–78% · n = 40
Does the judge’s model family matter?
blind review panel · AI pull requests vs merged human pull requests, judged blind
Speedtime per call or decision
Not measured in any study yet.
CostUS dollars per call, pass or decision
Not measured in any study yet.
Where it sits
Every measured value, grouped by study. Each row puts the value on its own track, with the other configurations of the same chart as muted dots. A range is the fastest to slowest recorded run and p50–p95 is the median to the 95th percentile; neither is a confidence interval. Use Table for the plain values.
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Does the judge’s model family matter? | 65% (26/40) | 40 | 50%–78% (95% CI) | blind review panel |
Whiskers: 95% Wilson intervaln beside each valueMuted dots: the other configurations on the same chart
Claude Fable 5 in AI pull requests vs merged human pull requests, judged blind: 1 value, first Does the judge’s model family matter? 65% (26/40).
Watch
AI pull requests vs merged human pull requests, judged blind
Blind critics preferred the AI change on 9 of 12 tasks at the latest attempt and 6 of 12 at the first. Same-family bias disclosed.
Transcript
- Blind review · 12 real merged pull requests. AI change vs the merged human change. Critic models review both, unlabelled, once in each order.
- Latest attempt: the panel preferred the AI change on 9 of 12 tasks. First attempt: 6 of 12. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Latest attempts are not independent first tries: they had lessons, review replays and some operator answers from earlier attempts on the same task.
- The first attempt is the cleaner estimate: 50%, with a 95% interval from 25% to 75%. Chart: First attempt vs latest attempt (n = 5–12 each). Caveat: n = 12 tasks: intervals are wide.
- Vote by vote: 9 of 12 tasks went unanimously to the AI change, 2 unanimously to the human one. Chart: Blind panel votes per task (latest attempt) (n = 6–8 each). Caveat: 7 of the 12 tasks come from one private Go service. They carry neutral labels (private-go-a and so on), and only their language and kind are published.
- Cross-family check: OpenAI critics preferred the AI change in 10 of 12 verdicts. Small n, wide intervals. Chart: Does the judge’s model family matter? (n = 2–40 each). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
- Open benchmarks: intervals, sources and every failure kept.
Write-ups that use this data
All postsAI coding benchmarks roundup, October 2026: sixteen studies, every number in one place
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
Do blind AI critics prefer AI pull requests over human ones?
Blind Claude and GPT critics preferred an AI worker's change over the merged human change on 9 of 12 real tasks, and 6 of 12 on the first try. Caveats inside.
Harness vs model: where do AI coding agent gains really come from?
Is it the model or the harness? Our data shows the harness clearly moves speed, tokens and cost. Whether it moves accuracy, our samples cannot yet say.