Model · OpenAI

GPT 5.5

OpenAI model, measured as a blind code-review critic on a small number of pairs.

1 values from 1 study · Updated

At a glance

The best-supported value per category: a 95% interval first, then a run range, then the larger n. Three separate values from separate studies, never one score.

Qualitypass rates, accuracy and scores

100% (2/2)

95% CI 34%–100% · n = 2

Does the judge’s model family matter?

blind review panel · AI pull requests vs merged human pull requests, judged blind

Speedtime per call or decision

Not measured in any study yet.

CostUS dollars per call, pass or decision

Not measured in any study yet.

Where it sits

Every measured value, grouped by study. Each row puts the value on its own track, with the other configurations of the same chart as muted dots. A range is the fastest to slowest recorded run and p50–p95 is the median to the 95th percentile; neither is a confidence interval. Use Table for the plain values.

Does the judge’s model family matter?n = 2 · 95% CI 34%–100% · blind review panel
100% (2/2)

Whiskers: 95% Wilson intervaln beside each valueMuted dots: the other configurations on the same chart

GPT 5.5 in AI pull requests vs merged human pull requests, judged blind: 1 value, first Does the judge’s model family matter? 100% (2/2).

Watch

Live story · 34 sAI pull requests vs merged human pull requests, judged blind

AI pull requests vs merged human pull requests, judged blind

Blind critics preferred the AI change on 9 of 12 tasks at the latest attempt and 6 of 12 at the first. Same-family bias disclosed.

Transcript
  1. Blind review · 12 real merged pull requests. AI change vs the merged human change. Critic models review both, unlabelled, once in each order.
  2. Latest attempt: the panel preferred the AI change on 9 of 12 tasks. First attempt: 6 of 12. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Latest attempts are not independent first tries: they had lessons, review replays and some operator answers from earlier attempts on the same task.
  3. The first attempt is the cleaner estimate: 50%, with a 95% interval from 25% to 75%. Chart: First attempt vs latest attempt (n = 5–12 each). Caveat: n = 12 tasks: intervals are wide.
  4. Vote by vote: 9 of 12 tasks went unanimously to the AI change, 2 unanimously to the human one. Chart: Blind panel votes per task (latest attempt) (n = 6–8 each). Caveat: 7 of the 12 tasks come from one private Go service. They carry neutral labels (private-go-a and so on), and only their language and kind are published.
  5. Cross-family check: OpenAI critics preferred the AI change in 10 of 12 verdicts. Small n, wide intervals. Chart: Does the judge’s model family matter? (n = 2–40 each). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
  6. Open benchmarks: intervals, sources and every failure kept.

Write-ups that use this data

All posts

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.