- Benchmarks
- Model picker
Only measured rows · 17 configurations · updated October 6, 2026
Which AI model should I use? Ask the data.
Pick your job. Then pick what matters after quality: cost or speed. The picker answers from rows we measured. It adds no score and makes no guess. When the data cannot separate two options, it says so.
$0.014
n = 24
Lowest cost per strict pass · Claude Sonnet 5.5 · Claude Code
What the data says
Lowest cost per strict pass: Claude Sonnet 5.5 · Claude Code, $0.014 (list-price calculation). 6 of 7 configurations have no comparison row that puts them behind on strict pass rate.
Sorted by cost per strict pass, lowest first. Configurations behind on strict pass rate come last.
6 of 7 configurations scored 100% on hard tasks with a strict reply format. This task set cannot separate those 6.
| Configuration | Rate | 95% interval | n | Behind |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 100% (24/24) | 86% to 100% | 24 | No row puts this configuration behind |
| GPT-6.1 Sol (high) · Codex CLI | 100% (16/16) | 81% to 100% | 16 | No row puts this configuration behind |
| GPT-6.1 Sol (medium) · Codex CLI | 100% (16/16) | 81% to 100% | 16 | No row puts this configuration behind |
| Claude Opus 5.5 · Claude Code | 100% (24/24) | 86% to 100% | 24 | No row puts this configuration behind |
| Claude Opus 5.5 (high) · Claude Code | 100% (24/24) | 86% to 100% | 24 | No row puts this configuration behind |
| Claude Fable 5.1 · Claude Code | 100% (24/24) | 86% to 100% | 24 | No row puts this configuration behind |
| Claude Haiku 4.5 · Claude Code | 46% (11/24) | 28% to 65% | 24 | Claude Sonnet 5.5 · Claude Code: The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%). |
7 rows. Highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 at 100%: this task set cannot separate them.
Strict pass. Sorted by cost per strict pass, lowest first. Configurations behind on strict pass rate come last.
Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the lowest value): a ratio of list-price calculations, not a measurement.
| Configuration | cost per strict pass | Range | n | Basis |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.014 | — | 24 | List-price calculation · Cost per strict pass |
| GPT-6.1 Sol (high) · Codex CLI | $0.015 | — | 16 | List-price calculation · Cost per strict pass |
| GPT-6.1 Sol (medium) · Codex CLI | $0.026 | — | 16 | List-price calculation · Cost per strict pass |
| Claude Opus 5.5 · Claude Code | $0.028 | — | 24 | List-price calculation · Cost per strict pass |
| Claude Opus 5.5 (high) · Claude Code | $0.033 | — | 24 | List-price calculation · Cost per strict pass |
| Claude Fable 5.1 · Claude Code | $0.093 | — | 24 | List-price calculation · Cost per strict pass |
| Claude Haiku 4.5 · Claude Code | $0.067 | — | 24 | List-price calculation · Cost per strict pass |
List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).
Notesn 16–24 per row
Sorted by cost per strict pass, lowest first. Configurations behind on strict pass rate come last.
Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.
Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
Why this answer
Each line comes from a dataset row or a count. Open the links to check it.
- 6 of 7 configurations scored 100% on hard tasks with a strict reply format. This task set cannot separate those 6.
- Claude Haiku 4.5 · Claude Code is behind: 46% (11/24) against Claude Sonnet 5.5 · Claude Code at 100% (24/24). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).
- Cost has no interval here. This order is for this run, not a tested ranking.
- Cost is a list-price calculation: reported tokens times list price. It is not a bill.
- Sample size: n = 16 to 24 calls per configuration.
Check the data
Questions and answers
These are the answers the picker gives with its first choices.
Which AI model should I use for hard coding tasks with a strict reply format?
Lowest cost per strict pass: Claude Sonnet 5.5 · Claude Code, $0.014 (list-price calculation). 6 of 7 configurations have no comparison row that puts them behind on strict pass rate. 6 of 7 configurations scored 100% on hard tasks with a strict reply format. This task set cannot separate those 6. Claude Haiku 4.5 · Claude Code is behind: 46% (11/24) against Claude Sonnet 5.5 · Claude Code at 100% (24/24). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).
Which AI model should I use for short coding tasks?
Lowest cost per pass: Claude Sonnet 5.5 · Claude Code, $0.0062 (list-price calculation). No comparison row puts any of the 9 configurations behind on pass rate. 8 of 9 configurations scored 100% on short tasks. This task set cannot separate those 8. Claude Sonnet 5.5 · Claude Code has the lowest cost but not the highest pass rate: 80% (12/15) against Claude Haiku 4.5 · Claude Code at 100% (15/15). The dataset row says: tie. The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Sonnet 5.5 55% to 93%), so this sample cannot separate them.
Which model and effort level should I use for hard coding tasks?
Lowest cost per strict pass: Claude Sonnet 5.5 (low) · Claude Code, $0.012 (list-price calculation). No comparison row puts any of the 11 configurations behind on strict pass rate. 11 of 11 configurations scored 100% on hard tasks at each effort level. This task set cannot separate those 11. Cost has no interval here. This order is for this run, not a tested ranking.
Which AI model should I use when I send the same prompt many times?
Highest strict pass rate on the "Exact number" prompt, tied at 100%: Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI. 2 of 3 configurations have no comparison row that puts them behind on strict pass rate on the "Exact number" prompt. Claude Haiku 4.5 · Claude Code is behind: 0% (0/10) against Claude Sonnet 5.5 · Claude Code at 100% (10/10). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 72% to 100%). This study has no cost chart, so the list follows strict pass rate on the "Exact number" prompt.
Which model or router should I use to route requests?
Lowest cost per 1,000 decisions: Jev 1.13 (TypeSafe), $0.034 (list-price calculation). No comparison row puts any of the 3 configurations behind on exact-decision rate. Jev 1.13 (TypeSafe) has the lowest cost but not the highest exact-decision rate: 90% against Claude Sonnet 5.5 at 94% (77/82). The dataset row says: tie. The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them. No comparison row names a winner on exact-decision rate, so quality does not rank these 3 configurations.
Which coding agent should I use on a real repository?
Lowest cost per passing session: Claude Sonnet 5.5 · Claude Code, $0.085 (list-price calculation). No comparison row puts any of the 3 configurations behind on hidden-check pass rate. 3 of 3 configurations scored 100% on small repository tasks graded by hidden tests. This task set cannot separate those 3. Cost has no interval here. This order is for this run, not a tested ranking.
How the picker decides
- Each job reads one study: its quality chart, and its cost and speed charts.
- A configuration is a model with its CLI and effort. We measured the pair, not the model alone.
- A configuration is behind only when a comparison row names it as the loser. The row says why.
- Cost or speed only orders the configurations that nothing puts behind.
- The order is for this run. It is not a tested ranking unless the page says so.
- There is no score and no weighted sum. Each value keeps its n and its interval or range.
What to keep in mind
- Each study ran on one host in one session. Times can differ on your machine.
- Samples are small, so intervals are wide. Every answer shows its n.
- Many jobs hit a ceiling: most configurations passed every call. Then quality cannot separate them.
- Cost is a list-price calculation: reported tokens times list price. The calls ran on flat subscriptions, so it is not a bill.
- The list holds the configurations we measured. A model that is not on it was not measured.