Only measured rows · 17 configurations · updated October 6, 2026

Which AI model should I use? Ask the data.

Pick your job. Then pick what matters after quality: cost or speed. The picker answers from rows we measured. It adds no score and makes no guess. When the data cannot separate two options, it says so.

1. What is the job?

Hard tasks with strict checks. A reply in the wrong format is not a pass.

2. After quality, what matters most?
3. Which route do you use?

A route is the CLI that ran the model.

Which replies count as a pass?

The reply must match the format exactly.

$0.014

n = 24

Lowest cost per strict pass · Claude Sonnet 5.5 · Claude Code

What the data says

Lowest cost per strict pass: Claude Sonnet 5.5 · Claude Code, $0.014 (list-price calculation). 6 of 7 configurations have no comparison row that puts them behind on strict pass rate.

Sorted by cost per strict pass, lowest first. Configurations behind on strict pass rate come last.

6 of 7 configurations scored 100% on hard tasks with a strict reply format. This task set cannot separate those 6.

0%100%
Claude Sonnet 5.5 · Claude Code100%n = 24
GPT-6.1 Sol (high) · Codex CLI100%n = 16
GPT-6.1 Sol (medium) · Codex CLI100%n = 16
Claude Opus 5.5 · Claude Code100%n = 24
Claude Opus 5.5 (high) · Claude Code100%n = 24
Claude Fable 5.1 · Claude Code100%n = 24
Claude Haiku 4.5 · Claude Code46%n = 24

7 rows. Highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 at 100%: this task set cannot separate them.

Strict pass. Sorted by cost per strict pass, lowest first. Configurations behind on strict pass rate come last.

Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Calculation
Claude Sonnet 5.5 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
Claude Fable 5.1 · Claude Code
Claude Haiku 4.5 · Claude Code

Hover or focus a bar for its ratio to Claude Sonnet 5.5 (the lowest value): a ratio of list-price calculations, not a measurement.

List-price calculation, not a run. 7 rows. Highest Claude Fable 5.1 · Claude Code $0.093 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).

Notesn 16–24 per row

Sorted by cost per strict pass, lowest first. Configurations behind on strict pass rate come last.

Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.

Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices

Why this answer

Each line comes from a dataset row or a count. Open the links to check it.

  • 6 of 7 configurations scored 100% on hard tasks with a strict reply format. This task set cannot separate those 6.
  • Claude Haiku 4.5 · Claude Code is behind: 46% (11/24) against Claude Sonnet 5.5 · Claude Code at 100% (24/24). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).
  • Cost has no interval here. This order is for this run, not a tested ranking.
  • Cost is a list-price calculation: reported tokens times list price. It is not a bill.
  • Sample size: n = 16 to 24 calls per configuration.

Check the data

Questions and answers

These are the answers the picker gives with its first choices.

Which AI model should I use for hard coding tasks with a strict reply format?

Lowest cost per strict pass: Claude Sonnet 5.5 · Claude Code, $0.014 (list-price calculation). 6 of 7 configurations have no comparison row that puts them behind on strict pass rate. 6 of 7 configurations scored 100% on hard tasks with a strict reply format. This task set cannot separate those 6. Claude Haiku 4.5 · Claude Code is behind: 46% (11/24) against Claude Sonnet 5.5 · Claude Code at 100% (24/24). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).

Which AI model should I use for short coding tasks?

Lowest cost per pass: Claude Sonnet 5.5 · Claude Code, $0.0062 (list-price calculation). No comparison row puts any of the 9 configurations behind on pass rate. 8 of 9 configurations scored 100% on short tasks. This task set cannot separate those 8. Claude Sonnet 5.5 · Claude Code has the lowest cost but not the highest pass rate: 80% (12/15) against Claude Haiku 4.5 · Claude Code at 100% (15/15). The dataset row says: tie. The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Sonnet 5.5 55% to 93%), so this sample cannot separate them.

Which model and effort level should I use for hard coding tasks?

Lowest cost per strict pass: Claude Sonnet 5.5 (low) · Claude Code, $0.012 (list-price calculation). No comparison row puts any of the 11 configurations behind on strict pass rate. 11 of 11 configurations scored 100% on hard tasks at each effort level. This task set cannot separate those 11. Cost has no interval here. This order is for this run, not a tested ranking.

Which AI model should I use when I send the same prompt many times?

Highest strict pass rate on the "Exact number" prompt, tied at 100%: Claude Sonnet 5.5 · Claude Code and GPT-6.1 Sol (medium) · Codex CLI. 2 of 3 configurations have no comparison row that puts them behind on strict pass rate on the "Exact number" prompt. Claude Haiku 4.5 · Claude Code is behind: 0% (0/10) against Claude Sonnet 5.5 · Claude Code at 100% (10/10). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 72% to 100%). This study has no cost chart, so the list follows strict pass rate on the "Exact number" prompt.

Which model or router should I use to route requests?

Lowest cost per 1,000 decisions: Jev 1.13 (TypeSafe), $0.034 (list-price calculation). No comparison row puts any of the 3 configurations behind on exact-decision rate. Jev 1.13 (TypeSafe) has the lowest cost but not the highest exact-decision rate: 90% against Claude Sonnet 5.5 at 94% (77/82). The dataset row says: tie. The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them. No comparison row names a winner on exact-decision rate, so quality does not rank these 3 configurations.

Which coding agent should I use on a real repository?

Lowest cost per passing session: Claude Sonnet 5.5 · Claude Code, $0.085 (list-price calculation). No comparison row puts any of the 3 configurations behind on hidden-check pass rate. 3 of 3 configurations scored 100% on small repository tasks graded by hidden tests. This task set cannot separate those 3. Cost has no interval here. This order is for this run, not a tested ranking.

How the picker decides

  • Each job reads one study: its quality chart, and its cost and speed charts.
  • A configuration is a model with its CLI and effort. We measured the pair, not the model alone.
  • A configuration is behind only when a comparison row names it as the loser. The row says why.
  • Cost or speed only orders the configurations that nothing puts behind.
  • The order is for this run. It is not a tested ranking unless the page says so.
  • There is no score and no weighted sum. Each value keeps its n and its interval or range.

What to keep in mind

  • Each study ran on one host in one session. Times can differ on your machine.
  • Samples are small, so intervals are wide. Every answer shows its n.
  • Many jobs hit a ceiling: most configurations passed every call. Then quality cannot separate them.
  • Cost is a list-price calculation: reported tokens times list price. The calls ran on flat subscriptions, so it is not a bill.
  • The list holds the configurations we measured. A model that is not on it was not measured.

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.