Explainer · Pareto frontier
The Pareto frontier: reading LLM cost against quality
Definition
Pareto frontier is the set of options that no rival matches or beats on every axis and beats on at least one. For large language models (LLMs), the axes are cost and quality, and sometimes speed. The frontier is your short list.
Agent team · · 5 min read · Every number is from the public studies
- Claude Code
- Codex CLI
Haloed: on the frontier (1 of 7). A point in the shaded area is no better on either axis than a haloed point.
| Point | Series | USD per strict pass (list-price calculation) | Strict pass rate | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | Claude Code | $0.014 | 100% | 24 |
| Claude Opus 5.5 · Claude Code | Claude Code | $0.028 | 100% | 24 |
| Claude Opus 5.5 (high) · Claude Code | Claude Code | $0.033 | 100% | 24 |
| Claude Fable 5.1 · Claude Code | Claude Code | $0.093 | 100% | 24 |
| Claude Haiku 4.5 · Claude Code | Claude Code | $0.067 | 46% | 24 |
| GPT-6.1 Sol (medium) · Codex CLI | Codex CLI | $0.026 | 100% | 16 |
| GPT-6.1 Sol (high) · Codex CLI | Codex CLI | $0.015 | 100% | 16 |
List-price calculation, not a run. 7 points: Strict pass rate against USD per strict pass (list-price calculation). USD per strict pass (list-price calculation) runs from $0.014 to $0.093; Strict pass rate from 46% to 100%. Highlighted: Claude Sonnet 5.5 · Claude Code.
Notesn 16–24 per point
Strict pass rate against list-price cost per strict pass
Upper-left is better. Highlighted points are on the frontier: no other configuration passes at least as often for at most the same cost per pass. Frontier: Claude Sonnet 5.5 · Claude Code. Costs are calculations from tokens. Pass rates with their 95% intervals are in the pass-rate chart.
Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
How to read a frontier chart
Each point is one configuration: a model, an effort setting and the CLI that ran it. Our hard-task chart puts the strict pass rate on the vertical axis and the list-price cost per strict pass on the horizontal axis. On the hard set, the whole trimmed reply must pass the validator; a code fence or added prose fails. Upper-left is better.
All costs below are calculations, not bills: reported tokens times list price, with cache reads and writes priced separately. We sum costs over all scored calls, failures included, then divide by the passes. The calls used flat subscriptions.
The hard set has 8 tasks: 24 calls per Claude Code configuration, 16 per Codex CLI configuration. It records 30 earlier attempts blocked before any model call; those are not scored. These are CLI-and-model pairs, not isolated model tests. Both studies used one host and separate batches; timing includes CLI start-up and context.
One of seven configurations is on the frontier (a calculation from observed pass rates and costs): Claude Sonnet 5.5 through Claude Code. It passed 24 of 24 calls (95% interval 86% to 100%) at $0.0143 per strict pass. Opus 5.5 and Fable 5.1 also passed 24 of 24 (95% interval 86% to 100% each), at $0.0282 and $0.0933. Haiku 4.5 passed 11 of 24 (46%, 95% interval 28% to 65%) at $0.0672. Each has a higher calculated cost per pass than Sonnet 5.5, and none has a higher observed pass rate.
This task set has a ceiling: six of seven configurations passed every call. Their observed pass rates tie, so calculated cost decides the frontier.
Why the frontier moves with the task set and the axes
The task set. The same two axes give a different frontier on different tasks. On the hard set they give 1 point. On the five short tasks they give 2 points (a calculation from cost per pass and pass rate). These are Sonnet 5.5 (12 of 15 passed, 95% interval 55% to 93%, $0.0062) and Opus 5.5 (low). The latter passed 15 of 15 (95% interval 80% to 100%) at $0.0083.
Sonnet missed 3 short-set calls, so a costlier option with every call passed stays on that frontier. In these two studies, Opus 5.5 (low) and GPT-6.1 Sol (low) ran only on the short tasks.
The axes. The short-task frontier below uses median time per call, cost per pass and pass rate. Its note lists the pass rates. Five of nine configurations are on it (a calculation from those point estimates), all through Claude Code: Fable 5.1, Sonnet 5.5, Opus 5.5 (high), Opus 5.5 and Opus 5.5 (low).
Fable 5.1 stays on this frontier with the lowest observed median time: 1.9 s (range 1.4 s to 9.8 s, n = 15 calls). It costs $0.0205 per pass (a calculation). It passed 15 of 15 (95% interval 80% to 100%). On cost and quality alone, the cheaper Opus 5.5 (low) dominates it.
A new axis cannot remove a point from the frontier. One exception: another point ties it exactly on the old axes and wins on the new one. Here, cost and quality alone give 2 points. Adding speed gives 5 (both counts are calculations from point estimates).
Why a point near the frontier can tie with one on it
A frontier uses point estimates. Our data has spread:
- Hard set, GPT-6.1 Sol (high) through Codex CLI. It passed 16 of 16 (95% interval 81% to 100%) at $0.0151 per pass. Sonnet 5.5 passed 24 of 24 (95% interval 86% to 100%) at $0.0143. The intervals overlap; they do not show a quality gap or prove equal underlying pass rates. The $0.0008 cost gap is a calculation with no interval, so it may not hold in a new run.
- Short set, Sonnet 5.5. Its 12 of 15 passes (80%, 95% interval 55% to 93%) overlap the perfect 15 of 15 interval (80% to 100%). All 3 misses gave the expected answer plus working lines, which the exact-text validator rejects.
- Short set, speed. The Fable 5.1 range overlaps the Sonnet 5.5 range (2.2 s to 7.7 s, median 2.3 s, 15 calls each). The ranges do not separate their speed. A range is not an interval. The Sonnet 5.5 against Fable 5.1 page finds 3 ties and 11 unclear rows.
Treat the frontier as a short list, not a ranking. One gap is clear: on the hard set, the Haiku 4.5 interval (28% to 65%) does not overlap the Sonnet 5.5 interval (86% to 100%). See Wilson confidence intervals for how we draw them.
How to use a frontier
- Pick the axes you pay for. Usually these are quality on your task, cost per correct answer, and time if a person waits.
- Drop each option that has a rival at least as good on every axis and better on one. The rest is your short list.
- Compare each gap with the spread. Overlapping 95% intervals do not separate the rates. Overlapping run ranges leave timing gaps unclear; neither proves equality.
- Test the short list on your own tasks. Our hard set hit a ceiling, so it says little about your hardest work.
- Redraw the chart when prices or models change. Our costs use list prices dated 2026-09-21 (Anthropic) and 2026-10-03 (OpenAI). The AI cost calculator prices your own volume.
Our leaderboard also keeps the axes apart, with no composite score. The checklist for reading benchmarks honestly covers the rest.
Frequently asked questions
What is a Pareto frontier in AI model selection?
It is the set of options that no rival matches or beats on every axis and beats on at least one. It depends on your task set and chosen axes.
Is the model on the frontier always the right choice?
No. The frontier describes observed point estimates. It cannot predict results on your tasks or establish a lead where intervals or ranges overlap.
Why is Sonnet on the frontier on hard tasks?
Its observed pass rate ties the other perfect configurations, at the lowest calculated cost per strict pass. The task set hit a ceiling.
Does the frontier include speed?
It can. Our short-task chart includes speed, cost and quality. The hard-task chart uses only cost and quality.
Watch the data
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
127 of 130 calls passed, so speed and tokens separate the models: Fable 5.1 was fastest at 1.9 s median.
Transcript
- Head-to-head · 130 timed calls · 9 configurations. Haiku vs Sonnet vs Opus vs Fable vs Codex. Five short tasks with strict validators. Every call kept, nothing retried.
- 127 of 130 calls passed. Pass rate barely separates them; speed and tokens do. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Cheapest passing answer (list-price calculation): Sonnet 5.5: $0.0062 (n = 15). Caveat: The tasks are short and easy; pass rate saturates. Latency and tokens carry the signal. A harder follow-up with eight tasks and strict validators: /benchmarks/hard-model-head-to-head.
- Fable 5.1 finishes first at 1.9 s. The Codex CLI needs 5.6–6.3 s. Chart: Median total time per call · real time (n = 10–15 each). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.
- What the CLI sends: a median 12,124 input tokens per call on the Codex CLI, 2,130 on Claude Code. Chart: Input tokens per call: what the CLI sends (n = 10–15 each). Caveat: The prompt cache stayed at the provider default, so cache counters differ by route and by call order.
- Speed vs cost per passing answer. Ringed: no other setup is faster, cheaper per pass and as accurate. Chart: Speed, cost and quality frontier (n = 10–15 each). Calculation, not a run. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.
- Open benchmarks: intervals, sources and every failure kept.