Recorded study values
AI model efficiency: one study, two metrics.
Choose one study and two recorded metrics. Each point joins the same entity and exact configuration context. Different studies never share a plot.
These are aggregate values, not paired raw runs or a causal correlation. Each plot shows one entity kind. Model and CLI records can describe the same route and are not independent runs. Each axis keeps its own sample size and interval scope. Separate axis intervals are not a joint 95% confidence region. Ranges are not confidence intervals; missing spans are not invented. List-price calculations are not vendor bills. No composite score, frontier or rank is computed.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
Updated 2026-10-06. Read the study and its limits.
7 exact matching entity/context records; 0 missing or ambiguous; 0 unqualified source facts.
Linear axes. Hollow marks include a calculation. Whiskers keep each axis’s own recorded interval or range. No frontier or rank is calculated.
x: List-price cost per strict pass on hard tasks (calculation) (usd) · y: Pass rate on eight hard tasks (Strict pass) (rate)
Coordinate tick labels use the existing chart formatter; exact recorded displays, bounds and sample sizes are in the table.Coordinates and source spans
| Entity and context | List-price cost per strict pass on hard tasks (calculation) | Pass rate on eight hard tasks (Strict pass) | Join |
|---|---|---|---|
| Claude Sonnet 5.5Claude Code · eight hard validated tasks | $0.014CalculationNo interval or range recorded · n 24hard-h2h-cost-per-pass | 100% (24/24)95% interval (as recorded): 86%–100% · n 24hard-h2h-pass-rate | paired |
| Claude Opus 5.5Claude Code · eight hard validated tasks | $0.028CalculationNo interval or range recorded · n 24hard-h2h-cost-per-pass | 100% (24/24)95% interval (as recorded): 86%–100% · n 24hard-h2h-pass-rate | paired |
| Claude Opus 5.5Claude Code · effort high · eight hard validated tasks | $0.033CalculationNo interval or range recorded · n 24hard-h2h-cost-per-pass | 100% (24/24)95% interval (as recorded): 86%–100% · n 24hard-h2h-pass-rate | paired |
| Claude Haiku 4.5Claude Code · eight hard validated tasks | $0.067CalculationNo interval or range recorded · n 24hard-h2h-cost-per-pass | 46% (11/24)95% interval (as recorded): 28%–65% · n 24hard-h2h-pass-rate | paired |
| Claude Fable 5.1Claude Code · eight hard validated tasks | $0.093CalculationNo interval or range recorded · n 24hard-h2h-cost-per-pass | 100% (24/24)95% interval (as recorded): 86%–100% · n 24hard-h2h-pass-rate | paired |
| GPT-6.1 Sol (Codex CLI)Codex CLI · effort medium · eight hard validated tasks | $0.026CalculationNo interval or range recorded · n 16hard-h2h-cost-per-pass | 100% (16/16)95% interval (as recorded): 81%–100% · n 16hard-h2h-pass-rate | paired |
| GPT-6.1 Sol (Codex CLI)Codex CLI · effort high · eight hard validated tasks | $0.015CalculationNo interval or range recorded · n 16hard-h2h-cost-per-pass | 100% (16/16)95% interval (as recorded): 81%–100% · n 16hard-h2h-pass-rate | paired |
Original study caveats
- 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.
- Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
- Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
- The Claude and Codex batches ran on different days on the same host, one call at a time per account. Each route used its own subscription.
- Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.
- CLI timings include CLI start-up and the CLI’s own system prompt. One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled.
- Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.
- List-price costs are calculations; the calls used a flat subscription.