Coding CLI · Anthropic
Claude Code
Anthropic’s coding CLI. Each measurement pairs it with one Claude model; the context names the model.
412 values from 16 studies (129 are list-price calculations) · Updated
At a glance
The best-supported value per category: a 95% interval first, then a run range, then the larger n. Three separate values from separate studies, never one score.
Qualitypass rates, accuracy and scores
97% (189/194)
95% CI 94%–99% · n = 194
Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy)
Claude Sonnet 5.5 · effort low · typed routing decisions, thinking on vs off · Does thinking pay for Claude Haiku 4.5? Thinking on vs off
- 92% (115/125)Per-question accuracy on unseen decisions
- 100% (24/24)Pass rate on eight hard tasks (Strict pass)
Speedtime per call or decision
2.60s
median to p95 2.6 s–4.3 s · n = 82
Haiku thinking study: time per routing decision (Wall time (CLI))
Claude Sonnet 5.5 · effort low · typed routing decisions, thinking on vs off · Does thinking pay for Claude Haiku 4.5? Thinking on vs off
CostUS dollars per call, pass or decision
$0.0036
range $0.0034–$0.01 · n = 15 · list-price calculation
List-price cost per call (calculation)
Claude Sonnet 5.5 · five short validated tasks · Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
Where it sits
Every measured value, grouped by study. Each row puts the value on its own track, with the other configurations of the same chart as muted dots. A range is the fastest to slowest recorded run and p50–p95 is the median to the 95th percentile; neither is a confidence interval. Use Table for the plain values.
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Pass rate on five validated tasks | 100% (15/15) | 15 | 80%–100% (95% CI) | Claude Fable 5.1 · five short validated tasks |
| Pass rate on five validated tasks | 80% (12/15) | 15 | 55%–93% (95% CI) | Claude Sonnet 5.5 · five short validated tasks |
| Pass rate on five validated tasks | 100% (15/15) | 15 | 80%–100% (95% CI) | Claude Opus 5.5 · effort high · five short validated tasks |
| Pass rate on five validated tasks | 100% (15/15) | 15 | 80%–100% (95% CI) | Claude Opus 5.5 · five short validated tasks |
| Pass rate on five validated tasks | 100% (15/15) | 15 | 80%–100% (95% CI) | Claude Opus 5.5 · effort low · five short validated tasks |
| Pass rate on five validated tasks | 100% (15/15) | 15 | 80%–100% (95% CI) | Claude Haiku 4.5 · five short validated tasks |
| Total time per call | 1.94 s | 15 | 1.4 s–9.8 s (range) | Claude Fable 5.1 · five short validated tasks |
| Total time per call | 2.31 s | 15 | 2.2 s–7.7 s (range) | Claude Sonnet 5.5 · five short validated tasks |
| Total time per call | 2.71 s | 15 | 2.5 s–11.8 s (range) | Claude Opus 5.5 · effort high · five short validated tasks |
| Total time per call | 2.75 s | 15 | 2.5 s–8.9 s (range) | Claude Opus 5.5 · five short validated tasks |
| Total time per call | 2.83 s | 15 | 2.4 s–6.6 s (range) | Claude Opus 5.5 · effort low · five short validated tasks |
| Total time per call | 4.43 s | 15 | 3.2 s–23.6 s (range) | Claude Haiku 4.5 · five short validated tasks |
| Time to first useful output | 1.20 s | 15 | 1 s–7.9 s (range) | Claude Fable 5.1 · five short validated tasks |
| Time to first useful output | 1.56 s | 15 | 1 s–6.4 s (range) | Claude Sonnet 5.5 · five short validated tasks |
| Time to first useful output | 2.04 s | 15 | 1.4 s–9.9 s (range) | Claude Opus 5.5 · effort high · five short validated tasks |
| Time to first useful output | 1.92 s | 15 | 1.6 s–7.2 s (range) | Claude Opus 5.5 · five short validated tasks |
| Time to first useful output | 2.39 s | 15 | 1.5 s–4.9 s (range) | Claude Opus 5.5 · effort low · five short validated tasks |
| Time to first useful output | 3.63 s | 15 | 2.8 s–22.3 s (range) | Claude Haiku 4.5 · five short validated tasks |
| Input tokens per call: what the CLI sends (Cache read) | 2,760 | 15 | — | Claude Fable 5.1 · five short validated tasks |
| Input tokens per call: what the CLI sends (Cache read) | 1,401 | 15 | — | Claude Sonnet 5.5 · five short validated tasks |
| Input tokens per call: what the CLI sends (Cache read) | 1,463 | 15 | — | Claude Opus 5.5 · effort high · five short validated tasks |
| Input tokens per call: what the CLI sends (Cache read) | 1,401 | 15 | — | Claude Opus 5.5 · five short validated tasks |
| Input tokens per call: what the CLI sends (Cache read) | 1,463 | 15 | — | Claude Opus 5.5 · effort low · five short validated tasks |
| Input tokens per call: what the CLI sends (Cache read) | 0 | 15 | — | Claude Haiku 4.5 · five short validated tasks |
| Input tokens per call: what the CLI sends (Other input) | 473 | 15 | — | Claude Fable 5.1 · five short validated tasks |
| Input tokens per call: what the CLI sends (Other input) | 685 | 15 | — | Claude Sonnet 5.5 · five short validated tasks |
| Input tokens per call: what the CLI sends (Other input) | 619 | 15 | — | Claude Opus 5.5 · effort high · five short validated tasks |
| Input tokens per call: what the CLI sends (Other input) | 680 | 15 | — | Claude Opus 5.5 · five short validated tasks |
| Input tokens per call: what the CLI sends (Other input) | 618 | 15 | — | Claude Opus 5.5 · effort low · five short validated tasks |
| Input tokens per call: what the CLI sends (Other input) | 3,790 | 15 | — | Claude Haiku 4.5 · five short validated tasks |
| Output tokens per call (Output tokens) | 64 | 15 | — | Claude Fable 5.1 · five short validated tasks |
| Output tokens per call (Output tokens) | 107 | 15 | — | Claude Sonnet 5.5 · five short validated tasks |
| Output tokens per call (Output tokens) | 78 | 15 | — | Claude Opus 5.5 · effort high · five short validated tasks |
| Output tokens per call (Output tokens) | 64 | 15 | — | Claude Opus 5.5 · five short validated tasks |
| Output tokens per call (Output tokens) | 64 | 15 | — | Claude Opus 5.5 · effort low · five short validated tasks |
| Output tokens per call (Output tokens) | 367 | 15 | — | Claude Haiku 4.5 · five short validated tasks |
| List-price cost per call (calculation) Calculation | $0.0099 | 15 | $0.0049–$0.058 (range) | Claude Fable 5.1 · five short validated tasks |
| List-price cost per call (calculation) Calculation | $0.0036 | 15 | $0.0034–$0.01 (range) | Claude Sonnet 5.5 · five short validated tasks |
| List-price cost per call (calculation) Calculation | $0.0069 | 15 | $0.0059–$0.027 (range) | Claude Opus 5.5 · effort high · five short validated tasks |
| List-price cost per call (calculation) Calculation | $0.0069 | 15 | $0.0059–$0.022 (range) | Claude Opus 5.5 · five short validated tasks |
| List-price cost per call (calculation) Calculation | $0.0069 | 15 | $0.0058–$0.018 (range) | Claude Opus 5.5 · effort low · five short validated tasks |
| List-price cost per call (calculation) Calculation | $0.0057 | 15 | $0.0051–$0.018 (range) | Claude Haiku 4.5 · five short validated tasks |
| List-price cost per passing answer (calculation) Calculation | $0.0062 | 15 | — | Claude Sonnet 5.5 · five short validated tasks |
| List-price cost per passing answer (calculation) Calculation | $0.0083 | 15 | — | Claude Opus 5.5 · effort low · five short validated tasks |
| List-price cost per passing answer (calculation) Calculation | $0.0084 | 15 | — | Claude Haiku 4.5 · five short validated tasks |
| List-price cost per passing answer (calculation) Calculation | $0.010 | 15 | — | Claude Opus 5.5 · five short validated tasks |
| List-price cost per passing answer (calculation) Calculation | $0.010 | 15 | — | Claude Opus 5.5 · effort high · five short validated tasks |
| List-price cost per passing answer (calculation) Calculation | $0.021 | 15 | — | Claude Fable 5.1 · five short validated tasks |
Whiskers: 95% Wilson intervalLines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Claude Code in Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head: 48 values, first Pass rate on five validated tasks 100% (15/15).
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Pass rate on eight hard tasks (Strict pass) | 100% (24/24) | 24 | 86%–100% (95% CI) | Claude Sonnet 5.5 · eight hard validated tasks |
| Pass rate on eight hard tasks (Strict pass) | 100% (24/24) | 24 | 86%–100% (95% CI) | Claude Opus 5.5 · eight hard validated tasks |
| Pass rate on eight hard tasks (Strict pass) | 100% (24/24) | 24 | 86%–100% (95% CI) | Claude Opus 5.5 · effort high · eight hard validated tasks |
| Pass rate on eight hard tasks (Strict pass) | 100% (24/24) | 24 | 86%–100% (95% CI) | Claude Fable 5.1 · eight hard validated tasks |
| Pass rate on eight hard tasks (Strict pass) | 46% (11/24) | 24 | 28%–65% (95% CI) | Claude Haiku 4.5 · eight hard validated tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 100% (24/24) | 24 | 86%–100% (95% CI) | Claude Sonnet 5.5 · eight hard validated tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 100% (24/24) | 24 | 86%–100% (95% CI) | Claude Opus 5.5 · eight hard validated tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 100% (24/24) | 24 | 86%–100% (95% CI) | Claude Opus 5.5 · effort high · eight hard validated tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 100% (24/24) | 24 | 86%–100% (95% CI) | Claude Fable 5.1 · eight hard validated tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 67% (16/24) | 24 | 47%–82% (95% CI) | Claude Haiku 4.5 · eight hard validated tasks |
| Total time per call on hard tasks (separate batches) | 7.75 s | 24 | 2.3 s–34.8 s (range) | Claude Sonnet 5.5 · eight hard validated tasks |
| Total time per call on hard tasks (separate batches) | 9.18 s | 24 | 4.2 s–27.2 s (range) | Claude Opus 5.5 · eight hard validated tasks |
| Total time per call on hard tasks (separate batches) | 11.0 s | 24 | 3.6 s–63 s (range) | Claude Opus 5.5 · effort high · eight hard validated tasks |
| Total time per call on hard tasks (separate batches) | 16.1 s | 24 | 4.5 s–90 s (range) | Claude Fable 5.1 · eight hard validated tasks |
| Total time per call on hard tasks (separate batches) | 39.0 s | 24 | 15.3 s–75.1 s (range) | Claude Haiku 4.5 · eight hard validated tasks |
| Time to first useful output on hard tasks | 5.95 s | 24 | 0.9 s–30.6 s (range) | Claude Sonnet 5.5 · eight hard validated tasks |
| Time to first useful output on hard tasks | 6.78 s | 24 | 2.4 s–21.8 s (range) | Claude Opus 5.5 · eight hard validated tasks |
| Time to first useful output on hard tasks | 7.13 s | 24 | 2.2 s–56.2 s (range) | Claude Opus 5.5 · effort high · eight hard validated tasks |
| Time to first useful output on hard tasks | 11.6 s | 24 | 2 s–85.3 s (range) | Claude Fable 5.1 · eight hard validated tasks |
| Time to first useful output on hard tasks | 35.5 s | 24 | 12.9 s–70.3 s (range) | Claude Haiku 4.5 · eight hard validated tasks |
| Output tokens per call on hard tasks (Output tokens) | 1,050 | 24 | — | Claude Sonnet 5.5 · eight hard validated tasks |
| Output tokens per call on hard tasks (Output tokens) | 945 | 24 | — | Claude Opus 5.5 · eight hard validated tasks |
| Output tokens per call on hard tasks (Output tokens) | 1,052 | 24 | — | Claude Opus 5.5 · effort high · eight hard validated tasks |
| Output tokens per call on hard tasks (Output tokens) | 1,366 | 24 | — | Claude Fable 5.1 · eight hard validated tasks |
| Output tokens per call on hard tasks (Output tokens) | 5,064 | 24 | — | Claude Haiku 4.5 · eight hard validated tasks |
| List-price cost per strict pass on hard tasks (calculation) Calculation | $0.014 | 24 | — | Claude Sonnet 5.5 · eight hard validated tasks |
| List-price cost per strict pass on hard tasks (calculation) Calculation | $0.028 | 24 | — | Claude Opus 5.5 · eight hard validated tasks |
| List-price cost per strict pass on hard tasks (calculation) Calculation | $0.033 | 24 | — | Claude Opus 5.5 · effort high · eight hard validated tasks |
| List-price cost per strict pass on hard tasks (calculation) Calculation | $0.067 | 24 | — | Claude Haiku 4.5 · eight hard validated tasks |
| List-price cost per strict pass on hard tasks (calculation) Calculation | $0.093 | 24 | — | Claude Fable 5.1 · eight hard validated tasks |
Whiskers: 95% Wilson intervalLines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Claude Code in Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks: 30 values, first Pass rate on eight hard tasks (Strict pass) 100% (24/24).
Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks
8 values · open the study
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Coding sessions that passed every hidden check | 100% (12/12) | 12 | 76%–100% (95% CI) | Claude Sonnet 5.5 · six small repository tasks with hidden tests |
| Coding sessions that passed every hidden check | 100% (12/12) | 12 | 76%–100% (95% CI) | Claude Opus 5.5 · six small repository tasks with hidden tests |
| Time per coding session | 23.1 s | 12 | 18.7 s–44.5 s (range) | Claude Sonnet 5.5 · six small repository tasks with hidden tests |
| Time per coding session | 56.9 s | 12 | 29.8 s–186 s (range) | Claude Opus 5.5 · six small repository tasks with hidden tests |
| Tool calls per coding session | 7.5 | 12 | 3–14 (range) | Claude Sonnet 5.5 · six small repository tasks with hidden tests |
| Tool calls per coding session | 7.5 | 12 | 5–14 (range) | Claude Opus 5.5 · six small repository tasks with hidden tests |
| List-price cost per passing coding session (calculation) Calculation | $0.085 | 12 | — | Claude Sonnet 5.5 · six small repository tasks with hidden tests |
| List-price cost per passing coding session (calculation) Calculation | $0.22 | 12 | — | Claude Opus 5.5 · six small repository tasks with hidden tests |
Whiskers: 95% Wilson intervalLines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Claude Code in Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks: 8 values, first Coding sessions that passed every hidden check 100% (12/12).
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks
32 values · open the study
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Strict pass rate by effort on eight hard tasks | 100% (16/16) | 16 | 81%–100% (95% CI) | Claude Sonnet 5.5 · effort low · eight hard validated tasks, effort ladder |
| Strict pass rate by effort on eight hard tasks | 100% (16/16) | 16 | 81%–100% (95% CI) | Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder |
| Strict pass rate by effort on eight hard tasks | 100% (16/16) | 16 | 81%–100% (95% CI) | Claude Sonnet 5.5 · effort high · eight hard validated tasks, effort ladder |
| Strict pass rate by effort on eight hard tasks | 100% (16/16) | 16 | 81%–100% (95% CI) | Claude Sonnet 5.5 · eight hard validated tasks, effort ladder |
| Strict pass rate by effort on eight hard tasks | 100% (16/16) | 16 | 81%–100% (95% CI) | Claude Opus 5.5 · effort low · eight hard validated tasks, effort ladder |
| Strict pass rate by effort on eight hard tasks | 100% (16/16) | 16 | 81%–100% (95% CI) | Claude Opus 5.5 · effort medium · eight hard validated tasks, effort ladder |
| Strict pass rate by effort on eight hard tasks | 100% (16/16) | 16 | 81%–100% (95% CI) | Claude Opus 5.5 · effort high · eight hard validated tasks, effort ladder |
| Strict pass rate by effort on eight hard tasks | 100% (16/16) | 16 | 81%–100% (95% CI) | Claude Opus 5.5 · eight hard validated tasks, effort ladder |
| Total time per call by effort on hard tasks | 5.82 s | 16 | 2.8 s–20 s (range) | Claude Sonnet 5.5 · effort low · eight hard validated tasks, effort ladder |
| Total time per call by effort on hard tasks | 7.63 s | 16 | 2.7 s–24 s (range) | Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder |
| Total time per call by effort on hard tasks | 8.81 s | 16 | 2.9 s–35.8 s (range) | Claude Sonnet 5.5 · effort high · eight hard validated tasks, effort ladder |
| Total time per call by effort on hard tasks | 7.97 s | 16 | 2.3 s–21.6 s (range) | Claude Sonnet 5.5 · eight hard validated tasks, effort ladder |
| Total time per call by effort on hard tasks | 7.50 s | 16 | 3.3 s–15.8 s (range) | Claude Opus 5.5 · effort low · eight hard validated tasks, effort ladder |
| Total time per call by effort on hard tasks | 9.72 s | 16 | 4.8 s–31.4 s (range) | Claude Opus 5.5 · effort medium · eight hard validated tasks, effort ladder |
| Total time per call by effort on hard tasks | 10.1 s | 16 | 3.6 s–63 s (range) | Claude Opus 5.5 · effort high · eight hard validated tasks, effort ladder |
| Total time per call by effort on hard tasks | 9.18 s | 16 | 4.2 s–27.2 s (range) | Claude Opus 5.5 · eight hard validated tasks, effort ladder |
| Output tokens per call by effort on hard tasks (Output tokens) | 667 | 16 | — | Claude Sonnet 5.5 · effort low · eight hard validated tasks, effort ladder |
| Output tokens per call by effort on hard tasks (Output tokens) | 770 | 16 | — | Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder |
| Output tokens per call by effort on hard tasks (Output tokens) | 1,192 | 16 | — | Claude Sonnet 5.5 · effort high · eight hard validated tasks, effort ladder |
| Output tokens per call by effort on hard tasks (Output tokens) | 1,054 | 16 | — | Claude Sonnet 5.5 · eight hard validated tasks, effort ladder |
| Output tokens per call by effort on hard tasks (Output tokens) | 594 | 16 | — | Claude Opus 5.5 · effort low · eight hard validated tasks, effort ladder |
| Output tokens per call by effort on hard tasks (Output tokens) | 853 | 16 | — | Claude Opus 5.5 · effort medium · eight hard validated tasks, effort ladder |
| Output tokens per call by effort on hard tasks (Output tokens) | 1,052 | 16 | — | Claude Opus 5.5 · effort high · eight hard validated tasks, effort ladder |
| Output tokens per call by effort on hard tasks (Output tokens) | 945 | 16 | — | Claude Opus 5.5 · eight hard validated tasks, effort ladder |
| List-price cost per strict pass by effort (calculation) Calculation | $0.012 | 16 | — | Claude Sonnet 5.5 · effort low · eight hard validated tasks, effort ladder |
| List-price cost per strict pass by effort (calculation) Calculation | $0.014 | 16 | — | Claude Sonnet 5.5 · effort medium · eight hard validated tasks, effort ladder |
| List-price cost per strict pass by effort (calculation) Calculation | $0.017 | 16 | — | Claude Sonnet 5.5 · effort high · eight hard validated tasks, effort ladder |
| List-price cost per strict pass by effort (calculation) Calculation | $0.014 | 16 | — | Claude Sonnet 5.5 · eight hard validated tasks, effort ladder |
| List-price cost per strict pass by effort (calculation) Calculation | $0.021 | 16 | — | Claude Opus 5.5 · effort low · eight hard validated tasks, effort ladder |
| List-price cost per strict pass by effort (calculation) Calculation | $0.029 | 16 | — | Claude Opus 5.5 · effort medium · eight hard validated tasks, effort ladder |
| List-price cost per strict pass by effort (calculation) Calculation | $0.034 | 16 | — | Claude Opus 5.5 · effort high · eight hard validated tasks, effort ladder |
| List-price cost per strict pass by effort (calculation) Calculation | $0.029 | 16 | — | Claude Opus 5.5 · eight hard validated tasks, effort ladder |
Whiskers: 95% Wilson intervalLines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Claude Code in Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks: 32 values, first Strict pass rate by effort on eight hard tasks 100% (16/16).
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| List-price cost of 5-question sessions with and without the cache (calculation) (With the cache, as recorded) Calculation | $0.14 | 15 | — | Claude Sonnet 5.5 · calculation: 5-turn cached sessions over a fixed ledger |
| List-price cost of 5-question sessions with and without the cache (calculation) (With the cache, as recorded) Calculation | $0.26 | 15 | — | Claude Opus 5.5 · calculation: 5-turn cached sessions over a fixed ledger |
| List-price cost of 5-question sessions with and without the cache (calculation) (Without a cache: every input token at the input price) Calculation | $0.27 | 15 | — | Claude Sonnet 5.5 · calculation: 5-turn cached sessions over a fixed ledger |
| List-price cost of 5-question sessions with and without the cache (calculation) (Without a cache: every input token at the input price) Calculation | $0.54 | 15 | — | Claude Opus 5.5 · calculation: 5-turn cached sessions over a fixed ledger |
| Time per turn: first turn vs later turns in a cached session (Turn 1 (writes the ledger to the cache)) | 1.64 s | 3 | 1.6 s–1.8 s (range) | Claude Sonnet 5.5 · 5-turn cached sessions over a fixed ledger |
| Time per turn: first turn vs later turns in a cached session (Turn 1 (writes the ledger to the cache)) | 1.90 s | 3 | 1.8 s–4.4 s (range) | Claude Opus 5.5 · 5-turn cached sessions over a fixed ledger |
| Time per turn: first turn vs later turns in a cached session (Turns 2-5 (read the ledger from the cache)) | 1.61 s | 12 | 1.4 s–5.6 s (range) | Claude Sonnet 5.5 · 5-turn cached sessions over a fixed ledger |
| Time per turn: first turn vs later turns in a cached session (Turns 2-5 (read the ledger from the cache)) | 2.40 s | 12 | 1.6 s–12.7 s (range) | Claude Opus 5.5 · 5-turn cached sessions over a fixed ledger |
| Same prompt, 10 times: strict pass rate (Exact number) | 0% (0/10) | 10 | 0%–28% (95% CI) | Claude Haiku 4.5 · same prompt repeated 10 times |
| Same prompt, 10 times: strict pass rate (Exact number) | 100% (10/10) | 10 | 72%–100% (95% CI) | Claude Sonnet 5.5 · same prompt repeated 10 times |
| Same prompt, 10 times: strict pass rate (JSON object) | 10% (1/10) | 10 | 1.8%–40% (95% CI) | Claude Haiku 4.5 · same prompt repeated 10 times |
| Same prompt, 10 times: strict pass rate (JSON object) | 100% (10/10) | 10 | 72%–100% (95% CI) | Claude Sonnet 5.5 · same prompt repeated 10 times |
| Same prompt, 10 times: strict pass rate (Code fix) | 100% (10/10) | 10 | 72%–100% (95% CI) | Claude Haiku 4.5 · same prompt repeated 10 times |
| Same prompt, 10 times: strict pass rate (Code fix) | 100% (10/10) | 10 | 72%–100% (95% CI) | Claude Sonnet 5.5 · same prompt repeated 10 times |
| Same prompt, 10 times: how many different answers (Exact number) | 1 | 10 | — | Claude Haiku 4.5 · same prompt repeated 10 times |
| Same prompt, 10 times: how many different answers (Exact number) | 1 | 10 | — | Claude Sonnet 5.5 · same prompt repeated 10 times |
| Same prompt, 10 times: how many different answers (JSON object) | 1 | 10 | — | Claude Haiku 4.5 · same prompt repeated 10 times |
| Same prompt, 10 times: how many different answers (JSON object) | 1 | 10 | — | Claude Sonnet 5.5 · same prompt repeated 10 times |
| Same prompt, 10 times: how many different answers (Code fix) | 6 | 10 | — | Claude Haiku 4.5 · same prompt repeated 10 times |
| Same prompt, 10 times: how many different answers (Code fix) | 3 | 10 | — | Claude Sonnet 5.5 · same prompt repeated 10 times |
| Same prompt, 10 times: time per call (Exact number) | 5.06 s | 10 | 4.4 s–6.2 s (range) | Claude Haiku 4.5 · same prompt repeated 10 times |
| Same prompt, 10 times: time per call (Exact number) | 6.89 s | 10 | 5.8 s–7.8 s (range) | Claude Sonnet 5.5 · same prompt repeated 10 times |
| Same prompt, 10 times: time per call (JSON object) | 7.03 s | 10 | 5.3 s–12.3 s (range) | Claude Haiku 4.5 · same prompt repeated 10 times |
| Same prompt, 10 times: time per call (JSON object) | 2.89 s | 10 | 2.7 s–5.3 s (range) | Claude Sonnet 5.5 · same prompt repeated 10 times |
| Same prompt, 10 times: time per call (Code fix) | 5.95 s | 10 | 4.9 s–7.3 s (range) | Claude Haiku 4.5 · same prompt repeated 10 times |
| Same prompt, 10 times: time per call (Code fix) | 2.67 s | 10 | 2.3 s–4.3 s (range) | Claude Sonnet 5.5 · same prompt repeated 10 times |
Lines: fastest–slowest run (not an interval)Whiskers: 95% Wilson intervaln beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Claude Code in Prompt caching and run-to-run consistency in Claude Code and Codex CLI: 26 values, first List-price cost of 5-question sessions with and without the cache (calculation) (With the cache, as recorded) $0.14.
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Claude Code sessions, every one graded (120 Sonnet 5.5, 80 Haiku 4.5) | 200 | 200 | — | — |
n beside each valueMuted dots: the other configurations on the same chart
Claude Code in Does memory help Claude Code? 8 kinds of agent memory, tested: 1 value, first Claude Code sessions, every one graded (120 Sonnet 5.5, 80 Haiku 4.5) 200.
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| CLI start-up tax on a one-word answer (First output event) | 563 ms | 5 | 519 ms–726 ms (range) | Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs |
| CLI start-up tax on a one-word answer (First model output) | 1,461 ms | 5 | 1.21 s–2.31 s (range) | Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs |
| CLI start-up tax on a one-word answer (Total wall time) | 2,529 ms | 5 | 2.27 s–3.38 s (range) | Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs |
| Input tokens a CLI sends for a one-word answer | 6,761 | 5 | — | Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs |
| Claude Code time outside the model on a one-word answer | 1,690 ms median (1,533 to 1,811) | 5 | — | routing overhead per decision |
Lines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chart
Claude Code in Routing overhead: deterministic policy vs LLM routers vs Jev: 5 values, first CLI start-up tax on a one-word answer (First output event) 563 ms.
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Repairing a scheduler: Claude Code vs Codex vs API (Total time) | 15.0 s | 3 | 13.9 s–15.9 s (range) | Sonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs |
| Repairing a scheduler: Claude Code vs Codex vs API (First useful output) | 7.55 s | 3 | 6.8 s–7.6 s (range) | Sonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs |
| Output tokens to repair the scheduler (Output tokens) | 2,227 | 3 | — | Sonnet 5.5 · effort medium · scheduler repair, 296 checks, 3 runs |
Lines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chart
Claude Code in Claude Code CLI vs Codex CLI vs the API: latency and tokens: 3 values, first Repairing a scheduler: Claude Code vs Codex vs API (Total time) 15.0 s.
Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
54 values · open the study
Whiskers: 95% Wilson intervalLines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Claude Code in Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks: 54 values, first Strict pass rate: single call vs agent loop on eight hard tasks 46% (11/24).
| Metric | Value | n | Interval or range | Configuration |
|---|---|---|---|---|
| Haiku thinking study: typed routing decisions answered exactly right (Exact decisions (every scored question right)) | 87% (71/82) | 82 | 78%–92% (95% CI) | Claude Haiku 4.5 · thinking off · typed routing decisions, thinking on vs off |
| Haiku thinking study: typed routing decisions answered exactly right (Exact decisions (every scored question right)) | 89% (73/82) | 82 | 80%–94% (95% CI) | Claude Haiku 4.5 · thinking on · typed routing decisions, thinking on vs off |
| Haiku thinking study: typed routing decisions answered exactly right (Exact decisions (every scored question right)) | 94% (77/82) | 82 | 87%–97% (95% CI) | Claude Sonnet 5.5 · effort low · typed routing decisions, thinking on vs off |
| Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy) | 91% (177/194) | 194 | 86%–94% (95% CI) | Claude Haiku 4.5 · thinking off · typed routing decisions, thinking on vs off |
| Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy) | 94% (183/194) | 194 | 90%–97% (95% CI) | Claude Haiku 4.5 · thinking on · typed routing decisions, thinking on vs off |
| Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy) | 97% (189/194) | 194 | 94%–99% (95% CI) | Claude Sonnet 5.5 · effort low · typed routing decisions, thinking on vs off |
| Haiku thinking study: time per routing decision (Wall time (CLI)) | 4.66 s | 82 | 4.7 s–8.2 s (p50–p95) | Claude Haiku 4.5 · thinking off · typed routing decisions, thinking on vs off |
| Haiku thinking study: time per routing decision (Wall time (CLI)) | 12.5 s | 82 | 12.5 s–34.5 s (p50–p95) | Claude Haiku 4.5 · thinking on · typed routing decisions, thinking on vs off |
| Haiku thinking study: time per routing decision (Wall time (CLI)) | 2.60 s | 82 | 2.6 s–4.3 s (p50–p95) | Claude Sonnet 5.5 · effort low · typed routing decisions, thinking on vs off |
| Haiku thinking study: time per routing decision (Model time (API)) | 3.79 s | 82 | 3.8 s–7.4 s (p50–p95) | Claude Haiku 4.5 · thinking off · typed routing decisions, thinking on vs off |
| Haiku thinking study: time per routing decision (Model time (API)) | 10.5 s | 82 | 10.5 s–32.1 s (p50–p95) | Claude Haiku 4.5 · thinking on · typed routing decisions, thinking on vs off |
| Haiku thinking study: time per routing decision (Model time (API)) | 1.60 s | 82 | 1.6 s–2.6 s (p50–p95) | Claude Sonnet 5.5 · effort low · typed routing decisions, thinking on vs off |
| Haiku thinking study: thinking and visible output tokens per routing decision (Thinking tokens) | 0 | 82 | — | Claude Haiku 4.5 · thinking off · typed routing decisions, thinking on vs off |
| Haiku thinking study: thinking and visible output tokens per routing decision (Thinking tokens) | 1,101 | 82 | — | Claude Haiku 4.5 · thinking on · typed routing decisions, thinking on vs off |
| Haiku thinking study: thinking and visible output tokens per routing decision (Thinking tokens) | 2 | 82 | — | Claude Sonnet 5.5 · effort low · typed routing decisions, thinking on vs off |
| Haiku thinking study: thinking and visible output tokens per routing decision (Visible output tokens) | 366 | 82 | — | Claude Haiku 4.5 · thinking off · typed routing decisions, thinking on vs off |
| Haiku thinking study: thinking and visible output tokens per routing decision (Visible output tokens) | 318 | 82 | — | Claude Haiku 4.5 · thinking on · typed routing decisions, thinking on vs off |
| Haiku thinking study: thinking and visible output tokens per routing decision (Visible output tokens) | 105 | 82 | — | Claude Sonnet 5.5 · effort low · typed routing decisions, thinking on vs off |
| Haiku thinking study: list-price cost per 1,000 routing decisions (calculation) Calculation | $3.36 | 82 | — | Claude Haiku 4.5 · thinking off · typed routing decisions, thinking on vs off |
| Haiku thinking study: list-price cost per 1,000 routing decisions (calculation) Calculation | $8.92 | 82 | — | Claude Haiku 4.5 · thinking on · typed routing decisions, thinking on vs off |
| Haiku thinking study: list-price cost per 1,000 routing decisions (calculation) Calculation | $7.32 | 82 | — | Claude Sonnet 5.5 · effort low · typed routing decisions, thinking on vs off |
| Haiku thinking study: pass rate on eight hard tasks (Strict pass) | 17% (4/24) | 24 | 6.7%–36% (95% CI) | Claude Haiku 4.5 · thinking off · eight hard validated tasks, thinking on vs off |
| Haiku thinking study: pass rate on eight hard tasks (Strict pass) | 46% (11/24) | 24 | 28%–65% (95% CI) | Claude Haiku 4.5 · thinking on · eight hard validated tasks, thinking on vs off |
| Haiku thinking study: pass rate on eight hard tasks (Lenient (format misses counted)) | 17% (4/24) | 24 | 6.7%–36% (95% CI) | Claude Haiku 4.5 · thinking off · eight hard validated tasks, thinking on vs off |
| Haiku thinking study: pass rate on eight hard tasks (Lenient (format misses counted)) | 67% (16/24) | 24 | 47%–82% (95% CI) | Claude Haiku 4.5 · thinking on · eight hard validated tasks, thinking on vs off |
| Haiku thinking study: total time per call on hard tasks | 2.95 s | 24 | 1.7 s–13 s (range) | Claude Haiku 4.5 · thinking off · eight hard validated tasks, thinking on vs off |
| Haiku thinking study: total time per call on hard tasks | 39.0 s | 24 | 15.3 s–75.1 s (range) | Claude Haiku 4.5 · thinking on · eight hard validated tasks, thinking on vs off |
Whiskers: 95% Wilson intervalLines: median to p95 (not an interval)Lines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Claude Code in Does thinking pay for Claude Haiku 4.5? Thinking on vs off: 27 values, first Haiku thinking study: typed routing decisions answered exactly right (Exact decisions (every scored question right)) 87% (71/82).
Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
32 values · open the study
Whiskers: 95% Wilson intervalLines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Claude Code in Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI: 32 values, first Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON) 0% (0/24).
Whiskers: 95% Wilson intervalLines: median to p95 (not an interval)n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Claude Code in Jev vs Claude routers on unseen decisions: a blind holdout: 22 values, first Unseen routing decisions answered exactly right 79% (44/56).
Lines: lowest–highest run (not an interval)n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Claude Code in How much of an AI bill is thinking? Reasoning tokens by model and effort: 46 values, first Reasoning share of output tokens per call on hard tasks (calculation) 91.7%.
Lines: fastest–slowest run (not an interval)Whiskers: 95% Wilson intervaln beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Claude Code in Where the seconds go: first text, output speed and prompt size for 6 LLMs: 33 values, first Time to first text: a 250-line answer, six models 4.00 s.
n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Claude Code in Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts: 16 values, first List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Interval merge fix $0.018.
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
29 values · open the study
Whiskers: 95% Wilson intervalLines: fastest–slowest run (not an interval)n beside each valueMuted dots: the other configurations on the same chartHollow: list-price calculation
Claude Code in GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks: 29 values, first Pass rate on 4 harder tasks (Strict pass) 42% (5/12).
Effort profile
The effort ladder study ran the same eight hard tasks at each effort setting. Claude Code is emphasized; the other models stay on the chart for scale.
- Claude Sonnet 5.5 · Claude Code
- Claude Opus 5.5 · Claude Code
- GPT-6.1 Sol · Codex CLI
Seconds (median)
* default: the effort flag was not passed; its level is not known, so no line joins it.
| Effort | Claude Sonnet 5.5 · Claude Code | Claude Opus 5.5 · Claude Code | GPT-6.1 Sol · Codex CLI | n |
|---|---|---|---|---|
| low | 5.8 s | 7.5 s | 13.6 s | 16 |
| medium | 7.6 s | 9.7 s | 13.1 s | 16 |
| high | 8.8 s | 10.1 s | 18.1 s | 16 |
| default | 8 s | 9.2 s | — | 16 |
4 efforts, 3 series: Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol · Codex CLI. Claude Sonnet 5.5 · Claude Code: slowest high 8.8 s (n 16). Fastest low 5.8 s (n 16). Claude Opus 5.5 · Claude Code: slowest high 10.1 s (n 16). Fastest low 7.5 s (n 16).
Notesn = 16 per row
One line per model and route; default = the effort flag was not passed
Medians only; the per-call ranges are in the total-time chart and they overlap. "default" is placed last because its level is not known: the CLI chose it.
Source: Effort ladder: the hard task set at each effort level
Compare Claude Code
Each bar counts the rows of one comparison: a side ahead only where its interval or range is apart, otherwise a tie or unclear.
vs coding cli
Claude Code vs Codex CLI
Claude Code ahead on 7 · Codex CLI ahead on 1 · 31 ties · 49 unclear
Watch
Agent Benchmarks, October 2026: 28 studies in one film
The headline of each of our 28 open benchmark studies, with sample sizes and intervals. Calculations labelled; failures counted.
Transcript
- October 2026 roundup · 28 studies. Agent Benchmarks: every headline. Recorded runs, 95% intervals where they exist, every failure counted. Calculations labelled.
- SWE-bench Verified: Agent resolved 25 of 33. The public panel averaged 74.1%; the intervals overlap, so no rank. Agent resolved (one attempt each): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Model cost per resolved instance (calculation, notional): $3.71 (n = 25). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
- SWE-bench, interim (3 of 8 pairs graded): Opus 5.5 as Agent’s brain resolved 2 of 3, Sonnet 5.5 1 of 3. Exact McNemar p = 1.0: no difference yet. Opus cost 2.6× as much, a calculation. Opus 5.5 in Agent: resolved (interim): 67% (2/3) (n = 3, 95% CI 21–94%). Sonnet 5.5 in Agent: same instances: 33% (1/3) (n = 3, 95% CI 6–79%). List-price cost, Opus vs Sonnet (calculation): 2.6× (n = 3). Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
- Blind critics preferred the AI pull request to the merged human one on 9 of 12 tasks at the latest attempt, 6 of 12 at the first. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
- Five short tasks, nine setups: 127 of 130 calls passed, so speed and tokens separate them. Fastest median: Fable 5.1 at 1.9 s. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Fastest median total time (Fable 5.1): 1.9 s (n = 15). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.
- Eight hard tasks: Sonnet 5.5, Opus 5.5, Opus 5.5 high, GPT-6.1 Sol medium, Fable 5.1 and GPT-6.1 Sol high passed every call (24/24 or 16/16). Haiku 4.5 passed 11/24. Chart: Eight hard tasks · strict pass rate · 95% intervals (n = 16–24 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
- Coding agents on hidden-test tasks: 36 of 36 sessions passed, so time separates them. Claude Code with Sonnet 5.5: median 23.1 s; Codex CLI took 4.9× as long, with extra standing instructions. Sessions that passed every hidden check: 100% (36/36) (n = 36, 95% CI 90–100%). Median time, Sonnet 5.5 in Claude Code: 23.1 s (n = 12). Median time, Codex CLI vs Claude Code with Sonnet: 4.9× (n = 12). Caveat: Codex CLI ran with the tester’s global AGENTS.md: --ignore-user-config does not switch that file off. 12 of 12 Codex sessions read the tester’s notes, 10 wrote a WORKLOG.md and 9 reported a commit attempt. Its time, tool calls and lines changed include that work. Claude Code ran with setting sources off, and no Claude session did any of it.
- Effort ladder: 11 model and effort settings passed 176 of 176 calls strictly (16/16 each, 81%–100%). No effort level wins any of 172 rows. Calls that passed strictly, low to high effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference). Effort comparison rows where one level wins: 0 (of 172 rows · 16 pairs). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Caching, a list-price calculation: 50% less on Sonnet 5.5, 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every time. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- Agent memory: Sonnet 5.5 followed team-only rules 40% (6/15) of the time without memory and 100% (15/15) with an 11-line file. Without the late-fee rate it asked 7/9 times. Team knowledge, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). Asked for the missing rate: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
- System One arena: on 1,085 checkable decisions, Jev 1.13 answered 76.8% right; the best open model, Clef 27B, 69.9%. Jev 1.13: 76.8% (805/1048) (n = 1048, 95% CI 74–79%). Clef 27B: 69.9% (733/1048) (n = 1048, 95% CI 67–73%). Checkable decisions: 1,085. Caveat: Local models ran quantized on one Mac; the vendors measured full-precision weights on data-centre GPUs. Latency on other hardware will differ, and Jev latency includes the network.
- Routing: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82 exact, and the intervals overlap. List-price cost is where they differ. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
- Routing overhead: an in-process policy decides in 1.42 µs; Sonnet 5.5 as a router takes 2.60 s through the CLI, about 1.8 million times longer. Chart: Time per routing decision · log scale · median to p95 (n = 82–20000 each). Caveat: The policy is timed in process and the LLM routers through a CLI: this compares the two ways of routing as deployed, not two models on equal footing. A direct API call would skip the CLI time (shown separately).
- Provider prices, reported by OpenRouter: open-weight models vary up to 12.6x across providers. Closed models: one price, plus a 5.5% credit fee. Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Gateway price equals the first-party price: 10 of 10. Endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
- Recorded SWE-bench tokens, repriced: $87.23 on Sonnet 5.5, $143.83 on Opus 5.5, $343.33 with no prompt cache. At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Sonnet 5.5 without the prompt cache: $343.33 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
- What a CLI adds: for a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens. Codex CLI vs OpenAI API, one-line answer: 3.5x slower (n = 30). Input tokens the Codex CLI sends for one line: 19,551 (n = 15). Scheduler repair, median: Claude Code vs Codex CLI: 15.0 s vs 61.2 s (n = 3, models differ too). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
- Three real pull requests: the latest build verified 1 of 3, at $11.06 notional. Guardrail refusals went 26 → 19. Verified deliveries, latest build: 1 of 3 (n = 3). Notional cost, latest build, 3 tasks: $11.06 (n = 3). Guardrail refusals, first vs latest slice: 26 → 19 (n = 3). Caveat: One attempt per cell: these are defect-finding runs, not rates.
- Agent-loop strict passes: Haiku 54% (13/24); Sonnet 100% (16/16). Attempts excluded for outside reads: 2 of 56. Claude Haiku 4.5 strict pass rate, agent loop: 54% (13/24) (n = 24, 95% CI 35–72%). Claude Sonnet 5.5 strict pass rate, agent loop: 100% (16/16) (n = 16, 95% CI 81–100%). Agent-loop attempts left out for reading outside the work folder: 2 of 56 (n = 56). Caveat: The Claude single-call cells ran in another batch on 2026-10-06 (03:23 to 04:02 UTC), with the same CLI version, tasks and validators; provider load can differ by hour.
- Haiku exact routing: thinking off 87% (71/82), on 89% (73/82). Paired McNemar p = 0.754; day and account differ. Claude Haiku 4.5 (thinking off): exact routing decisions: 87% (71/82) (n = 82, 95% CI 78–92%). Claude Haiku 4.5 (thinking on): exact routing decisions: 89% (73/82) (n = 82, 95% CI 80–94%). Paired exact test, thinking off vs on (exact McNemar p): 0.754 (n = 82). Caveat: The thinking-on and Sonnet arms are the recorded routing run of 2026-10-05, reused. The thinking-off arm ran on a different day and on a different Claude subscription account, so this is a confounded comparison. Day, account, CLI version and input-token differences can affect the results. The data cannot isolate the effect of thinking.
- Format misses: instructions 35% (17/48), JSON schema 0% (0/48). The schema still gave wrong values: 13% (6/48). Format-miss rate with instructions only, all models: 35% (17/48) (n = 48, 95% CI 23–50%). Format-miss rate with a JSON schema, all models: 0% (0/48) (n = 48, 95% CI 0–7%). Wrong-values rate with a JSON schema, all models: 13% (6/48) (n = 48, 95% CI 6–25%). Caveat: Small samples: Haiku 24 calls per mode, Sonnet 12 calls per mode and GPT-6.1 Sol 12 calls per mode. A 12/12 result has a 95% interval of 76% to 100%. This is not a minimum detectable difference.
- Later Claude sessions with substantial cache reuse: new folders 0 of 2 (95% interval 0% to 66%); a fixed folder 2 of 2 (95% interval 34% to 100%). This analysis is exploratory. Later sessions with at least 50% of turn-1 input cached, A: new folder each time: 0 of 2 (95% interval 0% to 66%) (n = 2, 95% CI 0–66%). Later sessions with at least 50% of turn-1 input cached, B: fixed folder: 2 of 2 (95% interval 34% to 100%) (n = 2, 95% CI 34–100%). Caveat: The surviving protocol file was created after all counted calls. Its claimed 00:32 UTC declaration is not supported by its file birth time. Amendment 1 and 2 state 00:36 and 00:37 UTC, but separate pre-edit copies do not verify those times. The current summary was regenerated at 07:41 UTC. Treat the analysis as exploratory.
- Cache break-even, a calculation: a new prefix with a one-hour write needs 2 reuses (the 3rd request). Recorded session payback: 2 turns (6 of 6 sessions). Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation): 2 reuses (the 3rd request). Turn at which a recorded Claude Code session’s total input cost with the cache first fell below its cost with no cache (calculation on recorded tokens): 2 turns (6 of 6 sessions) (n = 6). Calculation, not a run. Caveat: The retained Claude protocol file was created at 14:43:55 UTC, after the first counted session at 14:35:39 UTC on 2026-10-06. Its declaration says 14:25 UTC, but file times do not verify that claim. Treat this as a retrospective protocol record.
- Exact unseen routing decisions: Jev 82% (46/56), Haiku 79% (44/56), Sonnet 88% (49/56). The intervals overlap; no rank. Jev 1.13 (TypeSafe): exact on unseen decisions: 82% (46/56) (n = 56, 95% CI 70–90%). Claude Haiku 4.5 · Claude Code: exact on unseen decisions: 79% (44/56) (n = 56, 95% CI 66–87%). Claude Sonnet 5.5 (low) · Claude Code: exact on unseen decisions: 88% (49/56) (n = 56, 95% CI 76–94%). Caveat: The case author and two of the three routers are Claude models, so a same-family label bias is possible. A second labeller (GPT-6.1 Sol) labelled every case blind; the secondary scoring keeps only the keys where its label is inside our acceptable set.
- Reasoning cost at high vs low effort, a calculation from recorded calls: Sonnet 2.2x ($0.0043 at low, $0.0094 at high); Opus 3.6x ($0.0050 at low, $0.0180 at high). Reasoning cost per call, high ÷ low effort, Claude Sonnet 5.5 · Claude Code (calculation, means): 2.2x ($0.0043 at low, $0.0094 at high) (n = 16). Reasoning cost per call, high ÷ low effort, Claude Opus 5.5 · Claude Code (calculation, means): 3.6x ($0.0050 at low, $0.0180 at high) (n = 16). Calculation, not a run. Caveat: The hard-set protocol file was created after its first counted call. Both effort-ladder protocol files were created after their batches ended. The short-set file predates its first call, but its top-up amendment timing is unverified. Batch receipts preserve protocol text, but we cannot verify all rules were written before inference. Treat these as exploratory calculations, not preregistered tests.
- Speed anatomy, a calculation: same-text token count ratio 1.8x. Extra time to first text with the larger prompt: Haiku +0.9 s; Sonnet +1.6 s. The same 4,327-character reply: most tokens ÷ fewest tokens across models (calculation): 1.8x (n = 23). Haiku: extra time to first text at 64k vs 1k (calculation): +0.9 s (n = 6). Sonnet: extra time to first text at 64k vs 1k (calculation): +1.6 s (n = 6). Calculation, not a run. Caveat: 4 calls per model in part A and 3 per size in part B. Medians of so few calls move with one slow call, and the ranges are not confidence intervals. The 95% Wilson intervals on the lookup rates are wide.
- Cost per correct answer, a calculation: Sonnet every time $0.0132; Haiku with one retry then Sonnet $0.0566. No policy ran. Cost per correct answer, Sonnet 5.5 every time (calculation): $0.0132 (n = 24). Cost per correct answer, Haiku with one retry then Sonnet (calculation): $0.0566 (n = 48). Calculation, not a run. Caveat: Advance registration is not verified. The hard-set protocol file birth time is 2026-10-06 04:03:37 UTC; its first counted call started at 03:23:59 UTC. The ladder protocol file birth time is 14:35:11 UTC; its first Claude call started at 14:21:50 UTC. Both files claim advance declaration, but the available file times do not support that claim. Copying could explain the times; we cannot establish it.
- Strict passes on four selected harder tasks: Sol 69% (11/16); Opus 42% (5/12); Sonnet 38% (6/16). Tasks were selected with a Sonnet pilot. GPT-6.1 Sol (medium) (Codex CLI): strict pass rate on the harder tasks: 69% (11/16) (n = 16, 95% CI 44–86%). Opus 5.5 (Claude Code): strict pass rate on the harder tasks: 42% (5/12) (n = 12, 95% CI 19–68%). Sonnet 5.5 (Claude Code): strict pass rate on the harder tasks: 38% (6/16) (n = 16, 95% CI 18–61%). Caveat: Selection effect: the study picked tasks that Sonnet did not pass twice in the pilot. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls (calculation: 25% and 38%; the intervals overlap). Selection can produce this pattern, but the run does not establish its cause.
- Latency budget, a calculation: steps within an assumed 300 ms budget 3 of 14 (median: 3 of 14); within 1,500 ms 4 of 14 (median: 8 of 14), at p95 or the observed maximum. No voice turn ran. Steps that fit a 300 ms budget at the slow end (calculation): 3 of 14 (median: 3 of 14) (n = 14). Steps that fit a 1,500 ms budget at the slow end (calculation): 4 of 14 (median: 8 of 14) (n = 14). Calculation, not a run. Caveat: Timing samples omit 0 failed or untimed matched explorer calls and 0 failed or untimed Claude head-to-head calls. A failure is not a fast successful step. No pass rate is estimated here.
- Routing cost at an assumed million decisions a day, a calculation: Jev $33.70; Sonnet $7,324. No load test ran. Daily cost at 1 million decisions, Jev (calculation): $33.70 (n = 82). Daily cost at 1 million decisions, Sonnet 5.5 router (calculation): $7,324 (n = 82). Calculation, not a run. Caveat: The routing and live Jev protocols predate the first counted calls by file birth time. The overhead protocol predates its microbenchmark output. Claude ran 82 calls per router, below its 120-call cap. Jev ran 246 counted calls, below its 300-call cap. One Haiku pilot and one Jev probe are excluded. All counted calls completed without call errors. The policy microbenchmark used 5,000 warm-up iterations and 64 synthetic contexts; the minimum and individual timing samples are unavailable, so we cannot rebuild its quantiles or full range. No separate validator-control receipt or Sonnet pilot is retained.
- 28 studies, one place. Intervals, sources and every failure kept.
Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks
All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.
Transcript
- Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.
- 176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.
- All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.
- More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.
- List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.
- 16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Start at low effort and measure. Every call, interval and cost online.
Write-ups that use this data
All postsA latency budget for voice agents: which LLM steps fit in one turn?
136.5 ms for Jev, 0.82 s for a small-model API, 2.79 s to 3.79 s for Codex CLI: which steps fit a voice agent latency budget? A thought experiment.
A voice agent latency budget, with measured times: what fits in one turn?
Rules and Jev 1.13 fit every budget we assumed; a Claude router through a CLI fits none. 14 measured steps vs 300, 800 and 1,500 ms. A thought experiment.
AI coding agent best practices: 12 rules, each backed by a measurement
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.