113 measured metrics · 40 calculated · 14 studies
Claude Haiku 4.5vsClaude Sonnet 5.5
Claude Sonnet 5.5 ahead on 21; 58 ties, 74 unclear. A side is ahead only where the intervals or ranges do not overlap.
The verdict
Claude Haiku 4.5 and Claude Sonnet 5.5 share 113 measured metrics and 40 list-price calculations from 14 studies. Claude Sonnet 5.5 leads on 21 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); and 18 more. On those rows the 95% intervals, run ranges and p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 58 ties and 74 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 2 at the smallest).
Headline metrics
How far apart the two sides are on the headline metrics. Length is the ratio of the two values; it is not a winner.
Bar length is the ratio of the two values on a log scale, pointing to the larger one. Larger is not better for time, tokens or cost. A bar has a side’s color only when that side is ahead in the data; gray means the data does not separate them.
| Metric | Claude Haiku 4.5 | Claude Sonnet 5.5 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 100% (15/15)Claude Code · five short validated tasks | 80% (12/15)Claude Code · five short validated tasks | 15 | 95% CI: 80%–100% vs 55%–93% | Tie | The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Sonnet 5.5 55% to 93%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 4.43 sClaude Code · five short validated tasks | 2.31 sClaude Code · five short validated tasks | 15 | range: 3.2 s–23.6 s vs 2.2 s–7.7 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; Claude Sonnet 5.5 2.17 s to 7.73 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 3.63 sClaude Code · five short validated tasks | 1.56 sClaude Code · five short validated tasks | 15 | range: 2.8 s–22.3 s vs 1 s–6.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.78 s to 22.3 s; Claude Sonnet 5.5 0.99 s to 6.39 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Other input) | 3,790Claude Code · five short validated tasks | 685Claude Code · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 367Claude Code · five short validated tasks | 107Claude Code · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.0084Claude Code · five short validated tasks | $0.0062Claude Code · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.0084 vs $0.0062) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
6 headline metrics as the ratio of the two values. None of them separates the sides in the data. Widest ratio: Input tokens per call: what the CLI sends (Other input), 5.5x (Claude Haiku 4.5 larger).
Watch it build
A live story drawn from the same dataset in your browser. It shows the same numbers, notes and interval labels as this page.
Claude Haiku 4.5 vs Claude Sonnet 5.5: what the measurements say
153 comparison rows from 14 studies: 0 rows favour Haiku 4.5, 21 favour Sonnet 5.5, 132 are ties or unclear. Cost rows are calculations.
Transcript
- Comparison · 153 rows · 14 studies. Haiku 4.5 vs Sonnet 5.5. A winner only where the 95% intervals or run ranges do not overlap.
- 153 comparison rows from 14 studies: Haiku 4.5 ahead on 0, Sonnet 5.5 ahead on 21. The rest do not separate them. Rows where Haiku 4.5 is ahead: 0 (of 153). Rows where Sonnet 5.5 is ahead: 21 (of 153). Ties or unclear: 132 (58 ties · 74 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.
- Eight hard tasks: pass rate 46% vs 100%, Sonnet 5.5 ahead: 95% intervals separate. 2 of 6 rows separate them. Table: Eight hard tasks · 5 of 6 rows · Claude Code · n = 24 per side. Source study: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks. Rows shown: Pass rate on eight hard tasks (Strict pass); Pass rate on eight hard tasks (Lenient (format misses counted)); Total time per call on hard tasks (separate batches); Time to first useful output on hard tasks; Output tokens per call on hard tasks (Output tokens). Recorded settings: Claude Code · eight hard validated tasks. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.
- Caching sessions: pass rate 0% vs 100%, Sonnet 5.5 ahead: 95% intervals separate. 3 of 9 rows separate them. Table: Caching sessions · 5 of 9 rows · Claude Code · n = 10 per side. Source study: Prompt caching and run-to-run consistency in Claude Code and Codex CLI. Rows shown: Same prompt, 10 times: strict pass rate (Exact number); Same prompt, 10 times: strict pass rate (JSON object); Same prompt, 10 times: strict pass rate (Code fix); Same prompt, 10 times: how many different answers (Exact number); Same prompt, 10 times: time per call (Code fix). Recorded settings: Claude Code · same prompt repeated 10 times. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- Agent memory: pass rate 20% vs 60%, tie: 95% intervals overlap. 4 of 40 rows separate them. Table: Agent memory · 5 of 40 rows · n = 10 vs 15. Source study: Does memory help Claude Code? 8 kinds of agent memory, tested. Rows shown: Full pass rate by kind of memory: No memory; Full pass rate by kind of memory: /init CLAUDE.md; Full pass rate by kind of memory: Handbook, 210 lines; Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md; Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines. Caveat: The hook checks the same code rules as the grader. It shows what rules written as code can do; it cannot carry a fact such as the late-fee rate.
- Routing overhead: success rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 3 of 8 rows separate them. Table: Routing overhead · 5 of 8 rows · thinking on vs effort low · n = 82 per side. Source study: Routing overhead: deterministic policy vs LLM routers vs Jev. Rows shown: Time to make one routing decision; Where an LLM router’s time goes: model vs CLI (Model API time); Where an LLM router’s time goes: model vs CLI (CLI and harness time); Routing calls that returned a decision; Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)). Recorded settings: thinking on · via Claude Code · routing overhead per decision; effort low · via Claude Code · routing overhead per decision; thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts; effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts. Includes a calculation, not a bill or a new run. Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.
- No winner where the data shows none. Showing 20 of 153 rows; every row and its reason online.
Metric by metric
Both values of a row come from the same chart in the same study. Each row has its own axis. The shaded band is where the two intervals or ranges overlap: a side is ahead only when they do not.
- Claude Haiku 4.5
- Claude Sonnet 5.5
- 95% interval
- fastest–slowest run (not an interval)
- median to p95 (not an interval)
- where the two overlap
- hollow: list-price calculation
These calculations hold each rate fixed and assume independent samples. They are not a power calculation or a paired test. Open the sample-size planner.
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
- Pass rate on five validated tasks100% (15/15)n 1580% (12/15)n 15TiePass rate on five validated tasks: Claude Haiku 4.5 100% (15/15) (n 15, 95% interval 80%–100%); Claude Sonnet 5.5 80% (12/15) (n 15, 95% interval 55%–93%). Tie.
Calculation: at these rates, about 33 runs per side would separate them.
- Total time per call4.43 sn 152.31 sn 15UnclearTotal time per call: Claude Haiku 4.5 4.43 s (n 15, run range 3.2 s–23.6 s); Claude Sonnet 5.5 2.31 s (n 15, run range 2.2 s–7.7 s). Unclear.
- Time to first useful output3.63 sn 151.56 sn 15UnclearTime to first useful output: Claude Haiku 4.5 3.63 s (n 15, run range 2.8 s–22.3 s); Claude Sonnet 5.5 1.56 s (n 15, run range 1 s–6.4 s). Unclear.
- Input tokens per call: what the CLI sends (Cache read)0n 151,401n 15UnclearInput tokens per call: what the CLI sends (Cache read): Claude Haiku 4.5 0 (n 15); Claude Sonnet 5.5 1,401 (n 15). Unclear.
- Input tokens per call: what the CLI sends (Other input)3,790n 15685n 15UnclearInput tokens per call: what the CLI sends (Other input): Claude Haiku 4.5 3,790 (n 15); Claude Sonnet 5.5 685 (n 15). Unclear.
- Output tokens per call (Output tokens)367n 15107n 15UnclearOutput tokens per call (Output tokens): Claude Haiku 4.5 367 (n 15); Claude Sonnet 5.5 107 (n 15). Unclear.
- List-price cost per call (calculation)Calculation$0.0057n 15$0.0036n 15UnclearList-price cost per call (calculation), calculation: Claude Haiku 4.5 $0.0057 (n 15, run range $0.0051–$0.018); Claude Sonnet 5.5 $0.0036 (n 15, run range $0.0034–$0.01). Unclear.
- List-price cost per passing answer (calculation)Calculation$0.0084n 15$0.0062n 15UnclearList-price cost per passing answer (calculation), calculation: Claude Haiku 4.5 $0.0084 (n 15); Claude Sonnet 5.5 $0.0062 (n 15). Unclear.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
- Pass rate on eight hard tasks (Strict pass)46% (11/24)n 24100% (24/24)n 24Claude Sonnet 5.5 aheadPass rate on eight hard tasks (Strict pass): Claude Haiku 4.5 46% (11/24) (n 24, 95% interval 28%–65%); Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%). Claude Sonnet 5.5 ahead.
- Pass rate on eight hard tasks (Lenient (format misses counted))67% (16/24)n 24100% (24/24)n 24Claude Sonnet 5.5 aheadPass rate on eight hard tasks (Lenient (format misses counted)): Claude Haiku 4.5 67% (16/24) (n 24, 95% interval 47%–82%); Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%). Claude Sonnet 5.5 ahead.
- Total time per call on hard tasks (separate batches)39.0 sn 247.75 sn 24UnclearTotal time per call on hard tasks (separate batches): Claude Haiku 4.5 39.0 s (n 24, run range 15.3 s–75.1 s); Claude Sonnet 5.5 7.75 s (n 24, run range 2.3 s–34.8 s). Unclear.
- Time to first useful output on hard tasks35.5 sn 245.95 sn 24UnclearTime to first useful output on hard tasks: Claude Haiku 4.5 35.5 s (n 24, run range 12.9 s–70.3 s); Claude Sonnet 5.5 5.95 s (n 24, run range 0.9 s–30.6 s). Unclear.
- Output tokens per call on hard tasks (Output tokens)5,064n 241,050n 24UnclearOutput tokens per call on hard tasks (Output tokens): Claude Haiku 4.5 5,064 (n 24); Claude Sonnet 5.5 1,050 (n 24). Unclear.
- List-price cost per strict pass on hard tasks (calculation)Calculation$0.067n 24$0.014n 24UnclearList-price cost per strict pass on hard tasks (calculation), calculation: Claude Haiku 4.5 $0.067 (n 24); Claude Sonnet 5.5 $0.014 (n 24). Unclear.
Prompt caching and run-to-run consistency in Claude Code and Codex CLI
- Same prompt, 10 times: strict pass rate (Exact number)0% (0/10)n 10100% (10/10)n 10Claude Sonnet 5.5 aheadSame prompt, 10 times: strict pass rate (Exact number): Claude Haiku 4.5 0% (0/10) (n 10, 95% interval 0%–28%); Claude Sonnet 5.5 100% (10/10) (n 10, 95% interval 72%–100%). Claude Sonnet 5.5 ahead.
- Same prompt, 10 times: strict pass rate (JSON object)10% (1/10)n 10100% (10/10)n 10Claude Sonnet 5.5 aheadSame prompt, 10 times: strict pass rate (JSON object): Claude Haiku 4.5 10% (1/10) (n 10, 95% interval 1.8%–40%); Claude Sonnet 5.5 100% (10/10) (n 10, 95% interval 72%–100%). Claude Sonnet 5.5 ahead.
- Same prompt, 10 times: strict pass rate (Code fix)100% (10/10)n 10100% (10/10)n 10TieSame prompt, 10 times: strict pass rate (Code fix): Claude Haiku 4.5 100% (10/10) (n 10, 95% interval 72%–100%); Claude Sonnet 5.5 100% (10/10) (n 10, 95% interval 72%–100%). Tie.
- Same prompt, 10 times: how many different answers (Exact number)1n 101n 10TieSame prompt, 10 times: how many different answers (Exact number): Claude Haiku 4.5 1 (n 10); Claude Sonnet 5.5 1 (n 10). Tie.
- Same prompt, 10 times: how many different answers (JSON object)1n 101n 10TieSame prompt, 10 times: how many different answers (JSON object): Claude Haiku 4.5 1 (n 10); Claude Sonnet 5.5 1 (n 10). Tie.
- Same prompt, 10 times: how many different answers (Code fix)6n 103n 10UnclearSame prompt, 10 times: how many different answers (Code fix): Claude Haiku 4.5 6 (n 10); Claude Sonnet 5.5 3 (n 10). Unclear.
- Same prompt, 10 times: time per call (Exact number)5.06 sn 106.89 sn 10UnclearSame prompt, 10 times: time per call (Exact number): Claude Haiku 4.5 5.06 s (n 10, run range 4.4 s–6.2 s); Claude Sonnet 5.5 6.89 s (n 10, run range 5.8 s–7.8 s). Unclear.
- Same prompt, 10 times: time per call (JSON object)7.03 sn 102.89 sn 10UnclearSame prompt, 10 times: time per call (JSON object): Claude Haiku 4.5 7.03 s (n 10, run range 5.3 s–12.3 s); Claude Sonnet 5.5 2.89 s (n 10, run range 2.7 s–5.3 s). Unclear.
- Same prompt, 10 times: time per call (Code fix)5.95 sn 102.67 sn 10Claude Sonnet 5.5 aheadSame prompt, 10 times: time per call (Code fix): Claude Haiku 4.5 5.95 s (n 10, run range 4.9 s–7.3 s); Claude Sonnet 5.5 2.67 s (n 10, run range 2.3 s–4.3 s). Claude Sonnet 5.5 ahead.
Does memory help Claude Code? 8 kinds of agent memory, tested
- Full pass rate by kind of memory: No memory20% (2/10)n 1060% (9/15)n 15TieFull pass rate by kind of memory: No memory: Claude Haiku 4.5 20% (2/10) (n 10, 95% interval 5.7%–51%); Claude Sonnet 5.5 60% (9/15) (n 15, 95% interval 36%–80%). Tie.
Calculation: at these rates, about 21 runs per side would separate them.
- Full pass rate by kind of memory: /init CLAUDE.md20% (2/10)n 1060% (9/15)n 15TieFull pass rate by kind of memory: /init CLAUDE.md: Claude Haiku 4.5 20% (2/10) (n 10, 95% interval 5.7%–51%); Claude Sonnet 5.5 60% (9/15) (n 15, 95% interval 36%–80%). Tie.
Calculation: at these rates, about 21 runs per side would separate them.
- Full pass rate by kind of memory: Curated, 11 lines70% (7/10)n 10100% (15/15)n 15TieFull pass rate by kind of memory: Curated, 11 lines: Claude Haiku 4.5 70% (7/10) (n 10, 95% interval 40%–89%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
Calculation: at these rates, about 22 runs per side would separate them.
- Full pass rate by kind of memory: Raw notes, 60 lines60% (6/10)n 1093% (14/15)n 15TieFull pass rate by kind of memory: Raw notes, 60 lines: Claude Haiku 4.5 60% (6/10) (n 10, 95% interval 31%–83%); Claude Sonnet 5.5 93% (14/15) (n 15, 95% interval 70%–99%). Tie.
Calculation: at these rates, about 22 runs per side would separate them.
- Full pass rate by kind of memory: Dreamed notes70% (7/10)n 10100% (15/15)n 15TieFull pass rate by kind of memory: Dreamed notes: Claude Haiku 4.5 70% (7/10) (n 10, 95% interval 40%–89%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
Calculation: at these rates, about 22 runs per side would separate them.
- Full pass rate by kind of memory: Handbook, 210 lines30% (3/10)n 10100% (15/15)n 15Claude Sonnet 5.5 aheadFull pass rate by kind of memory: Handbook, 210 lines: Claude Haiku 4.5 30% (3/10) (n 10, 95% interval 11%–60%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Claude Sonnet 5.5 ahead.
- Full pass rate by kind of memory: Stop hook only80% (8/10)n 1080% (12/15)n 15TieFull pass rate by kind of memory: Stop hook only: Claude Haiku 4.5 80% (8/10) (n 10, 95% interval 49%–94%); Claude Sonnet 5.5 80% (12/15) (n 15, 95% interval 55%–93%). Tie.
- Full pass rate by kind of memory: Curated + hook90% (9/10)n 10100% (15/15)n 15TieFull pass rate by kind of memory: Curated + hook: Claude Haiku 4.5 90% (9/10) (n 10, 95% interval 60%–98%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
Calculation: at these rates, about 76 runs per side would separate them.
- Team knowledge followed, Sonnet vs Haiku: No memory0% (0/10)n 1040% (6/15)n 15TieTeam knowledge followed, Sonnet vs Haiku: No memory: Claude Haiku 4.5 0% (0/10) (n 10, 95% interval 0%–28%); Claude Sonnet 5.5 40% (6/15) (n 15, 95% interval 20%–64%). Tie.
Calculation: at these rates, about 17 runs per side would separate them.
- Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md10% (1/10)n 1067% (10/15)n 15Claude Sonnet 5.5 aheadTeam knowledge followed, Sonnet vs Haiku: /init CLAUDE.md: Claude Haiku 4.5 10% (1/10) (n 10, 95% interval 1.8%–40%); Claude Sonnet 5.5 67% (10/15) (n 15, 95% interval 42%–85%). Claude Sonnet 5.5 ahead.
- Team knowledge followed, Sonnet vs Haiku: Curated, 11 lines80% (8/10)n 10100% (15/15)n 15TieTeam knowledge followed, Sonnet vs Haiku: Curated, 11 lines: Claude Haiku 4.5 80% (8/10) (n 10, 95% interval 49%–94%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
Calculation: at these rates, about 33 runs per side would separate them.
- Team knowledge followed, Sonnet vs Haiku: Raw notes, 60 lines60% (6/10)n 10100% (15/15)n 15TieTeam knowledge followed, Sonnet vs Haiku: Raw notes, 60 lines: Claude Haiku 4.5 60% (6/10) (n 10, 95% interval 31%–83%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
Calculation: at these rates, about 17 runs per side would separate them.
- Team knowledge followed, Sonnet vs Haiku: Dreamed notes80% (8/10)n 10100% (15/15)n 15TieTeam knowledge followed, Sonnet vs Haiku: Dreamed notes: Claude Haiku 4.5 80% (8/10) (n 10, 95% interval 49%–94%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
Calculation: at these rates, about 33 runs per side would separate them.
- Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines30% (3/10)n 10100% (15/15)n 15Claude Sonnet 5.5 aheadTeam knowledge followed, Sonnet vs Haiku: Handbook, 210 lines: Claude Haiku 4.5 30% (3/10) (n 10, 95% interval 11%–60%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Claude Sonnet 5.5 ahead.
- Team knowledge followed, Sonnet vs Haiku: Stop hook only80% (8/10)n 1067% (10/15)n 15TieTeam knowledge followed, Sonnet vs Haiku: Stop hook only: Claude Haiku 4.5 80% (8/10) (n 10, 95% interval 49%–94%); Claude Sonnet 5.5 67% (10/15) (n 15, 95% interval 42%–85%). Tie.
Calculation: at these rates, about 167 runs per side would separate them.
- Team knowledge followed, Sonnet vs Haiku: Curated + hook100% (10/10)n 10100% (15/15)n 15TieTeam knowledge followed, Sonnet vs Haiku: Curated + hook: Claude Haiku 4.5 100% (10/10) (n 10, 95% interval 72%–100%); Claude Sonnet 5.5 100% (15/15) (n 15, 95% interval 80%–100%). Tie.
- A stale README command: who still ran it?: No memory100% (10/10)n 1080% (12/15)n 15TieA stale README command: who still ran it?: No memory: Claude Haiku 4.5 100% (10/10) (n 10, 95% interval 72%–100%); Claude Sonnet 5.5 80% (12/15) (n 15, 95% interval 55%–93%). Tie.
Calculation: at these rates, about 33 runs per side would separate them.
- A stale README command: who still ran it?: /init CLAUDE.md100% (10/10)n 1087% (13/15)n 15TieA stale README command: who still ran it?: /init CLAUDE.md: Claude Haiku 4.5 100% (10/10) (n 10, 95% interval 72%–100%); Claude Sonnet 5.5 87% (13/15) (n 15, 95% interval 62%–96%). Tie.
Calculation: at these rates, about 57 runs per side would separate them.
- A stale README command: who still ran it?: Curated, 11 lines0% (0/10)n 100% (0/15)n 15TieA stale README command: who still ran it?: Curated, 11 lines: Claude Haiku 4.5 0% (0/10) (n 10, 95% interval 0%–28%); Claude Sonnet 5.5 0% (0/15) (n 15, 95% interval 0%–20%). Tie.
- A stale README command: who still ran it?: Raw notes, 60 lines100% (10/10)n 100% (0/15)n 15Claude Sonnet 5.5 aheadA stale README command: who still ran it?: Raw notes, 60 lines: Claude Haiku 4.5 100% (10/10) (n 10, 95% interval 72%–100%); Claude Sonnet 5.5 0% (0/15) (n 15, 95% interval 0%–20%). Claude Sonnet 5.5 ahead.
- A stale README command: who still ran it?: Dreamed notes10% (1/10)n 100% (0/15)n 15TieA stale README command: who still ran it?: Dreamed notes: Claude Haiku 4.5 10% (1/10) (n 10, 95% interval 1.8%–40%); Claude Sonnet 5.5 0% (0/15) (n 15, 95% interval 0%–20%). Tie.
Calculation: at these rates, about 75 runs per side would separate them.
- A stale README command: who still ran it?: Handbook, 210 lines0% (0/10)n 100% (0/15)n 15TieA stale README command: who still ran it?: Handbook, 210 lines: Claude Haiku 4.5 0% (0/10) (n 10, 95% interval 0%–28%); Claude Sonnet 5.5 0% (0/15) (n 15, 95% interval 0%–20%). Tie.
- A stale README command: who still ran it?: Stop hook only100% (10/10)n 1060% (9/15)n 15TieA stale README command: who still ran it?: Stop hook only: Claude Haiku 4.5 100% (10/10) (n 10, 95% interval 72%–100%); Claude Sonnet 5.5 60% (9/15) (n 15, 95% interval 36%–80%). Tie.
Calculation: at these rates, about 17 runs per side would separate them.
- A stale README command: who still ran it?: Curated + hook0% (0/10)n 100% (0/15)n 15TieA stale README command: who still ran it?: Curated + hook: Claude Haiku 4.5 0% (0/10) (n 10, 95% interval 0%–28%); Claude Sonnet 5.5 0% (0/15) (n 15, 95% interval 0%–20%). Tie.
- List-price cost per fully correct result (calculation): No memoryCalculation$0.38n 2$0.14n 9UnclearList-price cost per fully correct result (calculation): No memory, calculation: Claude Haiku 4.5 $0.38 (n 2); Claude Sonnet 5.5 $0.14 (n 9). Unclear.
- List-price cost per fully correct result (calculation): /init CLAUDE.mdCalculation$0.43n 2$0.14n 9UnclearList-price cost per fully correct result (calculation): /init CLAUDE.md, calculation: Claude Haiku 4.5 $0.43 (n 2); Claude Sonnet 5.5 $0.14 (n 9). Unclear.
- List-price cost per fully correct result (calculation): Curated, 11 linesCalculation$0.11n 7$0.082n 15UnclearList-price cost per fully correct result (calculation): Curated, 11 lines, calculation: Claude Haiku 4.5 $0.11 (n 7); Claude Sonnet 5.5 $0.082 (n 15). Unclear.
- List-price cost per fully correct result (calculation): Raw notes, 60 linesCalculation$0.13n 6$0.10n 14UnclearList-price cost per fully correct result (calculation): Raw notes, 60 lines, calculation: Claude Haiku 4.5 $0.13 (n 6); Claude Sonnet 5.5 $0.10 (n 14). Unclear.
- List-price cost per fully correct result (calculation): Dreamed notesCalculation$0.12n 7$0.090n 15UnclearList-price cost per fully correct result (calculation): Dreamed notes, calculation: Claude Haiku 4.5 $0.12 (n 7); Claude Sonnet 5.5 $0.090 (n 15). Unclear.
- List-price cost per fully correct result (calculation): Handbook, 210 linesCalculation$0.26n 3$0.10n 15UnclearList-price cost per fully correct result (calculation): Handbook, 210 lines, calculation: Claude Haiku 4.5 $0.26 (n 3); Claude Sonnet 5.5 $0.10 (n 15). Unclear.
- List-price cost per fully correct result (calculation): Stop hook onlyCalculation$0.14n 8$0.13n 12UnclearList-price cost per fully correct result (calculation): Stop hook only, calculation: Claude Haiku 4.5 $0.14 (n 8); Claude Sonnet 5.5 $0.13 (n 12). Unclear.
- List-price cost per fully correct result (calculation): Curated + hookCalculation$0.096n 9$0.084n 15UnclearList-price cost per fully correct result (calculation): Curated + hook, calculation: Claude Haiku 4.5 $0.096 (n 9); Claude Sonnet 5.5 $0.084 (n 15). Unclear.
- Time per session: No memory54.2 sn 1018.0 sn 15UnclearTime per session: No memory: Claude Haiku 4.5 54.2 s (n 10); Claude Sonnet 5.5 18.0 s (n 15). Unclear.
- Time per session: /init CLAUDE.md52.9 sn 1019.0 sn 15UnclearTime per session: /init CLAUDE.md: Claude Haiku 4.5 52.9 s (n 10); Claude Sonnet 5.5 19.0 s (n 15). Unclear.
- Time per session: Curated, 11 lines51.7 sn 1021.9 sn 15UnclearTime per session: Curated, 11 lines: Claude Haiku 4.5 51.7 s (n 10); Claude Sonnet 5.5 21.9 s (n 15). Unclear.
- Time per session: Raw notes, 60 lines51.0 sn 1026.8 sn 15UnclearTime per session: Raw notes, 60 lines: Claude Haiku 4.5 51.0 s (n 10); Claude Sonnet 5.5 26.8 s (n 15). Unclear.
- Time per session: Dreamed notes51.8 sn 1027.2 sn 15UnclearTime per session: Dreamed notes: Claude Haiku 4.5 51.8 s (n 10); Claude Sonnet 5.5 27.2 s (n 15). Unclear.
- Time per session: Handbook, 210 lines49.9 sn 1023.6 sn 15UnclearTime per session: Handbook, 210 lines: Claude Haiku 4.5 49.9 s (n 10); Claude Sonnet 5.5 23.6 s (n 15). Unclear.
- Time per session: Stop hook only68.5 sn 1027.3 sn 15UnclearTime per session: Stop hook only: Claude Haiku 4.5 68.5 s (n 10); Claude Sonnet 5.5 27.3 s (n 15). Unclear.
- Time per session: Curated + hook52.8 sn 1022.0 sn 15UnclearTime per session: Curated + hook: Claude Haiku 4.5 52.8 s (n 10); Claude Sonnet 5.5 22.0 s (n 15). Unclear.
Jev vs Claude as a router: accuracy and cost
- Typed routing decisions answered exactly right89% (73/82)n 8294% (77/82)n 82TieTyped routing decisions answered exactly right: Claude Haiku 4.5 89% (73/82) (n 82, 95% interval 80%–94%); Claude Sonnet 5.5 94% (77/82) (n 82, 95% interval 87%–97%). Tie.
Calculation: at these rates, about 497 runs per side would separate them.
- Per-question accuracy94% (183/194)n 19497% (189/194)n 194TiePer-question accuracy: Claude Haiku 4.5 94% (183/194) (n 194, 95% interval 90%–97%); Claude Sonnet 5.5 97% (189/194) (n 194, 95% interval 94%–99%). Tie.
Calculation: at these rates, about 627 runs per side would separate them.
- Exact rate by decision type: Failure class94% (17/18)n 18100% (18/18)n 18TieExact rate by decision type: Failure class: Claude Haiku 4.5 94% (17/18) (n 18, 95% interval 74%–99%); Claude Sonnet 5.5 100% (18/18) (n 18, 95% interval 82%–100%). Tie.
Calculation: at these rates, about 135 runs per side would separate them.
- Exact rate by decision type: Message intent100% (20/20)n 20100% (20/20)n 20TieExact rate by decision type: Message intent: Claude Haiku 4.5 100% (20/20) (n 20, 95% interval 84%–100%); Claude Sonnet 5.5 100% (20/20) (n 20, 95% interval 84%–100%). Tie.
- Exact rate by decision type: Is it a rule?100% (12/12)n 12100% (12/12)n 12TieExact rate by decision type: Is it a rule?: Claude Haiku 4.5 100% (12/12) (n 12, 95% interval 76%–100%); Claude Sonnet 5.5 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
- Exact rate by decision type: Context shape75% (24/32)n 3284% (27/32)n 32TieExact rate by decision type: Context shape: Claude Haiku 4.5 75% (24/32) (n 32, 95% interval 58%–87%); Claude Sonnet 5.5 84% (27/32) (n 32, 95% interval 68%–93%). Tie.
Calculation: at these rates, about 271 runs per side would separate them.
- Cost per 1,000 routing decisionsCalculation$8.92n 82$5.00n 82UnclearCost per 1,000 routing decisions, calculation: Claude Haiku 4.5 $8.92 (n 82); Claude Sonnet 5.5 $5.00 (n 82). Unclear.
- Time per routing decision (Wall time (CLI))12,674 msn 822,598 msn 82Claude Sonnet 5.5 aheadTime per routing decision (Wall time (CLI)): Claude Haiku 4.5 12,674 ms (n 82, median to p95 12.67 s–34.41 s); Claude Sonnet 5.5 2,598 ms (n 82, median to p95 2.6 s–4.3 s). Claude Sonnet 5.5 ahead.
- Time per routing decision (Model time (API))10,734 msn 821,599 msn 82Claude Sonnet 5.5 aheadTime per routing decision (Model time (API)): Claude Haiku 4.5 10,734 ms (n 82, median to p95 10.73 s–32.07 s); Claude Sonnet 5.5 1,599 ms (n 82, median to p95 1.6 s–2.57 s). Claude Sonnet 5.5 ahead.
Routing overhead: deterministic policy vs LLM routers vs Jev
- Time to make one routing decision12,543 msn 822,597 msn 82Claude Sonnet 5.5 aheadTime to make one routing decision: Claude Haiku 4.5 12,543 ms (n 82, median to p95 12.54 s–34.48 s); Claude Sonnet 5.5 2,597 ms (n 82, median to p95 2.6 s–4.3 s). Claude Sonnet 5.5 ahead.
- Where an LLM router’s time goes: model vs CLI (Model API time)10,508 msn 821,596 msn 82Claude Sonnet 5.5 aheadWhere an LLM router’s time goes: model vs CLI (Model API time): Claude Haiku 4.5 10,508 ms (n 82, median to p95 10.51 s–32.13 s); Claude Sonnet 5.5 1,596 ms (n 82, median to p95 1.6 s–2.58 s). Claude Sonnet 5.5 ahead.
- Where an LLM router’s time goes: model vs CLI (CLI and harness time)1,698 msn 82973 msn 82Claude Sonnet 5.5 aheadWhere an LLM router’s time goes: model vs CLI (CLI and harness time): Claude Haiku 4.5 1,698 ms (n 82, median to p95 1.7 s–2.68 s); Claude Sonnet 5.5 973 ms (n 82, median to p95 973 ms–1.28 s). Claude Sonnet 5.5 ahead.
- Routing calls that returned a decision100% (82/82)n 82100% (82/82)n 82TieRouting calls that returned a decision: Claude Haiku 4.5 100% (82/82) (n 82, 95% interval 96%–100%); Claude Sonnet 5.5 100% (82/82) (n 82, 95% interval 96%–100%). Tie.
- Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))Calculation$441.74$247.30UnclearAdded routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)), calculation: Claude Haiku 4.5 $441.74; Claude Sonnet 5.5 $247.30. Unclear.
- Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))Calculation$62.47$34.97UnclearAdded routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)), calculation: Claude Haiku 4.5 $62.47; Claude Sonnet 5.5 $34.97. Unclear.
- Added routing delay per task (calculation) (Every model call routed (49.5 per task))Calculation620.9 s128.6 sUnclearAdded routing delay per task (calculation) (Every model call routed (49.5 per task)), calculation: Claude Haiku 4.5 620.9 s; Claude Sonnet 5.5 128.6 s. Unclear.
- Added routing delay per task (calculation) (Only System One decisions (7 per task))Calculation87.8 s18.2 sUnclearAdded routing delay per task (calculation) (Only System One decisions (7 per task)), calculation: Claude Haiku 4.5 87.8 s; Claude Sonnet 5.5 18.2 s. Unclear.
Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
- Strict pass rate: single call vs agent loop on eight hard tasks46% (11/24)n 24100% (24/24)n 24Claude Sonnet 5.5 aheadStrict pass rate: single call vs agent loop on eight hard tasks: Claude Haiku 4.5 46% (11/24) (n 24, 95% interval 28%–65%); Claude Sonnet 5.5 100% (24/24) (n 24, 95% interval 86%–100%). Claude Sonnet 5.5 ahead.
- Strict passes per task: single call vs agent loop: Interval merge fix100% (3/3)n 3100% (3/3)n 3TieStrict passes per task: single call vs agent loop: Interval merge fix: Claude Haiku 4.5 100% (3/3) (n 3, 95% interval 44%–100%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.
- Strict passes per task: single call vs agent loop: DST day-length fix33% (1/3)n 3100% (3/3)n 3TieStrict passes per task: single call vs agent loop: DST day-length fix: Claude Haiku 4.5 33% (1/3) (n 3, 95% interval 6.2%–79%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.
Calculation: at these rates, about 7 runs per side would separate them.
- Strict passes per task: single call vs agent loop: CSV parser67% (2/3)n 3100% (3/3)n 3TieStrict passes per task: single call vs agent loop: CSV parser: Claude Haiku 4.5 67% (2/3) (n 3, 95% interval 21%–94%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.
Calculation: at these rates, about 20 runs per side would separate them.
- Strict passes per task: single call vs agent loop: Event-loop order0% (0/3)n 3100% (3/3)n 3TieStrict passes per task: single call vs agent loop: Event-loop order: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.
Calculation: at these rates, about 4 runs per side would separate them.
- Strict passes per task: single call vs agent loop: Room schedule0% (0/3)n 3100% (3/3)n 3TieStrict passes per task: single call vs agent loop: Room schedule: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.
Calculation: at these rates, about 4 runs per side would separate them.
- Strict passes per task: single call vs agent loop: SemVer regex100% (3/3)n 3100% (3/3)n 3TieStrict passes per task: single call vs agent loop: SemVer regex: Claude Haiku 4.5 100% (3/3) (n 3, 95% interval 44%–100%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.
- Strict passes per task: single call vs agent loop: Money refactor67% (2/3)n 3100% (3/3)n 3TieStrict passes per task: single call vs agent loop: Money refactor: Claude Haiku 4.5 67% (2/3) (n 3, 95% interval 21%–94%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.
Calculation: at these rates, about 20 runs per side would separate them.
- Strict passes per task: single call vs agent loop: SQL report0% (0/3)n 3100% (3/3)n 3TieStrict passes per task: single call vs agent loop: SQL report: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 100% (3/3) (n 3, 95% interval 44%–100%). Tie.
Calculation: at these rates, about 4 runs per side would separate them.
- Total time per attempt: single call vs agent loop39.0 sn 247.75 sn 24UnclearTotal time per attempt: single call vs agent loop: Claude Haiku 4.5 39.0 s (n 24, run range 15.3 s–75.1 s); Claude Sonnet 5.5 7.75 s (n 24, run range 2.3 s–34.8 s). Unclear.
- Tokens per attempt: single call vs agent loop (Input tokens (cache reads included))3,941n 242,281n 24UnclearTokens per attempt: single call vs agent loop (Input tokens (cache reads included)): Claude Haiku 4.5 3,941 (n 24, run range 3,879–4,221); Claude Sonnet 5.5 2,281 (n 24, run range 2,234–2,669). Unclear.
- Tokens per attempt: single call vs agent loop (Output tokens)5,064n 241,050n 24UnclearTokens per attempt: single call vs agent loop (Output tokens): Claude Haiku 4.5 5,064 (n 24, run range 1,899–9,321); Claude Sonnet 5.5 1,050 (n 24, run range 176–3,895). Unclear.
- Tool calls per agent-loop attempt3n 240n 16UnclearTool calls per agent-loop attempt: Claude Haiku 4.5 3 (n 24, run range 2–18); Claude Sonnet 5.5 0 (n 16, run range 0–3). Unclear.
- List-price cost per strict pass: single call vs agent loop (calculation)Calculation$0.067n 24$0.014n 24UnclearList-price cost per strict pass: single call vs agent loop (calculation), calculation: Claude Haiku 4.5 $0.067 (n 24); Claude Sonnet 5.5 $0.014 (n 24). Unclear.
Does thinking pay for Claude Haiku 4.5? Thinking on vs off
- Haiku thinking study: typed routing decisions answered exactly right (Exact decisions (every scored question right))87% (71/82)n 8294% (77/82)n 82TieHaiku thinking study: typed routing decisions answered exactly right (Exact decisions (every scored question right)): Claude Haiku 4.5 87% (71/82) (n 82, 95% interval 78%–92%); Claude Sonnet 5.5 94% (77/82) (n 82, 95% interval 87%–97%). Tie.
Calculation: at these rates, about 235 runs per side would separate them.
- Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy)91% (177/194)n 19497% (189/194)n 194TieHaiku thinking study: typed routing decisions answered exactly right (Per-question accuracy): Claude Haiku 4.5 91% (177/194) (n 194, 95% interval 86%–94%); Claude Sonnet 5.5 97% (189/194) (n 194, 95% interval 94%–99%). Tie.
Calculation: at these rates, about 200 runs per side would separate them.
- Haiku thinking study: time per routing decision (Wall time (CLI))4.66 sn 822.60 sn 82Claude Sonnet 5.5 aheadHaiku thinking study: time per routing decision (Wall time (CLI)): Claude Haiku 4.5 4.66 s (n 82, median to p95 4.7 s–8.2 s); Claude Sonnet 5.5 2.60 s (n 82, median to p95 2.6 s–4.3 s). Claude Sonnet 5.5 ahead.
- Haiku thinking study: time per routing decision (Model time (API))3.79 sn 821.60 sn 82Claude Sonnet 5.5 aheadHaiku thinking study: time per routing decision (Model time (API)): Claude Haiku 4.5 3.79 s (n 82, median to p95 3.8 s–7.4 s); Claude Sonnet 5.5 1.60 s (n 82, median to p95 1.6 s–2.6 s). Claude Sonnet 5.5 ahead.
- Haiku thinking study: thinking and visible output tokens per routing decision (Thinking tokens)0n 822n 82UnclearHaiku thinking study: thinking and visible output tokens per routing decision (Thinking tokens): Claude Haiku 4.5 0 (n 82); Claude Sonnet 5.5 2 (n 82). Unclear.
- Haiku thinking study: thinking and visible output tokens per routing decision (Visible output tokens)366n 82105n 82UnclearHaiku thinking study: thinking and visible output tokens per routing decision (Visible output tokens): Claude Haiku 4.5 366 (n 82); Claude Sonnet 5.5 105 (n 82). Unclear.
- Haiku thinking study: list-price cost per 1,000 routing decisions (calculation)Calculation$3.36n 82$7.32n 82UnclearHaiku thinking study: list-price cost per 1,000 routing decisions (calculation), calculation: Claude Haiku 4.5 $3.36 (n 82); Claude Sonnet 5.5 $7.32 (n 82). Unclear.
Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
- Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)Calculation0% (0/24)n 24100% (12/12)n 12Claude Sonnet 5.5 aheadDoes a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON), calculation: Claude Haiku 4.5 0% (0/24) (n 24, 95% interval 0%–14%); Claude Sonnet 5.5 100% (12/12) (n 12, 95% interval 76%–100%). Claude Sonnet 5.5 ahead.
- Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))Calculation71% (17/24)n 24100% (12/12)n 12TieDoes a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss)), calculation: Claude Haiku 4.5 71% (17/24) (n 24, 95% interval 51%–85%); Claude Sonnet 5.5 100% (12/12) (n 12, 95% interval 76%–100%). Tie.
- What each call produced: strict pass, format miss, wrong values or error (Strict pass)0n 2412n 12UnclearWhat each call produced: strict pass, format miss, wrong values or error (Strict pass): Claude Haiku 4.5 0 (n 24); Claude Sonnet 5.5 12 (n 12). Unclear.
- What each call produced: strict pass, format miss, wrong values or error (Format miss)17n 240n 12UnclearWhat each call produced: strict pass, format miss, wrong values or error (Format miss): Claude Haiku 4.5 17 (n 24); Claude Sonnet 5.5 0 (n 12). Unclear.
- What each call produced: strict pass, format miss, wrong values or error (Wrong values)7n 240n 12UnclearWhat each call produced: strict pass, format miss, wrong values or error (Wrong values): Claude Haiku 4.5 7 (n 24); Claude Sonnet 5.5 0 (n 12). Unclear.
- What each call produced: strict pass, format miss, wrong values or error (Error)0n 240n 12TieWhat each call produced: strict pass, format miss, wrong values or error (Error): Claude Haiku 4.5 0 (n 24); Claude Sonnet 5.5 0 (n 12). Tie.
- Time per call, instructions vs schema mode9.52 sn 243.52 sn 12Claude Sonnet 5.5 aheadTime per call, instructions vs schema mode: Claude Haiku 4.5 9.52 s (n 24, run range 5.7 s–17 s); Claude Sonnet 5.5 3.52 s (n 12, run range 2.7 s–4.1 s). Claude Sonnet 5.5 ahead.
- Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call)1,128n 24368n 12UnclearOutput and reasoning tokens per call, instructions vs schema mode (Median output tokens per call): Claude Haiku 4.5 1,128 (n 24); Claude Sonnet 5.5 368 (n 12). Unclear.
Jev vs Claude routers on unseen decisions: a blind holdout
- Unseen routing decisions answered exactly right79% (44/56)n 5688% (49/56)n 56TieUnseen routing decisions answered exactly right: Claude Haiku 4.5 79% (44/56) (n 56, 95% interval 66%–87%); Claude Sonnet 5.5 88% (49/56) (n 56, 95% interval 76%–94%). Tie.
Calculation: at these rates, about 259 runs per side would separate them.
- Per-question accuracy on unseen decisions82% (102/125)n 12592% (115/125)n 125TiePer-question accuracy on unseen decisions: Claude Haiku 4.5 82% (102/125) (n 125, 95% interval 74%–87%); Claude Sonnet 5.5 92% (115/125) (n 125, 95% interval 86%–96%). Tie.
Calculation: at these rates, about 155 runs per side would separate them.
- Exact rate on unseen decisions, by decision type: Failure class93% (13/14)n 14100% (14/14)n 14TieExact rate on unseen decisions, by decision type: Failure class: Claude Haiku 4.5 93% (13/14) (n 14, 95% interval 69%–99%); Claude Sonnet 5.5 100% (14/14) (n 14, 95% interval 78%–100%). Tie.
Calculation: at these rates, about 106 runs per side would separate them.
- Exact rate on unseen decisions, by decision type: Message intent93% (13/14)n 14100% (14/14)n 14TieExact rate on unseen decisions, by decision type: Message intent: Claude Haiku 4.5 93% (13/14) (n 14, 95% interval 69%–99%); Claude Sonnet 5.5 100% (14/14) (n 14, 95% interval 78%–100%). Tie.
Calculation: at these rates, about 106 runs per side would separate them.
- Exact rate on unseen decisions, by decision type: Is it a rule?93% (13/14)n 1493% (13/14)n 14TieExact rate on unseen decisions, by decision type: Is it a rule?: Claude Haiku 4.5 93% (13/14) (n 14, 95% interval 69%–99%); Claude Sonnet 5.5 93% (13/14) (n 14, 95% interval 69%–99%). Tie.
- Exact rate on unseen decisions, by decision type: Context shape36% (5/14)n 1457% (8/14)n 14TieExact rate on unseen decisions, by decision type: Context shape: Claude Haiku 4.5 36% (5/14) (n 14, 95% interval 16%–61%); Claude Sonnet 5.5 57% (8/14) (n 14, 95% interval 33%–79%). Tie.
Calculation: at these rates, about 82 runs per side would separate them.
- Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm))89% (73/82)n 8294% (77/82)n 82TieTuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm)): Claude Haiku 4.5 89% (73/82) (n 82, 95% interval 80%–94%); Claude Sonnet 5.5 94% (77/82) (n 82, 95% interval 87%–97%). Tie.
Calculation: at these rates, about 497 runs per side would separate them.
- Tuned case set vs unseen holdout: exact rate per router (Unseen holdout)79% (44/56)n 5688% (49/56)n 56TieTuned case set vs unseen holdout: exact rate per router (Unseen holdout): Claude Haiku 4.5 79% (44/56) (n 56, 95% interval 66%–87%); Claude Sonnet 5.5 88% (49/56) (n 56, 95% interval 76%–94%). Tie.
Calculation: at these rates, about 259 runs per side would separate them.
- Time per routing decision, by route (Wall time)9.44 sn 562.36 sn 56Claude Sonnet 5.5 aheadTime per routing decision, by route (Wall time): Claude Haiku 4.5 9.44 s (n 56, median to p95 9.4 s–25.4 s); Claude Sonnet 5.5 2.36 s (n 56, median to p95 2.4 s–3.7 s). Claude Sonnet 5.5 ahead.
- Time per routing decision, by route (Model time (API, CLI-reported))7.52 sn 561.49 sn 56Claude Sonnet 5.5 aheadTime per routing decision, by route (Model time (API, CLI-reported)): Claude Haiku 4.5 7.52 s (n 56, median to p95 7.5 s–23.9 s); Claude Sonnet 5.5 1.49 s (n 56, median to p95 1.5 s–2.4 s). Claude Sonnet 5.5 ahead.
- Cost per 1,000 unseen routing decisionsCalculation$7.13n 56$7.24n 56UnclearCost per 1,000 unseen routing decisions, calculation: Claude Haiku 4.5 $7.13 (n 56); Claude Sonnet 5.5 $7.24 (n 56). Unclear.
How much of an AI bill is thinking? Reasoning tokens by model and effort
- Reasoning share of output tokens per call on hard tasks (calculation)Calculation91.7%n 2454.5%n 24UnclearReasoning share of output tokens per call on hard tasks (calculation), calculation: Claude Haiku 4.5 91.7% (n 24, run range 76%–99%); Claude Sonnet 5.5 54.5% (n 24, run range 0%–96%). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation$0.024n 24$0.0067n 24UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)), calculation: Claude Haiku 4.5 $0.024 (n 24); Claude Sonnet 5.5 $0.0067 (n 24). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation$0.0018n 24$0.0037n 24UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)), calculation: Claude Haiku 4.5 $0.0018 (n 24); Claude Sonnet 5.5 $0.0037 (n 24). Unclear.
- List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation$0.0045n 24$0.0040n 24UnclearList-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced)), calculation: Claude Haiku 4.5 $0.0045 (n 24); Claude Sonnet 5.5 $0.0040 (n 24). Unclear.
- Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation91.7%n 2454.5%n 24UnclearReasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks), calculation: Claude Haiku 4.5 91.7% (n 24, run range 76%–99%); Claude Sonnet 5.5 54.5% (n 24, run range 0%–96%). Unclear.
- Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation90.2%n 150%n 15UnclearReasoning share on short tasks vs hard tasks (calculation) (Five short tasks), calculation: Claude Haiku 4.5 90.2% (n 15, run range 73%–98%); Claude Sonnet 5.5 0% (n 15, run range 0%–73%). Unclear.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
- Time to first text: a 250-line answer, six models4.00 sn 41.96 sn 4UnclearTime to first text: a 250-line answer, six models: Claude Haiku 4.5 4.00 s (n 4, run range 2.8 s–6.4 s); Claude Sonnet 5.5 1.96 s (n 4, run range 0.9 s–4.1 s). Unclear.
- Output speed after the first text: visible tokens per second (calculation)Calculation153n 4232n 4UnclearOutput speed after the first text: visible tokens per second (calculation), calculation: Claude Haiku 4.5 153 (n 4, run range 153–216); Claude Sonnet 5.5 232 (n 4, run range 230–233). Unclear.
- Output speed in characters per second after the first text (calculation)Calculation547n 3517n 4UnclearOutput speed in characters per second after the first text (calculation), calculation: Claude Haiku 4.5 547 (n 3, run range 546–548); Claude Sonnet 5.5 517 (n 4, run range 513–519). Unclear.
- Time to first text as the prompt grows: 1kCalculation1.93 sn 31.45 sn 3UnclearTime to first text as the prompt grows: 1k, calculation: Claude Haiku 4.5 1.93 s (n 3, run range 1.9 s–2 s); Claude Sonnet 5.5 1.45 s (n 3, run range 1.2 s–1.7 s). Unclear.
- Time to first text as the prompt grows: 16kCalculation2.27 sn 31.78 sn 3UnclearTime to first text as the prompt grows: 16k, calculation: Claude Haiku 4.5 2.27 s (n 3, run range 2.2 s–2.5 s); Claude Sonnet 5.5 1.78 s (n 3, run range 1.6 s–2.1 s). Unclear.
- Time to first text as the prompt grows: 64kCalculation2.78 sn 33.07 sn 3UnclearTime to first text as the prompt grows: 64k, calculation: Claude Haiku 4.5 2.78 s (n 3, run range 2.5 s–2.9 s); Claude Sonnet 5.5 3.07 s (n 3, run range 1.4 s–3.6 s). Unclear.
- Total time per call by prompt size (1k prompt)2.34 sn 31.78 sn 3UnclearTotal time per call by prompt size (1k prompt): Claude Haiku 4.5 2.34 s (n 3, run range 2.2 s–2.5 s); Claude Sonnet 5.5 1.78 s (n 3, run range 1.6 s–2.1 s). Unclear.
- Total time per call by prompt size (16k prompt)2.79 sn 32.10 sn 3UnclearTotal time per call by prompt size (16k prompt): Claude Haiku 4.5 2.79 s (n 3, run range 2.6 s–2.8 s); Claude Sonnet 5.5 2.10 s (n 3, run range 2 s–2.5 s). Unclear.
- Total time per call by prompt size (64k prompt)3.13 sn 33.44 sn 3UnclearTotal time per call by prompt size (64k prompt): Claude Haiku 4.5 3.13 s (n 3, run range 2.8 s–3.3 s); Claude Sonnet 5.5 3.44 s (n 3, run range 1.7 s–4.4 s). Unclear.
- Exact lookup answers at the 1k, 16k and 64k prompt-size targets100% (9/9)n 9100% (9/9)n 9TieExact lookup answers at the 1k, 16k and 64k prompt-size targets: Claude Haiku 4.5 100% (9/9) (n 9, 95% interval 70%–100%); Claude Sonnet 5.5 100% (9/9) (n 9, 95% interval 70%–100%). Tie.
Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Interval merge fixCalculation$0.018n 3$0.0056n 3UnclearList-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Interval merge fix, calculation: Claude Haiku 4.5 $0.018 (n 3); Claude Sonnet 5.5 $0.0056 (n 3). Unclear.
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): DST day lengthCalculation$0.029n 3$0.025n 3UnclearList-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): DST day length, calculation: Claude Haiku 4.5 $0.029 (n 3); Claude Sonnet 5.5 $0.025 (n 3). Unclear.
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): CSV parserCalculation$0.029n 3$0.015n 3UnclearList-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): CSV parser, calculation: Claude Haiku 4.5 $0.029 (n 3); Claude Sonnet 5.5 $0.015 (n 3). Unclear.
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Event-loop orderCalculation$0.037n 3$0.016n 3UnclearList-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Event-loop order, calculation: Claude Haiku 4.5 $0.037 (n 3); Claude Sonnet 5.5 $0.016 (n 3). Unclear.
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Room scheduleCalculation$0.036n 3$0.012n 3UnclearList-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Room schedule, calculation: Claude Haiku 4.5 $0.036 (n 3); Claude Sonnet 5.5 $0.012 (n 3). Unclear.
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SemVer regexCalculation$0.042n 3$0.0051n 3UnclearList-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SemVer regex, calculation: Claude Haiku 4.5 $0.042 (n 3); Claude Sonnet 5.5 $0.0051 (n 3). Unclear.
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Money refactorCalculation$0.020n 3$0.0096n 3UnclearList-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Money refactor, calculation: Claude Haiku 4.5 $0.020 (n 3); Claude Sonnet 5.5 $0.0096 (n 3). Unclear.
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SQLite report queryCalculation$0.036n 3$0.017n 3UnclearList-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SQLite report query, calculation: Claude Haiku 4.5 $0.036 (n 3); Claude Sonnet 5.5 $0.017 (n 3). Unclear.
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
- Pass rate on 4 harder tasks (Strict pass)0% (0/12)n 1238% (6/16)n 16TiePass rate on 4 harder tasks (Strict pass): Claude Haiku 4.5 0% (0/12) (n 12, 95% interval 0%–24%); Claude Sonnet 5.5 38% (6/16) (n 16, 95% interval 18%–61%). Tie.
Calculation: at these rates, about 18 runs per side would separate them.
- Pass rate on 4 harder tasks (Lenient (format misses counted))0% (0/12)n 1238% (6/16)n 16TiePass rate on 4 harder tasks (Lenient (format misses counted)): Claude Haiku 4.5 0% (0/12) (n 12, 95% interval 0%–24%); Claude Sonnet 5.5 38% (6/16) (n 16, 95% interval 18%–61%). Tie.
Calculation: at these rates, about 18 runs per side would separate them.
- Calls that tried a tool although tools were off8% (1/12)n 1231% (5/16)n 16UnclearCalls that tried a tool although tools were off: Claude Haiku 4.5 8% (1/12) (n 12, 95% interval 1.5%–35%); Claude Sonnet 5.5 31% (5/16) (n 16, 95% interval 14%–56%). Unclear.
- Strict pass rate by task: 10x10 nonogram0% (0/3)n 3100% (4/4)n 4TieStrict pass rate by task: 10x10 nonogram: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 100% (4/4) (n 4, 95% interval 51%–100%). Tie.
Calculation: at these rates, about 4 runs per side would separate them.
- Strict pass rate by task: Sudoku, 22 givens0% (0/3)n 30% (0/4)n 4TieStrict pass rate by task: Sudoku, 22 givens: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 0% (0/4) (n 4, 95% interval 0%–49%). Tie.
- Strict pass rate by task: 6x6 Skyscrapers0% (0/3)n 30% (0/4)n 4TieStrict pass rate by task: 6x6 Skyscrapers: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 0% (0/4) (n 4, 95% interval 0%–49%). Tie.
- Strict pass rate by task: Seeded shuffle output0% (0/3)n 350% (2/4)n 4TieStrict pass rate by task: Seeded shuffle output: Claude Haiku 4.5 0% (0/3) (n 3, 95% interval 0%–56%); Claude Sonnet 5.5 50% (2/4) (n 4, 95% interval 15%–85%). Tie.
Calculation: at these rates, about 11 runs per side would separate them.
- Total time per call on harder tasks109.0 sn 1070.4 sn 12UnclearTotal time per call on harder tasks: Claude Haiku 4.5 109.0 s (n 10, run range 25.7 s–224 s); Claude Sonnet 5.5 70.4 s (n 12, run range 4.3 s–210 s). Unclear.
- Output tokens per call on harder tasks (Output tokens)12,508n 109,287n 12UnclearOutput tokens per call on harder tasks (Output tokens): Claude Haiku 4.5 12,508 (n 10, run range 2,965–26,532); Claude Sonnet 5.5 9,287 (n 12, run range 407–27,921). Unclear.
| Metric | Claude Haiku 4.5 | Claude Sonnet 5.5 | n | Interval or range | Outcome | Basis | Study |
|---|---|---|---|---|---|---|---|
| Pass rate on five validated tasks | 100% (15/15)Claude Code · five short validated tasks | 80% (12/15)Claude Code · five short validated tasks | 15 | 95% CI: 80%–100% vs 55%–93% | Tie | The 95% intervals overlap (Claude Haiku 4.5 80% to 100%; Claude Sonnet 5.5 55% to 93%), so this sample cannot separate them. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Total time per call | 4.43 sClaude Code · five short validated tasks | 2.31 sClaude Code · five short validated tasks | 15 | range: 3.2 s–23.6 s vs 2.2 s–7.7 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 3.16 s to 23.6 s; Claude Sonnet 5.5 2.17 s to 7.73 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Time to first useful output | 3.63 sClaude Code · five short validated tasks | 1.56 sClaude Code · five short validated tasks | 15 | range: 2.8 s–22.3 s vs 1 s–6.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.78 s to 22.3 s; Claude Sonnet 5.5 0.99 s to 6.39 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Cache read) | 0Claude Code · five short validated tasks | 1,401Claude Code · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Input tokens per call: what the CLI sends (Other input) | 3,790Claude Code · five short validated tasks | 685Claude Code · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Output tokens per call (Output tokens) | 367Claude Code · five short validated tasks | 107Claude Code · five short validated tasks | 15 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per call (calculation)Calculation | $0.0057Claude Code · five short validated tasks | $0.0036Claude Code · five short validated tasks | 15 | range: $0.0051–$0.018 vs $0.0034–$0.01 | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 $0.0051 to $0.018; Claude Sonnet 5.5 $0.0034 to $0.010); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| List-price cost per passing answer (calculation)Calculation | $0.0084Claude Code · five short validated tasks | $0.0062Claude Code · five short validated tasks | 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.0084 vs $0.0062) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head |
| Pass rate on eight hard tasks (Strict pass) | 46% (11/24)Claude Code · eight hard validated tasks | 100% (24/24)Claude Code · eight hard validated tasks | 24 | 95% CI: 28%–65% vs 86%–100% | Claude Sonnet 5.5 ahead | The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%). | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Pass rate on eight hard tasks (Lenient (format misses counted)) | 67% (16/24)Claude Code · eight hard validated tasks | 100% (24/24)Claude Code · eight hard validated tasks | 24 | 95% CI: 47%–82% vs 86%–100% | Claude Sonnet 5.5 ahead | The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Sonnet 5.5 86% to 100%). | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Total time per call on hard tasks (separate batches) | 39.0 sClaude Code · eight hard validated tasks | 7.75 sClaude Code · eight hard validated tasks | 24 | range: 15.3 s–75.1 s vs 2.3 s–34.8 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; Claude Sonnet 5.5 2.26 s to 34.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Time to first useful output on hard tasks | 35.5 sClaude Code · eight hard validated tasks | 5.95 sClaude Code · eight hard validated tasks | 24 | range: 12.9 s–70.3 s vs 0.9 s–30.6 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 12.9 s to 70.3 s; Claude Sonnet 5.5 0.86 s to 30.6 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Output tokens per call on hard tasks (Output tokens) | 5,064Claude Code · eight hard validated tasks | 1,050Claude Code · eight hard validated tasks | 24 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| List-price cost per strict pass on hard tasks (calculation)Calculation | $0.067Claude Code · eight hard validated tasks | $0.014Claude Code · eight hard validated tasks | 24 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.067 vs $0.014, 4.7x) is not tested against run-to-run variation. | Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks |
| Same prompt, 10 times: strict pass rate (Exact number) | 0% (0/10)Claude Code · same prompt repeated 10 times | 100% (10/10)Claude Code · same prompt repeated 10 times | 10 | 95% CI: 0%–28% vs 72%–100% | Claude Sonnet 5.5 ahead | The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 72% to 100%). | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: strict pass rate (JSON object) | 10% (1/10)Claude Code · same prompt repeated 10 times | 100% (10/10)Claude Code · same prompt repeated 10 times | 10 | 95% CI: 1.8%–40% vs 72%–100% | Claude Sonnet 5.5 ahead | The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 72% to 100%). | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: strict pass rate (Code fix) | 100% (10/10)Claude Code · same prompt repeated 10 times | 100% (10/10)Claude Code · same prompt repeated 10 times | 10 | 95% CI: 72%–100% vs 72%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 72% to 100%), so this sample cannot separate them. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (Exact number) | 1Claude Code · same prompt repeated 10 times | 1Claude Code · same prompt repeated 10 times | 10 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (JSON object) | 1Claude Code · same prompt repeated 10 times | 1Claude Code · same prompt repeated 10 times | 10 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: how many different answers (Code fix) | 6Claude Code · same prompt repeated 10 times | 3Claude Code · same prompt repeated 10 times | 10 | none recorded | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (Exact number) | 5.06 sClaude Code · same prompt repeated 10 times | 6.89 sClaude Code · same prompt repeated 10 times | 10 | range: 4.4 s–6.2 s vs 5.8 s–7.8 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 4.42 s to 6.20 s; Claude Sonnet 5.5 5.81 s to 7.81 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (JSON object) | 7.03 sClaude Code · same prompt repeated 10 times | 2.89 sClaude Code · same prompt repeated 10 times | 10 | range: 5.3 s–12.3 s vs 2.7 s–5.3 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 5.28 s to 12.3 s; Claude Sonnet 5.5 2.68 s to 5.30 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Same prompt, 10 times: time per call (Code fix) | 5.95 sClaude Code · same prompt repeated 10 times | 2.67 sClaude Code · same prompt repeated 10 times | 10 | range: 4.9 s–7.3 s vs 2.3 s–4.3 s | Claude Sonnet 5.5 ahead | The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; Claude Sonnet 5.5 2.32 s to 4.34 s). A range is not a confidence interval. | Prompt caching and run-to-run consistency in Claude Code and Codex CLI |
| Full pass rate by kind of memory: No memory | 20% (2/10) | 60% (9/15) | 10 / 15 | 95% CI: 5.7%–51% vs 36%–80% | Tie | The 95% intervals overlap (Claude Haiku 4.5 6% to 51%; Claude Sonnet 5.5 36% to 80%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Full pass rate by kind of memory: /init CLAUDE.md | 20% (2/10) | 60% (9/15) | 10 / 15 | 95% CI: 5.7%–51% vs 36%–80% | Tie | The 95% intervals overlap (Claude Haiku 4.5 6% to 51%; Claude Sonnet 5.5 36% to 80%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Full pass rate by kind of memory: Curated, 11 lines | 70% (7/10) | 100% (15/15) | 10 / 15 | 95% CI: 40%–89% vs 80%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 40% to 89%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Full pass rate by kind of memory: Raw notes, 60 lines | 60% (6/10) | 93% (14/15) | 10 / 15 | 95% CI: 31%–83% vs 70%–99% | Tie | The 95% intervals overlap (Claude Haiku 4.5 31% to 83%; Claude Sonnet 5.5 70% to 99%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Full pass rate by kind of memory: Dreamed notes | 70% (7/10) | 100% (15/15) | 10 / 15 | 95% CI: 40%–89% vs 80%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 40% to 89%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Full pass rate by kind of memory: Handbook, 210 lines | 30% (3/10) | 100% (15/15) | 10 / 15 | 95% CI: 11%–60% vs 80%–100% | Claude Sonnet 5.5 ahead | The 95% intervals do not overlap (Claude Haiku 4.5 11% to 60%; Claude Sonnet 5.5 80% to 100%). | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Full pass rate by kind of memory: Stop hook only | 80% (8/10) | 80% (12/15) | 10 / 15 | 95% CI: 49%–94% vs 55%–93% | Tie | The 95% intervals overlap (Claude Haiku 4.5 49% to 94%; Claude Sonnet 5.5 55% to 93%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Full pass rate by kind of memory: Curated + hook | 90% (9/10) | 100% (15/15) | 10 / 15 | 95% CI: 60%–98% vs 80%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 60% to 98%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Team knowledge followed, Sonnet vs Haiku: No memory | 0% (0/10) | 40% (6/15) | 10 / 15 | 95% CI: 0%–28% vs 20%–64% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 20% to 64%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md | 10% (1/10) | 67% (10/15) | 10 / 15 | 95% CI: 1.8%–40% vs 42%–85% | Claude Sonnet 5.5 ahead | The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 42% to 85%). | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Team knowledge followed, Sonnet vs Haiku: Curated, 11 lines | 80% (8/10) | 100% (15/15) | 10 / 15 | 95% CI: 49%–94% vs 80%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 49% to 94%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Team knowledge followed, Sonnet vs Haiku: Raw notes, 60 lines | 60% (6/10) | 100% (15/15) | 10 / 15 | 95% CI: 31%–83% vs 80%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 31% to 83%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Team knowledge followed, Sonnet vs Haiku: Dreamed notes | 80% (8/10) | 100% (15/15) | 10 / 15 | 95% CI: 49%–94% vs 80%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 49% to 94%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines | 30% (3/10) | 100% (15/15) | 10 / 15 | 95% CI: 11%–60% vs 80%–100% | Claude Sonnet 5.5 ahead | The 95% intervals do not overlap (Claude Haiku 4.5 11% to 60%; Claude Sonnet 5.5 80% to 100%). | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Team knowledge followed, Sonnet vs Haiku: Stop hook only | 80% (8/10) | 67% (10/15) | 10 / 15 | 95% CI: 49%–94% vs 42%–85% | Tie | The 95% intervals overlap (Claude Haiku 4.5 49% to 94%; Claude Sonnet 5.5 42% to 85%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Team knowledge followed, Sonnet vs Haiku: Curated + hook | 100% (10/10) | 100% (15/15) | 10 / 15 | 95% CI: 72%–100% vs 80%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 80% to 100%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| A stale README command: who still ran it?: No memory | 100% (10/10) | 80% (12/15) | 10 / 15 | 95% CI: 72%–100% vs 55%–93% | Tie | The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 55% to 93%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| A stale README command: who still ran it?: /init CLAUDE.md | 100% (10/10) | 87% (13/15) | 10 / 15 | 95% CI: 72%–100% vs 62%–96% | Tie | The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 62% to 96%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| A stale README command: who still ran it?: Curated, 11 lines | 0% (0/10) | 0% (0/15) | 10 / 15 | 95% CI: 0%–28% vs 0%–20% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 0% to 20%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| A stale README command: who still ran it?: Raw notes, 60 lines | 100% (10/10) | 0% (0/15) | 10 / 15 | 95% CI: 72%–100% vs 0%–20% | Claude Sonnet 5.5 ahead | The 95% intervals do not overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 0% to 20%). | Does memory help Claude Code? 8 kinds of agent memory, tested |
| A stale README command: who still ran it?: Dreamed notes | 10% (1/10) | 0% (0/15) | 10 / 15 | 95% CI: 1.8%–40% vs 0%–20% | Tie | The 95% intervals overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 0% to 20%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| A stale README command: who still ran it?: Handbook, 210 lines | 0% (0/10) | 0% (0/15) | 10 / 15 | 95% CI: 0%–28% vs 0%–20% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 0% to 20%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| A stale README command: who still ran it?: Stop hook only | 100% (10/10) | 60% (9/15) | 10 / 15 | 95% CI: 72%–100% vs 36%–80% | Tie | The 95% intervals overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 36% to 80%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| A stale README command: who still ran it?: Curated + hook | 0% (0/10) | 0% (0/15) | 10 / 15 | 95% CI: 0%–28% vs 0%–20% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 0% to 20%), so this sample cannot separate them. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| List-price cost per fully correct result (calculation): No memoryCalculation | $0.38 | $0.14 | 2 / 9 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.38 vs $0.14, 2.7x) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| List-price cost per fully correct result (calculation): /init CLAUDE.mdCalculation | $0.43 | $0.14 | 2 / 9 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.43 vs $0.14, 3.2x) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| List-price cost per fully correct result (calculation): Curated, 11 linesCalculation | $0.11 | $0.082 | 7 / 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.11 vs $0.082) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| List-price cost per fully correct result (calculation): Raw notes, 60 linesCalculation | $0.13 | $0.10 | 6 / 14 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.13 vs $0.10) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| List-price cost per fully correct result (calculation): Dreamed notesCalculation | $0.12 | $0.090 | 7 / 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.12 vs $0.090) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| List-price cost per fully correct result (calculation): Handbook, 210 linesCalculation | $0.26 | $0.10 | 3 / 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.26 vs $0.10, 2.6x) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| List-price cost per fully correct result (calculation): Stop hook onlyCalculation | $0.14 | $0.13 | 8 / 12 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.14 vs $0.13) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| List-price cost per fully correct result (calculation): Curated + hookCalculation | $0.096 | $0.084 | 9 / 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.096 vs $0.084) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Time per session: No memory | 54.2 s | 18.0 s | 10 / 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap (54.2 s vs 18.0 s, 3.0x) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Time per session: /init CLAUDE.md | 52.9 s | 19.0 s | 10 / 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap (52.9 s vs 19.0 s, 2.8x) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Time per session: Curated, 11 lines | 51.7 s | 21.9 s | 10 / 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap (51.7 s vs 21.9 s, 2.4x) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Time per session: Raw notes, 60 lines | 51.0 s | 26.8 s | 10 / 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap (51.0 s vs 26.8 s) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Time per session: Dreamed notes | 51.8 s | 27.2 s | 10 / 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap (51.8 s vs 27.2 s) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Time per session: Handbook, 210 lines | 49.9 s | 23.6 s | 10 / 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap (49.9 s vs 23.6 s, 2.1x) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Time per session: Stop hook only | 68.5 s | 27.3 s | 10 / 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap (68.5 s vs 27.3 s, 2.5x) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Time per session: Curated + hook | 52.8 s | 22.0 s | 10 / 15 | none recorded | Unclear | No interval or range was recorded for either side, so the gap (52.8 s vs 22.0 s, 2.4x) is not tested against run-to-run variation. | Does memory help Claude Code? 8 kinds of agent memory, tested |
| Typed routing decisions answered exactly right | 89% (73/82)typed routing decisions · via Claude Code | 94% (77/82)typed routing decisions · via Claude Code | 82 | 95% CI: 80%–94% vs 87%–97% | Tie | The 95% intervals overlap (Claude Haiku 4.5 80% to 94%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Per-question accuracy | 94% (183/194)typed routing decisions · via Claude Code | 97% (189/194)typed routing decisions · via Claude Code | 194 | 95% CI: 90%–97% vs 94%–99% | Tie | The 95% intervals overlap (Claude Haiku 4.5 90% to 97%; Claude Sonnet 5.5 94% to 99%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Failure class | 94% (17/18)typed routing decisions · via Claude Code | 100% (18/18)typed routing decisions · via Claude Code | 18 | 95% CI: 74%–99% vs 82%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 74% to 99%; Claude Sonnet 5.5 82% to 100%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Message intent | 100% (20/20)typed routing decisions · via Claude Code | 100% (20/20)typed routing decisions · via Claude Code | 20 | 95% CI: 84%–100% vs 84%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 84% to 100%; Claude Sonnet 5.5 84% to 100%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Is it a rule? | 100% (12/12)typed routing decisions · via Claude Code | 100% (12/12)typed routing decisions · via Claude Code | 12 | 95% CI: 76%–100% vs 76%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 76% to 100%; Claude Sonnet 5.5 76% to 100%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Exact rate by decision type: Context shape | 75% (24/32)typed routing decisions · via Claude Code | 84% (27/32)typed routing decisions · via Claude Code | 32 | 95% CI: 58%–87% vs 68%–93% | Tie | The 95% intervals overlap (Claude Haiku 4.5 58% to 87%; Claude Sonnet 5.5 68% to 93%), so this sample cannot separate them. | Jev vs Claude as a router: accuracy and cost |
| Cost per 1,000 routing decisionsCalculation | $8.92typed routing decisions · via Claude Code | $5.00typed routing decisions · via Claude Code | 82 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($8.92 vs $5.00) is not tested against run-to-run variation. | Jev vs Claude as a router: accuracy and cost |
| Time per routing decision (Wall time (CLI)) | 12,674 mstyped routing decisions · via Claude Code | 2,598 mstyped routing decisions · via Claude Code | 82 | p50–p95: 12.67 s–34.41 s vs 2.6 s–4.3 s | Claude Sonnet 5.5 ahead | Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 12,674 ms to 34,413 ms; Claude Sonnet 5.5 2,598 ms to 4,298 ms); not a confidence interval. | Jev vs Claude as a router: accuracy and cost |
| Time per routing decision (Model time (API)) | 10,734 mstyped routing decisions · via Claude Code | 1,599 mstyped routing decisions · via Claude Code | 82 | p50–p95: 10.73 s–32.07 s vs 1.6 s–2.57 s | Claude Sonnet 5.5 ahead | Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 10,734 ms to 32,072 ms; Claude Sonnet 5.5 1,599 ms to 2,574 ms); not a confidence interval. | Jev vs Claude as a router: accuracy and cost |
| Time to make one routing decision | 12,543 msthinking on · via Claude Code · routing overhead per decision | 2,597 mseffort low · via Claude Code · routing overhead per decision | 82 | p50–p95: 12.54 s–34.48 s vs 2.6 s–4.3 s | Claude Sonnet 5.5 ahead | Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 12,543 ms to 34,481 ms; Claude Sonnet 5.5 2,597 ms to 4,298 ms); not a confidence interval. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Where an LLM router’s time goes: model vs CLI (Model API time) | 10,508 msthinking on · via Claude Code · routing overhead per decision | 1,596 mseffort low · via Claude Code · routing overhead per decision | 82 | p50–p95: 10.51 s–32.13 s vs 1.6 s–2.58 s | Claude Sonnet 5.5 ahead | Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 10,508 ms to 32,132 ms; Claude Sonnet 5.5 1,596 ms to 2,583 ms); not a confidence interval. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Where an LLM router’s time goes: model vs CLI (CLI and harness time) | 1,698 msthinking on · via Claude Code · routing overhead per decision | 973 mseffort low · via Claude Code · routing overhead per decision | 82 | p50–p95: 1.7 s–2.68 s vs 973 ms–1.28 s | Claude Sonnet 5.5 ahead | Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 1,698 ms to 2,677 ms; Claude Sonnet 5.5 973 ms to 1,277 ms); not a confidence interval. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Routing calls that returned a decision | 100% (82/82)thinking on · via Claude Code · routing overhead per decision | 100% (82/82)effort low · via Claude Code · routing overhead per decision | 82 | 95% CI: 96%–100% vs 96%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 96% to 100%; Claude Sonnet 5.5 96% to 100%), so this sample cannot separate them. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task))Calculation | $441.74thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts | $247.30effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($441.74 vs $247.30) is not tested against run-to-run variation. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task))Calculation | $62.47thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts | $34.97effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($62.47 vs $34.97) is not tested against run-to-run variation. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Added routing delay per task (calculation) (Every model call routed (49.5 per task))Calculation | 620.9 sthinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line | 128.6 seffort low · via Claude Code · calculation per task from recorded decision counts, decisions in line | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap (620.9 s vs 128.6 s, 4.8x) is not tested against run-to-run variation. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Added routing delay per task (calculation) (Only System One decisions (7 per task))Calculation | 87.8 sthinking on · via Claude Code · calculation per task from recorded decision counts, decisions in line | 18.2 seffort low · via Claude Code · calculation per task from recorded decision counts, decisions in line | — | none recorded | Unclear | No interval or range was recorded for either side, so the gap (87.8 s vs 18.2 s, 4.8x) is not tested against run-to-run variation. | Routing overhead: deterministic policy vs LLM routers vs Jev |
| Strict pass rate: single call vs agent loop on eight hard tasks | 46% (11/24)Claude Code · single call | 100% (24/24)Claude Code · single call | 24 | 95% CI: 28%–65% vs 86%–100% | Claude Sonnet 5.5 ahead | The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%). | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Interval merge fix | 100% (3/3)Claude Code · single call | 100% (3/3)Claude Code · single call | 3 | 95% CI: 44%–100% vs 44%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: DST day-length fix | 33% (1/3)Claude Code · single call | 100% (3/3)Claude Code · single call | 3 | 95% CI: 6.2%–79% vs 44%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 6% to 79%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: CSV parser | 67% (2/3)Claude Code · single call | 100% (3/3)Claude Code · single call | 3 | 95% CI: 21%–94% vs 44%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Event-loop order | 0% (0/3)Claude Code · single call | 100% (3/3)Claude Code · single call | 3 | 95% CI: 0%–56% vs 44%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Room schedule | 0% (0/3)Claude Code · single call | 100% (3/3)Claude Code · single call | 3 | 95% CI: 0%–56% vs 44%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SemVer regex | 100% (3/3)Claude Code · single call | 100% (3/3)Claude Code · single call | 3 | 95% CI: 44%–100% vs 44%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 44% to 100%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: Money refactor | 67% (2/3)Claude Code · single call | 100% (3/3)Claude Code · single call | 3 | 95% CI: 21%–94% vs 44%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 21% to 94%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Strict passes per task: single call vs agent loop: SQL report | 0% (0/3)Claude Code · single call | 100% (3/3)Claude Code · single call | 3 | 95% CI: 0%–56% vs 44%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 44% to 100%), so this sample cannot separate them. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Total time per attempt: single call vs agent loop | 39.0 sClaude Code · single call | 7.75 sClaude Code · single call | 24 | range: 15.3 s–75.1 s vs 2.3 s–34.8 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 15.3 s to 75.1 s; Claude Sonnet 5.5 2.26 s to 34.8 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Input tokens (cache reads included)) | 3,941Claude Code · single call | 2,281Claude Code · single call | 24 | range: 3,879–4,221 vs 2,234–2,669 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tokens per attempt: single call vs agent loop (Output tokens) | 5,064Claude Code · single call | 1,050Claude Code · single call | 24 | range: 1,899–9,321 vs 176–3,895 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Tool calls per agent-loop attempt | 3Claude Code · agent loop | 0Claude Code · agent loop | 24 / 16 | range: 2–18 vs 0–3 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| List-price cost per strict pass: single call vs agent loop (calculation)Calculation | $0.067Claude Code · single call | $0.014Claude Code · single call | 24 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.067 vs $0.014, 4.7x) is not tested against run-to-run variation. | Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks |
| Haiku thinking study: typed routing decisions answered exactly right (Exact decisions (every scored question right)) | 87% (71/82)Claude Code · thinking off · typed routing decisions, thinking on vs off | 94% (77/82)Claude Code · effort low · typed routing decisions, thinking on vs off | 82 | 95% CI: 78%–92% vs 87%–97% | Tie | The 95% intervals overlap (Claude Haiku 4.5 78% to 92%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them. | Does thinking pay for Claude Haiku 4.5? Thinking on vs off |
| Haiku thinking study: typed routing decisions answered exactly right (Per-question accuracy) | 91% (177/194)Claude Code · thinking off · typed routing decisions, thinking on vs off | 97% (189/194)Claude Code · effort low · typed routing decisions, thinking on vs off | 194 | 95% CI: 86%–94% vs 94%–99% | Tie | The 95% intervals overlap (Claude Haiku 4.5 86% to 94%; Claude Sonnet 5.5 94% to 99%), so this sample cannot separate them. | Does thinking pay for Claude Haiku 4.5? Thinking on vs off |
| Haiku thinking study: time per routing decision (Wall time (CLI)) | 4.66 sClaude Code · thinking off · typed routing decisions, thinking on vs off | 2.60 sClaude Code · effort low · typed routing decisions, thinking on vs off | 82 | p50–p95: 4.7 s–8.2 s vs 2.6 s–4.3 s | Claude Sonnet 5.5 ahead | Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 4.66 s to 8.18 s; Claude Sonnet 5.5 2.60 s to 4.30 s); not a confidence interval. | Does thinking pay for Claude Haiku 4.5? Thinking on vs off |
| Haiku thinking study: time per routing decision (Model time (API)) | 3.79 sClaude Code · thinking off · typed routing decisions, thinking on vs off | 1.60 sClaude Code · effort low · typed routing decisions, thinking on vs off | 82 | p50–p95: 3.8 s–7.4 s vs 1.6 s–2.6 s | Claude Sonnet 5.5 ahead | Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 3.79 s to 7.43 s; Claude Sonnet 5.5 1.60 s to 2.58 s); not a confidence interval. | Does thinking pay for Claude Haiku 4.5? Thinking on vs off |
| Haiku thinking study: thinking and visible output tokens per routing decision (Thinking tokens) | 0Claude Code · thinking off · typed routing decisions, thinking on vs off | 2Claude Code · effort low · typed routing decisions, thinking on vs off | 82 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does thinking pay for Claude Haiku 4.5? Thinking on vs off |
| Haiku thinking study: thinking and visible output tokens per routing decision (Visible output tokens) | 366Claude Code · thinking off · typed routing decisions, thinking on vs off | 105Claude Code · effort low · typed routing decisions, thinking on vs off | 82 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does thinking pay for Claude Haiku 4.5? Thinking on vs off |
| Haiku thinking study: list-price cost per 1,000 routing decisions (calculation)Calculation | $3.36Claude Code · thinking off · typed routing decisions, thinking on vs off | $7.32Claude Code · effort low · typed routing decisions, thinking on vs off | 82 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($3.36 vs $7.32, 2.2x) is not tested against run-to-run variation. | Does thinking pay for Claude Haiku 4.5? Thinking on vs off |
| Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON)Calculation | 0% (0/24)Claude Code · instructions | 100% (12/12)Claude Code · instructions | 24 / 12 | 95% CI: 0%–14% vs 76%–100% | Claude Sonnet 5.5 ahead | The 95% intervals do not overlap (Claude Haiku 4.5 0% to 14%; Claude Sonnet 5.5 76% to 100%). | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Does a JSON schema raise the pass rate? Instructions vs schema mode (Right answer in any format (strict pass or format miss))Calculation | 71% (17/24)Claude Code · instructions | 100% (12/12)Claude Code · instructions | 24 / 12 | 95% CI: 51%–85% vs 76%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 51% to 85%; Claude Sonnet 5.5 76% to 100%), so this sample cannot separate them. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Strict pass) | 0Claude Code · instructions | 12Claude Code · instructions | 24 / 12 | none recorded | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Format miss) | 17Claude Code · instructions | 0Claude Code · instructions | 24 / 12 | none recorded | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Wrong values) | 7Claude Code · instructions | 0Claude Code · instructions | 24 / 12 | none recorded | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| What each call produced: strict pass, format miss, wrong values or error (Error) | 0Claude Code · instructions | 0Claude Code · instructions | 24 / 12 | none recorded | Tie | Same value. More or fewer is not better by itself for this metric. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Time per call, instructions vs schema mode | 9.52 sClaude Code · instructions | 3.52 sClaude Code · instructions | 24 / 12 | range: 5.7 s–17 s vs 2.7 s–4.1 s | Claude Sonnet 5.5 ahead | The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 5.67 s to 17.0 s; Claude Sonnet 5.5 2.67 s to 4.12 s). A range is not a confidence interval. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Output and reasoning tokens per call, instructions vs schema mode (Median output tokens per call) | 1,128Claude Code · instructions | 368Claude Code · instructions | 24 / 12 | none recorded | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI |
| Unseen routing decisions answered exactly right | 79% (44/56)Claude Code | 88% (49/56)Claude Code · effort low | 56 | 95% CI: 66%–87% vs 76%–94% | Tie | The 95% intervals overlap (Claude Haiku 4.5 66% to 87%; Claude Sonnet 5.5 76% to 94%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Per-question accuracy on unseen decisions | 82% (102/125)Claude Code | 92% (115/125)Claude Code · effort low | 125 | 95% CI: 74%–87% vs 86%–96% | Tie | The 95% intervals overlap (Claude Haiku 4.5 74% to 87%; Claude Sonnet 5.5 86% to 96%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Exact rate on unseen decisions, by decision type: Failure class | 93% (13/14)Claude Code | 100% (14/14)Claude Code · effort low | 14 | 95% CI: 69%–99% vs 78%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 69% to 99%; Claude Sonnet 5.5 78% to 100%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Exact rate on unseen decisions, by decision type: Message intent | 93% (13/14)Claude Code | 100% (14/14)Claude Code · effort low | 14 | 95% CI: 69%–99% vs 78%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 69% to 99%; Claude Sonnet 5.5 78% to 100%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Exact rate on unseen decisions, by decision type: Is it a rule? | 93% (13/14)Claude Code | 93% (13/14)Claude Code · effort low | 14 | 95% CI: 69%–99% vs 69%–99% | Tie | The 95% intervals overlap (Claude Haiku 4.5 69% to 99%; Claude Sonnet 5.5 69% to 99%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Exact rate on unseen decisions, by decision type: Context shape | 36% (5/14)Claude Code | 57% (8/14)Claude Code · effort low | 14 | 95% CI: 16%–61% vs 33%–79% | Tie | The 95% intervals overlap (Claude Haiku 4.5 16% to 61%; Claude Sonnet 5.5 33% to 79%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Tuned case set vs unseen holdout: exact rate per router (Tuned set (routing-jev-vs-llm)) | 89% (73/82)Claude Code | 94% (77/82)Claude Code · effort low | 82 | 95% CI: 80%–94% vs 87%–97% | Tie | The 95% intervals overlap (Claude Haiku 4.5 80% to 94%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Tuned case set vs unseen holdout: exact rate per router (Unseen holdout) | 79% (44/56)Claude Code | 88% (49/56)Claude Code · effort low | 56 | 95% CI: 66%–87% vs 76%–94% | Tie | The 95% intervals overlap (Claude Haiku 4.5 66% to 87%; Claude Sonnet 5.5 76% to 94%), so this sample cannot separate them. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Time per routing decision, by route (Wall time) | 9.44 sClaude Code | 2.36 sClaude Code · effort low | 56 | p50–p95: 9.4 s–25.4 s vs 2.4 s–3.7 s | Claude Sonnet 5.5 ahead | Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 9.44 s to 25.4 s; Claude Sonnet 5.5 2.36 s to 3.66 s); not a confidence interval. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Time per routing decision, by route (Model time (API, CLI-reported)) | 7.52 sClaude Code | 1.49 sClaude Code · effort low | 56 | p50–p95: 7.5 s–23.9 s vs 1.5 s–2.4 s | Claude Sonnet 5.5 ahead | Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 7.52 s to 23.9 s; Claude Sonnet 5.5 1.49 s to 2.38 s); not a confidence interval. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Cost per 1,000 unseen routing decisionsCalculation | $7.13Claude Code | $7.24Claude Code · effort low | 56 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($7.13 vs $7.24) is not tested against run-to-run variation. | Jev vs Claude routers on unseen decisions: a blind holdout |
| Reasoning share of output tokens per call on hard tasks (calculation)Calculation | 91.7%Claude Code | 54.5%Claude Code | 24 | range: 76%–99% vs 0%–96% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens))Calculation | $0.024Claude Code | $0.0067Claude Code | 24 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate))Calculation | $0.0018Claude Code | $0.0037Claude Code | 24 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| List-price cost per call: reasoning, remaining output and input (calculation) (Input (prompt, cache priced))Calculation | $0.0045Claude Code | $0.0040Claude Code | 24 | none recorded | Unclear | More or fewer usd is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Eight hard tasks)Calculation | 91.7%Claude Code | 54.5%Claude Code | 24 | range: 76%–99% vs 0%–96% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Reasoning share on short tasks vs hard tasks (calculation) (Five short tasks)Calculation | 90.2%Claude Code | 0%Claude Code | 15 | range: 73%–98% vs 0%–73% | Unclear | More or fewer percent is not better or worse by itself; this row describes behaviour, not a winner. | How much of an AI bill is thinking? Reasoning tokens by model and effort |
| Time to first text: a 250-line answer, six models | 4.00 sClaude Code | 1.96 sClaude Code | 4 | range: 2.8 s–6.4 s vs 0.9 s–4.1 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 6.38 s; Claude Sonnet 5.5 0.88 s to 4.09 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed after the first text: visible tokens per second (calculation)Calculation | 153Claude Code | 232Claude Code | 4 | range: 153–216 vs 230–233 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Output speed in characters per second after the first text (calculation)Calculation | 547Claude Code | 517Claude Code | 3 / 4 | range: 546–548 vs 513–519 | Unclear | More or fewer count is not better or worse by itself; this row describes behaviour, not a winner. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 1kCalculation | 1.93 sClaude Code | 1.45 sClaude Code | 3 | range: 1.9 s–2 s vs 1.2 s–1.7 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 1.85 s to 2.04 s; Claude Sonnet 5.5 1.23 s to 1.72 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 16kCalculation | 2.27 sClaude Code | 1.78 sClaude Code | 3 | range: 2.2 s–2.5 s vs 1.6 s–2.1 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.47 s; Claude Sonnet 5.5 1.64 s to 2.11 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Time to first text as the prompt grows: 64kCalculation | 2.78 sClaude Code | 3.07 sClaude Code | 3 | range: 2.5 s–2.9 s vs 1.4 s–3.6 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.45 s to 2.89 s; Claude Sonnet 5.5 1.38 s to 3.61 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (1k prompt) | 2.34 sClaude Code | 1.78 sClaude Code | 3 | range: 2.2 s–2.5 s vs 1.6 s–2.1 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.22 s to 2.46 s; Claude Sonnet 5.5 1.57 s to 2.12 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (16k prompt) | 2.79 sClaude Code | 2.10 sClaude Code | 3 | range: 2.6 s–2.8 s vs 2 s–2.5 s | Unclear | Only 3 runs per side; the run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 2.58 s to 2.84 s; Claude Sonnet 5.5 1.98 s to 2.48 s), but 3 runs cannot show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Total time per call by prompt size (64k prompt) | 3.13 sClaude Code | 3.44 sClaude Code | 3 | range: 2.8 s–3.3 s vs 1.7 s–4.4 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 2.84 s to 3.28 s; Claude Sonnet 5.5 1.74 s to 4.38 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| Exact lookup answers at the 1k, 16k and 64k prompt-size targets | 100% (9/9)Claude Code | 100% (9/9)Claude Code | 9 | 95% CI: 70%–100% vs 70%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 70% to 100%; Claude Sonnet 5.5 70% to 100%), so this sample cannot separate them. | Where the seconds go: first text, output speed and prompt size for 6 LLMs |
| List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Interval merge fixCalculation | $0.018Claude Code | $0.0056Claude Code | 3 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.018 vs $0.0056, 3.2x) is not tested against run-to-run variation. | Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts |
| List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): DST day lengthCalculation | $0.029Claude Code | $0.025Claude Code | 3 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.029 vs $0.025) is not tested against run-to-run variation. | Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts |
| List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): CSV parserCalculation | $0.029Claude Code | $0.015Claude Code | 3 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.029 vs $0.015) is not tested against run-to-run variation. | Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts |
| List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Event-loop orderCalculation | $0.037Claude Code | $0.016Claude Code | 3 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.037 vs $0.016, 2.3x) is not tested against run-to-run variation. | Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts |
| List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Room scheduleCalculation | $0.036Claude Code | $0.012Claude Code | 3 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.036 vs $0.012, 2.9x) is not tested against run-to-run variation. | Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts |
| List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SemVer regexCalculation | $0.042Claude Code | $0.0051Claude Code | 3 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.042 vs $0.0051, 8.2x) is not tested against run-to-run variation. | Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts |
| List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Money refactorCalculation | $0.020Claude Code | $0.0096Claude Code | 3 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.020 vs $0.0096, 2.1x) is not tested against run-to-run variation. | Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts |
| List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SQLite report queryCalculation | $0.036Claude Code | $0.017Claude Code | 3 | none recorded | Unclear | No interval or range was recorded for either side, so the gap ($0.036 vs $0.017, 2.1x) is not tested against run-to-run variation. | Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts |
| Pass rate on 4 harder tasks (Strict pass) | 0% (0/12)Claude Code | 38% (6/16)Claude Code | 12 / 16 | 95% CI: 0%–24% vs 18%–61% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 24%; Claude Sonnet 5.5 18% to 61%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Pass rate on 4 harder tasks (Lenient (format misses counted)) | 0% (0/12)Claude Code | 38% (6/16)Claude Code | 12 / 16 | 95% CI: 0%–24% vs 18%–61% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 24%; Claude Sonnet 5.5 18% to 61%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Calls that tried a tool although tools were off | 8% (1/12)Claude Code | 31% (5/16)Claude Code | 12 / 16 | 95% CI: 1.5%–35% vs 14%–56% | Unclear | More or fewer rate is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 10x10 nonogram | 0% (0/3)Claude Code | 100% (4/4)Claude Code | 3 / 4 | 95% CI: 0%–56% vs 51%–100% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 51% to 100%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Sudoku, 22 givens | 0% (0/3)Claude Code | 0% (0/4)Claude Code | 3 / 4 | 95% CI: 0%–56% vs 0%–49% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 0% to 49%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: 6x6 Skyscrapers | 0% (0/3)Claude Code | 0% (0/4)Claude Code | 3 / 4 | 95% CI: 0%–56% vs 0%–49% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 0% to 49%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Strict pass rate by task: Seeded shuffle output | 0% (0/3)Claude Code | 50% (2/4)Claude Code | 3 / 4 | 95% CI: 0%–56% vs 15%–85% | Tie | The 95% intervals overlap (Claude Haiku 4.5 0% to 56%; Claude Sonnet 5.5 15% to 85%), so this sample cannot separate them. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Total time per call on harder tasks | 109.0 sClaude Code | 70.4 sClaude Code | 10 / 12 | range: 25.7 s–224 s vs 4.3 s–210 s | Unclear | The run ranges (fastest to slowest) overlap (Claude Haiku 4.5 25.7 s to 223.9 s; Claude Sonnet 5.5 4.32 s to 210.1 s); the medians alone do not show a reliable difference. A range is not a confidence interval. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
| Output tokens per call on harder tasks (Output tokens) | 12,508Claude Code | 9,287Claude Code | 10 / 12 | range: 2,965–26,532 vs 407–27,921 | Unclear | More or fewer tokens is not better or worse by itself; this row describes behaviour, not a winner. | GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks |
Marks: 95% intervals (Wilson for rates); fastest–slowest run ranges (not intervals); median to p95 bands (not intervals)n is shown per side on every rowHollow marks: list-price calculations, not runs
153 rows from 14 studies. Claude Sonnet 5.5 ahead on 21; 58 ties, 74 unclear. A side is ahead only where the intervals or ranges do not overlap.
When to pick which
Only from the rows above. A tie is not a reason to pick either side.
When to pick Claude Haiku 4.5
- Haiku thinking study: list-price cost per 1,000 routing decisions (calculation): $3.36 vs $7.32. A list-price calculation, not a measured difference. Calculation
- List-price cost per call: reasoning, remaining output and input (calculation) (Remaining output (visible-answer estimate)): $0.0018 vs $0.0037. A list-price calculation, not a measured difference. Calculation
When to pick Claude Sonnet 5.5
- List-price cost per call (calculation): $0.0036 vs $0.0057. A list-price calculation, not a measured difference. Calculation
- Pass rate on eight hard tasks (Strict pass): 100% (24/24) vs 46% (11/24). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).
- Pass rate on eight hard tasks (Lenient (format misses counted)): 100% (24/24) vs 67% (16/24). The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Sonnet 5.5 86% to 100%).
- List-price cost per strict pass on hard tasks (calculation): $0.014 vs $0.067. A list-price calculation, not a measured difference. Calculation
- Same prompt, 10 times: strict pass rate (Exact number): 100% (10/10) vs 0% (0/10). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 72% to 100%).
- Same prompt, 10 times: strict pass rate (JSON object): 100% (10/10) vs 10% (1/10). The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 72% to 100%).
- Same prompt, 10 times: time per call (Code fix): 2.67 s vs 5.95 s. The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; Claude Sonnet 5.5 2.32 s to 4.34 s). A range is not a confidence interval.
- Full pass rate by kind of memory: Handbook, 210 lines: 100% (15/15) vs 30% (3/10). The 95% intervals do not overlap (Claude Haiku 4.5 11% to 60%; Claude Sonnet 5.5 80% to 100%).
- Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md: 67% (10/15) vs 10% (1/10). The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 42% to 85%).
- Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines: 100% (15/15) vs 30% (3/10). The 95% intervals do not overlap (Claude Haiku 4.5 11% to 60%; Claude Sonnet 5.5 80% to 100%).
- A stale README command: who still ran it?: Raw notes, 60 lines: 0% (0/15) vs 100% (10/10). The 95% intervals do not overlap (Claude Haiku 4.5 72% to 100%; Claude Sonnet 5.5 0% to 20%).
- List-price cost per fully correct result (calculation): No memory: $0.14 vs $0.38. A list-price calculation, not a measured difference. Calculation
- List-price cost per fully correct result (calculation): /init CLAUDE.md: $0.14 vs $0.43. A list-price calculation, not a measured difference. Calculation
- List-price cost per fully correct result (calculation): Handbook, 210 lines: $0.10 vs $0.26. A list-price calculation, not a measured difference. Calculation
- Cost per 1,000 routing decisions: $5.00 vs $8.92. A list-price calculation, not a measured difference. Calculation
- Time per routing decision (Wall time (CLI)): 2,598 ms vs 12,674 ms. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 12,674 ms to 34,413 ms; Claude Sonnet 5.5 2,598 ms to 4,298 ms); not a confidence interval.
- Time per routing decision (Model time (API)): 1,599 ms vs 10,734 ms. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 10,734 ms to 32,072 ms; Claude Sonnet 5.5 1,599 ms to 2,574 ms); not a confidence interval.
- Time to make one routing decision: 2,597 ms vs 12,543 ms. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 12,543 ms to 34,481 ms; Claude Sonnet 5.5 2,597 ms to 4,298 ms); not a confidence interval.
- Where an LLM router’s time goes: model vs CLI (Model API time): 1,596 ms vs 10,508 ms. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 10,508 ms to 32,132 ms; Claude Sonnet 5.5 1,596 ms to 2,583 ms); not a confidence interval.
- Where an LLM router’s time goes: model vs CLI (CLI and harness time): 973 ms vs 1,698 ms. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 1,698 ms to 2,677 ms; Claude Sonnet 5.5 973 ms to 1,277 ms); not a confidence interval.
- Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)): $247.30 vs $441.74. A list-price calculation, not a measured difference. Calculation
- Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)): $34.97 vs $62.47. A list-price calculation, not a measured difference. Calculation
- Added routing delay per task (calculation) (Every model call routed (49.5 per task)): 128.6 s vs 620.9 s. A list-price calculation, not a measured difference. Calculation
- Added routing delay per task (calculation) (Only System One decisions (7 per task)): 18.2 s vs 87.8 s. A list-price calculation, not a measured difference. Calculation
- Strict pass rate: single call vs agent loop on eight hard tasks: 100% (24/24) vs 46% (11/24). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).
- List-price cost per strict pass: single call vs agent loop (calculation): $0.014 vs $0.067. A list-price calculation, not a measured difference. Calculation
- Haiku thinking study: time per routing decision (Wall time (CLI)): 2.60 s vs 4.66 s. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 4.66 s to 8.18 s; Claude Sonnet 5.5 2.60 s to 4.30 s); not a confidence interval.
- Haiku thinking study: time per routing decision (Model time (API)): 1.60 s vs 3.79 s. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 3.79 s to 7.43 s; Claude Sonnet 5.5 1.60 s to 2.58 s); not a confidence interval.
- Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON): 100% (12/12) vs 0% (0/24). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 14%; Claude Sonnet 5.5 76% to 100%).
- Time per call, instructions vs schema mode: 3.52 s vs 9.52 s. The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 5.67 s to 17.0 s; Claude Sonnet 5.5 2.67 s to 4.12 s). A range is not a confidence interval.
- Time per routing decision, by route (Wall time): 2.36 s vs 9.44 s. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 9.44 s to 25.4 s; Claude Sonnet 5.5 2.36 s to 3.66 s); not a confidence interval.
- Time per routing decision, by route (Model time (API, CLI-reported)): 1.49 s vs 7.52 s. Claude Haiku 4.5’s median is above Claude Sonnet 5.5’s 95th percentile (p50–p95 bands: Claude Haiku 4.5 7.52 s to 23.9 s; Claude Sonnet 5.5 1.49 s to 2.38 s); not a confidence interval.
- List-price cost per call: reasoning, remaining output and input (calculation) (Reasoning (output tokens)): $0.0067 vs $0.024. A list-price calculation, not a measured difference. Calculation
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Interval merge fix: $0.0056 vs $0.018. A list-price calculation, not a measured difference. Calculation
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): CSV parser: $0.015 vs $0.029. A list-price calculation, not a measured difference. Calculation
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Event-loop order: $0.016 vs $0.037. A list-price calculation, not a measured difference. Calculation
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Room schedule: $0.012 vs $0.036. A list-price calculation, not a measured difference. Calculation
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SemVer regex: $0.0051 vs $0.042. A list-price calculation, not a measured difference. Calculation
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): Money refactor: $0.0096 vs $0.020. A list-price calculation, not a measured difference. Calculation
- List-price cost of one call, Haiku 4.5 vs Sonnet 5.5, on each hard task (calculation): SQLite report query: $0.017 vs $0.036. A list-price calculation, not a measured difference. Calculation
Side by side
The study charts, showing only these two. Open a study for every configuration.
| Item | Pass rate | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 80% | 55%–93% | 15 |
| Claude Haiku 4.5 · Claude Code | 100% | 80%–100% | 15 |
2 rows. Highest Claude Haiku 4.5 · Claude Code 100% (95% interval 80%–100%, n 15). Lowest Claude Sonnet 5.5 · Claude Code 80% (95% interval 55%–93%, n 15). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 15 per row
Every call counts; failures and timeouts are non-passes
Whiskers are 95% Wilson intervals. The tasks are short and most configurations pass nearly all of them, so pass rate does not separate the models here.
Source: Provider head-to-head: Claude Code models vs Codex efforts
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 2.3 s | 2.2 s–7.7 s | 15 |
| Claude Haiku 4.5 · Claude Code | 4.4 s | 3.2 s–23.6 s | 15 |
2 rows. Slowest Claude Haiku 4.5 · Claude Code 4.4 s (range 3.2 s–23.6 s, n 15). Fastest Claude Sonnet 5.5 · Claude Code 2.3 s (range 2.2 s–7.7 s, n 15). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 15 per row
Median per configuration; whiskers = fastest and slowest call
One host, one network, one day. Whiskers are a range, not a confidence interval.
Source: Provider head-to-head: Claude Code models vs Codex efforts
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first useful output | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1.6 s | 1 s–6.4 s | 15 |
| Claude Haiku 4.5 · Claude Code | 3.6 s | 2.8 s–22.3 s | 15 |
2 rows. Slowest Claude Haiku 4.5 · Claude Code 3.6 s (range 2.8 s–22.3 s, n 15). Fastest Claude Sonnet 5.5 · Claude Code 1.6 s (range 1 s–6.4 s, n 15). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 15 per row
Median per configuration; whiskers = fastest and slowest call
One host, one network, one day. Whiskers are a range, not a confidence interval.
Source: Provider head-to-head: Claude Code models vs Codex efforts
- Cache read
- Other input
Totals are the sum of the parts shown. Shares are calculated from the same values.
| Item | Cache read | Other input | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1,401 | 685 | 15 |
| Claude Haiku 4.5 · Claude Code | 0 | 3,790 | 15 |
2 rows, 2 series: Cache read, Other input. Cache read: highest Claude Sonnet 5.5 · Claude Code 1,401 (n 15). Lowest Claude Haiku 4.5 · Claude Code 0 (n 15). Other input: highest Claude Haiku 4.5 · Claude Code 3,790 (n 15). Lowest Claude Sonnet 5.5 · Claude Code 685 (n 15).
Notesn = 15 per row
Mean per call, split into prompt-cache reads and other input
The prompts are a few hundred tokens; most input is the CLI’s own system prompt and tool context. Input counts include cache reads, as the vendors report them.
Source: Provider head-to-head: Claude Code models vs Codex efforts
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 107 | 0 | 15 |
| Claude Haiku 4.5 · Claude Code | 367 | 297 | 15 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 367 (n 15). Lowest Claude Sonnet 5.5 · Claude Code 107 (n 15). Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 297 (n 15). Lowest Claude Sonnet 5.5 · Claude Code 0 (n 15).
Notesn = 15 per row
Median per configuration; reasoning tokens where the CLI reports them
Vendors count reasoning differently; compare within a vendor first. A median of 0 can mean the CLI did not report reasoning for most calls.
Source: Provider head-to-head: Claude Code models vs Codex efforts
| Item | Cost per call | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.0036 | $0.0034–$0.01 | 15 |
| Claude Haiku 4.5 · Claude Code | $0.0057 | $0.0051–$0.018 | 15 |
List-price calculation, not a run. 2 rows. Highest Claude Haiku 4.5 · Claude Code $0.0057 (range $0.0051–$0.018, n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.0036 (range $0.0034–$0.01, n 15). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 15 per row
Reported tokens × list price; the calls ran on subscriptions
Calculation, not a bill: the calls ran on flat subscriptions. Whiskers = cheapest and most expensive call.
Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
| Item | Cost per pass | n |
|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.0062 | 15 |
| Claude Haiku 4.5 · Claude Code | $0.0084 | 15 |
List-price calculation, not a run. 2 rows. Highest Claude Haiku 4.5 · Claude Code $0.0084 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.0062 (n 15).
Notesn = 15 per row
All calls in a configuration, failures included, divided by its passes
Calculation, not a bill: reported tokens × list price; the calls ran on flat subscriptions. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the speed/cost/quality frontier.
Sources: Provider head-to-head: Claude Code models vs Codex efforts, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Strict pass
- Lenient (format misses counted)
| Item | Strict pass | Lenient (format misses counted) | 95% interval | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 100% | 100% | Strict pass: 86%–100%; Lenient (format misses counted): 86%–100% | 24 |
| Claude Haiku 4.5 · Claude Code | 46% | 67% | Strict pass: 28%–65%; Lenient (format misses counted): 47%–82% | 24 |
2 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 67% (95% interval 47%–82%, n 24). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 24 per row
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call on hard tasks (separate batches) | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 7.8 s | 2.3 s–34.8 s | 24 |
| Claude Haiku 4.5 · Claude Code | 39 s | 15.3 s–75.1 s | 24 |
2 rows. Slowest Claude Haiku 4.5 · Claude Code 39 s (range 15.3 s–75.1 s, n 24). Fastest Claude Sonnet 5.5 · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 24 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first useful output on hard tasks | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 6 s | 0.9 s–30.6 s | 24 |
| Claude Haiku 4.5 · Claude Code | 35.5 s | 12.9 s–70.3 s | 24 |
2 rows. Slowest Claude Haiku 4.5 · Claude Code 35.5 s (range 12.9 s–70.3 s, n 24). Fastest Claude Sonnet 5.5 · Claude Code 6 s (range 0.9 s–30.6 s, n 24). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 24 per row
Median per configuration; whiskers = fastest and slowest call
One host and network; the counted Claude and Codex batches ran hours apart. Host load was not controlled. Whiskers are a range, not a confidence interval. Highlighted: configurations that passed every call.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 1,050 | 585 | 24 |
| Claude Haiku 4.5 · Claude Code | 5,064 | 4,556 | 24 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 5,064 (n 24). Lowest Claude Sonnet 5.5 · Claude Code 1,050 (n 24). Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 4,556 (n 24). Lowest Claude Sonnet 5.5 · Claude Code 585 (n 24).
Notesn = 24 per row
Median per configuration; reasoning tokens as the CLI reports them
Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. More tokens is not better or worse by itself.
Source: Provider head-to-head, hard set: eight hard tasks with strict validators
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Sonnet 5.5 · Claude Code | $0.014 | 24 |
| Claude Haiku 4.5 · Claude Code | $0.067 | 24 |
List-price calculation, not a run. 2 rows. Highest Claude Haiku 4.5 · Claude Code $0.067 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.014 (n 24).
Notesn = 24 per row
All calls in a configuration, failures and format misses included, divided by its strict passes
Calculation, not a bill: reported tokens × list price; the calls ran on a flat subscription. A failed call still costs, so a lower pass rate raises the cost per pass. Highlighted bars are on the quality-vs-cost frontier.
Sources: Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Exact number
- JSON object
- Code fix
| Item | Exact number | JSON object | Code fix | 95% interval | n |
|---|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 0% | 10% | 100% | Exact number: 0%–28%; JSON object: 1.8%–40%; Code fix: 72%–100% | 10 |
| Claude Sonnet 5.5 · Claude Code | 100% | 100% | 100% | Exact number: 72%–100%; JSON object: 72%–100%; Code fix: 72%–100% | 10 |
2 rows, 3 series: Exact number, JSON object, Code fix. Exact number: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–28%, n 10). Not all intervals overlap. JSON object: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 72%–100%, n 10). Lowest Claude Haiku 4.5 · Claude Code 10% (95% interval 1.8%–40%, n 10). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 10 per row
One series per prompt; whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals over 10 repetitions. Strict: the whole reply passes the validator. A correct answer in the wrong format (for example in a code fence) is a format miss and never a pass; the table lists format misses apart from wrong answers.
Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)
Exact number
JSON object
Code fix
One panel per series, all on the same axis.
| Item | Exact number | JSON object | Code fix | n |
|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 1 | 1 | 6 | 10 |
| Claude Sonnet 5.5 · Claude Code | 1 | 1 | 3 | 10 |
2 rows, 3 series: Exact number, JSON object, Code fix. Exact number: all at 1. JSON object: all at 1.
Notesn = 10 per row
Distinct normalized answers over 10 repetitions (1 = the same answer every time)
Normalized: the last number for the exact-number prompt; key-sorted JSON for the JSON prompt; for the code fix, the code without fences, comments, whitespace, quote style and line-end semicolons. One distinct answer is not the same as a correct answer: an answer can be the same and wrong every time. Different code can be equally correct.
Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)
- Exact number
- JSON object
- Code fix
| Item | Exact number | JSON object | Code fix | Range (lowest–highest run) | n |
|---|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 5.1 s | 7 s | 6 s | Exact number: 4.4 s–6.2 s; JSON object: 5.3 s–12.3 s; Code fix: 4.9 s–7.3 s | 10 |
| Claude Sonnet 5.5 · Claude Code | 6.9 s | 2.9 s | 2.7 s | Exact number: 5.8 s–7.8 s; JSON object: 2.7 s–5.3 s; Code fix: 2.3 s–4.3 s | 10 |
2 rows, 3 series: Exact number, JSON object, Code fix. Exact number: slowest Claude Sonnet 5.5 · Claude Code 6.9 s (range 5.8 s–7.8 s, n 10). Fastest Claude Haiku 4.5 · Claude Code 5.1 s (range 4.4 s–6.2 s, n 10). All run ranges overlap. JSON object: slowest Claude Haiku 4.5 · Claude Code 7 s (range 5.3 s–12.3 s, n 10). Fastest Claude Sonnet 5.5 · Claude Code 2.9 s (range 2.7 s–5.3 s, n 10). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 10 per row
Median; whiskers = fastest and slowest of 10 calls
Whiskers are a range (fastest and slowest call), not a confidence interval. The interquartile range and the coefficient of variation are in the table. Codex CLI timings include its start-up and its larger system prompt.
Source: Caching sessions and repeated prompts (Claude Code and Codex CLI)
- Claude Sonnet 5.5
- Claude Haiku 4.5
| Item | Claude Sonnet 5.5 | Claude Haiku 4.5 | 95% interval | n |
|---|---|---|---|---|
| No memory | 60% | 20% | Claude Sonnet 5.5: 36%–80%; Claude Haiku 4.5: 5.7%–51% | 15 |
| Curated, 11 lines | 100% | 70% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 40%–89% | 15 |
| Raw notes, 60 lines | 93% | 60% | Claude Sonnet 5.5: 70%–99%; Claude Haiku 4.5: 31%–83% | 15 |
| Handbook, 210 lines | 100% | 30% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 11%–60% | 15 |
| Stop hook only | 80% | 80% | Claude Sonnet 5.5: 55%–93%; Claude Haiku 4.5: 49%–94% | 15 |
| Curated + hook | 100% | 90% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 60%–98% | 15 |
6 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest No memory 60% (95% interval 36%–80%, n 15). All intervals overlap. Claude Haiku 4.5: highest Curated + hook 90% (95% interval 60%–98%, n 10). Lowest No memory 20% (95% interval 5.7%–51%, n 10). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 10–15 per row3 of 6 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.
Hidden tests pass and every convention check passes · 95% Wilson intervals
Claude Code 2.1.286, 5 tasks in one small repository. Sonnet: 3 repetitions per cell (n = 15 per condition); Haiku: 2 (n = 10). A condition is better only when its interval does not overlap the other's.
Source: Agent memory study: 8 kinds of project memory on Claude Code
- Claude Sonnet 5.5
- Claude Haiku 4.5
| Item | Claude Sonnet 5.5 | Claude Haiku 4.5 | 95% interval | n |
|---|---|---|---|---|
| No memory | 40% | 0% | Claude Sonnet 5.5: 20%–64%; Claude Haiku 4.5: 0%–28% | 15 |
| /init CLAUDE.md | 67% | 10% | Claude Sonnet 5.5: 42%–85%; Claude Haiku 4.5: 1.8%–40% | 15 |
| Curated, 11 lines | 100% | 80% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 49%–94% | 15 |
| Raw notes, 60 lines | 100% | 60% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 31%–83% | 15 |
| Handbook, 210 lines | 100% | 30% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 11%–60% | 15 |
| Curated + hook | 100% | 100% | Claude Sonnet 5.5: 80%–100%; Claude Haiku 4.5: 72%–100% | 15 |
6 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest Curated, 11 lines 100% (95% interval 80%–100%, n 15). Lowest No memory 40% (95% interval 20%–64%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest Curated + hook 100% (95% interval 72%–100%, n 10). Lowest No memory 0% (95% interval 0%–28%, n 10). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 10–15 per row4 of 6 (Claude Sonnet 5.5) at 100%: this task set cannot separate them.
Changelog rule and late-fee rate, pooled · 95% Wilson intervals
Both models had the same memory files. The smaller model followed the team rules less often when the facts sat in long or messy files.
Source: Agent memory study: 8 kinds of project memory on Claude Code
- Claude Sonnet 5.5
- Claude Haiku 4.5
| Item | Claude Sonnet 5.5 | Claude Haiku 4.5 | 95% interval | n |
|---|---|---|---|---|
| No memory | 80% | 100% | Claude Sonnet 5.5: 55%–93%; Claude Haiku 4.5: 72%–100% | 15 |
| /init CLAUDE.md | 87% | 100% | Claude Sonnet 5.5: 62%–96%; Claude Haiku 4.5: 72%–100% | 15 |
| Curated, 11 lines | 0% | 0% | Claude Sonnet 5.5: 0%–20%; Claude Haiku 4.5: 0%–28% | 15 |
| Dreamed notes | 0% | 10% | Claude Sonnet 5.5: 0%–20%; Claude Haiku 4.5: 1.8%–40% | 15 |
| Stop hook only | 60% | 100% | Claude Sonnet 5.5: 36%–80%; Claude Haiku 4.5: 72%–100% | 15 |
5 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest /init CLAUDE.md 87% (95% interval 62%–96%, n 15). Lowest Dreamed notes 0% (95% interval 0%–20%, n 15). Not all intervals overlap. Claude Haiku 4.5: highest No memory 100% (95% interval 72%–100%, n 10). Lowest Curated, 11 lines 0% (95% interval 0%–28%, n 10). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln 10–15 per row
Sessions that ran the README test command, which fails on Node 25 · 95% Wilson intervals
The README and npm test give a command that fails on Node 25. The curated, dreamed and handbook files name the right one. The raw notes hold both: an old "run npm test" note and a later correction. The /init file repeats the README.
Source: Agent memory study: 8 kinds of project memory on Claude Code
- Claude Sonnet 5.5
- Claude Haiku 4.5 (square)
Gap labels, Claude Haiku 4.5 vs Claude Sonnet 5.5: Claude Haiku 4.5 is x% higher (+) or lower (−) than Claude Sonnet 5.5, calculated from the two values shown (the change counted from Claude Sonnet 5.5’s value).
| Item | Claude Sonnet 5.5 | Claude Haiku 4.5 | n |
|---|---|---|---|
| No memory | $0.14 | $0.38 | 9 |
| /init CLAUDE.md | $0.14 | $0.43 | 9 |
| Curated, 11 lines | $0.082 | $0.11 | 15 |
| Raw notes, 60 lines | $0.1 | $0.13 | 14 |
| Dreamed notes | $0.09 | $0.12 | 15 |
| Handbook, 210 lines | $0.1 | $0.26 | 15 |
| Stop hook only | $0.13 | $0.14 | 12 |
| Curated + hook | $0.084 | $0.096 | 15 |
List-price calculation, not a run. 8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: highest No memory $0.14 (n 9). Lowest Curated, 11 lines $0.082 (n 15). Claude Haiku 4.5: highest /init CLAUDE.md $0.43 (n 2). Lowest Curated + hook $0.096 (n 9).
Notesn 2–15 per row
Sum of the CLI's cost estimates for a condition, divided by its full passes
Sessions ran on a subscription; these are the CLI's list-price estimates, not bills. A failed session still costs money, so cost per correct result falls when fewer sessions fail.
Source: Agent memory study: 8 kinds of project memory on Claude Code
Time per session
Median wall time in seconds
- Claude Sonnet 5.5
- Claude Haiku 4.5 (square)
Gap labels, Claude Haiku 4.5 vs Claude Sonnet 5.5: Claude Haiku 4.5 is x% higher (+) or lower (−) than Claude Sonnet 5.5, calculated from the two values shown (the change counted from Claude Sonnet 5.5’s value).
| Item | Claude Sonnet 5.5 | Claude Haiku 4.5 | n |
|---|---|---|---|
| No memory | 18 s | 54.2 s | 15 |
| /init CLAUDE.md | 19 s | 52.9 s | 15 |
| Curated, 11 lines | 21.9 s | 51.7 s | 15 |
| Raw notes, 60 lines | 26.8 s | 51 s | 15 |
| Dreamed notes | 27.2 s | 51.8 s | 15 |
| Handbook, 210 lines | 23.6 s | 49.9 s | 15 |
| Stop hook only | 27.3 s | 68.5 s | 15 |
| Curated + hook | 22 s | 52.8 s | 15 |
8 rows, 2 series: Claude Sonnet 5.5, Claude Haiku 4.5. Claude Sonnet 5.5: slowest Stop hook only 27.3 s (n 15). Fastest No memory 18 s (n 15). Claude Haiku 4.5: slowest Stop hook only 68.5 s (n 10). Fastest Handbook, 210 lines 49.9 s (n 10).
Notesn 10–15 per row
Up to four sessions ran at a time on one machine. Sessions without memory were often shorter because they stopped to ask or skipped the changelog.
Source: Agent memory study: 8 kinds of project memory on Claude Code
| Item | Exact rate | 95% interval | n |
|---|---|---|---|
| Claude Haiku 4.5 | 89% | 80%–94% | 82 |
| Claude Sonnet 5.5 | 94% | 87%–97% | 82 |
2 rows. Highest Claude Sonnet 5.5 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 89% (95% interval 80%–94%, n 82). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 82 per row
Share of asked cases where every scored question was acceptable
Whiskers are 95% Wilson intervals on the 82 decisions. Jev is the live run: 3 repeats of the same 82 decisions. Repeats of one decision are not independent, so its interval is taken at n = 82, not 246. 221 of 246 Jev calls were exact (74, 73 and 74 of 82 per repeat). Each Claude router made one pass through the Claude Code CLI. The recorded production run of Jev scored 74 of 82. The case sets were revised against Jev answers, so Jev has a home advantage.
Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
| Item | Key accuracy | 95% interval | n |
|---|---|---|---|
| Claude Haiku 4.5 | 94% | 90%–97% | 194 |
| Claude Sonnet 5.5 | 97% | 94%–99% | 194 |
2 rows. Highest Claude Sonnet 5.5 97% (95% interval 94%–99%, n 194). Lowest Claude Haiku 4.5 94% (95% interval 90%–97%, n 194). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 194 per row
Each open question the router was asked; an unanswered question counts as wrong
Whiskers are 95% Wilson intervals on the 194 scored questions. Jev: 552 of 582 answers over 3 repeats, with the interval taken at n = 194 because the repeats are not independent. Each Claude router made one pass.
Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
Jev 1.13 (TypeSafe)
Claude Haiku 4.5
Claude Sonnet 5.5
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | Jev 1.13 (TypeSafe) | Claude Haiku 4.5 | Claude Sonnet 5.5 | 95% interval | n |
|---|---|---|---|---|---|
| Failure class | 100% | 94% | 100% | Jev 1.13 (TypeSafe): 82%–100%; Claude Haiku 4.5: 74%–99%; Claude Sonnet 5.5: 82%–100% | 18 |
| Message intent | 100% | 100% | 100% | Jev 1.13 (TypeSafe): 84%–100%; Claude Haiku 4.5: 84%–100%; Claude Sonnet 5.5: 84%–100% | 20 |
| Context shape | 74% | 75% | 84% | Jev 1.13 (TypeSafe): 58%–87%; Claude Haiku 4.5: 58%–87%; Claude Sonnet 5.5: 68%–93% | 32 |
3 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5, Claude Sonnet 5.5. Jev 1.13 (TypeSafe): highest Failure class 100% (95% interval 82%–100%, n 18). Lowest Context shape 74% (95% interval 58%–87%, n 32). All intervals overlap. Claude Haiku 4.5: highest Message intent 100% (95% interval 84%–100%, n 20). Lowest Context shape 75% (95% interval 58%–87%, n 32). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 18–32 per row
A missing bar means that router was not run on that decision type. Whiskers are 95% Wilson intervals. Jev: the pooled rate of the live repeats, with the interval taken at the number of decisions of that type.
Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
| Item | Cost | n |
|---|---|---|
| Claude Haiku 4.5 | $8.92 | 82 |
| Claude Sonnet 5.5 | $5.00 | 82 |
List-price calculation, not a run. 2 rows. Highest Claude Haiku 4.5 $8.92 (n 82). Lowest Claude Sonnet 5.5 $5.00 (n 82).
Notesn = 82 per row
List price × reported tokens per decision
List-price calculation from tokens. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free); the provider-reported cost of its recorded run is the same figure. The LLM routers ran through a subscription CLI, so CLI tool-schema and thinking tokens are included because the CLI reports them.
Sources: Routing runs: Jev router vs LLM routing, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models), Jev live run: 246 timed calls on the 82 routing decisions
- Wall time (CLI)
- Model time (API)
| Item | Wall time (CLI) | Model time (API) | Median to p95 | n |
|---|---|---|---|---|
| Claude Haiku 4.5 | 12.67 s | 10.73 s | Wall time (CLI): 12.67 s–34.41 s; Model time (API): 10.73 s–32.07 s | 82 |
| Claude Sonnet 5.5 | 2.6 s | 1.6 s | Wall time (CLI): 2.6 s–4.3 s; Model time (API): 1.6 s–2.57 s | 82 |
2 rows, 2 series: Wall time (CLI), Model time (API). Wall time (CLI): slowest Claude Haiku 4.5 12.67 s (median to p95 12.67 s–34.41 s, n 82). Fastest Claude Sonnet 5.5 2.6 s (median to p95 2.6 s–4.3 s, n 82). Not all run ranges overlap. Model time (API): slowest Claude Haiku 4.5 10.73 s (median to p95 10.73 s–32.07 s, n 82). Fastest Claude Sonnet 5.5 1.6 s (median to p95 1.6 s–2.57 s, n 82). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n = 82 per row
Median wall time, whisker to the 95th percentile
Whiskers run from p50 to p95. The Claude routers ran through the Claude Code CLI, so their wall time includes CLI start-up and the tool schema; one pass of 82 decisions each. Jev was called directly over HTTPS from one Mac on a home network: 246 calls in a 35-second window, client wall time with the network inside it. Its API reports no server time, so Jev has no model-time point. These are different routes: the chart shows what a caller waits per decision, not model compute time.
Sources: Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.
| Item | Decision time | Median to p95 | n |
|---|---|---|---|
| Claude Sonnet 5.5 (effort low, via Claude Code) | 2.6 s | 2.6 s–4.3 s | 82 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 12.54 s | 12.54 s–34.48 s | 82 |
2 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 2.6 s (median to p95 2.6 s–4.3 s, n 82). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n = 82 per row
Median; whiskers = median to 95th percentile
The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
- Model API time
- CLI and harness time
| Item | Model API time | CLI and harness time | Median to p95 | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 (effort low, via Claude Code) | 1.6 s | 973 ms | Model API time: 1.6 s–2.58 s; CLI and harness time: 973 ms–1.28 s | 82 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 10.51 s | 1.7 s | Model API time: 10.51 s–32.13 s; CLI and harness time: 1.7 s–2.68 s | 82 |
2 rows, 2 series: Model API time, CLI and harness time. Model API time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 10.51 s (median to p95 10.51 s–32.13 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 1.6 s (median to p95 1.6 s–2.58 s, n 82). Not all run ranges overlap. CLI and harness time: slowest Claude Haiku 4.5 (thinking on, via Claude Code) 1.7 s (median to p95 1.7 s–2.68 s, n 82). Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 973 ms (median to p95 973 ms–1.28 s, n 82). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n = 82 per row
Median per call; whiskers = median to 95th percentile
Model API time is the API duration the CLI reports; CLI and harness time is wall time minus that, per call. Medians of the parts do not add up to the median of the whole. Jev is not split: its API reports no server time. The whisker is the median to the 95th percentile, not a confidence interval.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing
| Item | Completed | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 (effort low, via Claude Code) | 100% | 96%–100% | 82 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 100% | 96%–100% | 82 |
2 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 82 per row
Completed calls ÷ calls; whiskers = 95% Wilson interval
A completed call returned a decision, right or wrong (accuracy is in the routing study). Whiskers are 95% Wilson intervals.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
- Every model call routed (49.5 per task)
- Only System One decisions (7 per task) (square)
Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).
| Item | Every model call routed (49.5 per task) | Only System One decisions (7 per task) |
|---|---|---|
| Claude Sonnet 5.5 (effort low, via Claude Code) | $247 | $34.97 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | $442 | $62.47 |
List-price calculation, not a run. 2 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $442. Lowest Claude Sonnet 5.5 (effort low, via Claude Code) $247. Only System One decisions (7 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $62.47. Lowest Claude Sonnet 5.5 (effort low, via Claude Code) $34.97.
Notes
Decisions per task from recorded runs × cost per decision
A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions
- Every model call routed (49.5 per task)
- Only System One decisions (7 per task) (square)
Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).
| Item | Every model call routed (49.5 per task) | Only System One decisions (7 per task) |
|---|---|---|
| Claude Sonnet 5.5 (effort low, via Claude Code) | 129 s | 18.2 s |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 621 s | 87.8 s |
List-price calculation, not a run. 2 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 621 s. Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 129 s. Only System One decisions (7 per task): slowest Claude Haiku 4.5 (thinking on, via Claude Code) 87.8 s. Fastest Claude Sonnet 5.5 (effort low, via Claude Code) 18.2 s.
Notes
Decisions per task × median decision time, if every decision waits in line
A calculation and an upper bound: it assumes each decision waits for the one before. Median recorded task wall time: 10.3 min. Jev’s delay uses its live median over the API from one Mac (network included); the Claude routers’ includes the CLI.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions
| Item | Strict pass | 95% interval | n |
|---|---|---|---|
| Claude Haiku 4.5 (single call) · Claude Code | 46% | 28%–65% | 24 |
| Claude Sonnet 5.5 (single call) · Claude Code | 100% | 86%–100% | 24 |
2 rows. Highest Claude Sonnet 5.5 (single call) · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 (single call) · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 24 per row
Same tasks and validators. Whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals. A format miss never counts as a pass; an attempt with no final answer (error, time-out, turn limit) counts as a fail. Contaminated attempts (a read outside the work folder) are left out. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
Claude Haiku 4.5 (single call) · Claude Code
Claude Haiku 4.5 (agent loop) · Claude Code
Claude Sonnet 5.5 (single call) · Claude Code
Claude Sonnet 5.5 (agent loop) · Claude Code
GPT-6 Luna (single call) · Codex CLI
GPT-6 Luna (agent loop) · Codex CLI
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | Claude Haiku 4.5 (single call) · Claude Code | Claude Haiku 4.5 (agent loop) · Claude Code | Claude Sonnet 5.5 (single call) · Claude Code | Claude Sonnet 5.5 (agent loop) · Claude Code | GPT-6 Luna (single call) · Codex CLI | GPT-6 Luna (agent loop) · Codex CLI | 95% interval | n |
|---|---|---|---|---|---|---|---|---|
| Interval merge fix | 100% | 67% | 100% | 100% | 100% | 100% | Claude Haiku 4.5 (single call) · Claude Code: 44%–100%; Claude Haiku 4.5 (agent loop) · Claude Code: 21%–94%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 34%–100%; GPT-6 Luna (agent loop) · Codex CLI: 34%–100% | 3 |
| DST day-length fix | 33% | 67% | 100% | 100% | 100% | — | Claude Haiku 4.5 (single call) · Claude Code: 6.2%–79%; Claude Haiku 4.5 (agent loop) · Claude Code: 21%–94%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 34%–100% | 3 |
| CSV parser | 67% | 67% | 100% | 100% | 100% | 50% | Claude Haiku 4.5 (single call) · Claude Code: 21%–94%; Claude Haiku 4.5 (agent loop) · Claude Code: 21%–94%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 34%–100%; GPT-6 Luna (agent loop) · Codex CLI: 9.4%–91% | 3 |
| Event-loop order | 0% | 100% | 100% | 100% | 0% | 100% | Claude Haiku 4.5 (single call) · Claude Code: 0%–56%; Claude Haiku 4.5 (agent loop) · Claude Code: 44%–100%; Claude Sonnet 5.5 (single call) · Claude Code: 44%–100%; Claude Sonnet 5.5 (agent loop) · Claude Code: 34%–100%; GPT-6 Luna (single call) · Codex CLI: 0%–66%; GPT-6 Luna (agent loop) · Codex CLI: 34%–100% | 3 |
4 rows, 6 series: Claude Haiku 4.5 (single call) · Claude Code, Claude Haiku 4.5 (agent loop) · Claude Code, Claude Sonnet 5.5 (single call) · Claude Code, Claude Sonnet 5.5 (agent loop) · Claude Code, GPT-6 Luna (single call) · Codex CLI, GPT-6 Luna (agent loop) · Codex CLI. Claude Haiku 4.5 (single call) · Claude Code: highest Interval merge fix 100% (95% interval 44%–100%, n 3). Lowest Event-loop order 0% (95% interval 0%–56%, n 3). All intervals overlap. Claude Haiku 4.5 (agent loop) · Claude Code: highest Event-loop order 100% (95% interval 44%–100%, n 3). Lowest CSV parser 67% (95% interval 21%–94%, n 3). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 2–3 per row
Share of attempts per task that passed strictly; 2 or 3 attempts per task and configuration
Whiskers are 95% Wilson intervals on 2 or 3 attempts, so they are very wide: read this chart for where a configuration failed, not for a ranking. Contaminated attempts (a read outside the work folder) are left out.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per attempt | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 (single call) · Claude Code | 39 s | 15.3 s–75.1 s | 24 |
| Claude Sonnet 5.5 (single call) · Claude Code | 7.8 s | 2.3 s–34.8 s | 24 |
2 rows. Slowest Claude Haiku 4.5 (single call) · Claude Code 39 s (range 15.3 s–75.1 s, n 24). Fastest Claude Sonnet 5.5 (single call) · Claude Code 7.8 s (range 2.3 s–34.8 s, n 24). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 24 per row
Median per configuration; whiskers = fastest and slowest attempt
Whiskers are a range (fastest and slowest attempt), not a confidence interval. Agent loop: wall time of the whole session, failures and time-outs included. One host, one network. The Claude single-call cells are reference cells from the hard head-to-head (8 tasks × 3 repetitions), reused, not rerun. The reference cells ran in a different hour.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
- Input tokens (cache reads included)
- Output tokens
| Item | Input tokens (cache reads included) | Output tokens | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Haiku 4.5 (single call) · Claude Code | 3,941 | 5,064 | Input tokens (cache reads included): 3,879–4,221; Output tokens: 1,899–9,321 | 24 |
| Claude Sonnet 5.5 (single call) · Claude Code | 2,281 | 1,050 | Input tokens (cache reads included): 2,234–2,669; Output tokens: 176–3,895 | 24 |
2 rows, 2 series: Input tokens (cache reads included), Output tokens. Input tokens (cache reads included): highest Claude Haiku 4.5 (single call) · Claude Code 3,941 (range 3,879–4,221, n 24). Lowest Claude Sonnet 5.5 (single call) · Claude Code 2,281 (range 2,234–2,669, n 24). Not all run ranges overlap. Output tokens: highest Claude Haiku 4.5 (single call) · Claude Code 5,064 (range 1,899–9,321, n 24). Lowest Claude Sonnet 5.5 (single call) · Claude Code 1,050 (range 176–3,895, n 24). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 24 per row
Median per configuration; whiskers = fewest and most
Whiskers are a range (fewest and most), not a confidence interval. Input counts the whole prompt of every model request in the attempt, cache reads and writes included; an agent loop re-sends its growing context each turn. Codex input includes its own system prompt and tool schemas. More tokens is not better or worse by itself.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators
| Item | Tool calls per attempt | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 (agent loop) · Claude Code | 3 | 2–18 | 24 |
| Claude Sonnet 5.5 (agent loop) · Claude Code | 0 | 0–3 | 16 |
2 rows. Highest Claude Haiku 4.5 (agent loop) · Claude Code 3 (range 2–18, n 24). Lowest Claude Sonnet 5.5 (agent loop) · Claude Code 0 (range 0–3, n 16). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 16–24 per row
Median per configuration; whiskers = fewest and most. A single call makes none
Whiskers are a range (fewest and most), not a confidence interval. Claude Code tools: shell, read, edit, write, glob, grep. Codex CLI: shell commands and file changes. The model chose whether to test its answer; the prompt allowed it but did not require it.
Source: Single call vs agent loop
| Item | Cost per strict pass | n |
|---|---|---|
| Claude Haiku 4.5 (single call) · Claude Code | $0.067 | 24 |
| Claude Sonnet 5.5 (single call) · Claude Code | $0.014 | 24 |
List-price calculation, not a run. 2 rows. Highest Claude Haiku 4.5 (single call) · Claude Code $0.067 (n 24). Lowest Claude Sonnet 5.5 (single call) · Claude Code $0.014 (n 24).
Notesn = 24 per row
All attempts in a configuration divided by its strict passes
Calculation, not a bill: reported tokens × list price for every attempt (per model when a session used several), cache reads at the cache-read price and cache writes at the one-hour write price, divided by the strict passes. The calls ran on flat subscriptions.
Sources: Single call vs agent loop, Provider head-to-head, hard set: eight hard tasks with strict validators, Repricing calculation, Anthropic list prices (Claude models), OpenAI list prices
- Exact decisions (every scored question right)
- Per-question accuracy
| Item | Exact decisions (every scored question right) | Per-question accuracy | 95% interval | n |
|---|---|---|---|---|
| Claude Haiku 4.5 (thinking off) · Claude Code | 87% | 91% | Exact decisions (every scored question right): 78%–92%; Per-question accuracy: 86%–94% | 82 |
| Claude Sonnet 5.5 (low) · Claude Code | 94% | 97% | Exact decisions (every scored question right): 87%–97%; Per-question accuracy: 94%–99% | 82 |
2 rows, 2 series: Exact decisions (every scored question right), Per-question accuracy. Exact decisions (every scored question right): highest Claude Sonnet 5.5 (low) · Claude Code 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 (thinking off) · Claude Code 87% (95% interval 78%–92%, n 82). All intervals overlap. Per-question accuracy: highest Claude Sonnet 5.5 (low) · Claude Code 97% (95% interval 94%–99%, n 194). Lowest Claude Haiku 4.5 (thinking off) · Claude Code 91% (95% interval 86%–94%, n 194). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 82–194 per row
Claude Haiku 4.5 and Claude Sonnet 5.5 (low effort); 82 decisions, the same cases for every arm
Whiskers are 95% Wilson intervals. Questions within a decision are related; per-question intervals are descriptive, not an independent-question test. The thinking-on and Sonnet arms are the recorded 2026-10-05 routing run, reused, not rerun; the thinking-off arm ran later on another account. An unanswered question counts as wrong.
Sources: Haiku thinking on vs off, Routing runs: Jev router vs LLM routing
- Wall time (CLI)
- Model time (API)
| Item | Wall time (CLI) | Model time (API) | Median to p95 | n |
|---|---|---|---|---|
| Claude Haiku 4.5 (thinking off) · Claude Code | 4.7 s | 3.8 s | Wall time (CLI): 4.7 s–8.2 s; Model time (API): 3.8 s–7.4 s | 82 |
| Claude Sonnet 5.5 (low) · Claude Code | 2.6 s | 1.6 s | Wall time (CLI): 2.6 s–4.3 s; Model time (API): 1.6 s–2.6 s | 82 |
2 rows, 2 series: Wall time (CLI), Model time (API). Wall time (CLI): slowest Claude Haiku 4.5 (thinking off) · Claude Code 4.7 s (median to p95 4.7 s–8.2 s, n 82). Fastest Claude Sonnet 5.5 (low) · Claude Code 2.6 s (median to p95 2.6 s–4.3 s, n 82). Not all run ranges overlap. Model time (API): slowest Claude Haiku 4.5 (thinking off) · Claude Code 3.8 s (median to p95 3.8 s–7.4 s, n 82). Fastest Claude Sonnet 5.5 (low) · Claude Code 1.6 s (median to p95 1.6 s–2.6 s, n 82). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n = 82 per row
Median wall time and model (API) time; whisker to the 95th percentile
Whiskers run from p50 to p95, not a confidence interval. Percentiles use the nearest-rank rule of the routing-overhead study, so the thinking-on and Sonnet medians match that study (the Jev-vs-LLM page interpolates between ranks and shows slightly different values). One call at a time through the Claude Code CLI; the thinking-on and Sonnet arms ran on another day.
Sources: Haiku thinking on vs off, Routing runs: Jev router vs LLM routing
- Thinking tokens
- Visible output tokens
| Item | Thinking tokens | Visible output tokens | n |
|---|---|---|---|
| Claude Haiku 4.5 (thinking off) · Claude Code | 0 | 366 | 82 |
| Claude Sonnet 5.5 (low) · Claude Code | 2 | 105 | 82 |
2 rows, 2 series: Thinking tokens, Visible output tokens. Thinking tokens: highest Claude Sonnet 5.5 (low) · Claude Code 2 (n 82). Lowest Claude Haiku 4.5 (thinking off) · Claude Code 0 (n 82). Visible output tokens: highest Claude Haiku 4.5 (thinking off) · Claude Code 366 (n 82). Lowest Claude Sonnet 5.5 (low) · Claude Code 105 (n 82).
Notesn = 82 per row
Mean per decision, as the Claude Code CLI reports them
Visible output = output tokens minus thinking tokens. It includes the structured answer the CLI asks for. Thinking tokens are counted by the CLI; their content is never captured. More tokens is not better or worse by itself.
Sources: Haiku thinking on vs off, Routing runs: Jev router vs LLM routing
| Item | Cost per 1,000 decisions | n |
|---|---|---|
| Claude Haiku 4.5 (thinking off) · Claude Code | $3.36 | 82 |
| Claude Sonnet 5.5 (low) · Claude Code | $7.32 | 82 |
List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 (low) · Claude Code $7.32 (n 82). Lowest Claude Haiku 4.5 (thinking off) · Claude Code $3.36 (n 82).
Notesn = 82 per row
Reported tokens × list price, per 1,000 decisions
Calculation, not a bill: the calls ran on a flat subscription. Reported input, cache and output tokens (thinking tokens are part of output) × list price, with the price table of the Jev-vs-LLM study. Sonnet cache writes use the one-hour rate ($4 per million tokens), as recorded in its receipts. The calculation matches the CLI-reported cost.
Sources: Haiku thinking on vs off, Routing runs: Jev router vs LLM routing, Repricing calculation, Anthropic list prices (Claude models)
- Strict pass: the whole reply is the right JSON
- Right answer in any format (strict pass or format miss)
| Item | Strict pass: the whole reply is the right JSON | Right answer in any format (strict pass or format miss) | 95% interval | n |
|---|---|---|---|---|
| Claude Haiku 4.5 (instructions) · Claude Code | 0% | 71% | Strict pass: the whole reply is the right JSON: 0%–14%; Right answer in any format (strict pass or format miss): 51%–85% | 24 |
| Claude Sonnet 5.5 (instructions) · Claude Code | 100% | 100% | Strict pass: the whole reply is the right JSON: 76%–100%; Right answer in any format (strict pass or format miss): 76%–100% | 12 |
List-price calculation, not a run. 2 rows, 2 series: Strict pass: the whole reply is the right JSON, Right answer in any format (strict pass or format miss). Strict pass: the whole reply is the right JSON: highest Claude Sonnet 5.5 (instructions) · Claude Code 100% (95% interval 76%–100%, n 12). Lowest Claude Haiku 4.5 (instructions) · Claude Code 0% (95% interval 0%–14%, n 24). Not all intervals overlap. Right answer in any format (strict pass or format miss): highest Claude Sonnet 5.5 (instructions) · Claude Code 100% (95% interval 76%–100%, n 12). Lowest Claude Haiku 4.5 (instructions) · Claude Code 71% (95% interval 51%–85%, n 24). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 12–24 per row
Three extraction prompts pooled; whiskers are 95% Wilson intervals
Whiskers are 95% Wilson intervals (a calculation) over 24 calls and 12 calls per configuration; every error counts as a fail. Strict: the whole reply parses as JSON and matches the expected answer exactly. A format miss is a right answer inside a code fence or prose, so it is never a strict pass.
Source: JSON schema vs instructions
- Strict pass
- Format miss
- Wrong values
- Error
One square per call; counts at the right are exact and in legend order.
| Item | Strict pass | Format miss | Wrong values | Error | n |
|---|---|---|---|---|---|
| Claude Haiku 4.5 (instructions) · Claude Code | 0 | 17 | 7 | 0 | 24 |
| Claude Sonnet 5.5 (instructions) · Claude Code | 12 | 0 | 0 | 0 | 12 |
2 rows, 4 series: Strict pass, Format miss, Wrong values, Error. Strict pass: highest Claude Sonnet 5.5 (instructions) · Claude Code 12 (n 12). Lowest Claude Haiku 4.5 (instructions) · Claude Code 0 (n 24). Format miss: highest Claude Haiku 4.5 (instructions) · Claude Code 17 (n 24). Lowest Claude Sonnet 5.5 (instructions) · Claude Code 0 (n 12).
Notesn 12–24 per row
Counts of calls per configuration; the three prompts pooled
Counts of calls, not rates; the pass-rate chart carries the same results with 95% intervals. Format miss: the right answer inside a code fence or prose. Wrong values: any other completed reply, with a wrong value, key or type (a reply that sits in a code fence and also has a wrong value is counted here). Error: the call did not complete.
Source: JSON schema vs instructions
1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.
| Item | Median time per call (the three prompts pooled) | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 (instructions) · Claude Code | 9.5 s | 5.7 s–17 s | 24 |
| Claude Sonnet 5.5 (instructions) · Claude Code | 3.5 s | 2.7 s–4.1 s | 12 |
2 rows. Slowest Claude Haiku 4.5 (instructions) · Claude Code 9.5 s (range 5.7 s–17 s, n 24). Fastest Claude Sonnet 5.5 (instructions) · Claude Code 3.5 s (range 2.7 s–4.1 s, n 12). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 12–24 per row
Median; whiskers = fastest and slowest completed call
Whiskers are a range (fastest and slowest call), not a confidence interval. Wall time from process start to exit, so it includes CLI start-up; Codex CLI timings include its larger system prompt. The three prompts differ in length, which widens every range.
Source: JSON schema vs instructions
- Median output tokens per call
- of which median reasoning tokens per call (thinking) (inner bar)
| Item | Median output tokens per call | Median reasoning tokens per call (thinking) | n |
|---|---|---|---|
| Claude Haiku 4.5 (instructions) · Claude Code | 1,128 | 934 | 24 |
| Claude Sonnet 5.5 (instructions) · Claude Code | 368 | 182 | 12 |
2 rows, 2 series: Median output tokens per call, Median reasoning tokens per call (thinking). Median output tokens per call: highest Claude Haiku 4.5 (instructions) · Claude Code 1,128 (n 24). Lowest Claude Sonnet 5.5 (instructions) · Claude Code 368 (n 12). Median reasoning tokens per call (thinking): highest Claude Haiku 4.5 (instructions) · Claude Code 934 (n 24). Lowest Claude Sonnet 5.5 (instructions) · Claude Code 182 (n 12).
Notesn 12–24 per row
Median per call; ranges and sample sizes are in the note
Output tokens include reasoning tokens. Claude schema-mode output includes the CLI’s structured-output tool call. Input counts include CLI context and are not compared. Ranges (not intervals): Claude Haiku 4.5 (instructions) · Claude Code: output 698 to 2059 (n = 24); reasoning 559 to 1830 (n = 24); Claude Haiku 4.5 (JSON schema) · Claude Code: output 716 to 1462 (n = 24); reasoning 522 to 1187 (n = 24); Claude Sonnet 5.5 (instructions) · Claude Code: output 154 to 456 (n = 12); reasoning 54 to 278 (n = 12); Claude Sonnet 5.5 (JSON schema) · Claude Code: output 273 to 541 (n = 12); reasoning 0 to 260 (n = 12); GPT-6.1 Sol (low, instructions) · Codex CLI: output 69 to 259 (n = 12); reasoning 0 to 60 (n = 12); GPT-6.1 Sol (low, JSON schema) · Codex CLI: output 98 to 181 (n = 12); reasoning 0 to 49 (n = 12).
Source: JSON schema vs instructions
| Item | Exact rate | 95% interval | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 79% | 66%–87% | 56 |
| Claude Sonnet 5.5 (low) · Claude Code | 88% | 76%–94% | 56 |
2 rows. Highest Claude Sonnet 5.5 (low) · Claude Code 88% (95% interval 76%–94%, n 56). Lowest Claude Haiku 4.5 · Claude Code 79% (95% interval 66%–87%, n 56). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 56 per row
Share of the 56 holdout cases where every scored question was acceptable
Whiskers are 95% Wilson intervals. Cases written blind to every router answer, then frozen. Jev: its first of three repetitions.
Source: Routing on unseen holdout decisions
| Item | Key accuracy | 95% interval | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 82% | 74%–87% | 125 |
| Claude Sonnet 5.5 (low) · Claude Code | 92% | 86%–96% | 125 |
2 rows. Highest Claude Sonnet 5.5 (low) · Claude Code 92% (95% interval 86%–96%, n 125). Lowest Claude Haiku 4.5 · Claude Code 82% (95% interval 74%–87%, n 125). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 125 per row
Each open question a router was asked; an unanswered question counts as wrong
Whiskers are nominal 95% Wilson intervals. Context-shape cases ask up to 8 questions each, the other decision types one. Questions in one case are not independent; these intervals do not adjust for that grouping.
Source: Routing on unseen holdout decisions
Jev 1.13 (TypeSafe)
Claude Haiku 4.5 · Claude Code
Claude Sonnet 5.5 (low) · Claude Code
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | Jev 1.13 (TypeSafe) | Claude Haiku 4.5 · Claude Code | Claude Sonnet 5.5 (low) · Claude Code | 95% interval | n |
|---|---|---|---|---|---|
| Failure class | 93% | 93% | 100% | Jev 1.13 (TypeSafe): 69%–99%; Claude Haiku 4.5 · Claude Code: 69%–99%; Claude Sonnet 5.5 (low) · Claude Code: 78%–100% | 14 |
| Is it a rule? | 93% | 93% | 93% | Jev 1.13 (TypeSafe): 69%–99%; Claude Haiku 4.5 · Claude Code: 69%–99%; Claude Sonnet 5.5 (low) · Claude Code: 69%–99% | 14 |
| Context shape | 57% | 36% | 57% | Jev 1.13 (TypeSafe): 33%–79%; Claude Haiku 4.5 · Claude Code: 16%–61%; Claude Sonnet 5.5 (low) · Claude Code: 33%–79% | 14 |
3 rows, 3 series: Jev 1.13 (TypeSafe), Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 (low) · Claude Code. Jev 1.13 (TypeSafe): highest Failure class 93% (95% interval 69%–99%, n 14). Lowest Context shape 57% (95% interval 33%–79%, n 14). All intervals overlap. Claude Haiku 4.5 · Claude Code: highest Failure class 93% (95% interval 69%–99%, n 14). Lowest Context shape 36% (95% interval 16%–61%, n 14). Not all intervals overlap.
NotesWhiskers: 95% Wilson intervaln = 14 per row
14 cases per decision type
Whiskers are 95% Wilson intervals. With 14 cases a perfect score has an interval of 78% to 100%, so a decision type where every router scores 14 of 14 is at its ceiling and cannot rank them.
Source: Routing on unseen holdout decisions
- Tuned set (routing-jev-vs-llm)
- Unseen holdout (square)
Gap labels, Unseen holdout vs Tuned set (routing-jev-vs-llm): Unseen holdout is x percentage points higher (+) or lower (−) than Tuned set (routing-jev-vs-llm), calculated from the two values shown; lines are the 95% Wilson interval.
| Item | Tuned set (routing-jev-vs-llm) | Unseen holdout | 95% interval | n |
|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 89% | 79% | Tuned set (routing-jev-vs-llm): 80%–94%; Unseen holdout: 66%–87% | 82 |
| Claude Sonnet 5.5 (low) · Claude Code | 94% | 88% | Tuned set (routing-jev-vs-llm): 87%–97%; Unseen holdout: 76%–94% | 82 |
2 rows, 2 series: Tuned set (routing-jev-vs-llm), Unseen holdout. Tuned set (routing-jev-vs-llm): highest Claude Sonnet 5.5 (low) · Claude Code 94% (95% interval 87%–97%, n 82). Lowest Claude Haiku 4.5 · Claude Code 89% (95% interval 80%–94%, n 82). All intervals overlap. Unseen holdout: highest Claude Sonnet 5.5 (low) · Claude Code 88% (95% interval 76%–94%, n 56). Lowest Claude Haiku 4.5 · Claude Code 79% (95% interval 66%–87%, n 56). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 56–82 per row
Tuned set: the routing study’s 82 cases, revised against Jev answers. Holdout: 56 new cases, frozen before any router call
Whiskers are 95% Wilson intervals. The two case sets differ in mix and size, so a gap mixes a change of case set with any change in the router, and the data cannot separate them. A gap counts only when the two intervals do not overlap.
Sources: Routing on unseen holdout decisions, Routing runs: Jev router vs LLM routing
- Wall time
- Model time (API, CLI-reported)
| Item | Wall time | Model time (API, CLI-reported) | Median to p95 | n |
|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 9.4 s | 7.5 s | Wall time: 9.4 s–25.4 s; Model time (API, CLI-reported): 7.5 s–23.9 s | 56 |
| Claude Sonnet 5.5 (low) · Claude Code | 2.4 s | 1.5 s | Wall time: 2.4 s–3.7 s; Model time (API, CLI-reported): 1.5 s–2.4 s | 56 |
2 rows, 2 series: Wall time, Model time (API, CLI-reported). Wall time: slowest Claude Haiku 4.5 · Claude Code 9.4 s (median to p95 9.4 s–25.4 s, n 56). Fastest Claude Sonnet 5.5 (low) · Claude Code 2.4 s (median to p95 2.4 s–3.7 s, n 56). Not all run ranges overlap. Model time (API, CLI-reported): slowest Claude Haiku 4.5 · Claude Code 7.5 s (median to p95 7.5 s–23.9 s, n 56). Fastest Claude Sonnet 5.5 (low) · Claude Code 1.5 s (median to p95 1.5 s–2.4 s, n 56). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n = 56 per row
Median, whisker to the 95th percentile
The routes differ. Jev: one HTTPS call to the vendor API from one Mac, all three repetitions. Claude: the Claude Code CLI from process start to exit, which adds start-up time a direct API call would not. The whisker runs from the median to the 95th percentile; it is not a confidence interval. The shared Mac also ran local arena model servers from 15:18 local time (20:18 UTC), before all holdout calls. Host contention may affect wall times, especially CLI times. We did not rerun the timing.
Source: Routing on unseen holdout decisions
| Item | Cost | n |
|---|---|---|
| Claude Haiku 4.5 · Claude Code | $7.13 | 56 |
| Claude Sonnet 5.5 (low) · Claude Code | $7.24 | 56 |
List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 (low) · Claude Code $7.24 (n 56). Lowest Claude Haiku 4.5 · Claude Code $7.13 (n 56).
Notesn = 56 per row
Reported tokens per decision × list price
Calculation, not a bill. Jev: reported input tokens × $0.042 per million, output free. Claude: CLI-reported tokens × list price; the CLI wrote its prompt cache as 1-hour writes, priced at 2× input, and adds its own system prompt and tool-schema tokens. The Claude routers ran on a subscription. At the tuned-set study’s convention (every cache write at 1.25× input), Sonnet 5.5 (low) would be $4.877 here. That study shows $4.996 for it on its own cases, where the CLI reported $7.324, so its cost row and this one differ by convention and by case mix.
Sources: Routing on unseen holdout decisions, Repricing calculation, Jev 1.13 list price, Anthropic list prices (Claude models)
- Reasoning (output tokens)
- Remaining output (visible-answer estimate)
- Input (prompt, cache priced)
Totals are the sum of the parts shown. Shares are calculated from the same values.
| Item | Reasoning (output tokens) | Remaining output (visible-answer estimate) | Input (prompt, cache priced) | n |
|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | $0.024 | $0.0018 | $0.0045 | 24 |
| Claude Sonnet 5.5 · Claude Code | $0.0067 | $0.0037 | $0.004 | 24 |
List-price calculation, not a run. 2 rows, 3 series: Reasoning (output tokens), Remaining output (visible-answer estimate), Input (prompt, cache priced). Reasoning (output tokens): highest Claude Haiku 4.5 · Claude Code $0.024 (n 24). Lowest Claude Sonnet 5.5 · Claude Code $0.0067 (n 24). Remaining output (visible-answer estimate): highest Claude Sonnet 5.5 · Claude Code $0.0037 (n 24). Lowest Claude Haiku 4.5 · Claude Code $0.0018 (n 24).
Notesn = 24 per row
Mean per call on the hard tasks; the three parts add up to the call
Calculation, not a bill: reported tokens × list price (per million output tokens: Haiku 4.5 $5, Sonnet 5.5 $10, Opus 5.5 $20, Fable 5.1 $50 and GPT-6.1 Sol $10); the calls ran on flat subscriptions. Reasoning and remaining output both use the output price. The remainder estimates visible-answer tokens; its exact meaning depends on the CLI counters. Input is the whole prompt, with cache reads and writes priced as in the hard head-to-head. It includes CLI context; these receipts do not separate task tokens from CLI context. It changes with cache counters. Batch timing and cache behavior were not controlled. Means, not medians, so the parts add up.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
- Eight hard tasks
- Five short tasks (square)
Gap labels, Five short tasks vs Eight hard tasks: Five short tasks is x percentage points higher (+) or lower (−) than Eight hard tasks, calculated from the two values shown; lines are the lowest–highest run (not an interval).
| Item | Eight hard tasks | Five short tasks | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 92% | 90% | Eight hard tasks: 76%–99%; Five short tasks: 73%–98% | 24 |
| Claude Sonnet 5.5 · Claude Code | 55% | 0% | Eight hard tasks: 0%–96%; Five short tasks: 0%–73% | 24 |
List-price calculation, not a run. 2 rows, 2 series: Eight hard tasks, Five short tasks. Eight hard tasks: highest Claude Haiku 4.5 · Claude Code 92% (range 76%–99%, n 24). Lowest Claude Sonnet 5.5 · Claude Code 55% (range 0%–96%, n 24). All run ranges overlap. Five short tasks: highest Claude Haiku 4.5 · Claude Code 90% (range 73%–98%, n 15). Lowest Claude Sonnet 5.5 · Claude Code 0% (range 0%–73%, n 15). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 15–24 per row
Median call per configuration; five short tasks and eight hard tasks
Calculation from reported tokens, not a run: the median of per-call reasoning ÷ output. A median of 0% means at least half of the calls reported 0 reasoning tokens in the receipt. On a route that reports reasoning, we retain a numeric 0 as the recorded counter. It does not prove the model did no internal reasoning. The table says how many calls reported 0. The short tasks are five small validated tasks (a bug fix, a JSON extraction, an arithmetic problem, a refactor and a ticket classification). Per-call ranges are in the tables.
Sources: Reasoning token bill (calculation), Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Provider head-to-head: Claude Code models vs Codex efforts, Anthropic list prices (Claude models), OpenAI list prices
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Time to first text | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 4 s | 2.8 s–6.4 s | 4 |
| Claude Sonnet 5.5 · Claude Code | 2 s | 0.9 s–4.1 s | 4 |
2 rows. Slowest Claude Haiku 4.5 · Claude Code 4 s (range 2.8 s–6.4 s, n 4). Fastest Claude Sonnet 5.5 · Claude Code 2 s (range 0.9 s–4.1 s, n 4). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = fastest and slowest call
Dot = median; whiskers = fastest and slowest call (a range, not a confidence interval). The clock starts when the CLI starts, so first text includes CLI start-up (see chart cli-startup-tax) and any reasoning before the first word. One Mac, one network, two sessions on one night.
Source: LLM speed anatomy
| Item | Visible tokens per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 153 | 153–216 | 4 |
| Claude Sonnet 5.5 · Claude Code | 232 | 230–233 | 4 |
List-price calculation, not a run. 2 rows. Highest Claude Sonnet 5.5 · Claude Code 232 (range 230–233, n 4). Lowest Claude Haiku 4.5 · Claude Code 153 (range 153–216, n 4). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n = 4 per row
Median of 4 calls per model; whiskers = slowest and fastest call
Calculation, not a measurement: visible output tokens (the CLI’s output token count minus its reported reasoning tokens) divided by the time from the first text to the end of the call. The denominator includes the CLI’s exit overhead; its size is not measured here. Whiskers are a range of calls, not a confidence interval. The calculation removes reported reasoning tokens. These receipts do not locate all reasoning in time.
Source: LLM speed anatomy
| Item | Characters per second | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 547 | 546–548 | 3 |
| Claude Sonnet 5.5 · Claude Code | 517 | 513–519 | 4 |
List-price calculation, not a run. 2 rows. Highest Claude Haiku 4.5 · Claude Code 547 (range 546–548, n 3). Lowest Claude Sonnet 5.5 · Claude Code 517 (range 513–519, n 4). Not all run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 3–4 per row
Only replies that matched all 250 lines; median per model, whiskers = slowest and fastest call
Calculation, not a measurement: the characters of a correct reply (the same text for every model) divided by the time from the first text to the end of the call. Each vendor counts the same text as a different number of tokens, so characters compare across models where tokens do not. The CLI’s exit time is inside that time. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
- Claude Haiku 4.5 · Claude Code
- Claude Sonnet 5.5 · Claude Code
- Claude Opus 5.5 · Claude Code
- GPT-6.1 Sol (low) · Codex CLI
| Prompt-size target (approximate Haiku tokens; calibration calculation) | Claude Haiku 4.5 · Claude Code | Claude Sonnet 5.5 · Claude Code | Claude Opus 5.5 · Claude Code | GPT-6.1 Sol (low) · Codex CLI | Range (lowest–highest run) | n |
|---|---|---|---|---|---|---|
| 1k | 1.9 s | 1.5 s | 1.5 s | 3.4 s | Claude Haiku 4.5 · Claude Code: 1.9 s–2 s; Claude Sonnet 5.5 · Claude Code: 1.2 s–1.7 s; Claude Opus 5.5 · Claude Code: 1.5 s–2 s; GPT-6.1 Sol (low) · Codex CLI: 3.4 s–4.8 s | 3 |
| 16k | 2.3 s | 1.8 s | 1.7 s | 4 s | Claude Haiku 4.5 · Claude Code: 2.2 s–2.5 s; Claude Sonnet 5.5 · Claude Code: 1.6 s–2.1 s; Claude Opus 5.5 · Claude Code: 1.7 s–3 s; GPT-6.1 Sol (low) · Codex CLI: 3.3 s–4.3 s | 3 |
| 64k | 2.8 s | 3.1 s | 1.8 s | 3.9 s | Claude Haiku 4.5 · Claude Code: 2.5 s–2.9 s; Claude Sonnet 5.5 · Claude Code: 1.4 s–3.6 s; Claude Opus 5.5 · Claude Code: 1.7 s–3.7 s; GPT-6.1 Sol (low) · Codex CLI: 3.4 s–4.4 s | 3 |
List-price calculation, not a run. 3 rows, 4 series: Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Opus 5.5 · Claude Code, GPT-6.1 Sol (low) · Codex CLI. Claude Haiku 4.5 · Claude Code: slowest 64k 2.8 s (range 2.5 s–2.9 s, n 3). Fastest 1k 1.9 s (range 1.9 s–2 s, n 3). Not all run ranges overlap. Claude Sonnet 5.5 · Claude Code: slowest 64k 3.1 s (range 1.4 s–3.6 s, n 3). Fastest 1k 1.5 s (range 1.2 s–1.7 s, n 3). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Median of 3 calls per size; whiskers = fastest and slowest call
Each call used a new ledger seed. Cache-read counts stayed within the short-prompt baseline (see the cache table). This does not identify which tokens were cached. Sizes name the text we send; each model’s reported input tokens are in the table and include the CLI’s own prefix. The size calibration subtracts estimated prefixes from probe input counts; these are calculations, not measured prefix counts for each call. Whiskers are a range of calls, not a confidence interval.
Source: LLM speed anatomy
1k prompt
16k prompt
64k prompt
One panel per series, all on the same axis; whiskers are the fastest–slowest run (not an interval).
| Item | 1k prompt | 16k prompt | 64k prompt | Range (lowest–highest run) | n |
|---|---|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 2.3 s | 2.8 s | 3.1 s | 1k prompt: 2.2 s–2.5 s; 16k prompt: 2.6 s–2.8 s; 64k prompt: 2.8 s–3.3 s | 3 |
| Claude Sonnet 5.5 · Claude Code | 1.8 s | 2.1 s | 3.4 s | 1k prompt: 1.6 s–2.1 s; 16k prompt: 2 s–2.5 s; 64k prompt: 1.7 s–4.4 s | 3 |
2 rows, 3 series: 1k prompt, 16k prompt, 64k prompt. 1k prompt: slowest Claude Haiku 4.5 · Claude Code 2.3 s (range 2.2 s–2.5 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 1.8 s (range 1.6 s–2.1 s, n 3). Not all run ranges overlap. 16k prompt: slowest Claude Haiku 4.5 · Claude Code 2.8 s (range 2.6 s–2.8 s, n 3). Fastest Claude Sonnet 5.5 · Claude Code 2.1 s (range 2 s–2.5 s, n 3). Not all run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n = 3 per row
Median of 3 calls per bar; whiskers = fastest and slowest call
Whole call: CLI start-up, first text and the one-line answer. Whiskers are a range of calls, not a confidence interval. Each call used a new ledger.
Source: LLM speed anatomy
| Item | Exact answer | 95% interval | n |
|---|---|---|---|
| Claude Haiku 4.5 · Claude Code | 100% | 70%–100% | 9 |
| Claude Sonnet 5.5 · Claude Code | 100% | 70%–100% | 9 |
2 rows. All at 100%.
NotesWhiskers: 95% Wilson intervaln = 9 per row
All sizes together per model; whiskers = 95% Wilson intervals
Whiskers are 95% Wilson intervals. One lookup question per call; a reply with extra words is a format miss, not a pass. With 9 calls per model, a perfect score still has a wide interval.
Source: LLM speed anatomy
- Claude Haiku 4.5 · Claude Code
- Claude Sonnet 5.5 · Claude Code (square)
Gap labels, Claude Sonnet 5.5 · Claude Code vs Claude Haiku 4.5 · Claude Code: Claude Sonnet 5.5 · Claude Code is x% higher (+) or lower (−) than Claude Haiku 4.5 · Claude Code, calculated from the two values shown (the change counted from Claude Haiku 4.5 · Claude Code’s value).
| Item | Claude Haiku 4.5 · Claude Code | Claude Sonnet 5.5 · Claude Code | n |
|---|---|---|---|
| Interval merge fix | $0.018 | $0.0056 | 3 |
| DST day length | $0.029 | $0.025 | 3 |
| CSV parser | $0.029 | $0.015 | 3 |
| Event-loop order | $0.037 | $0.016 | 3 |
| Room schedule | $0.036 | $0.012 | 3 |
| SemVer regex | $0.042 | $0.0051 | 3 |
| Money refactor | $0.02 | $0.0096 | 3 |
| SQLite report query | $0.036 | $0.017 | 3 |
List-price calculation, not a run. 8 rows, 2 series: Claude Haiku 4.5 · Claude Code, Claude Sonnet 5.5 · Claude Code. Claude Haiku 4.5 · Claude Code: highest SemVer regex $0.042 (n 3). Lowest Interval merge fix $0.018 (n 3). Claude Sonnet 5.5 · Claude Code: highest DST day length $0.025 (n 3). Lowest SemVer regex $0.0051 (n 3).
Notesn = 3 per row
Median of each task’s calls, default effort, Claude Code
Calculation: reported tokens × list price for each call, then the median per task; the calls ran on a flat subscription. Haiku’s list price per token is lower. Its median calls cost more and contained more output tokens. This does not isolate the effect of effort or thinking. 3 calls per task and configuration; the table shows each call-cost range, not a confidence interval.
Sources: Repricing calculation, Provider head-to-head, hard set: eight hard tasks with strict validators, Effort ladder: the hard task set at each effort level, Anthropic list prices (Claude models)
- Strict pass
- Lenient (format misses counted)
| Item | Strict pass | Lenient (format misses counted) | 95% interval | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 38% | 38% | Strict pass: 18%–61%; Lenient (format misses counted): 18%–61% | 16 |
| Claude Haiku 4.5 · Claude Code | 0% | 0% | Strict pass: 0%–24%; Lenient (format misses counted): 0%–24% | 12 |
2 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 38% (95% interval 18%–61%, n 16). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–24%, n 12). All intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 38% (95% interval 18%–61%, n 16). Lowest Claude Haiku 4.5 · Claude Code 0% (95% interval 0%–24%, n 12). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 12–16 per row
Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts
Whiskers are 95% Wilson intervals. Strict passes decide the result. The lenient reading shows which failures contain a correct answer in the wrong format. A format miss never counts as a pass. The study picked tasks that Sonnet did not pass twice in a pilot. This selection can lower a Sonnet row. Counted calls are new calls.
Source: Harder tasks head-to-head
| Item | Tool attempt | 95% interval | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 31% | 14%–56% | 16 |
| Claude Haiku 4.5 · Claude Code | 8.3% | 1.5%–35% | 12 |
2 rows. Highest Claude Sonnet 5.5 · Claude Code 31% (95% interval 14%–56%, n 16). Lowest Claude Haiku 4.5 · Claude Code 8.3% (95% interval 1.5%–35%, n 12). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 12–16 per row
Share of calls whose reply held tool-call markup or whose tool call the CLI could not parse
Every prompt asked for the answer only. Both CLIs disabled tools. The Codex runner adds a developer instruction not to call tools. The Claude Code runner adds no such instruction. The validator decides the score. Tool-call markup is a separate flag; a flagged call can still pass if its whole reply meets the validator. It is a behaviour, not a quality score: more is not better. No call with a tool attempt passed. Whiskers are 95% Wilson intervals. The flag comes from the reply text, which is not published.
Source: Harder tasks head-to-head
GPT-6.1 Sol (medium) · Codex CLI
Claude Opus 5.5 · Claude Code
Claude Sonnet 5.5 · Claude Code
Claude Haiku 4.5 · Claude Code
One panel per series, all on the same axis; whiskers are the 95% Wilson interval.
| Item | GPT-6.1 Sol (medium) · Codex CLI | Claude Opus 5.5 · Claude Code | Claude Sonnet 5.5 · Claude Code | Claude Haiku 4.5 · Claude Code | 95% interval | n |
|---|---|---|---|---|---|---|
| 10x10 nonogram | 75% | 100% | 100% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 30%–95%; Claude Opus 5.5 · Claude Code: 44%–100%; Claude Sonnet 5.5 · Claude Code: 51%–100%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
| Sudoku, 22 givens | 25% | 0% | 0% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 4.6%–70%; Claude Opus 5.5 · Claude Code: 0%–56%; Claude Sonnet 5.5 · Claude Code: 0%–49%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
| Seeded shuffle output | 75% | 33% | 50% | 0% | GPT-6.1 Sol (medium) · Codex CLI: 30%–95%; Claude Opus 5.5 · Claude Code: 6.2%–79%; Claude Sonnet 5.5 · Claude Code: 15%–85%; Claude Haiku 4.5 · Claude Code: 0%–56% | 4 |
3 rows, 4 series: GPT-6.1 Sol (medium) · Codex CLI, Claude Opus 5.5 · Claude Code, Claude Sonnet 5.5 · Claude Code, Claude Haiku 4.5 · Claude Code. GPT-6.1 Sol (medium) · Codex CLI: highest 10x10 nonogram 75% (95% interval 30%–95%, n 4). Lowest Sudoku, 22 givens 25% (95% interval 4.6%–70%, n 4). All intervals overlap. Claude Opus 5.5 · Claude Code: highest 10x10 nonogram 100% (95% interval 44%–100%, n 3). Lowest Sudoku, 22 givens 0% (95% interval 0%–56%, n 3). All intervals overlap.
NotesWhiskers: 95% Wilson intervaln 3–4 per row
One bar per configuration and task; each bar rests on only a few calls, so the 95% intervals are wide
Each bar is 3 to 4 calls, so one call moves a bar by a quarter or a third. Whiskers are 95% Wilson intervals; with this few calls they overlap almost everywhere.
Source: Harder tasks head-to-head
Every lane’s run range overlaps another lane’s, so the chart shows no finish order.
| Item | Total time per call | Range (lowest–highest run) | n |
|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 70.4 s | 4.3 s–210 s | 12 |
| Claude Haiku 4.5 · Claude Code | 109 s | 25.7 s–224 s | 10 |
2 rows. Slowest Claude Haiku 4.5 · Claude Code 109 s (range 25.7 s–224 s, n 10). Fastest Claude Sonnet 5.5 · Claude Code 70.4 s (range 4.3 s–210 s, n 12). All run ranges overlap.
NotesLines: fastest–slowest run (not an interval)n 10–12 per row
Median per configuration; whiskers = fastest and slowest call
Median and range over the calls that completed. Completed calls include wrong answers and format misses. Only timeouts and tool-call parse errors are excluded from this run’s timings. Both count as non-passes in the outcomes chart. One Mac, one network, one session. Arena servers shared the Mac during part of the run. Host load was not controlled, so these times cannot isolate model speed. Whiskers are a range, not a confidence interval. Times include the CLI start-up and the CLI’s own system prompt. Highlighted: configurations that passed every call.
Source: Harder tasks head-to-head
- Output tokens
- of which reasoning tokens (inner bar)
| Item | Output tokens | Reasoning tokens | Range (lowest–highest run) | n |
|---|---|---|---|---|
| Claude Sonnet 5.5 · Claude Code | 9,287 | 6,557 | Output tokens: 407–27,921; Reasoning tokens: 63–27,902 | 12 |
| Claude Haiku 4.5 · Claude Code | 12,508 | 12,483 | Output tokens: 2,965–26,532; Reasoning tokens: 2,924–26,510 | 10 |
2 rows, 2 series: Output tokens, Reasoning tokens. Output tokens: highest Claude Haiku 4.5 · Claude Code 12,508 (range 2,965–26,532, n 10). Lowest Claude Sonnet 5.5 · Claude Code 9,287 (range 407–27,921, n 12). All run ranges overlap. Reasoning tokens: highest Claude Haiku 4.5 · Claude Code 12,483 (range 2,924–26,510, n 10). Lowest Claude Sonnet 5.5 · Claude Code 6,557 (range 63–27,902, n 12). All run ranges overlap.
NotesLines: lowest–highest run (not an interval)n 10–12 per row
Median per configuration; reasoning tokens as the CLI reports them
Medians and minimum-to-maximum token ranges cover completed calls only. Ranges are not confidence intervals. The chart omits unknown reasoning counts. Reasoning tokens are part of the output tokens where the CLI reports them. Their content is never captured. Claude Code used an output-token cap setting of 16,000. Some reported totals exceeded it. Codex CLI had no cap. More tokens is not better or worse by itself.
Source: Harder tasks head-to-head
How a row is called
- AheadThe 95% intervals do not overlap, or the run ranges or p50–p95 bands do not overlap with at least 5 runs per side.
- TieThe values match, or both sit at the same ceiling.
- UnclearThe intervals or ranges overlap, too few runs were recorded, no interval was recorded, or more is not better (token counts are never a win).
- CalculationDerived from list prices and recorded counts. Not a bill and not a run.
The dataset compiler makes every call; this page only draws it. No composite score, no rank.
Questions
- Which is better, Claude Haiku 4.5 or Claude Sonnet 5.5?
- Claude Haiku 4.5 and Claude Sonnet 5.5 share 113 measured metrics and 40 list-price calculations from 14 studies. Claude Sonnet 5.5 leads on 21 rows: Pass rate on eight hard tasks (Strict pass), 100% (24/24) vs 46% (11/24); Pass rate on eight hard tasks (Lenient (format misses counted)), 100% (24/24) vs 67% (16/24); Same prompt, 10 times: strict pass rate (Exact number), 100% (10/10) vs 0% (0/10); and 18 more. On those rows the 95% intervals, run ranges and p50–p95 bands do not overlap; only a 95% interval is a confidence interval. The other rows are 58 ties and 74 unclear; each row says why. Calculation rows are derived from list prices and recorded counts; they are not bills or runs. Some rows rest on small samples (n = 2 at the smallest).
- How were Claude Haiku 4.5 and Claude Sonnet 5.5 measured?
- They share 113 measured metrics and 40 list-price calculations from 14 public studies: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head; Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks; Prompt caching and run-to-run consistency in Claude Code and Codex CLI; Does memory help Claude Code? 8 kinds of agent memory, tested; Jev vs Claude as a router: accuracy and cost; Routing overhead: deterministic policy vs LLM routers vs Jev; Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks; Does thinking pay for Claude Haiku 4.5? Thinking on vs off; Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI; Jev vs Claude routers on unseen decisions: a blind holdout; How much of an AI bill is thinking? Reasoning tokens by model and effort; Where the seconds go: first text, output speed and prompt size for 6 LLMs; Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts; GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Every row names its configuration, its sample size and its interval or range.
- How do Claude Haiku 4.5 and Claude Sonnet 5.5 compare on pass rate on eight hard tasks (Strict pass)?
- Claude Haiku 4.5: 46% (11/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 28% to 65%). Claude Sonnet 5.5: 100% (24/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 86% to 100%). The 95% intervals do not overlap (Claude Haiku 4.5 28% to 65%; Claude Sonnet 5.5 86% to 100%).
- How do Claude Haiku 4.5 and Claude Sonnet 5.5 compare on pass rate on eight hard tasks (Lenient (format misses counted))?
- Claude Haiku 4.5: 67% (16/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 47% to 82%). Claude Sonnet 5.5: 100% (24/24) (Claude Code · eight hard validated tasks; n = 24; 95% interval 86% to 100%). The 95% intervals do not overlap (Claude Haiku 4.5 47% to 82%; Claude Sonnet 5.5 86% to 100%).
- How do Claude Haiku 4.5 and Claude Sonnet 5.5 compare on same prompt, 10 times: strict pass rate (Exact number)?
- Claude Haiku 4.5: 0% (0/10) (Claude Code · same prompt repeated 10 times; n = 10; 95% interval 0% to 28%). Claude Sonnet 5.5: 100% (10/10) (Claude Code · same prompt repeated 10 times; n = 10; 95% interval 72% to 100%). The 95% intervals do not overlap (Claude Haiku 4.5 0% to 28%; Claude Sonnet 5.5 72% to 100%).
- How do Claude Haiku 4.5 and Claude Sonnet 5.5 compare on same prompt, 10 times: strict pass rate (JSON object)?
- Claude Haiku 4.5: 10% (1/10) (Claude Code · same prompt repeated 10 times; n = 10; 95% interval 1.8% to 40%). Claude Sonnet 5.5: 100% (10/10) (Claude Code · same prompt repeated 10 times; n = 10; 95% interval 72% to 100%). The 95% intervals do not overlap (Claude Haiku 4.5 2% to 40%; Claude Sonnet 5.5 72% to 100%).
- How do Claude Haiku 4.5 and Claude Sonnet 5.5 compare on same prompt, 10 times: time per call (Code fix)?
- Claude Haiku 4.5: 5.95 s (Claude Code · same prompt repeated 10 times; n = 10; run range 4.9 s to 7.3 s). Claude Sonnet 5.5: 2.67 s (Claude Code · same prompt repeated 10 times; n = 10; run range 2.3 s to 4.3 s). The run ranges (fastest to slowest) do not overlap (Claude Haiku 4.5 4.89 s to 7.33 s; Claude Sonnet 5.5 2.32 s to 4.34 s). A range is not a confidence interval.
The studies behind this page
Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head
130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.
Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks
152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.
Prompt caching and run-to-run consistency in Claude Code and Codex CLI
135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.
Does memory help Claude Code? 8 kinds of agent memory, tested
200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.
Jev vs Claude as a router: accuracy and cost
Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.
Routing overhead: deterministic policy vs LLM routers vs Jev
How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.
Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks
118 attempts on 8 hard tasks: one call vs an agent loop that runs code in a sandbox. Pass rate, time, tokens and cost.
Does thinking pay for Claude Haiku 4.5? Thinking on vs off
Claude Haiku 4.5 with extended thinking on and off: 82 routing decisions and 8 hard tasks. Accuracy with 95% intervals, time and cost.
Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
96 calls. Strict passes, schema vs instructions: Haiku 18/24 vs 0/24, Sonnet 12/12 vs 12/12, GPT-6.1 Sol 12/12 vs 12/12. With intervals.
Jev vs Claude routers on unseen decisions: a blind holdout
Jev 1.13, Claude Haiku 4.5 and Sonnet 5.5 on 56 routing decisions written blind and frozen first: exact rate with 95% intervals, tuned vs unseen.
How much of an AI bill is thinking? Reasoning tokens by model and effort
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.
Where the seconds go: first text, output speed and prompt size for 6 LLMs
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.
Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts
Haiku 4.5 first, then retry or escalate to Sonnet 5.5? A calculation on 64 recorded calls across 8 hard tasks: cost and time per correct answer.
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.