AI coding benchmarks roundup, October 2026: sixteen studies, every number in one place
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
TL;DR
- Sixteen studies, one dataset file, n and a 95% interval for every rate. Each section below links the study, its posts and the comparison and model pages. The leaderboard lists each option's best-supported facts, with no composite score.
- The strongest result is a cost result, not a quality result. Most quality rows tie within their intervals. Price and speed separate the options far more often.
- Hard tasks: Claude Sonnet 5.5, Opus 5.5 and Fable 5.1 each passed 24/24, and GPT-6.1 Sol in the Codex CLI passed 16/16 at medium and at high effort. Haiku 4.5 passed 11/24. That gap is the clearest quality difference in the dataset.
- New: routing overhead. Rules decide in 1.42 µs for $0; Jev's direct API call took a median 137 ms; Sonnet as a router takes 2.60 s through its CLI (a different route). Routing every call adds $1.67 per 1,000 tasks with Jev and $247.30 with Sonnet (a calculation).
- New: provider index. OpenRouter listed the vendor's per-token price for 10 of 10 models we checked; the cost is a 5.5% credit fee. Open-weight models differ by up to 12.6x between providers (snapshot 2026-10-06, third-party-reported).
- Routes matter: the same GPT-6.1 Sol model repaired a scheduler in a median 17.3 s through its API and 61.2 s through the Codex CLI.
- New: effort ladder. Sonnet, Opus and GPT-6.1 Sol at low, medium and high effort: all 11 cells passed 16/16 on the hard set (81% to 100% each). More effort raised output tokens and list-price cost, not passes.
- New: caching and consistency. Claude Code read 97% of later-turn input from the cache; at list price that halved a 5-turn session (a calculation). 7 of 9 repeated-prompt cells passed 10/10; Haiku gave the same wrong number 10 times.
- Caching matters: our agent's recorded work would cost about 3.9 times as much at Sonnet prices without prompt caching (a calculation).
- New: coding agents on hidden tests. Claude Code with Sonnet 5.5 or Opus 5.5 and Codex CLI with GPT-6.1 Sol passed 36 of 36 sessions. Time separated them: medians 23.1 s, 56.9 s and 113.4 s. Codex also read the tester's global
AGENTS.mdin 12 of 12 sessions, a confound we disclose. - New, interim: Opus vs Sonnet on SWE-bench Verified. 3 of 8 declared pairs graded: Opus 5.5 resolved 2, Sonnet 5.5 1; exact McNemar p = 1.0. Opus cost 2.6x as much (a calculation).
- New: agent memory. In 200 Claude Code sessions, memory carried what the repository cannot show: team-knowledge checks went from 6 of 15 with no memory to 15 of 15 with an 11-line file.
- New: System One arena. On 1,048 graded decisions with a checkable answer, Jev 1.13 was right 76.8% of the time; the best open model, Clef 27B, 69.9%, at a median 1.9 s on one Mac Studio against 137 ms for Jev over the network. In Pong, where an answer counts only once it arrives, Jev won 13 of 16 games and Clef 27B 4.
- New: an AI cost calculator built on the same dataset.
Agent Benchmarks, October 2026: 28 studies in one film
The headline of each of our 28 open benchmark studies, with sample sizes and intervals. Calculations labelled; failures counted.
Transcript
- October 2026 roundup · 28 studies. Agent Benchmarks: every headline. Recorded runs, 95% intervals where they exist, every failure counted. Calculations labelled.
- SWE-bench Verified: Agent resolved 25 of 33. The public panel averaged 74.1%; the intervals overlap, so no rank. Agent resolved (one attempt each): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Model cost per resolved instance (calculation, notional): $3.71 (n = 25). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
- SWE-bench, interim (3 of 8 pairs graded): Opus 5.5 as Agent’s brain resolved 2 of 3, Sonnet 5.5 1 of 3. Exact McNemar p = 1.0: no difference yet. Opus cost 2.6× as much, a calculation. Opus 5.5 in Agent: resolved (interim): 67% (2/3) (n = 3, 95% CI 21–94%). Sonnet 5.5 in Agent: same instances: 33% (1/3) (n = 3, 95% CI 6–79%). List-price cost, Opus vs Sonnet (calculation): 2.6× (n = 3). Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
- Blind critics preferred the AI pull request to the merged human one on 9 of 12 tasks at the latest attempt, 6 of 12 at the first. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
- Five short tasks, nine setups: 127 of 130 calls passed, so speed and tokens separate them. Fastest median: Fable 5.1 at 1.9 s. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Fastest median total time (Fable 5.1): 1.9 s (n = 15). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.
- Eight hard tasks: Sonnet 5.5, Opus 5.5, Opus 5.5 high, GPT-6.1 Sol medium, Fable 5.1 and GPT-6.1 Sol high passed every call (24/24 or 16/16). Haiku 4.5 passed 11/24. Chart: Eight hard tasks · strict pass rate · 95% intervals (n = 16–24 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
- Coding agents on hidden-test tasks: 36 of 36 sessions passed, so time separates them. Claude Code with Sonnet 5.5: median 23.1 s; Codex CLI took 4.9× as long, with extra standing instructions. Sessions that passed every hidden check: 100% (36/36) (n = 36, 95% CI 90–100%). Median time, Sonnet 5.5 in Claude Code: 23.1 s (n = 12). Median time, Codex CLI vs Claude Code with Sonnet: 4.9× (n = 12). Caveat: Codex CLI ran with the tester’s global AGENTS.md: --ignore-user-config does not switch that file off. 12 of 12 Codex sessions read the tester’s notes, 10 wrote a WORKLOG.md and 9 reported a commit attempt. Its time, tool calls and lines changed include that work. Claude Code ran with setting sources off, and no Claude session did any of it.
- Effort ladder: 11 model and effort settings passed 176 of 176 calls strictly (16/16 each, 81%–100%). No effort level wins any of 172 rows. Calls that passed strictly, low to high effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference). Effort comparison rows where one level wins: 0 (of 172 rows · 16 pairs). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
- Caching, a list-price calculation: 50% less on Sonnet 5.5, 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every time. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
- Agent memory: Sonnet 5.5 followed team-only rules 40% (6/15) of the time without memory and 100% (15/15) with an 11-line file. Without the late-fee rate it asked 7/9 times. Team knowledge, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). Asked for the missing rate: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
- System One arena: on 1,085 checkable decisions, Jev 1.13 answered 76.8% right; the best open model, Clef 27B, 69.9%. Jev 1.13: 76.8% (805/1048) (n = 1048, 95% CI 74–79%). Clef 27B: 69.9% (733/1048) (n = 1048, 95% CI 67–73%). Checkable decisions: 1,085. Caveat: Local models ran quantized on one Mac; the vendors measured full-precision weights on data-centre GPUs. Latency on other hardware will differ, and Jev latency includes the network.
- Routing: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82 exact, and the intervals overlap. List-price cost is where they differ. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
- Routing overhead: an in-process policy decides in 1.42 µs; Sonnet 5.5 as a router takes 2.60 s through the CLI, about 1.8 million times longer. Chart: Time per routing decision · log scale · median to p95 (n = 82–20000 each). Caveat: The policy is timed in process and the LLM routers through a CLI: this compares the two ways of routing as deployed, not two models on equal footing. A direct API call would skip the CLI time (shown separately).
- Provider prices, reported by OpenRouter: open-weight models vary up to 12.6x across providers. Closed models: one price, plus a 5.5% credit fee. Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Gateway price equals the first-party price: 10 of 10. Endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
- Recorded SWE-bench tokens, repriced: $87.23 on Sonnet 5.5, $143.83 on Opus 5.5, $343.33 with no prompt cache. At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Sonnet 5.5 without the prompt cache: $343.33 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
- What a CLI adds: for a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens. Codex CLI vs OpenAI API, one-line answer: 3.5x slower (n = 30). Input tokens the Codex CLI sends for one line: 19,551 (n = 15). Scheduler repair, median: Claude Code vs Codex CLI: 15.0 s vs 61.2 s (n = 3, models differ too). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
- Three real pull requests: the latest build verified 1 of 3, at $11.06 notional. Guardrail refusals went 26 → 19. Verified deliveries, latest build: 1 of 3 (n = 3). Notional cost, latest build, 3 tasks: $11.06 (n = 3). Guardrail refusals, first vs latest slice: 26 → 19 (n = 3). Caveat: One attempt per cell: these are defect-finding runs, not rates.
- Agent-loop strict passes: Haiku 54% (13/24); Sonnet 100% (16/16). Attempts excluded for outside reads: 2 of 56. Claude Haiku 4.5 strict pass rate, agent loop: 54% (13/24) (n = 24, 95% CI 35–72%). Claude Sonnet 5.5 strict pass rate, agent loop: 100% (16/16) (n = 16, 95% CI 81–100%). Agent-loop attempts left out for reading outside the work folder: 2 of 56 (n = 56). Caveat: The Claude single-call cells ran in another batch on 2026-10-06 (03:23 to 04:02 UTC), with the same CLI version, tasks and validators; provider load can differ by hour.
- Haiku exact routing: thinking off 87% (71/82), on 89% (73/82). Paired McNemar p = 0.754; day and account differ. Claude Haiku 4.5 (thinking off): exact routing decisions: 87% (71/82) (n = 82, 95% CI 78–92%). Claude Haiku 4.5 (thinking on): exact routing decisions: 89% (73/82) (n = 82, 95% CI 80–94%). Paired exact test, thinking off vs on (exact McNemar p): 0.754 (n = 82). Caveat: The thinking-on and Sonnet arms are the recorded routing run of 2026-10-05, reused. The thinking-off arm ran on a different day and on a different Claude subscription account, so this is a confounded comparison. Day, account, CLI version and input-token differences can affect the results. The data cannot isolate the effect of thinking.
- Format misses: instructions 35% (17/48), JSON schema 0% (0/48). The schema still gave wrong values: 13% (6/48). Format-miss rate with instructions only, all models: 35% (17/48) (n = 48, 95% CI 23–50%). Format-miss rate with a JSON schema, all models: 0% (0/48) (n = 48, 95% CI 0–7%). Wrong-values rate with a JSON schema, all models: 13% (6/48) (n = 48, 95% CI 6–25%). Caveat: Small samples: Haiku 24 calls per mode, Sonnet 12 calls per mode and GPT-6.1 Sol 12 calls per mode. A 12/12 result has a 95% interval of 76% to 100%. This is not a minimum detectable difference.
- Later Claude sessions with substantial cache reuse: new folders 0 of 2 (95% interval 0% to 66%); a fixed folder 2 of 2 (95% interval 34% to 100%). This analysis is exploratory. Later sessions with at least 50% of turn-1 input cached, A: new folder each time: 0 of 2 (95% interval 0% to 66%) (n = 2, 95% CI 0–66%). Later sessions with at least 50% of turn-1 input cached, B: fixed folder: 2 of 2 (95% interval 34% to 100%) (n = 2, 95% CI 34–100%). Caveat: The surviving protocol file was created after all counted calls. Its claimed 00:32 UTC declaration is not supported by its file birth time. Amendment 1 and 2 state 00:36 and 00:37 UTC, but separate pre-edit copies do not verify those times. The current summary was regenerated at 07:41 UTC. Treat the analysis as exploratory.
- Cache break-even, a calculation: a new prefix with a one-hour write needs 2 reuses (the 3rd request). Recorded session payback: 2 turns (6 of 6 sessions). Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation): 2 reuses (the 3rd request). Turn at which a recorded Claude Code session’s total input cost with the cache first fell below its cost with no cache (calculation on recorded tokens): 2 turns (6 of 6 sessions) (n = 6). Calculation, not a run. Caveat: The retained Claude protocol file was created at 14:43:55 UTC, after the first counted session at 14:35:39 UTC on 2026-10-06. Its declaration says 14:25 UTC, but file times do not verify that claim. Treat this as a retrospective protocol record.
- Exact unseen routing decisions: Jev 82% (46/56), Haiku 79% (44/56), Sonnet 88% (49/56). The intervals overlap; no rank. Jev 1.13 (TypeSafe): exact on unseen decisions: 82% (46/56) (n = 56, 95% CI 70–90%). Claude Haiku 4.5 · Claude Code: exact on unseen decisions: 79% (44/56) (n = 56, 95% CI 66–87%). Claude Sonnet 5.5 (low) · Claude Code: exact on unseen decisions: 88% (49/56) (n = 56, 95% CI 76–94%). Caveat: The case author and two of the three routers are Claude models, so a same-family label bias is possible. A second labeller (GPT-6.1 Sol) labelled every case blind; the secondary scoring keeps only the keys where its label is inside our acceptable set.
- Reasoning cost at high vs low effort, a calculation from recorded calls: Sonnet 2.2x ($0.0043 at low, $0.0094 at high); Opus 3.6x ($0.0050 at low, $0.0180 at high). Reasoning cost per call, high ÷ low effort, Claude Sonnet 5.5 · Claude Code (calculation, means): 2.2x ($0.0043 at low, $0.0094 at high) (n = 16). Reasoning cost per call, high ÷ low effort, Claude Opus 5.5 · Claude Code (calculation, means): 3.6x ($0.0050 at low, $0.0180 at high) (n = 16). Calculation, not a run. Caveat: The hard-set protocol file was created after its first counted call. Both effort-ladder protocol files were created after their batches ended. The short-set file predates its first call, but its top-up amendment timing is unverified. Batch receipts preserve protocol text, but we cannot verify all rules were written before inference. Treat these as exploratory calculations, not preregistered tests.
- Speed anatomy, a calculation: same-text token count ratio 1.8x. Extra time to first text with the larger prompt: Haiku +0.9 s; Sonnet +1.6 s. The same 4,327-character reply: most tokens ÷ fewest tokens across models (calculation): 1.8x (n = 23). Haiku: extra time to first text at 64k vs 1k (calculation): +0.9 s (n = 6). Sonnet: extra time to first text at 64k vs 1k (calculation): +1.6 s (n = 6). Calculation, not a run. Caveat: 4 calls per model in part A and 3 per size in part B. Medians of so few calls move with one slow call, and the ranges are not confidence intervals. The 95% Wilson intervals on the lookup rates are wide.
- Cost per correct answer, a calculation: Sonnet every time $0.0132; Haiku with one retry then Sonnet $0.0566. No policy ran. Cost per correct answer, Sonnet 5.5 every time (calculation): $0.0132 (n = 24). Cost per correct answer, Haiku with one retry then Sonnet (calculation): $0.0566 (n = 48). Calculation, not a run. Caveat: Advance registration is not verified. The hard-set protocol file birth time is 2026-10-06 04:03:37 UTC; its first counted call started at 03:23:59 UTC. The ladder protocol file birth time is 14:35:11 UTC; its first Claude call started at 14:21:50 UTC. Both files claim advance declaration, but the available file times do not support that claim. Copying could explain the times; we cannot establish it.
- Strict passes on four selected harder tasks: Sol 69% (11/16); Opus 42% (5/12); Sonnet 38% (6/16). Tasks were selected with a Sonnet pilot. GPT-6.1 Sol (medium) (Codex CLI): strict pass rate on the harder tasks: 69% (11/16) (n = 16, 95% CI 44–86%). Opus 5.5 (Claude Code): strict pass rate on the harder tasks: 42% (5/12) (n = 12, 95% CI 19–68%). Sonnet 5.5 (Claude Code): strict pass rate on the harder tasks: 38% (6/16) (n = 16, 95% CI 18–61%). Caveat: Selection effect: the study picked tasks that Sonnet did not pass twice in the pilot. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls (calculation: 25% and 38%; the intervals overlap). Selection can produce this pattern, but the run does not establish its cause.
- Latency budget, a calculation: steps within an assumed 300 ms budget 3 of 14 (median: 3 of 14); within 1,500 ms 4 of 14 (median: 8 of 14), at p95 or the observed maximum. No voice turn ran. Steps that fit a 300 ms budget at the slow end (calculation): 3 of 14 (median: 3 of 14) (n = 14). Steps that fit a 1,500 ms budget at the slow end (calculation): 4 of 14 (median: 8 of 14) (n = 14). Calculation, not a run. Caveat: Timing samples omit 0 failed or untimed matched explorer calls and 0 failed or untimed Claude head-to-head calls. A failure is not a fast successful step. No pass rate is estimated here.
- Routing cost at an assumed million decisions a day, a calculation: Jev $33.70; Sonnet $7,324. No load test ran. Daily cost at 1 million decisions, Jev (calculation): $33.70 (n = 82). Daily cost at 1 million decisions, Sonnet 5.5 router (calculation): $7,324 (n = 82). Calculation, not a run. Caveat: The routing and live Jev protocols predate the first counted calls by file birth time. The overhead protocol predates its microbenchmark output. Claude ran 82 calls per router, below its 120-call cap. Jev ran 246 counted calls, below its 300-call cap. One Haiku pilot and one Jev probe are excluded. All counted calls completed without call errors. The policy microbenchmark used 5,000 warm-up iterations and 64 synthetic contexts; the minimum and individual timing samples are unavailable, so we cannot rebuild its quantiles or full range. No separate validator-control receipt or Sonnet pilot is retained.
- 28 studies, one place. Intervals, sources and every failure kept.
1. SWE-bench Verified: an agent pipeline vs 11 public runs
Question: how does Agent, a full coding pipeline on Claude Sonnet 5.5, compare with 11 public runs on the same instances?
Answer: Agent resolved 25 of 33 (76%, 95% interval 59% to 87%). On the same instances, the public runs resolved between 21 and 28 (panel mean 74.1%). Every interval overlaps, so this sample cannot rank Agent above or below any panel model. Agent spent a notional $2.81 and 49 model calls per attempt.
Study: /benchmarks/swe-bench-verified. Post: Agent vs 11 public models. Compare: Agent vs Claude Sonnet 4.5.
2. Blind review: AI pull requests vs merged human ones
Question: when critics do not know which change is which, do they prefer the AI change or the merged human change?
Answer: on the latest attempt per task, the panel preferred the AI change on 9 of 12 tasks (75%, 95% interval 47% to 91%). On the first scored attempt, the cleaner estimate, it was 6 of 12 (25% to 75%). Most critics are Anthropic models, so same-family preference is possible.
Study: /benchmarks/blind-review-head-to-head. Post: Do blind AI critics prefer AI pull requests?
3. Five-task model head-to-head
Question: on five short tasks, how do Claude Code models compare with GPT-6.1 Sol in the Codex CLI?
Answer: 127 of 130 calls passed (98%, 93% to 99%), so pass rate barely separates them. Speed does more: Claude Fable 5.1 had the fastest median at 1.9 s; Codex CLI medians ran 5.6 to 6.3 s. The Codex CLI sent a median 12,124 input tokens per call against 2,130 for Claude Code.
Study: /benchmarks/model-head-to-head. Post: Claude Haiku vs Sonnet vs Opus vs Fable vs Codex. Compare: Sonnet 5.5 vs GPT-6.1 Sol in Codex CLI.
4. Hard model head-to-head: Claude vs Codex
Question: on eight hard tasks with strict validators, does pass rate separate the Claude Code models and GPT-6.1 Sol in the Codex CLI?
Answer: 139 of 152 calls passed strictly (91%, 86% to 95%). Sonnet 5.5, Opus 5.5, Opus 5.5 (high) and Fable 5.1 each passed 24/24 (86% to 100%), and GPT-6.1 Sol passed 16/16 at medium and at high effort (81% to 100%). Pass rate does not separate these six. Haiku 4.5 passed 11/24 (46%, 28% to 65%); 5 more of its replies were right but in the wrong format, and 8 were wrong. Median time per call: Sonnet 7.7 s, Sol (medium) 13.1 s, Sol (high) 18.1 s; ranges overlap. The lowest cost per strict pass was Sonnet at $0.0143, with Sol (high) at $0.0151 (calculations). 30 earlier Codex attempts were blocked before any model call; they stay in the record, not scored.
Study: /benchmarks/hard-model-head-to-head. Posts: Claude vs Codex on hard tasks, When tasks get hard and Sonnet vs Opus: when is Opus worth the price?. Compare: Sonnet vs GPT-6.1 Sol, Haiku vs GPT-6.1 Sol, Opus vs Fable.
5. Routing: Jev vs Claude as a router
Question: can a small dedicated routing model match Claude on typed routing decisions, and at what cost?
Answer: exact decisions: Jev 1.13 74/82 (90%, 82% to 95%), Claude Haiku 4.5 73/82 (89%, 80% to 94%), Claude Sonnet 5.5 77/82 (94%, 87% to 97%). The intervals overlap. Cost does differ: $0.0337 per 1,000 decisions for Jev against $8.92 for Haiku and $5.00 for Sonnet. The case sets were tuned against Jev answers, which gives Jev a home advantage.
Study: /benchmarks/routing-jev-vs-llm. Posts: Jev vs Claude Haiku and Sonnet as a router and an honest comparison, row by row. Compare: Jev vs Sonnet 5.5, Jev vs Haiku 4.5.
6. Routing overhead: what a router adds (new)
Question: what delay and what cost does each kind of router add before the real work starts?
Answer: the deterministic routing policy decided in a median 1.42 µs (p95 2.33 µs, n = 20,000) for $0. Claude Sonnet 5.5 as a router through the Claude Code CLI took 2.60 s (p95 4.30 s, n = 82), about 1.8 million times longer; 973 ms of that was CLI time. Haiku 4.5 with default thinking took 12.54 s. Jev 1.13, called directly over HTTPS from the same Mac, took a median 137 ms per decision (p95 196 ms, n = 246 calls, network included) and cost $0.0337 per 1,000 decisions (a calculation from its reported input tokens); that is a different route from the CLI, so it is not a model-against-model comparison. As a calculation over 48 recorded tasks (a median 49.5 model calls each), routing every call adds $1.67 per 1,000 tasks with Jev and $247.30 with Sonnet, plus up to 6.76 s of waiting per task with Jev and 129 s with Sonnet. Routing only the 7 System One decisions cuts Sonnet to $34.97 and 18 s.
Study: /benchmarks/routing-overhead. Post: What does a router cost you? Compare: rules vs Sonnet, Jev vs rules.
7. Inference provider index: OpenRouter vs direct (new)
Question: for the same model, how much do providers differ in price, and what does a gateway add?
Answer, third-party-reported (OpenRouter's public API, snapshot 2026-10-06): 265 endpoints from 52 providers for 27 models. For 10 of 10 models with a first-party list price, OpenRouter's per-token price equals the vendor's; the gateway's cost is the 5.5% credit-purchase fee on the Standard plan. All 15 closed models (Claude, GPT, Gemini) with 2 or more providers had one standard-tier price at every provider. Open-weight models differ a lot (blended 3:1, a calculation): DeepSeek V4 Flash 12.6x, DeepSeek V4 Pro 11.2x, gpt-oss-120b 6.9x, Llama 3.3 70B 6.7x. The cheapest endpoints often report fp4 or fp8 precision or a shorter context. Latency was returned for 0 of 265 endpoints, and gateway delay is not measured: no key in the environment.
Study: /benchmarks/inference-provider-index. Posts: OpenRouter vs going direct and the cheapest place to run open models. Compare: OpenRouter vs Anthropic, Groq vs Cerebras, Amazon Bedrock vs Google Vertex.
8. Cost thought experiments
Question: what would the same recorded tokens cost at other list prices?
Answer, a calculation, not a run: Agent's SWE-bench work used 162.9M input tokens (94.0% cache reads) and 1.8M output tokens. At list prices that is $43.61 on Haiku 4.5, $87.23 on Sonnet 5.5, $143.83 on Opus 5.5 and $321.31 on Fable 5.1. Without caching, the Sonnet figure is $343.33.
Study: /benchmarks/cost-thought-experiments. Posts: What if every call ran on Opus? and How to estimate your AI coding bill. Tool: AI cost calculator.
9. CLI vs API: latency and hidden tokens
Question: how much time and context does a coding CLI add?
Answer: for a one-line answer, the Codex CLI took a median 3.5 times as long as the OpenAI API with the same model and effort, and sent about 19,551 input tokens instead of 17. On a scheduler repair with 296 checks, all 9 runs passed: Claude Code with Sonnet 5.5 15.0 s, OpenAI API with GPT-6.1 Sol 17.3 s, Codex CLI with GPT-6.1 Sol 61.2 s.
Study: /benchmarks/cli-model-latency-tokens. Posts: Claude Code vs Codex CLI vs the API and the hidden context tax. Compare: GPT-6.1 Sol: Codex CLI vs API.
10. Coding calibration on three real pull requests
Question: what broke on three real open-source pull requests across four platform builds?
Answer: verified delivery went from 0 of 3 to 1 of 3. h3 regressed to an empty patch; uvicorn passed its suite but stayed unverified. The latest slice cost $11.06 for the three tasks (notional). One attempt per cell finds defects; it does not measure a rate.
Study: /benchmarks/coding-calibration. Post: Why we count every failed attempt.
11. Effort ladder: does more effort buy quality? (new)
Question: on the 8 hard tasks, does a higher effort setting buy a higher pass rate for Sonnet 5.5, Opus 5.5 and GPT-6.1 Sol (Codex CLI)?
Answer: not on this set. All 11 configurations passed 16 of 16 strictly (95% interval 81% to 100% each), with 0 format misses and 0 wrong answers: 176 of 176 calls. The set has a ceiling; it cannot rule out a gap of up to about 19 points. More effort changed time and tokens. Median time per call, low to high: Sonnet 5.8 s to 8.8 s, Opus 7.5 s to 10.1 s, GPT-6.1 Sol 13.6 s to 18.1 s; the ranges overlap. List-price cost per strict pass (a calculation), low to high: Sonnet $0.0122 to $0.0167, Opus $0.0212 to $0.0337, GPT-6.1 Sol $0.0128 to $0.0151. All 16 effort comparisons in the dataset are tie or unclear.
Study: /benchmarks/effort-ladder. Post: Does reasoning effort buy quality?
12. Prompt caching and run-to-run consistency (new)
Question: in a CLI session that reuses a fixed context, how much input comes from the cache and what does that save? When the same prompt runs 10 times, how much do the result and the time vary?
Answer, caching: in Claude Code, turns 2 to 5 read 97% of their input from the cache on average (turn 1: 19%). At list price, a calculation, 15 turns cost Sonnet $0.1350 vs $0.2698 without the cache (50% less) and Opus $0.2551 vs $0.5442 (53% less). Turn 1 costs more with the cache, because a 1-hour write is priced at 2x input. 0 of 4 later sessions reused an earlier session's cache; the cause was not tested. No clear latency effect. The Codex CLI read 99% of later-turn input from its cache, but reports no cache writes, so it is not priced.
Answer, consistency: 7 of 9 model-and-prompt cells passed 10 of 10 (72% to 100%). Haiku 4.5 passed 0/10 on the exact-number prompt, with the same wrong number every time, and 1/10 on the JSON prompt (9 format misses). Those are the only pass-rate rows where a model wins: Sonnet and GPT-6.1 Sol beat Haiku. Some time rows also separate on run ranges; see the consistency post. On the code fix, all passed 10/10 with 6, 3 and 6 distinct code bodies (Haiku, Sonnet, GPT-6.1 Sol).
Study: /benchmarks/caching-consistency. Posts: How much does prompt caching save? and Same prompt, ten answers. Compare: Haiku vs Sonnet, Claude Code vs Codex CLI.
13. Agent memory: does CLAUDE.md help? (new)
Question: does project memory make Claude Code do better work, and which kind of memory?
Answer: 200 graded sessions (120 with Sonnet 5.5, 80 with Haiku 4.5) and 8 kinds of memory, from none to a 211-line handbook and a Stop hook. With no memory, Sonnet already followed 95% of the rules that the code shows and 100% of the rules that the folders hint at, but only 6 of 15 team-knowledge checks. With an 11-line curated file it passed 15 of 15. When a memory file held the late-fee rate, Sonnet used it 15 of 15 times; without it, Sonnet stopped and asked 7 of 9 times, and Haiku 4.5 invented a rate in 6 of 6 sessions without a flag. Messy memory misled Haiku: with raw notes it ran a stale test command 10 of 10 times, and after one "dreaming" pass 1 of 10. A Stop hook enforced every code rule but cannot carry a fact, and it used 1.6x the input tokens of no memory. Full-pass intervals overlap for most pairs (n = 15 per condition).
Study: /benchmarks/agent-memory. Post: Does CLAUDE.md help? We tested 8 kinds of agent memory.
14. Coding agents on hidden tests: Claude Code vs Codex CLI (new)
Question: on small real repository tasks graded by hidden tests, how do coding-agent CLIs compare with their normal file and shell tools?
Answer: 36 of 36 sessions passed: 12 of 12 each for Claude Code with Sonnet 5.5, Claude Code with Opus 5.5 and Codex CLI with GPT-6.1 Sol at medium effort (95% interval 76% to 100% each). The task set has a ceiling, so it cannot rank them on quality. Median time per session: Sonnet 23.1 s, Opus 56.9 s, Codex 113.4 s. The run ranges of Sonnet (18.7 s to 44.5 s) and Codex (78.5 s to 221.9 s) do not overlap, so the comparison pages mark Claude Code with Sonnet faster on that row (4.9x by medians). A confound: Codex read the tester's global AGENTS.md in 12 of 12 sessions despite --ignore-user-config; 10 wrote a work log nobody asked for and 9 reported a commit attempt, so its time, tool calls and diffs include that work. The agents wrote about the same amount of task code; the median diff was 45, 77.5 and 94 lines, and most of the difference is tests. List-price cost per pass (a calculation): Sonnet $0.085, Codex $0.098, Opus $0.22. Gemini CLI was not run: it asked for a browser login.
Study: /benchmarks/coding-agents-head-to-head. Post: Claude Code vs Codex CLI on hidden tests. Compare: Claude Code vs Codex CLI, Sonnet 5.5 vs GPT-6.1 Sol, Opus 5.5 vs GPT-6.1 Sol.
15. Opus vs Sonnet as the agent's brain on SWE-bench Verified (new, interim)
Question: does Agent resolve more SWE-bench Verified instances with Claude Opus 5.5 as its brain than with Claude Sonnet 5.5, and at what cost and time?
Answer, interim: 3 of 8 declared pairs are graded; a usage gate stopped the run, and the other 5 resume after the reset on 2026-10-09. With Opus, Agent resolved 2 of 3 (95% interval 21% to 94%); with Sonnet, 1 of 3 of the same instances (6% to 79%). The one pair that differs, django__django-10554, went to Opus, and the exact McNemar p is 1.0: no evidence of a difference. On the same instances, Opus cost 2.6x as much ($22.76 vs $8.64, a list-price calculation) and used 1.9x the worker minutes. The two arms ran on different platform builds, and the instances were chosen to hold Sonnet misses and resolves in equal numbers.
Study: /benchmarks/swe-bench-opus-vs-sonnet. Post: Opus 5.5 vs Sonnet 5.5 on real GitHub issues: an early look. Compare: Sonnet vs Opus.
16. System One arena: Jev vs Clef and five open decision models (new)
Question: how good are typed-decision models at decisions with a checkable answer, and which one wins when they play each other?
Answer: Jev 1.13 answered 76.8% of 1,048 graded decisions correctly (805/1,048), ahead of every other model (exact McNemar p < 0.001 for each pair). The best open model, Clef 27B, reached 69.9%, at a median 1.9 s per decision on one Mac Studio against 137 ms for Jev 1.13 over the network. Clef-Flash 9B (65.2%) and Kev 4B (64.0%) were statistically tied. Laya (27.4%) and Julia-1 (25.9%) were tied, close to the 22.4% that picking an option at random scores on the same items (a calculation). In 1,512 round-robin games, Jev 1.13 rated highest, but only Jev 1.13 (30-11 of 42) and Clef-Flash 9B (29-11 of 42) beat the random player clearly. In Pong, where an answer counts only once it arrives, Jev 1.13 won 13 of 16 games and Kev 4B 12, while Clef 27B, the most accurate open model on the exam, won 4 at 1.8 s per decision. The items are synthetic, the local models ran quantized on one Mac, and Jev 1.13 latency includes the network.
Study: /benchmarks/system-one-arena. Posts: Jev vs Clef and five open decision models and Decision models play Connect Four, Nim and Pong.
What the sixteen studies say together
- Quality ties are the norm. Across all comparison pages, 147 of 162 pass-rate, accuracy and behaviour-rate rows are ties: the 95% intervals overlap (our count). The other 15 all separate Claude Haiku 4.5 from a larger model: 7 on the hard set, 4 on repeated prompts and 4 in the memory study. In all 15, Haiku had the worse result. Effort did not separate any model from itself, and neither the coding-agent ceiling nor the interim SWE-bench pairs separate Sonnet from Opus.
- Cost gaps are large and real as arithmetic. Sonnet 5.5 had the lowest cost per pass on both model head-to-heads and in the coding-agent sessions. Opus cost 2.6x as much as Sonnet on the coding-agent tasks and on the interim SWE-bench pairs. A dedicated router cost about 148 times less per decision than Sonnet, and rules cost nothing. Caching moved the agent bill by about 3.9 times, and halved a 5-turn session. For open-weight models, the provider moved the price by up to 12.6 times; for closed models, not at all.
- The route is part of the result. CLI context and start-up can matter more than the model. Name both when you quote a number. A CLI's standing instructions count too: Codex read the tester's
AGENTS.mdalthough user config was off. - Accuracy is one axis, and time counts in games. On the arena exam, Clef 27B was the most accurate open model (69.9%) at a median 1.9 s per decision. In Pong, where each answer moves the paddle only after its measured latency, it won 4 of 16 games at 1.8 s per decision; Jev 1.13 won 13 of 16.
- Small samples need humble claims. Many cells hold 3 to 24 runs. The SWE-bench Opus probe holds 3 graded pairs so far.
- Consistent is not correct. A model can repeat the same wrong answer 10 times. Validate against the right answer.
Not in this roundup: an Opus-vs-Sonnet or Claude-vs-GPT-6.1 Sol quality difference, or a quality gain from higher effort (our tasks did not find one), the 5 SWE-bench Opus pairs that resume after 2026-10-09, a Codex rerun without the tester's notes, Gemini CLI (it asked for a browser login), a single composite score (why), gateway delay (no key in the environment), and provider speed or quality (no data).
Browse by model or pair
- Models and tools: Sonnet 5.5, Opus 5.5, Haiku 4.5, Fable 5.1, Claude Code, Codex CLI, Jev 1.13, Agent. All: /models.
- Routers: Jev 1.13 and the deterministic routing policy.
- Providers: every price is on the provider index study; pairs such as Anthropic vs Amazon Bedrock and DeepInfra vs Parasail line them up.
- Pairs: Sonnet vs Opus, Opus vs GPT-6.1 Sol, Haiku vs Sonnet, Claude Code vs Codex CLI, OpenRouter vs OpenAI. All: /compare.
What to read next
- Claude Code vs Codex CLI on hidden tests: when every agent passes, what differs?
- Opus 5.5 vs Sonnet 5.5 on real GitHub issues: an early look
- Does CLAUDE.md help? We tested 8 kinds of agent memory
- Does reasoning effort buy quality?
- How much does prompt caching save?
- Same prompt, ten answers: how consistent are Claude and Codex?
- An AI model leaderboard without a composite score
- Claude vs Codex on hard tasks
- What does a router cost you?
- OpenRouter vs going direct
- When tasks get hard: Haiku vs Sonnet vs Opus vs Fable
- Claude Sonnet vs Opus: when is Opus worth the price?
- How to estimate your AI coding bill
Run the benchmark on your own work
Agent keeps the same receipts for your tasks: the model, the route, the tokens, the time, the cost and the validation result. Try Agent and see your own numbers.