• Roundup
  • Head to head
  • SWE-bench
  • Coding Agents

AI coding benchmarks roundup, October 2026: sixteen studies, every number in one place

Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.

TL;DR

  • Sixteen studies, one dataset file, n and a 95% interval for every rate. Each section below links the study, its posts and the comparison and model pages. The leaderboard lists each option's best-supported facts, with no composite score.
  • The strongest result is a cost result, not a quality result. Most quality rows tie within their intervals. Price and speed separate the options far more often.
  • Hard tasks: Claude Sonnet 5.5, Opus 5.5 and Fable 5.1 each passed 24/24, and GPT-6.1 Sol in the Codex CLI passed 16/16 at medium and at high effort. Haiku 4.5 passed 11/24. That gap is the clearest quality difference in the dataset.
  • New: routing overhead. Rules decide in 1.42 µs for $0; Jev's direct API call took a median 137 ms; Sonnet as a router takes 2.60 s through its CLI (a different route). Routing every call adds $1.67 per 1,000 tasks with Jev and $247.30 with Sonnet (a calculation).
  • New: provider index. OpenRouter listed the vendor's per-token price for 10 of 10 models we checked; the cost is a 5.5% credit fee. Open-weight models differ by up to 12.6x between providers (snapshot 2026-10-06, third-party-reported).
  • Routes matter: the same GPT-6.1 Sol model repaired a scheduler in a median 17.3 s through its API and 61.2 s through the Codex CLI.
  • New: effort ladder. Sonnet, Opus and GPT-6.1 Sol at low, medium and high effort: all 11 cells passed 16/16 on the hard set (81% to 100% each). More effort raised output tokens and list-price cost, not passes.
  • New: caching and consistency. Claude Code read 97% of later-turn input from the cache; at list price that halved a 5-turn session (a calculation). 7 of 9 repeated-prompt cells passed 10/10; Haiku gave the same wrong number 10 times.
  • Caching matters: our agent's recorded work would cost about 3.9 times as much at Sonnet prices without prompt caching (a calculation).
  • New: coding agents on hidden tests. Claude Code with Sonnet 5.5 or Opus 5.5 and Codex CLI with GPT-6.1 Sol passed 36 of 36 sessions. Time separated them: medians 23.1 s, 56.9 s and 113.4 s. Codex also read the tester's global AGENTS.md in 12 of 12 sessions, a confound we disclose.
  • New, interim: Opus vs Sonnet on SWE-bench Verified. 3 of 8 declared pairs graded: Opus 5.5 resolved 2, Sonnet 5.5 1; exact McNemar p = 1.0. Opus cost 2.6x as much (a calculation).
  • New: agent memory. In 200 Claude Code sessions, memory carried what the repository cannot show: team-knowledge checks went from 6 of 15 with no memory to 15 of 15 with an 11-line file.
  • New: System One arena. On 1,048 graded decisions with a checkable answer, Jev 1.13 was right 76.8% of the time; the best open model, Clef 27B, 69.9%, at a median 1.9 s on one Mac Studio against 137 ms for Jev over the network. In Pong, where an answer counts only once it arrives, Jev won 13 of 16 games and Clef 27B 4.
  • New: an AI cost calculator built on the same dataset.
Live story · 211 sAgent Benchmarks, October 2026: 28 studies in one film

Agent Benchmarks, October 2026: 28 studies in one film

The headline of each of our 28 open benchmark studies, with sample sizes and intervals. Calculations labelled; failures counted.

Transcript
  1. October 2026 roundup · 28 studies. Agent Benchmarks: every headline. Recorded runs, 95% intervals where they exist, every failure counted. Calculations labelled.
  2. SWE-bench Verified: Agent resolved 25 of 33. The public panel averaged 74.1%; the intervals overlap, so no rank. Agent resolved (one attempt each): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Model cost per resolved instance (calculation, notional): $3.71 (n = 25). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.
  3. SWE-bench, interim (3 of 8 pairs graded): Opus 5.5 as Agent’s brain resolved 2 of 3, Sonnet 5.5 1 of 3. Exact McNemar p = 1.0: no difference yet. Opus cost 2.6× as much, a calculation. Opus 5.5 in Agent: resolved (interim): 67% (2/3) (n = 3, 95% CI 21–94%). Sonnet 5.5 in Agent: same instances: 33% (1/3) (n = 3, 95% CI 6–79%). List-price cost, Opus vs Sonnet (calculation): 2.6× (n = 3). Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.
  4. Blind critics preferred the AI pull request to the merged human one on 9 of 12 tasks at the latest attempt, 6 of 12 at the first. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.
  5. Five short tasks, nine setups: 127 of 130 calls passed, so speed and tokens separate them. Fastest median: Fable 5.1 at 1.9 s. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Fastest median total time (Fable 5.1): 1.9 s (n = 15). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.
  6. Eight hard tasks: Sonnet 5.5, Opus 5.5, Opus 5.5 high, GPT-6.1 Sol medium, Fable 5.1 and GPT-6.1 Sol high passed every call (24/24 or 16/16). Haiku 4.5 passed 11/24. Chart: Eight hard tasks · strict pass rate · 95% intervals (n = 16–24 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.
  7. Coding agents on hidden-test tasks: 36 of 36 sessions passed, so time separates them. Claude Code with Sonnet 5.5: median 23.1 s; Codex CLI took 4.9× as long, with extra standing instructions. Sessions that passed every hidden check: 100% (36/36) (n = 36, 95% CI 90–100%). Median time, Sonnet 5.5 in Claude Code: 23.1 s (n = 12). Median time, Codex CLI vs Claude Code with Sonnet: 4.9× (n = 12). Caveat: Codex CLI ran with the tester’s global AGENTS.md: --ignore-user-config does not switch that file off. 12 of 12 Codex sessions read the tester’s notes, 10 wrote a WORKLOG.md and 9 reported a commit attempt. Its time, tool calls and lines changed include that work. Claude Code ran with setting sources off, and no Claude session did any of it.
  8. Effort ladder: 11 model and effort settings passed 176 of 176 calls strictly (16/16 each, 81%–100%). No effort level wins any of 172 rows. Calls that passed strictly, low to high effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference). Effort comparison rows where one level wins: 0 (of 172 rows · 16 pairs). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.
  9. Caching, a list-price calculation: 50% less on Sonnet 5.5, 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every time. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.
  10. Agent memory: Sonnet 5.5 followed team-only rules 40% (6/15) of the time without memory and 100% (15/15) with an 11-line file. Without the late-fee rate it asked 7/9 times. Team knowledge, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). Asked for the missing rate: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.
  11. System One arena: on 1,085 checkable decisions, Jev 1.13 answered 76.8% right; the best open model, Clef 27B, 69.9%. Jev 1.13: 76.8% (805/1048) (n = 1048, 95% CI 74–79%). Clef 27B: 69.9% (733/1048) (n = 1048, 95% CI 67–73%). Checkable decisions: 1,085. Caveat: Local models ran quantized on one Mac; the vendors measured full-precision weights on data-centre GPUs. Latency on other hardware will differ, and Jev latency includes the network.
  12. Routing: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82 exact, and the intervals overlap. List-price cost is where they differ. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.
  13. Routing overhead: an in-process policy decides in 1.42 µs; Sonnet 5.5 as a router takes 2.60 s through the CLI, about 1.8 million times longer. Chart: Time per routing decision · log scale · median to p95 (n = 82–20000 each). Caveat: The policy is timed in process and the LLM routers through a CLI: this compares the two ways of routing as deployed, not two models on equal footing. A direct API call would skip the CLI time (shown separately).
  14. Provider prices, reported by OpenRouter: open-weight models vary up to 12.6x across providers. Closed models: one price, plus a 5.5% credit fee. Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Gateway price equals the first-party price: 10 of 10. Endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.
  15. Recorded SWE-bench tokens, repriced: $87.23 on Sonnet 5.5, $143.83 on Opus 5.5, $343.33 with no prompt cache. At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Sonnet 5.5 without the prompt cache: $343.33 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.
  16. What a CLI adds: for a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens. Codex CLI vs OpenAI API, one-line answer: 3.5x slower (n = 30). Input tokens the Codex CLI sends for one line: 19,551 (n = 15). Scheduler repair, median: Claude Code vs Codex CLI: 15.0 s vs 61.2 s (n = 3, models differ too). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.
  17. Three real pull requests: the latest build verified 1 of 3, at $11.06 notional. Guardrail refusals went 26 → 19. Verified deliveries, latest build: 1 of 3 (n = 3). Notional cost, latest build, 3 tasks: $11.06 (n = 3). Guardrail refusals, first vs latest slice: 26 → 19 (n = 3). Caveat: One attempt per cell: these are defect-finding runs, not rates.
  18. Agent-loop strict passes: Haiku 54% (13/24); Sonnet 100% (16/16). Attempts excluded for outside reads: 2 of 56. Claude Haiku 4.5 strict pass rate, agent loop: 54% (13/24) (n = 24, 95% CI 35–72%). Claude Sonnet 5.5 strict pass rate, agent loop: 100% (16/16) (n = 16, 95% CI 81–100%). Agent-loop attempts left out for reading outside the work folder: 2 of 56 (n = 56). Caveat: The Claude single-call cells ran in another batch on 2026-10-06 (03:23 to 04:02 UTC), with the same CLI version, tasks and validators; provider load can differ by hour.
  19. Haiku exact routing: thinking off 87% (71/82), on 89% (73/82). Paired McNemar p = 0.754; day and account differ. Claude Haiku 4.5 (thinking off): exact routing decisions: 87% (71/82) (n = 82, 95% CI 78–92%). Claude Haiku 4.5 (thinking on): exact routing decisions: 89% (73/82) (n = 82, 95% CI 80–94%). Paired exact test, thinking off vs on (exact McNemar p): 0.754 (n = 82). Caveat: The thinking-on and Sonnet arms are the recorded routing run of 2026-10-05, reused. The thinking-off arm ran on a different day and on a different Claude subscription account, so this is a confounded comparison. Day, account, CLI version and input-token differences can affect the results. The data cannot isolate the effect of thinking.
  20. Format misses: instructions 35% (17/48), JSON schema 0% (0/48). The schema still gave wrong values: 13% (6/48). Format-miss rate with instructions only, all models: 35% (17/48) (n = 48, 95% CI 23–50%). Format-miss rate with a JSON schema, all models: 0% (0/48) (n = 48, 95% CI 0–7%). Wrong-values rate with a JSON schema, all models: 13% (6/48) (n = 48, 95% CI 6–25%). Caveat: Small samples: Haiku 24 calls per mode, Sonnet 12 calls per mode and GPT-6.1 Sol 12 calls per mode. A 12/12 result has a 95% interval of 76% to 100%. This is not a minimum detectable difference.
  21. Later Claude sessions with substantial cache reuse: new folders 0 of 2 (95% interval 0% to 66%); a fixed folder 2 of 2 (95% interval 34% to 100%). This analysis is exploratory. Later sessions with at least 50% of turn-1 input cached, A: new folder each time: 0 of 2 (95% interval 0% to 66%) (n = 2, 95% CI 0–66%). Later sessions with at least 50% of turn-1 input cached, B: fixed folder: 2 of 2 (95% interval 34% to 100%) (n = 2, 95% CI 34–100%). Caveat: The surviving protocol file was created after all counted calls. Its claimed 00:32 UTC declaration is not supported by its file birth time. Amendment 1 and 2 state 00:36 and 00:37 UTC, but separate pre-edit copies do not verify those times. The current summary was regenerated at 07:41 UTC. Treat the analysis as exploratory.
  22. Cache break-even, a calculation: a new prefix with a one-hour write needs 2 reuses (the 3rd request). Recorded session payback: 2 turns (6 of 6 sessions). Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation): 2 reuses (the 3rd request). Turn at which a recorded Claude Code session’s total input cost with the cache first fell below its cost with no cache (calculation on recorded tokens): 2 turns (6 of 6 sessions) (n = 6). Calculation, not a run. Caveat: The retained Claude protocol file was created at 14:43:55 UTC, after the first counted session at 14:35:39 UTC on 2026-10-06. Its declaration says 14:25 UTC, but file times do not verify that claim. Treat this as a retrospective protocol record.
  23. Exact unseen routing decisions: Jev 82% (46/56), Haiku 79% (44/56), Sonnet 88% (49/56). The intervals overlap; no rank. Jev 1.13 (TypeSafe): exact on unseen decisions: 82% (46/56) (n = 56, 95% CI 70–90%). Claude Haiku 4.5 · Claude Code: exact on unseen decisions: 79% (44/56) (n = 56, 95% CI 66–87%). Claude Sonnet 5.5 (low) · Claude Code: exact on unseen decisions: 88% (49/56) (n = 56, 95% CI 76–94%). Caveat: The case author and two of the three routers are Claude models, so a same-family label bias is possible. A second labeller (GPT-6.1 Sol) labelled every case blind; the secondary scoring keeps only the keys where its label is inside our acceptable set.
  24. Reasoning cost at high vs low effort, a calculation from recorded calls: Sonnet 2.2x ($0.0043 at low, $0.0094 at high); Opus 3.6x ($0.0050 at low, $0.0180 at high). Reasoning cost per call, high ÷ low effort, Claude Sonnet 5.5 · Claude Code (calculation, means): 2.2x ($0.0043 at low, $0.0094 at high) (n = 16). Reasoning cost per call, high ÷ low effort, Claude Opus 5.5 · Claude Code (calculation, means): 3.6x ($0.0050 at low, $0.0180 at high) (n = 16). Calculation, not a run. Caveat: The hard-set protocol file was created after its first counted call. Both effort-ladder protocol files were created after their batches ended. The short-set file predates its first call, but its top-up amendment timing is unverified. Batch receipts preserve protocol text, but we cannot verify all rules were written before inference. Treat these as exploratory calculations, not preregistered tests.
  25. Speed anatomy, a calculation: same-text token count ratio 1.8x. Extra time to first text with the larger prompt: Haiku +0.9 s; Sonnet +1.6 s. The same 4,327-character reply: most tokens ÷ fewest tokens across models (calculation): 1.8x (n = 23). Haiku: extra time to first text at 64k vs 1k (calculation): +0.9 s (n = 6). Sonnet: extra time to first text at 64k vs 1k (calculation): +1.6 s (n = 6). Calculation, not a run. Caveat: 4 calls per model in part A and 3 per size in part B. Medians of so few calls move with one slow call, and the ranges are not confidence intervals. The 95% Wilson intervals on the lookup rates are wide.
  26. Cost per correct answer, a calculation: Sonnet every time $0.0132; Haiku with one retry then Sonnet $0.0566. No policy ran. Cost per correct answer, Sonnet 5.5 every time (calculation): $0.0132 (n = 24). Cost per correct answer, Haiku with one retry then Sonnet (calculation): $0.0566 (n = 48). Calculation, not a run. Caveat: Advance registration is not verified. The hard-set protocol file birth time is 2026-10-06 04:03:37 UTC; its first counted call started at 03:23:59 UTC. The ladder protocol file birth time is 14:35:11 UTC; its first Claude call started at 14:21:50 UTC. Both files claim advance declaration, but the available file times do not support that claim. Copying could explain the times; we cannot establish it.
  27. Strict passes on four selected harder tasks: Sol 69% (11/16); Opus 42% (5/12); Sonnet 38% (6/16). Tasks were selected with a Sonnet pilot. GPT-6.1 Sol (medium) (Codex CLI): strict pass rate on the harder tasks: 69% (11/16) (n = 16, 95% CI 44–86%). Opus 5.5 (Claude Code): strict pass rate on the harder tasks: 42% (5/12) (n = 12, 95% CI 19–68%). Sonnet 5.5 (Claude Code): strict pass rate on the harder tasks: 38% (6/16) (n = 16, 95% CI 18–61%). Caveat: Selection effect: the study picked tasks that Sonnet did not pass twice in the pilot. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls (calculation: 25% and 38%; the intervals overlap). Selection can produce this pattern, but the run does not establish its cause.
  28. Latency budget, a calculation: steps within an assumed 300 ms budget 3 of 14 (median: 3 of 14); within 1,500 ms 4 of 14 (median: 8 of 14), at p95 or the observed maximum. No voice turn ran. Steps that fit a 300 ms budget at the slow end (calculation): 3 of 14 (median: 3 of 14) (n = 14). Steps that fit a 1,500 ms budget at the slow end (calculation): 4 of 14 (median: 8 of 14) (n = 14). Calculation, not a run. Caveat: Timing samples omit 0 failed or untimed matched explorer calls and 0 failed or untimed Claude head-to-head calls. A failure is not a fast successful step. No pass rate is estimated here.
  29. Routing cost at an assumed million decisions a day, a calculation: Jev $33.70; Sonnet $7,324. No load test ran. Daily cost at 1 million decisions, Jev (calculation): $33.70 (n = 82). Daily cost at 1 million decisions, Sonnet 5.5 router (calculation): $7,324 (n = 82). Calculation, not a run. Caveat: The routing and live Jev protocols predate the first counted calls by file birth time. The overhead protocol predates its microbenchmark output. Claude ran 82 calls per router, below its 120-call cap. Jev ran 246 counted calls, below its 300-call cap. One Haiku pilot and one Jev probe are excluded. All counted calls completed without call errors. The policy microbenchmark used 5,000 warm-up iterations and 64 synthetic contexts; the minimum and individual timing samples are unavailable, so we cannot rebuild its quantiles or full range. No separate validator-control receipt or Sonnet pilot is retained.
  30. 28 studies, one place. Intervals, sources and every failure kept.

1. SWE-bench Verified: an agent pipeline vs 11 public runs

GPT 5.2 (high)
Gemini 3 Flash (high)
GLM 5 (high)
Agent (Sonnet 5.5, full pipeline)
Claude 4.5 Sonnet (high)
Claude 4.5 Haiku (high)
Claude 4.5 Opus (high)
DeepSeek V3.2 (high)
MiniMax M2.5 (high)
Claude 4.6 Opus
Kimi K2.5 (high)
GPT 5 mini

Every interval overlaps every other: this chart does not order these rows.

12 rows. Highest GPT 5.2 (high) 85% (95% interval 69%–93%, n 33). Lowest GPT 5 mini 64% (95% interval 47%–78%, n 33). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 33 per row

Agent vs 11 public mini-SWE-agent v2 runs, one attempt each

Dots show the rate; whiskers show the 95% Wilson interval for n = 33. The panel compares models under one harness; Agent is a full system on one model.

Sources: Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances), SWE-bench Verified leaderboard, mini-SWE-agent v2 runs

Question: how does Agent, a full coding pipeline on Claude Sonnet 5.5, compare with 11 public runs on the same instances?

Answer: Agent resolved 25 of 33 (76%, 95% interval 59% to 87%). On the same instances, the public runs resolved between 21 and 28 (panel mean 74.1%). Every interval overlaps, so this sample cannot rank Agent above or below any panel model. Agent spent a notional $2.81 and 49 model calls per attempt.

Study: /benchmarks/swe-bench-verified. Post: Agent vs 11 public models. Compare: Agent vs Claude Sonnet 4.5.

2. Blind review: AI pull requests vs merged human ones

Question: when critics do not know which change is which, do they prefer the AI change or the merged human change?

Answer: on the latest attempt per task, the panel preferred the AI change on 9 of 12 tasks (75%, 95% interval 47% to 91%). On the first scored attempt, the cleaner estimate, it was 6 of 12 (25% to 75%). Most critics are Anthropic models, so same-family preference is possible.

Study: /benchmarks/blind-review-head-to-head. Post: Do blind AI critics prefer AI pull requests?

3. Five-task model head-to-head

Question: on five short tasks, how do Claude Code models compare with GPT-6.1 Sol in the Codex CLI?

Answer: 127 of 130 calls passed (98%, 93% to 99%), so pass rate barely separates them. Speed does more: Claude Fable 5.1 had the fastest median at 1.9 s; Codex CLI medians ran 5.6 to 6.3 s. The Codex CLI sent a median 12,124 input tokens per call against 2,130 for Claude Code.

Study: /benchmarks/model-head-to-head. Post: Claude Haiku vs Sonnet vs Opus vs Fable vs Codex. Compare: Sonnet 5.5 vs GPT-6.1 Sol in Codex CLI.

4. Hard model head-to-head: Claude vs Codex

  • Strict pass
  • Lenient (format misses counted)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
Claude Opus 5.5 (high) · Claude Code
GPT-6.1 Sol (medium) · Codex CLI
Claude Fable 5.1 · Claude Code
GPT-6.1 Sol (high) · Codex CLI
Claude Haiku 4.5 · Claude Code

7 rows, 2 series: Strict pass, Lenient (format misses counted). Strict pass: highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 46% (95% interval 28%–65%, n 24). Not all intervals overlap. Lenient (format misses counted): highest Claude Sonnet 5.5 · Claude Code 100% (95% interval 86%–100%, n 24). Lowest Claude Haiku 4.5 · Claude Code 67% (95% interval 47%–82%, n 24). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln 16–24 per row6 of 7 (Strict pass) at 100%: this task set cannot separate them.

Strict: the reply passes as given. Lenient: a correct answer in the wrong format also counts

Whiskers are 95% Wilson intervals. The strict pass is the result; the lenient reading is shown so a reader can see how much is format and how much is a wrong answer. A format miss never counts as a pass.

Source: Provider head-to-head, hard set: eight hard tasks with strict validators

Question: on eight hard tasks with strict validators, does pass rate separate the Claude Code models and GPT-6.1 Sol in the Codex CLI?

Answer: 139 of 152 calls passed strictly (91%, 86% to 95%). Sonnet 5.5, Opus 5.5, Opus 5.5 (high) and Fable 5.1 each passed 24/24 (86% to 100%), and GPT-6.1 Sol passed 16/16 at medium and at high effort (81% to 100%). Pass rate does not separate these six. Haiku 4.5 passed 11/24 (46%, 28% to 65%); 5 more of its replies were right but in the wrong format, and 8 were wrong. Median time per call: Sonnet 7.7 s, Sol (medium) 13.1 s, Sol (high) 18.1 s; ranges overlap. The lowest cost per strict pass was Sonnet at $0.0143, with Sol (high) at $0.0151 (calculations). 30 earlier Codex attempts were blocked before any model call; they stay in the record, not scored.

Study: /benchmarks/hard-model-head-to-head. Posts: Claude vs Codex on hard tasks, When tasks get hard and Sonnet vs Opus: when is Opus worth the price?. Compare: Sonnet vs GPT-6.1 Sol, Haiku vs GPT-6.1 Sol, Opus vs Fable.

5. Routing: Jev vs Claude as a router

Question: can a small dedicated routing model match Claude on typed routing decisions, and at what cost?

Answer: exact decisions: Jev 1.13 74/82 (90%, 82% to 95%), Claude Haiku 4.5 73/82 (89%, 80% to 94%), Claude Sonnet 5.5 77/82 (94%, 87% to 97%). The intervals overlap. Cost does differ: $0.0337 per 1,000 decisions for Jev against $8.92 for Haiku and $5.00 for Sonnet. The case sets were tuned against Jev answers, which gives Jev a home advantage.

Study: /benchmarks/routing-jev-vs-llm. Posts: Jev vs Claude Haiku and Sonnet as a router and an honest comparison, row by row. Compare: Jev vs Sonnet 5.5, Jev vs Haiku 4.5.

6. Routing overhead: what a router adds (new)

Calculation
  • Every model call routed (49.5 per task)
  • Only System One decisions (7 per task) (square)
In chart order.
Deterministic routing policy (Agent, in process)
Jev 1.13 (TypeSafe)
Claude Sonnet 5.5 (effort low, via Claude Code)
Claude Haiku 4.5 (thinking on, via Claude Code)

Gap labels, Only System One decisions (7 per task) vs Every model call routed (49.5 per task): Only System One decisions (7 per task) is x% higher (+) or lower (−) than Every model call routed (49.5 per task), calculated from the two values shown (the change counted from Every model call routed (49.5 per task)’s value).

List-price calculation, not a run. 4 rows, 2 series: Every model call routed (49.5 per task), Only System One decisions (7 per task). Every model call routed (49.5 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $442. Lowest Deterministic routing policy (Agent, in process) $0. Only System One decisions (7 per task): highest Claude Haiku 4.5 (thinking on, via Claude Code) $62.47. Lowest Deterministic routing policy (Agent, in process) $0.

Notes

Decisions per task from recorded runs × cost per decision

A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).

Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing overhead per 1,000 tasks (calculation), Routing runs: Jev router vs LLM routing, Anthropic list prices (Claude models), Jev 1.13 list price, Jev live run: 246 timed calls on the 82 routing decisions

Question: what delay and what cost does each kind of router add before the real work starts?

Answer: the deterministic routing policy decided in a median 1.42 µs (p95 2.33 µs, n = 20,000) for $0. Claude Sonnet 5.5 as a router through the Claude Code CLI took 2.60 s (p95 4.30 s, n = 82), about 1.8 million times longer; 973 ms of that was CLI time. Haiku 4.5 with default thinking took 12.54 s. Jev 1.13, called directly over HTTPS from the same Mac, took a median 137 ms per decision (p95 196 ms, n = 246 calls, network included) and cost $0.0337 per 1,000 decisions (a calculation from its reported input tokens); that is a different route from the CLI, so it is not a model-against-model comparison. As a calculation over 48 recorded tasks (a median 49.5 model calls each), routing every call adds $1.67 per 1,000 tasks with Jev and $247.30 with Sonnet, plus up to 6.76 s of waiting per task with Jev and 129 s with Sonnet. Routing only the 7 System One decisions cuts Sonnet to $34.97 and 18 s.

Study: /benchmarks/routing-overhead. Post: What does a router cost you? Compare: rules vs Sonnet, Jev vs rules.

7. Inference provider index: OpenRouter vs direct (new)

Calculation
ModelSpread
  1. DeepSeek V4 Flash 0423n = 15 providers
  2. DeepSeek V4 Pro 0423n = 15 providers
  3. gpt-oss-120bn = 20 providers
  4. Llama 3.3 70B Instructn = 10 providers
  5. GLM 5.3n = 32 providers
  6. Kimi K3n = 19 providers
  7. Llama 4 Maverickn = 3 providers
  8. 15 more models: the same price at every providerClaude Haiku 4.5, Claude Sonnet 5, Claude Sonnet 5.5, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Fable 5.1, GPT-6 Sol, GPT-6 Luna, GPT-6 Astra, GPT-5.5, Gemini 3.8 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash Lite, Gemini 3.1 Pro Preview
    1x

List-price calculation, not a run. 22 rows. Highest DeepSeek V4 Flash 0423 12.6x (n 15). Lowest Gemini 3.1 Pro Preview 1x (n 2).

Notesn 2–32 per row

Standard tier, blended price (3 input : 1 output), most expensive provider ÷ cheapest provider

A calculation on prices reported by OpenRouter’s public API, snapshot 2026-10-06. n = providers with a standard-tier endpoint. 1x means every provider charges the same. Highlighted: a spread of 4x or more.

Source: OpenRouter public API: models and provider endpoints (snapshot)

Question: for the same model, how much do providers differ in price, and what does a gateway add?

Answer, third-party-reported (OpenRouter's public API, snapshot 2026-10-06): 265 endpoints from 52 providers for 27 models. For 10 of 10 models with a first-party list price, OpenRouter's per-token price equals the vendor's; the gateway's cost is the 5.5% credit-purchase fee on the Standard plan. All 15 closed models (Claude, GPT, Gemini) with 2 or more providers had one standard-tier price at every provider. Open-weight models differ a lot (blended 3:1, a calculation): DeepSeek V4 Flash 12.6x, DeepSeek V4 Pro 11.2x, gpt-oss-120b 6.9x, Llama 3.3 70B 6.7x. The cheapest endpoints often report fp4 or fp8 precision or a shorter context. Latency was returned for 0 of 265 endpoints, and gateway delay is not measured: no key in the environment.

Study: /benchmarks/inference-provider-index. Posts: OpenRouter vs going direct and the cheapest place to run open models. Compare: OpenRouter vs Anthropic, Groq vs Cerebras, Amazon Bedrock vs Google Vertex.

8. Cost thought experiments

Question: what would the same recorded tokens cost at other list prices?

Answer, a calculation, not a run: Agent's SWE-bench work used 162.9M input tokens (94.0% cache reads) and 1.8M output tokens. At list prices that is $43.61 on Haiku 4.5, $87.23 on Sonnet 5.5, $143.83 on Opus 5.5 and $321.31 on Fable 5.1. Without caching, the Sonnet figure is $343.33.

Study: /benchmarks/cost-thought-experiments. Posts: What if every call ran on Opus? and How to estimate your AI coding bill. Tool: AI cost calculator.

9. CLI vs API: latency and hidden tokens

Question: how much time and context does a coding CLI add?

Answer: for a one-line answer, the Codex CLI took a median 3.5 times as long as the OpenAI API with the same model and effort, and sent about 19,551 input tokens instead of 17. On a scheduler repair with 296 checks, all 9 runs passed: Claude Code with Sonnet 5.5 15.0 s, OpenAI API with GPT-6.1 Sol 17.3 s, Codex CLI with GPT-6.1 Sol 61.2 s.

Study: /benchmarks/cli-model-latency-tokens. Posts: Claude Code vs Codex CLI vs the API and the hidden context tax. Compare: GPT-6.1 Sol: Codex CLI vs API.

10. Coding calibration on three real pull requests

Question: what broke on three real open-source pull requests across four platform builds?

Answer: verified delivery went from 0 of 3 to 1 of 3. h3 regressed to an empty patch; uvicorn passed its suite but stayed unverified. The latest slice cost $11.06 for the three tasks (notional). One attempt per cell finds defects; it does not measure a rate.

Study: /benchmarks/coding-calibration. Post: Why we count every failed attempt.

11. Effort ladder: does more effort buy quality? (new)

Every rate is 95% or more
Claude Sonnet 5.5 (low)
Claude Sonnet 5.5 (medium)
Claude Sonnet 5.5 (high)
Claude Sonnet 5.5
Claude Opus 5.5 (low)
Claude Opus 5.5 (medium)
Claude Opus 5.5 (high)
Claude Opus 5.5
GPT-6.1 Sol (low)
GPT-6.1 Sol (medium)
GPT-6.1 Sol (high)

Every interval overlaps every other: this chart does not order these rows.

11 rows. All at 100%.

NotesWhiskers: 95% Wilson intervaln = 16 per row11 of 11 at 100%: this task set cannot separate them.

Every configuration: 8 tasks × 2 repetitions. Whiskers are 95% Wilson intervals

Whiskers are 95% Wilson intervals. A format miss never counts as a pass. Reference cells (efforts the hard head-to-head already ran) are reused, not rerun; Claude reference cells keep repetitions 1-2, so every cell is 8 tasks × 2 repetitions.

Source: Effort ladder: the hard task set at each effort level

Question: on the 8 hard tasks, does a higher effort setting buy a higher pass rate for Sonnet 5.5, Opus 5.5 and GPT-6.1 Sol (Codex CLI)?

Answer: not on this set. All 11 configurations passed 16 of 16 strictly (95% interval 81% to 100% each), with 0 format misses and 0 wrong answers: 176 of 176 calls. The set has a ceiling; it cannot rule out a gap of up to about 19 points. More effort changed time and tokens. Median time per call, low to high: Sonnet 5.8 s to 8.8 s, Opus 7.5 s to 10.1 s, GPT-6.1 Sol 13.6 s to 18.1 s; the ranges overlap. List-price cost per strict pass (a calculation), low to high: Sonnet $0.0122 to $0.0167, Opus $0.0212 to $0.0337, GPT-6.1 Sol $0.0128 to $0.0151. All 16 effort comparisons in the dataset are tie or unclear.

Study: /benchmarks/effort-ladder. Post: Does reasoning effort buy quality?

12. Prompt caching and run-to-run consistency (new)

Calculation
  • With the cache, as recorded
  • Without a cache: every input token at the input price (square)
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code

Gap labels, Without a cache: every input token at the input price vs With the cache, as recorded: Without a cache: every input token at the input price is x% higher (+) or lower (−) than With the cache, as recorded, calculated from the two values shown (the change counted from With the cache, as recorded’s value).

List-price calculation, not a run. 2 rows, 2 series: With the cache, as recorded, Without a cache: every input token at the input price. With the cache, as recorded: highest Claude Opus 5.5 · Claude Code $0.26 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.14 (n 15). Without a cache: every input token at the input price: highest Claude Opus 5.5 · Claude Code $0.54 (n 15). Lowest Claude Sonnet 5.5 · Claude Code $0.27 (n 15).

Notesn = 15 per row

All recorded turns per model; the same reported tokens priced two ways

Calculation, not a bill: the calls ran on a subscription. Cache reads at the cache-read price, 1-hour cache writes at 2× the input price (every write in this run was a 1-hour write). Codex CLI is not priced here: it reports no cache-write count.

Sources: Caching sessions and repeated prompts (Claude Code and Codex CLI), Cost with and without the prompt cache (calculation), Anthropic list prices (Claude models)

Question: in a CLI session that reuses a fixed context, how much input comes from the cache and what does that save? When the same prompt runs 10 times, how much do the result and the time vary?

Answer, caching: in Claude Code, turns 2 to 5 read 97% of their input from the cache on average (turn 1: 19%). At list price, a calculation, 15 turns cost Sonnet $0.1350 vs $0.2698 without the cache (50% less) and Opus $0.2551 vs $0.5442 (53% less). Turn 1 costs more with the cache, because a 1-hour write is priced at 2x input. 0 of 4 later sessions reused an earlier session's cache; the cause was not tested. No clear latency effect. The Codex CLI read 99% of later-turn input from its cache, but reports no cache writes, so it is not priced.

Answer, consistency: 7 of 9 model-and-prompt cells passed 10 of 10 (72% to 100%). Haiku 4.5 passed 0/10 on the exact-number prompt, with the same wrong number every time, and 1/10 on the JSON prompt (9 format misses). Those are the only pass-rate rows where a model wins: Sonnet and GPT-6.1 Sol beat Haiku. Some time rows also separate on run ranges; see the consistency post. On the code fix, all passed 10/10 with 6, 3 and 6 distinct code bodies (Haiku, Sonnet, GPT-6.1 Sol).

Study: /benchmarks/caching-consistency. Posts: How much does prompt caching save? and Same prompt, ten answers. Compare: Haiku vs Sonnet, Claude Code vs Codex CLI.

13. Agent memory: does CLAUDE.md help? (new)

Rules the code already shows

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Rules the folders hint at

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

Team knowledge only

No memory
/init CLAUDE.md
Curated, 11 lines
Raw notes, 60 lines
Dreamed notes
Handbook, 210 lines
Stop hook only
Curated + hook

One panel per series, all on the same axis; whiskers are the 95% Wilson interval.

8 rows, 3 series: Rules the code already shows, Rules the folders hint at, Team knowledge only. Rules the code already shows: highest Curated, 11 lines 100% (95% interval 94%–100%, n 60). Lowest /init CLAUDE.md 95% (95% interval 86%–98%, n 60). All intervals overlap. Rules the folders hint at: all at 100%.

NotesWhiskers: 95% Wilson intervaln 15–60 per row6 of 8 (Rules the code already shows) at 100%: this task set cannot separate them.

Share of convention checks passed, pooled by kind of knowledge (Claude Sonnet 5.5)

Code shows: money in cents, coded errors, the clock helper, the log helper. Folders hint: append-only migrations, a new migration, the generated report. Team knowledge only: the changelog rule and the late-fee rate.

Source: Agent memory study: 8 kinds of project memory on Claude Code

Question: does project memory make Claude Code do better work, and which kind of memory?

Answer: 200 graded sessions (120 with Sonnet 5.5, 80 with Haiku 4.5) and 8 kinds of memory, from none to a 211-line handbook and a Stop hook. With no memory, Sonnet already followed 95% of the rules that the code shows and 100% of the rules that the folders hint at, but only 6 of 15 team-knowledge checks. With an 11-line curated file it passed 15 of 15. When a memory file held the late-fee rate, Sonnet used it 15 of 15 times; without it, Sonnet stopped and asked 7 of 9 times, and Haiku 4.5 invented a rate in 6 of 6 sessions without a flag. Messy memory misled Haiku: with raw notes it ran a stale test command 10 of 10 times, and after one "dreaming" pass 1 of 10. A Stop hook enforced every code rule but cannot carry a fact, and it used 1.6x the input tokens of no memory. Full-pass intervals overlap for most pairs (n = 15 per condition).

Study: /benchmarks/agent-memory. Post: Does CLAUDE.md help? We tested 8 kinds of agent memory.

14. Coding agents on hidden tests: Claude Code vs Codex CLI (new)

Entrance: medians race at 81× real time
Claude Sonnet 5.5 · Claude Code
Claude Opus 5.5 · Claude Code
GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI

Every lane’s run range overlaps another lane’s, so the chart shows no finish order.

3 rows. Slowest GPT-6.1 Sol (medium, tester’s AGENTS.md) · Codex CLI 113 s (range 78.5 s–222 s, n 12). Fastest Claude Sonnet 5.5 · Claude Code 23.1 s (range 18.7 s–44.5 s, n 12). Not all run ranges overlap.

NotesLines: fastest–slowest run (not an interval)n = 12 per row

Median wall time; whiskers = fastest and slowest of 12 sessions (not an interval)

CLI process start to exit. One session at a time per lane; the Claude Code and Codex CLI lanes ran at the same time on one machine with different accounts. Codex time includes reading the tester’s notes, writing a work log and trying to commit. A range is not a confidence interval.

Source: Coding agents head-to-head: Claude Code and Codex CLI on 6 hidden-test tasks

Question: on small real repository tasks graded by hidden tests, how do coding-agent CLIs compare with their normal file and shell tools?

Answer: 36 of 36 sessions passed: 12 of 12 each for Claude Code with Sonnet 5.5, Claude Code with Opus 5.5 and Codex CLI with GPT-6.1 Sol at medium effort (95% interval 76% to 100% each). The task set has a ceiling, so it cannot rank them on quality. Median time per session: Sonnet 23.1 s, Opus 56.9 s, Codex 113.4 s. The run ranges of Sonnet (18.7 s to 44.5 s) and Codex (78.5 s to 221.9 s) do not overlap, so the comparison pages mark Claude Code with Sonnet faster on that row (4.9x by medians). A confound: Codex read the tester's global AGENTS.md in 12 of 12 sessions despite --ignore-user-config; 10 wrote a work log nobody asked for and 9 reported a commit attempt, so its time, tool calls and diffs include that work. The agents wrote about the same amount of task code; the median diff was 45, 77.5 and 94 lines, and most of the difference is tests. List-price cost per pass (a calculation): Sonnet $0.085, Codex $0.098, Opus $0.22. Gemini CLI was not run: it asked for a browser login.

Study: /benchmarks/coding-agents-head-to-head. Post: Claude Code vs Codex CLI on hidden tests. Compare: Claude Code vs Codex CLI, Sonnet 5.5 vs GPT-6.1 Sol, Opus 5.5 vs GPT-6.1 Sol.

15. Opus vs Sonnet as the agent's brain on SWE-bench Verified (new, interim)

Claude Opus 5.5 (Agent, new build)
Claude Sonnet 5.5 (Agent, older builds)

2 rows. Highest Claude Opus 5.5 (Agent, new build) 67% (95% interval 21%–94%, n 3). Lowest Claude Sonnet 5.5 (Agent, older builds) 33% (95% interval 6.2%–79%, n 3). All intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 3 per row

One attempt per arm per instance, official grader · 95% Wilson intervals

Interim: 3 of 8 declared pairs are graded; 5 were never started; no resumed attempts are included here. With n = 3 the intervals span most of the axis, so this chart supports no ranking. The instances were chosen to hold Sonnet misses and resolves in equal numbers, so neither rate estimates SWE-bench Verified as a whole.

Sources: Agent on SWE-bench Verified with Opus 5.5 as the brain, paired with Sonnet 5.5 (interim), Agent on SWE-bench Verified, campaign 1 (25 instances), Agent on SWE-bench Verified, campaign 2 (8 compiled-extension instances)

Question: does Agent resolve more SWE-bench Verified instances with Claude Opus 5.5 as its brain than with Claude Sonnet 5.5, and at what cost and time?

Answer, interim: 3 of 8 declared pairs are graded; a usage gate stopped the run, and the other 5 resume after the reset on 2026-10-09. With Opus, Agent resolved 2 of 3 (95% interval 21% to 94%); with Sonnet, 1 of 3 of the same instances (6% to 79%). The one pair that differs, django__django-10554, went to Opus, and the exact McNemar p is 1.0: no evidence of a difference. On the same instances, Opus cost 2.6x as much ($22.76 vs $8.64, a list-price calculation) and used 1.9x the worker minutes. The two arms ran on different platform builds, and the instances were chosen to hold Sonnet misses and resolves in equal numbers.

Study: /benchmarks/swe-bench-opus-vs-sonnet. Post: Opus 5.5 vs Sonnet 5.5 on real GitHub issues: an early look. Compare: Sonnet vs Opus.

16. System One arena: Jev vs Clef and five open decision models (new)

Jev 1.13
Clef 27B
Clef-Flash 9B
Kev 4B
lev 4B
Laya
Julia-1

7 rows. Highest Jev 1.13 77% (95% interval 74%–79%, n 1048). Lowest Julia-1 26% (95% interval 23%–29%, n 1048). Not all intervals overlap.

NotesWhiskers: 95% Wilson intervaln = 1048 per row

Share of graded items answered correctly, first presentation · 95% Wilson intervals

1048 graded items in five suites. An error, a timeout or a label outside the option set counts as wrong. Two models differ clearly only where the paired McNemar test says so (table below).

Source: System One arena: typed-decision models on checkable decisions and in head-to-head games

Question: how good are typed-decision models at decisions with a checkable answer, and which one wins when they play each other?

Answer: Jev 1.13 answered 76.8% of 1,048 graded decisions correctly (805/1,048), ahead of every other model (exact McNemar p < 0.001 for each pair). The best open model, Clef 27B, reached 69.9%, at a median 1.9 s per decision on one Mac Studio against 137 ms for Jev 1.13 over the network. Clef-Flash 9B (65.2%) and Kev 4B (64.0%) were statistically tied. Laya (27.4%) and Julia-1 (25.9%) were tied, close to the 22.4% that picking an option at random scores on the same items (a calculation). In 1,512 round-robin games, Jev 1.13 rated highest, but only Jev 1.13 (30-11 of 42) and Clef-Flash 9B (29-11 of 42) beat the random player clearly. In Pong, where an answer counts only once it arrives, Jev 1.13 won 13 of 16 games and Kev 4B 12, while Clef 27B, the most accurate open model on the exam, won 4 at 1.8 s per decision. The items are synthetic, the local models ran quantized on one Mac, and Jev 1.13 latency includes the network.

Study: /benchmarks/system-one-arena. Posts: Jev vs Clef and five open decision models and Decision models play Connect Four, Nim and Pong.

What the sixteen studies say together

  • Quality ties are the norm. Across all comparison pages, 147 of 162 pass-rate, accuracy and behaviour-rate rows are ties: the 95% intervals overlap (our count). The other 15 all separate Claude Haiku 4.5 from a larger model: 7 on the hard set, 4 on repeated prompts and 4 in the memory study. In all 15, Haiku had the worse result. Effort did not separate any model from itself, and neither the coding-agent ceiling nor the interim SWE-bench pairs separate Sonnet from Opus.
  • Cost gaps are large and real as arithmetic. Sonnet 5.5 had the lowest cost per pass on both model head-to-heads and in the coding-agent sessions. Opus cost 2.6x as much as Sonnet on the coding-agent tasks and on the interim SWE-bench pairs. A dedicated router cost about 148 times less per decision than Sonnet, and rules cost nothing. Caching moved the agent bill by about 3.9 times, and halved a 5-turn session. For open-weight models, the provider moved the price by up to 12.6 times; for closed models, not at all.
  • The route is part of the result. CLI context and start-up can matter more than the model. Name both when you quote a number. A CLI's standing instructions count too: Codex read the tester's AGENTS.md although user config was off.
  • Accuracy is one axis, and time counts in games. On the arena exam, Clef 27B was the most accurate open model (69.9%) at a median 1.9 s per decision. In Pong, where each answer moves the paddle only after its measured latency, it won 4 of 16 games at 1.8 s per decision; Jev 1.13 won 13 of 16.
  • Small samples need humble claims. Many cells hold 3 to 24 runs. The SWE-bench Opus probe holds 3 graded pairs so far.
  • Consistent is not correct. A model can repeat the same wrong answer 10 times. Validate against the right answer.

Not in this roundup: an Opus-vs-Sonnet or Claude-vs-GPT-6.1 Sol quality difference, or a quality gain from higher effort (our tasks did not find one), the 5 SWE-bench Opus pairs that resume after 2026-10-09, a Codex rerun without the tester's notes, Gemini CLI (it asked for a browser login), a single composite score (why), gateway delay (no key in the environment), and provider speed or quality (no data).

Browse by model or pair

Run the benchmark on your own work

Agent keeps the same receipts for your tasks: the model, the route, the tokens, the time, the cost and the validation result. Try Agent and see your own numbers.

The data behind this post

  • SWE-bench
  • Coding Agents

Agent on SWE-bench Verified vs 11 public models

Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.

76% (25/33)Agent resolved, all 33 attempted instances · n = 33

6 chartsUpdated October 5, 2026

  • Code Review
  • AI vs human

AI pull requests vs merged human pull requests, judged blind

A blind panel of Claude and GPT critics preferred Agent's change over the merged human change on 9 of 12 real tasks. Votes, scores, caveats.

75% (9/12)Tasks where the panel preferred the AI change (latest attempt) · n = 12

4 chartsUpdated October 5, 2026

  • Head to head
  • Claude Haiku

Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head

130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.

98% (127/130)Calls that passed their validator · n = 130

9 chartsUpdated October 5, 2026

  • Prompt Caching
  • Consistency

Prompt caching and run-to-run consistency in Claude Code and Codex CLI

135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.

50%($0.1350 vs $0.2698) · List-price saving from the cache over 15 turns, Claude Sonnet 5.5 · Claude Code (calculation)

6 chartsUpdated October 6, 2026

Includes calculations
  • Routing
  • Jev

Jev vs Claude as a router: accuracy and cost

Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.

90% (74/82)Jev 1.13 (TypeSafe): exact decisions · n = 82

7 chartsUpdated October 6, 2026

  • Routing
  • Latency

Routing overhead: deterministic policy vs LLM routers vs Jev

How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.

1.42 µsp95 2.33 µs, p99 3.04 µs · Deterministic routing policy: median decision time · n = 20,000

9 chartsUpdated October 6, 2026

  • Inference
  • Providers

Inference provider index: 27 models, 52 providers

Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.

265endpoints, 52 providers, 27 models · Provider endpoints in the snapshot · n = 27

34 chartsUpdated October 6, 2026

Includes calculations
  • Thought experiment
  • LLM pricing

What if every call ran on Opus? Repricing real agent tokens

Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.

162.9MInput tokens recorded · n = 33

5 chartsUpdated October 5, 2026

  • Claude Code
  • Codex

Claude Code CLI vs Codex CLI vs the API: latency and tokens

194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.

100% (194/194)Evaluated runs that passed their validator · n = 194

5 chartsUpdated October 5, 2026

  • Calibration
  • Coding Agents

Coding calibration: what broke on three real pull requests

One attempt per task, four platform builds, failures kept: how an AI worker did on real fastify/session, h3 and uvicorn issues, and what broke.

1 of 3Verified deliveries, latest build · n = 3

4 chartsUpdated October 5, 2026

  • Claude Code
  • Agent Memory

Does memory help Claude Code? 8 kinds of agent memory, tested

200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.

40% (6/15)Team-knowledge checks passed with no memory (Sonnet 5.5) · n = 15

10 chartsUpdated October 6, 2026

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.