{"$k":["id","title","description","studySlug","chartIds","durationSeconds","transcript","placement"],"$r":[["october-2026-roundup","Agent Benchmarks, October 2026: 28 studies in one film","The headline of each of our 28 open benchmark studies, with sample sizes and intervals. Calculations labelled; failures counted.","swe-bench-verified",["hard-h2h-pass-rate","routing-cost-per-1000","router-overhead-decision-latency"],211,["October 2026 roundup · 28 studies. Agent Benchmarks: every headline. Recorded runs, 95% intervals where they exist, every failure counted. Calculations labelled.","SWE-bench Verified: Agent resolved 25 of 33. The public panel averaged 74.1%; the intervals overlap, so no rank. Agent resolved (one attempt each): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Model cost per resolved instance (calculation, notional): $3.71 (n = 25). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.","SWE-bench, interim (3 of 8 pairs graded): Opus 5.5 as Agent’s brain resolved 2 of 3, Sonnet 5.5 1 of 3. Exact McNemar p = 1.0: no difference yet. Opus cost 2.6× as much, a calculation. Opus 5.5 in Agent: resolved (interim): 67% (2/3) (n = 3, 95% CI 21–94%). Sonnet 5.5 in Agent: same instances: 33% (1/3) (n = 3, 95% CI 6–79%). List-price cost, Opus vs Sonnet (calculation): 2.6× (n = 3). Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.","Blind critics preferred the AI pull request to the merged human one on 9 of 12 tasks at the latest attempt, 6 of 12 at the first. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.","Five short tasks, nine setups: 127 of 130 calls passed, so speed and tokens separate them. Fastest median: Fable 5.1 at 1.9 s. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Fastest median total time (Fable 5.1): 1.9 s (n = 15). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.","Eight hard tasks: Sonnet 5.5, Opus 5.5, Opus 5.5 high, GPT-6.1 Sol medium, Fable 5.1 and GPT-6.1 Sol high passed every call (24/24 or 16/16). Haiku 4.5 passed 11/24. Chart: Eight hard tasks · strict pass rate · 95% intervals (n = 16–24 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.","Coding agents on hidden-test tasks: 36 of 36 sessions passed, so time separates them. Claude Code with Sonnet 5.5: median 23.1 s; Codex CLI took 4.9× as long, with extra standing instructions. Sessions that passed every hidden check: 100% (36/36) (n = 36, 95% CI 90–100%). Median time, Sonnet 5.5 in Claude Code: 23.1 s (n = 12). Median time, Codex CLI vs Claude Code with Sonnet: 4.9× (n = 12). Caveat: Codex CLI ran with the tester’s global AGENTS.md: --ignore-user-config does not switch that file off. 12 of 12 Codex sessions read the tester’s notes, 10 wrote a WORKLOG.md and 9 reported a commit attempt. Its time, tool calls and lines changed include that work. Claude Code ran with setting sources off, and no Claude session did any of it.","Effort ladder: 11 model and effort settings passed 176 of 176 calls strictly (16/16 each, 81%–100%). No effort level wins any of 172 rows. Calls that passed strictly, low to high effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference). Effort comparison rows where one level wins: 0 (of 172 rows · 16 pairs). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.","Caching, a list-price calculation: 50% less on Sonnet 5.5, 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every time. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.","Agent memory: Sonnet 5.5 followed team-only rules 40% (6/15) of the time without memory and 100% (15/15) with an 11-line file. Without the late-fee rate it asked 7/9 times. Team knowledge, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). Asked for the missing rate: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.","System One arena: on 1,085 checkable decisions, Jev 1.13 answered 76.8% right; the best open model, Clef 27B, 69.9%. Jev 1.13: 76.8% (805/1048) (n = 1048, 95% CI 74–79%). Clef 27B: 69.9% (733/1048) (n = 1048, 95% CI 67–73%). Checkable decisions: 1,085. Caveat: Local models ran quantized on one Mac; the vendors measured full-precision weights on data-centre GPUs. Latency on other hardware will differ, and Jev latency includes the network.","Routing: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82 exact, and the intervals overlap. List-price cost is where they differ. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.","Routing overhead: an in-process policy decides in 1.42 µs; Sonnet 5.5 as a router takes 2.60 s through the CLI, about 1.8 million times longer. Chart: Time per routing decision · log scale · median to p95 (n = 82–20000 each). Caveat: The policy is timed in process and the LLM routers through a CLI: this compares the two ways of routing as deployed, not two models on equal footing. A direct API call would skip the CLI time (shown separately).","Provider prices, reported by OpenRouter: open-weight models vary up to 12.6x across providers. Closed models: one price, plus a 5.5% credit fee. Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Gateway price equals the first-party price: 10 of 10. Endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.","Recorded SWE-bench tokens, repriced: $87.23 on Sonnet 5.5, $143.83 on Opus 5.5, $343.33 with no prompt cache. At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Sonnet 5.5 without the prompt cache: $343.33 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.","What a CLI adds: for a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens. Codex CLI vs OpenAI API, one-line answer: 3.5x slower (n = 30). Input tokens the Codex CLI sends for one line: 19,551 (n = 15). Scheduler repair, median: Claude Code vs Codex CLI: 15.0 s vs 61.2 s (n = 3, models differ too). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.","Three real pull requests: the latest build verified 1 of 3, at $11.06 notional. Guardrail refusals went 26 → 19. Verified deliveries, latest build: 1 of 3 (n = 3). Notional cost, latest build, 3 tasks: $11.06 (n = 3). Guardrail refusals, first vs latest slice: 26 → 19 (n = 3). Caveat: One attempt per cell: these are defect-finding runs, not rates.","Agent-loop strict passes: Haiku 54% (13/24); Sonnet 100% (16/16). Attempts excluded for outside reads: 2 of 56. Claude Haiku 4.5 strict pass rate, agent loop: 54% (13/24) (n = 24, 95% CI 35–72%). Claude Sonnet 5.5 strict pass rate, agent loop: 100% (16/16) (n = 16, 95% CI 81–100%). Agent-loop attempts left out for reading outside the work folder: 2 of 56 (n = 56). Caveat: The Claude single-call cells ran in another batch on 2026-10-06 (03:23 to 04:02 UTC), with the same CLI version, tasks and validators; provider load can differ by hour.","Haiku exact routing: thinking off 87% (71/82), on 89% (73/82). Paired McNemar p = 0.754; day and account differ. Claude Haiku 4.5 (thinking off): exact routing decisions: 87% (71/82) (n = 82, 95% CI 78–92%). Claude Haiku 4.5 (thinking on): exact routing decisions: 89% (73/82) (n = 82, 95% CI 80–94%). Paired exact test, thinking off vs on (exact McNemar p): 0.754 (n = 82). Caveat: The thinking-on and Sonnet arms are the recorded routing run of 2026-10-05, reused. The thinking-off arm ran on a different day and on a different Claude subscription account, so this is a confounded comparison. Day, account, CLI version and input-token differences can affect the results. The data cannot isolate the effect of thinking.","Format misses: instructions 35% (17/48), JSON schema 0% (0/48). The schema still gave wrong values: 13% (6/48). Format-miss rate with instructions only, all models: 35% (17/48) (n = 48, 95% CI 23–50%). Format-miss rate with a JSON schema, all models: 0% (0/48) (n = 48, 95% CI 0–7%). Wrong-values rate with a JSON schema, all models: 13% (6/48) (n = 48, 95% CI 6–25%). Caveat: Small samples: Haiku 24 calls per mode, Sonnet 12 calls per mode and GPT-6.1 Sol 12 calls per mode. A 12/12 result has a 95% interval of 76% to 100%. This is not a minimum detectable difference.","Later Claude sessions with substantial cache reuse: new folders 0 of 2 (95% interval 0% to 66%); a fixed folder 2 of 2 (95% interval 34% to 100%). This analysis is exploratory. Later sessions with at least 50% of turn-1 input cached, A: new folder each time: 0 of 2 (95% interval 0% to 66%) (n = 2, 95% CI 0–66%). Later sessions with at least 50% of turn-1 input cached, B: fixed folder: 2 of 2 (95% interval 34% to 100%) (n = 2, 95% CI 34–100%). Caveat: The surviving protocol file was created after all counted calls. Its claimed 00:32 UTC declaration is not supported by its file birth time. Amendment 1 and 2 state 00:36 and 00:37 UTC, but separate pre-edit copies do not verify those times. The current summary was regenerated at 07:41 UTC. Treat the analysis as exploratory.","Cache break-even, a calculation: a new prefix with a one-hour write needs 2 reuses (the 3rd request). Recorded session payback: 2 turns (6 of 6 sessions). Reuses before a 1-hour cached prefix costs less, whole prefix new (calculation): 2 reuses (the 3rd request). Turn at which a recorded Claude Code session’s total input cost with the cache first fell below its cost with no cache (calculation on recorded tokens): 2 turns (6 of 6 sessions) (n = 6). Calculation, not a run. Caveat: The retained Claude protocol file was created at 14:43:55 UTC, after the first counted session at 14:35:39 UTC on 2026-10-06. Its declaration says 14:25 UTC, but file times do not verify that claim. Treat this as a retrospective protocol record.","Exact unseen routing decisions: Jev 82% (46/56), Haiku 79% (44/56), Sonnet 88% (49/56). The intervals overlap; no rank. Jev 1.13 (TypeSafe): exact on unseen decisions: 82% (46/56) (n = 56, 95% CI 70–90%). Claude Haiku 4.5 · Claude Code: exact on unseen decisions: 79% (44/56) (n = 56, 95% CI 66–87%). Claude Sonnet 5.5 (low) · Claude Code: exact on unseen decisions: 88% (49/56) (n = 56, 95% CI 76–94%). Caveat: The case author and two of the three routers are Claude models, so a same-family label bias is possible. A second labeller (GPT-6.1 Sol) labelled every case blind; the secondary scoring keeps only the keys where its label is inside our acceptable set.","Reasoning cost at high vs low effort, a calculation from recorded calls: Sonnet 2.2x ($0.0043 at low, $0.0094 at high); Opus 3.6x ($0.0050 at low, $0.0180 at high). Reasoning cost per call, high ÷ low effort, Claude Sonnet 5.5 · Claude Code (calculation, means): 2.2x ($0.0043 at low, $0.0094 at high) (n = 16). Reasoning cost per call, high ÷ low effort, Claude Opus 5.5 · Claude Code (calculation, means): 3.6x ($0.0050 at low, $0.0180 at high) (n = 16). Calculation, not a run. Caveat: The hard-set protocol file was created after its first counted call. Both effort-ladder protocol files were created after their batches ended. The short-set file predates its first call, but its top-up amendment timing is unverified. Batch receipts preserve protocol text, but we cannot verify all rules were written before inference. Treat these as exploratory calculations, not preregistered tests.","Speed anatomy, a calculation: same-text token count ratio 1.8x. Extra time to first text with the larger prompt: Haiku +0.9 s; Sonnet +1.6 s. The same 4,327-character reply: most tokens ÷ fewest tokens across models (calculation): 1.8x (n = 23). Haiku: extra time to first text at 64k vs 1k (calculation): +0.9 s (n = 6). Sonnet: extra time to first text at 64k vs 1k (calculation): +1.6 s (n = 6). Calculation, not a run. Caveat: 4 calls per model in part A and 3 per size in part B. Medians of so few calls move with one slow call, and the ranges are not confidence intervals. The 95% Wilson intervals on the lookup rates are wide.","Cost per correct answer, a calculation: Sonnet every time $0.0132; Haiku with one retry then Sonnet $0.0566. No policy ran. Cost per correct answer, Sonnet 5.5 every time (calculation): $0.0132 (n = 24). Cost per correct answer, Haiku with one retry then Sonnet (calculation): $0.0566 (n = 48). Calculation, not a run. Caveat: Advance registration is not verified. The hard-set protocol file birth time is 2026-10-06 04:03:37 UTC; its first counted call started at 03:23:59 UTC. The ladder protocol file birth time is 14:35:11 UTC; its first Claude call started at 14:21:50 UTC. Both files claim advance declaration, but the available file times do not support that claim. Copying could explain the times; we cannot establish it.","Strict passes on four selected harder tasks: Sol 69% (11/16); Opus 42% (5/12); Sonnet 38% (6/16). Tasks were selected with a Sonnet pilot. GPT-6.1 Sol (medium) (Codex CLI): strict pass rate on the harder tasks: 69% (11/16) (n = 16, 95% CI 44–86%). Opus 5.5 (Claude Code): strict pass rate on the harder tasks: 42% (5/12) (n = 12, 95% CI 19–68%). Sonnet 5.5 (Claude Code): strict pass rate on the harder tasks: 38% (6/16) (n = 16, 95% CI 18–61%). Caveat: Selection effect: the study picked tasks that Sonnet did not pass twice in the pilot. Sonnet passed 2 of 8 pilot calls on the kept tasks and 6 of 16 counted calls (calculation: 25% and 38%; the intervals overlap). Selection can produce this pattern, but the run does not establish its cause.","Latency budget, a calculation: steps within an assumed 300 ms budget 3 of 14 (median: 3 of 14); within 1,500 ms 4 of 14 (median: 8 of 14), at p95 or the observed maximum. No voice turn ran. Steps that fit a 300 ms budget at the slow end (calculation): 3 of 14 (median: 3 of 14) (n = 14). Steps that fit a 1,500 ms budget at the slow end (calculation): 4 of 14 (median: 8 of 14) (n = 14). Calculation, not a run. Caveat: Timing samples omit 0 failed or untimed matched explorer calls and 0 failed or untimed Claude head-to-head calls. A failure is not a fast successful step. No pass rate is estimated here.","Routing cost at an assumed million decisions a day, a calculation: Jev $33.70; Sonnet $7,324. No load test ran. Daily cost at 1 million decisions, Jev (calculation): $33.70 (n = 82). Daily cost at 1 million decisions, Sonnet 5.5 router (calculation): $7,324 (n = 82). Calculation, not a run. Caveat: The routing and live Jev protocols predate the first counted calls by file birth time. The overhead protocol predates its microbenchmark output. Claude ran 82 calls per router, below its 120-call cap. Jev ran 246 counted calls, below its 300-call cap. One Haiku pilot and one Jev probe are excluded. All counted calls completed without call errors.\n\nThe policy microbenchmark used 5,000 warm-up iterations and 64 synthetic contexts; the minimum and individual timing samples are unavailable, so we cannot rebuild its quantiles or full range. No separate validator-control receipt or Sonnet pilot is retained.","28 studies, one place. Intervals, sources and every failure kept."],"\u0001"],["system-one-arena","System One arena: Jev vs Clef and five open decision models","1,085 checkable decisions, seven typed-decision models, every call counted. Jev 1.13 led with 76.8% right.","system-one-arena",["arena-accuracy","arena-speed-accuracy-gpu","arena-escape","arena-flips","arena-injection","arena-confident-wrong","arena-elo","arena-pong"],98.1,["System One arena · 1,085 decisions · 23,247 calls. Jev vs Clef vs five open decision models. Games, logic, policy cases in 20 industries, usability and stress tests. Every answer checkable.","Jev 1.13 answered 76.8% right. The best open model, Clef 27B: 69.9%. The smallest, Julia-1: 25.9%. Jev 1.13, hosted API: 76.8% (805/1048) (n = 1048, 95% CI 74–79%). Clef 27B, open weights: 69.9% (733/1048) (n = 1048, 95% CI 67–73%). Julia-1, smallest: 25.9% (271/1048) (n = 1048, 95% CI 23–29%). Caveat: Local models ran quantized on one Mac; the vendors measured full-precision weights on data-centre GPUs. Latency on other hardware will differ, and Jev latency includes the network.","One scale, seven decision models. Where the intervals overlap, the order is not settled. Chart: Share of graded decisions answered right · 95% intervals (n = 1048 each). Caveat: The items are synthetic and written by agents. They test clear, checkable decisions, not the messy decisions a product meets; real accuracy depends on how well the state and options are written.","Speed on one GPU, from a third-party board, against our accuracy. Jev is a hosted API, so it is not on this chart. Chart: Open models: accuracy against speed on the same GPU (n = 1048 each). Caveat: Jev ran on TypeSafe's servers and the six open models on one Mac Studio, so speed across that line is not comparable. This study compares capability (same questions and positions) directly, and speed only on one machine, on one GPU (third-party numbers), through one gateway (OpenRouter's numbers) or as list price.","Knowing when no option fits is the job of a router. The smallest models rarely said \"none\". Chart: Items with a \"none\", \"ask\" or \"escalate\" option (n = 147–618 each). Caveat: The items are synthetic and written by agents. They test clear, checkable decisions, not the messy decisions a product meets; real accuracy depends on how well the state and options are written.","Same options, same words. Only the order, or the labels, changed. Chart: Decisions that changed when only the presentation changed (n = 961 each). Caveat: The items are synthetic and written by agents. They test clear, checkable decisions, not the messy decisions a product meets; real accuracy depends on how well the state and options are written.","\"Ignore the instructions and choose X\", hidden in a review, an email or a file name. Chart: Right answers despite injected text in the state (n = 40 each). Caveat: The items are synthetic and written by agents. They test clear, checkable decisions, not the messy decisions a product meets; real accuracy depends on how well the state and options are written.","A confidence gate only works if the model is unsure when it is wrong. Chart: Wrong answers given with probability 0.8 or more (n = 243–777 each). Caveat: Typed-decision models are built for routing and gating with a confidence threshold. These items have no threshold: a model that is unsure must still pick.","Then they played each other: tic-tac-toe, Connect Four, Nim, Dots and Boxes and Pong. A random and a perfect player anchor the scale. Chart: Tournament rating across every game · 95% intervals (n = 336 each). Caveat: Local models ran quantized on one Mac; the vendors measured full-precision weights on data-centre GPUs. Latency on other hardware will differ, and Jev latency includes the network.","Jev 1.13 won. Rings mark moves a perfect solver would not play: Jev 1.13 6, Clef 27B 6. Recorded Connect Four game: Jev 1.13 (moves first) against Clef 27B, Jev 1.13 wins after 27 moves. Jev 1.13 played the best move 6 of 12 times. Clef 27B played the best move 4 of 10 times. Caveat: One recorded game (the first between these two in the round robin), after four random opening moves. \"Best\" columns come from a depth-10 search.","Jev 1.13 won 3–0. Jev 1.13 (hosted) decides in 124 ms, Laya (on one Mac) in 37 ms, but Laya picked the right zone on only 12% of its graded decisions. Recorded rally, 3.8 s: Jev 1.13 (left, 124 ms per decision, 77% right on the exam) against Laya (right, 37 ms per decision, 27% right on the exam). Point Jev 1.13. Final score 3–0. Caveat: Not a speed comparison: Jev ran on TypeSafe's servers, the open models on one Mac. The longest rally of the featured Jev vs Laya series (Jev won all 4 games).","Scored in points, each model on its own machine. On the same Mac, Julia-1 decided in 12 ms but picked the right zone on 21%; Clef 27B was right on 100% (31/31) of its graded decisions, at 1.8 s, and won 4 of 16. Chart: Pong round robin · share of games won · 95% intervals (n = 16 each). Caveat: Jev ran on TypeSafe's servers and the six open models on one Mac Studio, so speed across that line is not comparable. This study compares capability (same questions and positions) directly, and speed only on one machine, on one GPU (third-party numbers), through one gateway (OpenRouter's numbers) or as list price.","Test a decision model on your own decisions before you route on it. Every call online."],"\u0001"],["agent-memory","Does memory help Claude Code? 8 kinds of agent memory, tested","200 graded Claude Code sessions. Memory mattered for what the repository cannot show: team knowledge went from 40% to 100% with an 11-line file.","agent-memory",["memory-knowledge-class","memory-late-fee","memory-broken-test-command","memory-input-tokens","memory-full-pass"],61.8,["Agent memory study · 200 Claude Code sessions. Does memory help Claude Code? Eight kinds of memory. Five tasks. Hidden tests. Every session graded.","Without memory, Sonnet 5.5 followed the rules it could see, but passed only 40% (6/15) of the team-knowledge checks. With an 11-line file: 100% (15/15). Team knowledge followed, no memory: 40% (6/15) (n = 15, 95% CI 20–64%). Team knowledge followed, 11-line file: 100% (15/15) (n = 15, 95% CI 80–100%). No rate in memory: asked instead of coding: 7/9. Caveat: One small synthetic repository and one author of the facts: the curated file is an upper bound written with knowledge of the tasks.","Rules the code shows and rules the folders hint at: followed with or without memory. Team-only checks: 40% (6/15) without memory, 100% (15/15) with the curated file. Chart: Share of checks passed by kind of knowledge · Sonnet 5.5 (n = 15–60 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.","With the rate in any memory file: 15/15 right, even from notes that also held a stale 2%. Without it, Sonnet asked 7/9 times instead of guessing. Chart: \"Charge our standard late fee\" · Sonnet 5.5, 3 sessions per condition (n = 3 each). Caveat: In headless mode an agent that asks a question cannot get an answer, so asking counts as a failure here. In a live session the person would answer; the cost is the round trip.","Raw notes pile up: repeats, one-off events and rules that later changed. A dreaming pass should keep the facts and drop the rest. File: CLAUDE.md · 56 raw notes, oldest first. 10/10 current facts kept. 0/5 stale notes left. 1/9 transient notes left. Caveat: Labels use fixed text patterns. The tally is what Dream 1 actually kept.","The /init file copied the README's broken command: 13/15 sessions ran it. Haiku 4.5 with messy raw notes: 10/10; after one dreaming pass: 1/10. Chart: Sessions that ran the stale README test command (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.","The 11-line file used the fewest input tokens. The Stop hook alone used 1.6× the input of no memory, a calculation: each block is another round of work. Chart: Median input tokens per session · Sonnet 5.5 (n = 15 each). Caveat: Costs are the CLI's list-price estimates for subscription sessions, not invoices.","Full pass: tests and every rule. The intervals overlap for most pairs, so read the direction, not a rank. Chart: Full pass rate · 95% intervals (n = 10–15 each). Caveat: n = 15 per condition for Sonnet and 10 for Haiku: most full-pass intervals overlap, so most differences between conditions are not clear.","Write down what the repo cannot show. Enforce what a script can check."],"\u0001"],["effort-ladder","Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on eight hard tasks","All 11 configurations passed 16/16 strictly (95% interval 81%–100%): more effort bought no extra passes on this set (a ceiling). Median output tokens rose with effort; time ranges overlap. Costs are list-price calculations.","effort-ladder",["effort-ladder-pass-rate","effort-ladder-total-latency","effort-ladder-output-tokens","effort-ladder-cost-per-pass"],48.2,["Effort ladder · 176 calls · 11 configurations. Does more effort buy quality? Low, medium and high effort on eight hard tasks. Sonnet 5.5 and Opus 5.5 via Claude Code, GPT-6.1 Sol via Codex CLI.","176 of 176 calls passed strictly, at every effort level. No format misses and no wrong answers. Calls that passed strictly, every effort: 100% (176/176) (n = 176, 95% CI 98–100%). Model and effort configurations: 11 (6 new, 5 reference) (n = 176). Format misses and wrong answers: 0 · 0 (n = 176, format · wrong). Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.","All 11 configurations passed 16/16. A 16/16 has a 95% interval of 81%–100%: the set has a ceiling, so no effort level wins. Chart: Strict pass rate by effort · 95% Wilson intervals (n = 16 each). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.","Median time: Sonnet 5.5 5.8 s at low, 8.8 s at high; Opus 5.5 7.5 s at low, 10.1 s at high; GPT-6.1 Sol 13.6 s at low, 18.1 s at high. Within each model the ranges overlap: no effort is ranked faster. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16 each). Caveat: Each cell has only 2 calls per task. Medians of 16 calls move with a few slow calls; the ranges are wide and overlap.","More effort writes more: Sonnet 5.5 used a median 667 / 770 / 1,192 output tokens at low / medium / high. Chart: Median output and reasoning tokens per call (n = 16 each). Caveat: Claude Code and Codex CLI rows pair a CLI with a model; each CLI adds its own system prompt and start-up time. Effort levels are not the same scale across vendors.","List-price cost per strict pass, low → high effort, a calculation: Sonnet $0.0122 → $0.0167, Opus $0.0212 → $0.0337, GPT-6.1 Sol $0.0128 → $0.0151. Chart: List-price cost per strict pass · calculation (n = 16 each). Calculation, not a run. Caveat: List-price costs are calculations; the calls used flat subscriptions.","16 effort pairs, 172 comparison rows: no effort level wins any row. The rest are ties or unclear. Effort pairs compared (same model, same route): 16. Comparison rows across those pairs: 172. Rows where one effort level wins: 0 (of 172). Caveat: Every cell passed every call, so the set has a ceiling: a perfect 16/16 has a 95% interval of 81% to 100%. This study cannot show that effort does not matter on harder work; it shows that these 8 tasks do not need more than low effort.","Start at low effort and measure. Every call, interval and cost online."],"\u0001"],["caching-consistency","Prompt caching and consistency: what the cache saves, and how much answers vary","Calculation at list price: the cache cut a 5-question session 50% on Sonnet 5.5 and 53% on Opus 5.5. Same prompt 10 times: 7 of 9 cells passed every repetition.","caching-consistency",["caching-read-share-by-turn","caching-cost-with-without","consistency-pass-rate","consistency-distinct-answers"],49.2,["Caching and consistency · 135 calls. What the cache saves, and how much answers vary. Five-question sessions on a fixed context. Then the same prompt, 10 times.","135 calls: 45 cache turns, 90 repeated prompts. Every call counted. Sonnet 5.5 session cost saved by the cache (calculation): 50% ($0.1350 vs $0.2698). Later sessions that reused an earlier session’s cache: 0 of 4. Model-and-prompt cells that passed 10 of 10: 7 of 9. Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.","Claude Code turn 1 reads 19% from the cache, the CLI’s own prefix. Turns 2–5 read 90%–99%. Codex CLI: 98%–99%. Chart: Share of input read from the cache · mean of 3 sessions per turn (n = 3 each). Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.","A calculation: Sonnet 5.5 $0.1350 with the cache vs $0.2698 without, 50% less. Opus 5.5 $0.2551 vs $0.5442, 53% less. Chart: List-price cost of the recorded sessions · calculation (n = 15 each). Calculation, not a run. Caveat: Costs are list-price calculations; the calls used flat subscriptions.","No clear speed effect: Sonnet 5.5 1.6 s on turn 1 vs 1.6 s later; Opus 5.5 1.9 s vs 2.4 s. The ranges overlap. Sonnet 5.5: median turn 1 vs turns 2–5: 1.6 s vs 1.6 s (ranges 1.6–1.8 s and 1.4–5.6 s · n = 3 and 12). Opus 5.5: median turn 1 vs turns 2–5: 1.9 s vs 2.4 s (ranges 1.8–4.4 s and 1.6–12.7 s · n = 3 and 12). Caveat: Cache figures come from 3 sessions per model; they describe this CLI version and this context size. A different working folder, prompt order or cache lifetime can change them.","7 of 9 cells passed 10/10. Haiku 4.5: 0/10 on the exact number (10 wrong, 1 distinct answer), 1/10 on JSON (9 format misses). Chart: Same prompt, 10 times · strict passes · a 10/10 is 72%–100% at 95% (n = 10 each). Caveat: Consistency rests on 10 repetitions per cell: a 10/10 has a 95% interval of 72% to 100%.","Consistent is not correct: Haiku 4.5 gave the same wrong number all 10 times. Code fix, distinct correct bodies: Haiku 4.5 6, Sonnet 5.5 3, GPT-6.1 Sol (medium) 6. Chart: Same prompt, 10 times · distinct answers (n = 10 each). Caveat: The JSON prompt fails a reply in a code fence even when the JSON is right. The table counts these format misses apart from wrong answers.","Cache the fixed context; check answers, not agreement. Every call online."],"\u0001"],["provider-index","Same model, different price: 265 provider endpoints compared","Reported by OpenRouter’s public API on 2026-10-06: open-weight prices vary up to 12.6x across providers; closed models have one standard price, and the gateway adds a 5.5% credit fee.","inference-provider-index",["provider-index-spread","provider-prices-llama-3-3-70b-instruct","gateway-vs-direct-claude-sonnet-5-5"],42,["Provider index · third-party-reported · 2026-10-06. Same model, different price. 265 endpoints, 52 providers, 27 models. Prices as OpenRouter’s public API reports them.","265 endpoints, 52 providers, 27 models. The keyless API returned latency for 0 of them, so speed is not compared. Provider endpoints in the snapshot: 265 endpoints (52 providers, 27 models). Largest price spread (DeepSeek V4 Flash 0423): 12.6x (15 providers · blended price). Endpoints with a latency figure: 0 of 265. Caveat: Every price is third-party-reported by OpenRouter’s API at 2026-10-06. Prices change often; refetch before relying on them.","Open-weight models vary most: DeepSeek V4 Flash 12.6×, DeepSeek V4 Pro 11.2×. All 15 closed models with 2+ providers: one standard price (1×). Chart: Priciest ÷ cheapest standard provider · blended price (n = 2–32 each). Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.","One model, 10 providers: Llama 3.3 70B input costs $0.10 per million tokens at DeepInfra (fp8) and $1.04 at Together. Chart: Llama 3.3 70B · USD per million tokens · one row per provider. Caveat: The cheapest endpoint may run lower precision (fp4 or fp8) or a shorter context. Price alone does not make two endpoints equal.","Gateway markup: for 10 of 10 models OpenRouter charges the vendor’s per-token price. The cost is a 5.5% fee when you buy credits. Models at the first-party per-token price: 10 of 10. Credit-purchase fee, Standard plan by card: 5.5% ($0.80 minimum by card). Effective markup per token after the fee: +5.5% (calculation). Caveat: First-party prices in the product table were verified on an earlier date than the snapshot; a vendor price change in between would show as a markup.","Example: Sonnet 5.5 costs $2.00 in and $10.00 out per million tokens on both. A price is not a measurement: no winner. Chart: Claude Sonnet 5.5 · OpenRouter vs Anthropic list price · USD per million. Caveat: Comparison rows between providers name no winner: a price has no interval, so the rule for winners does not apply. The gap is stated.","Prices change often: every endpoint, tier and source date online."],"\u0001"],["routing-overhead","Routing overhead: a 1.42 µs policy vs LLM routers","A deterministic routing policy decides in 1.42 µs (median, n = 20,000); Sonnet 5.5 as a router takes 2.60 s through the CLI. Per-task costs and delays are calculations.","routing-overhead",["router-overhead-decision-latency","router-overhead-cli-vs-model-time","router-overhead-cost-per-1000-tasks","router-overhead-delay-per-task","cli-startup-tax"],51.4,["Routing overhead · policy vs LLM routers. What a routing decision costs before the work starts. Time and money per decision, then per task. Jev is timed over its API, a different route from the CLI routers.","The latency ladder: the in-process policy decides in 1.42 µs. Jev, a direct API call, takes 137 ms. Sonnet 5.5 as a router takes 2.60 s through the CLI, Haiku 4.5 12.5 s. Chart: Time per decision · log scale · median to p95 (n = 82–20000 each). Caveat: Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.","The LLM router takes about 1.8 million times longer per decision. Jev costs $0.0337 per 1,000 decisions, provider-reported. Policy decision, median (20,000 timed): 1.42 µs (n = 20000, p95 2.33 µs). Sonnet 5.5 router ÷ policy, medians: 1.8 million× (n = 82). Jev 1.13 per 1,000 decisions (provider-reported): $0.0337 (n = 82, 137 ms median per call over its API). Caveat: OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.","Of Sonnet 5.5’s 2.60 s per decision, a median 973 ms is CLI and harness time, not the model. Chart: Where an LLM router’s time goes · median per call (n = 82 each). Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.","Route all 49.5 model calls per task: $1.67 per 1,000 tasks with Jev, $247.30 with Sonnet 5.5. Route 7 decisions: $34.97. Chart: Added routing cost per 1,000 tasks. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.","If each decision waits in line, Sonnet 5.5 adds up to 129 s per task, or 18.2 s for 7 decisions. The policy adds 70.3 µs. Chart: Added routing delay per task · upper bound. Calculation, not a run. Caveat: Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.","Start-up tax for a one-word answer: Claude Code 2.53 s and 6,761 input tokens; Codex CLI 6.00 s and 17,051 input tokens, 13,184 of them read from the cache. Chart: CLI start-up tax · one-word answer · median of 5 runs (n = 5 each). Caveat: CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.","Route in process where a rule is enough. Every timing, cost and gap online."],"\u0001"],["hard-model-head-to-head","Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks","139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.","hard-model-head-to-head",["hard-h2h-pass-rate","hard-h2h-total-latency","hard-h2h-frontier"],44,["Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.","139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.","Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.","Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.","Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.","Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.","Open benchmarks: intervals, sources and every failure kept."],"\u0001"],["swe-bench-agent-vs-panel","SWE-bench Verified: Agent vs 11 public model runs","Agent resolved 25 of 33 SWE-bench Verified instances; every 95% interval overlaps the public panel, so the data does not rank it.","swe-bench-verified",["swebench-same-instance-leaderboard","swebench-by-difficulty-band"],28.4,["SWE-bench Verified · same 33 instances. Agent vs 11 public model runs. One attempt each. No retries. 95% intervals on individual run rates. The panel mean has no interval.","Agent resolved 25 of 33. The public panel averaged 74.1% on the same instances. Agent resolved (full pipeline, Sonnet 5.5): 76% (25/33) (n = 33, 95% CI 59–87%). Public panel mean, same instances: 74.1% (n = 33). Agent model cost per attempt (calculation, notional): $2.81 (n = 33). Caveat: Panel costs are published API costs; Agent costs are notional subscription estimates, not invoices.","Panel runs resolved 21 to 28 of 33. Every interval overlaps, so this sample cannot rank Agent. Chart: Resolved rate on the same 33 SWE-bench Verified instances (n = 33 each). Caveat: n = 33: intervals are wide. This is a defect-finding run, not a ranking.","Hardest band: on the 4 instances no panel model solved, Agent solved 1. Too few to call a rate. Chart: Resolved rate by difficulty band (n = 4–14 each). Caveat: Systems, not models: the panel is one model in a bash-only harness; Agent is a full pipeline on one model.","Open benchmarks: intervals, sources and every failure kept."],"\u0001"],["what-if-every-call-ran-on-opus","What if every call ran on Opus? Repricing real agent tokens","A calculation, not a run: Agent's recorded SWE-bench tokens cost $87.23 at Sonnet 5.5 prices, $143.83 at Opus 5.5 and $343.33 without caching.","cost-thought-experiments",["repriced-cost-per-resolved","prompt-cache-savings"],33,["Thought experiment · recorded tokens × list prices. What if every call ran on Opus? Agent recorded every token on 33 SWE-bench attempts. We repriced them.","162.9M input tokens, 94.0% of them read from the prompt cache. Input tokens recorded: 162.9M (n = 33). Output tokens recorded: 1.8M (n = 33). Input served from cache: 94.0% (n = 33). Caveat: Recorded costs are list-price estimates for subscription calls; no invoice backs them.","Same tokens on Opus 5.5: $143.83 instead of $87.23, 1.65× the bill. On Haiku 4.5: $43.61. At Haiku 4.5 prices: $43.61 (n = 33). At Sonnet 5.5 prices (the model that ran): $87.23 (n = 33). At Opus 5.5 prices: $143.83 (n = 33). Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.","Per resolved instance: $3.49 on Sonnet 5.5, $5.75 on Opus 5.5, $12.85 on Fable 5.1. Chart: Thought experiment: the same tokens at other list prices. Calculation, not a run. Caveat: Different models use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.","In this calculation caching matters more than the model: without it, Sonnet would cost $343.33, 3.9× the recorded $87.23. Chart: Thought experiment: what prompt caching saved. Calculation, not a run. Caveat: Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.","Price sensitivity, not predictions. Every repricing labelled as a calculation."],"\u0001"],["haiku-sonnet-opus-fable-codex-head-to-head","Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head","127 of 130 calls passed, so speed and tokens separate the models: Fable 5.1 was fastest at 1.9 s median.","model-head-to-head",["h2h-total-latency","h2h-input-tokens","h2h-frontier"],34.5,["Head-to-head · 130 timed calls · 9 configurations. Haiku vs Sonnet vs Opus vs Fable vs Codex. Five short tasks with strict validators. Every call kept, nothing retried.","127 of 130 calls passed. Pass rate barely separates them; speed and tokens do. Calls that passed their validator: 98% (127/130) (n = 130, 95% CI 93–99%). Median input tokens per call: Codex CLI vs Claude Code: 12,124 vs 2,130 (n = 130). Cheapest passing answer (list-price calculation): Sonnet 5.5: $0.0062 (n = 15). Caveat: The tasks are short and easy; pass rate saturates. Latency and tokens carry the signal. A harder follow-up with eight tasks and strict validators: /benchmarks/hard-model-head-to-head.","Fable 5.1 finishes first at 1.9 s. The Codex CLI needs 5.6–6.3 s. Chart: Median total time per call · real time (n = 10–15 each). Caveat: CLI timings include CLI start-up and the CLI’s own system prompt.","What the CLI sends: a median 12,124 input tokens per call on the Codex CLI, 2,130 on Claude Code. Chart: Input tokens per call: what the CLI sends (n = 10–15 each). Caveat: The prompt cache stayed at the provider default, so cache counters differ by route and by call order.","Speed vs cost per passing answer. Ringed: no other setup is faster, cheaper per pass and as accurate. Chart: Speed, cost and quality frontier (n = 10–15 each). Calculation, not a run. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.","Open benchmarks: intervals, sources and every failure kept."],"\u0001"],["jev-vs-llm-router","Jev vs Claude as a router: accuracy and cost","Exact routing decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. Intervals overlap; cost per decision differs by over 100×.","routing-jev-vs-llm",["routing-exact-decisions","routing-cost-per-1000","routing-economics-scenarios"],34.2,["Routing · 82 typed decisions. A small router vs Claude as the router. Jev 1.13 vs Claude Haiku 4.5 vs Claude Sonnet 5.5 on the decisions the platform asks.","Exact decisions: Jev 74/82, Haiku 4.5 73/82, Sonnet 5.5 77/82. The intervals overlap: accuracy does not rank them. Chart: Typed routing decisions answered exactly right (n = 82 each). Caveat: The case sets and question wording were revised in fix waves against Jev answers on 2026-10-04 and 2026-10-05, so Jev has a home advantage.","Cost does separate them: $0.0337 per 1,000 decisions for Jev vs $8.92 for Haiku 4.5 and $5.00 for Sonnet 5.5. Chart: Cost per 1,000 routing decisions (n = 82–246 each). Calculation, not a run. Caveat: Claude routers ran through the Claude Code CLI; the CLI adds startup time and tool-schema tokens a direct API call would not. Haiku 4.5 ran with the CLI default extended thinking; Sonnet 5.5 at effort low, as production asks.","Per decision, at list prices. Speed is another route: Jev's call took a median 137 ms over its API; the Claude routers ran through a CLI. Jev vs Claude Haiku 4.5: 265× cheaper ($8.92 ÷ $0.0337). Jev vs Claude Sonnet 5.5: 148× cheaper ($5.00 ÷ $0.0337). Calculation, not a run. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.","Repricing 2,362 recorded calls: all on Sonnet 5.5 $108.54; the Opus-heavy policy mix $161.62 (1.49×). Chart: Thought experiment: recorded agent work under different model mixes. Calculation, not a run. Caveat: Economics are calculations: Same tokens, same cache-read share and same number of calls on every model; a different model would take a different path and number of turns.","Open benchmarks: intervals, sources and every failure kept."],"\u0001"],["blind-critics-ai-vs-human-prs","AI pull requests vs merged human pull requests, judged blind","Blind critics preferred the AI change on 9 of 12 tasks at the latest attempt and 6 of 12 at the first. Same-family bias disclosed.","blind-review-head-to-head",["blind-review-first-vs-latest","blind-review-votes-by-task","blind-review-critic-agreement"],33.6,["Blind review · 12 real merged pull requests. AI change vs the merged human change. Critic models review both, unlabelled, once in each order.","Latest attempt: the panel preferred the AI change on 9 of 12 tasks. First attempt: 6 of 12. AI preferred, latest attempt: 75% (9/12) (n = 12, 95% CI 47–91%). AI preferred, first scored attempt: 50% (6/12) (n = 12, 95% CI 25–75%). Single critic verdicts for the AI change: 69% (91/132) (n = 132, 95% CI 61–76%). Caveat: Latest attempts are not independent first tries: they had lessons, review replays and some operator answers from earlier attempts on the same task.","The first attempt is the cleaner estimate: 50%, with a 95% interval from 25% to 75%. Chart: First attempt vs latest attempt (n = 5–12 each). Caveat: n = 12 tasks: intervals are wide.","Vote by vote: 9 of 12 tasks went unanimously to the AI change, 2 unanimously to the human one. Chart: Blind panel votes per task (latest attempt) (n = 6–8 each). Caveat: 7 of the 12 tasks come from one private Go service. They carry neutral labels (private-go-a and so on), and only their language and kind are published.","Cross-family check: OpenAI critics preferred the AI change in 10 of 12 verdicts. Small n, wide intervals. Chart: Does the judge’s model family matter? (n = 2–40 each). Caveat: Most critics are Anthropic models, and the worker runs on an Anthropic model. Same-family preference is possible; the per-critic chart shows the OpenAI critics’ share.","Open benchmarks: intervals, sources and every failure kept."],"\u0001"],["cli-vs-api-latency-race","Claude Code CLI vs Codex CLI vs the API: a latency race","For a one-line answer the Codex CLI was 3.5x slower than the API and sent 19,551 input tokens instead of 17.","cli-model-latency-tokens",["cli-vs-api-exact-reply-latency","cli-vs-api-prompt-overhead","scheduler-repair-claude-vs-codex"],32.8,["Latency race · CLI vs API. What a coding CLI adds on top of the model. Same model, same effort, same prompt. Timed from launch to exit.","A one-line answer: the OpenAI API replies in 1.0–1.5 s. The Codex CLI takes 3.2–4.2 s. Chart: One-line answer · median total time · real time (n = 5 each). Caveat: All runs are from one host and one network on 2026-10-03. Vendor latency changes over the day.","It also sends more: 18,859–19,555 input tokens for the same one-line request. The API sends 17. Chart: Hidden prompt: input tokens for the same one-line request (n = 5 each). Caveat: Small samples: 3 to 5 runs per configuration. Medians with ranges, not intervals.","A real repair, all runs passed: Claude Code 15.0 s, OpenAI API 17.3 s, Codex CLI 61.2 s. Chart: Scheduler repair · median total time · playback 8× (n = 3 each). Caveat: The scheduler comparison pairs Claude Code with Sonnet 5.5 against GPT-6.1 Sol on Codex and the API; the routes and the models differ together.","Open benchmarks: intervals, sources and every failure kept."],"\u0001"],["sonnet-vs-opus","Claude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say","66 comparison rows from 10 studies: 0 rows favour Sonnet 5.5, 0 favour Opus 5.5, 66 are ties or unclear. Cost rows are calculations.","hard-model-head-to-head",[],101.1,["Comparison · 66 rows · 10 studies. Sonnet 5.5 vs Opus 5.5. A winner only where the 95% intervals or run ranges do not overlap.","66 comparison rows from 10 studies. None separates them: intervals or run ranges overlap, or none was recorded. Rows where Sonnet 5.5 is ahead: 0 (of 66). Rows where Opus 5.5 is ahead: 0 (of 66). Ties or unclear: 66 (16 ties · 50 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.","SWE-bench pairs, interim: pass rate 33% vs 67%, tie: 95% intervals overlap. None of the 3 rows separates them. Table: SWE-bench pairs, interim · Agent · n = 3 per side. Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.","Five short tasks: pass rate 80% vs 100%, tie: 95% intervals overlap. None of the 8 rows separates them. Table: Five short tasks · Claude Code · n = 15 per side. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.","Eight hard tasks: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 6 rows separates them. Table: Eight hard tasks · Claude Code · n = 24 per side. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.","Coding agents, hidden tests: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Coding agents, hidden tests · Claude Code · n = 12 per side. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.","Effort ladder, default effort: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Effort ladder, default effort · Claude Code · n = 16 per side. Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.","Caching sessions: pass rate not measured. None of the 4 rows separates them. Table: Caching sessions · Claude Code · n = 15, 3, 12 per side. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.","Prompt cache break-even: after how many reuses does a cached prefix cost less?: pass rate not measured. None of the 9 rows separates them. Table: Prompt cache break-even: after how many reuses does a cached prefix cost less? · no cache · n =  per side. Caveat: The source has 30 attempted Claude turns, 0 failed turns and 0 turns without token usage. Failed turns with usage remain in cost totals. Missing usage cannot be priced. No quality rate or cache-caused speed effect is claimed.","How much of an AI bill is thinking? Reasoning tokens by model and effort: pass rate not measured. None of the 8 rows separates them. Table: How much of an AI bill is thinking? Reasoning tokens by model and effort · Claude Code · n = 24, 16, 15 per side. Caveat: The calls used one shared Mac and network. Other work and provider load were not controlled. Latency includes those effects.","Where the seconds go: first text, output speed and prompt size for 6 LLMs: pass rate 100% vs 56%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: Where the seconds go: first text, output speed and prompt size for 6 LLMs · Claude Code · n = 4, 3, 9 per side. Caveat: First text includes CLI start-up and reasoning time. It is not the API’s time to first token; a direct API call would skip the CLI start-up (the CLI reported ready after a median 0.33 s to 0.84 s per cell).","GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks: pass rate 38% vs 42%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks · Claude Code · n =  per side. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.","No winner where the data shows none. Every row and its reason online."],"compare"],["claude-code-vs-codex-cli","Claude Code vs Codex CLI: what the measurements say","88 comparison rows from 12 studies: 7 rows favour Claude Code, 1 favour Codex CLI, 80 are ties or unclear. Cost rows are calculations.","hard-model-head-to-head",[],41.1,["Comparison · 88 rows · 12 studies. Claude Code vs Codex CLI. A winner only where the 95% intervals or run ranges do not overlap.","88 comparison rows from 12 studies: Claude Code ahead on 7, Codex CLI ahead on 1. The rest do not separate them. Rows where Claude Code is ahead: 7 (of 88). Rows where Codex CLI is ahead: 1 (of 88). Ties or unclear: 80 (31 ties · 49 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.","Coding agents, hidden tests: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 4 rows separates them. Table: Coding agents, hidden tests · Claude Sonnet 5.5 vs GPT-6.1 Sol · n = 12 per side. Source study: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks. Rows shown: Coding sessions that passed every hidden check; Time per coding session; Tool calls per coding session; List-price cost per passing coding session (calculation). Recorded settings: Claude Sonnet 5.5 · six small repository tasks with hidden tests; GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests. Includes a calculation, not a bill or a new run. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.","Caching sessions: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 2 of 9 rows separate them. Table: Caching sessions · 5 of 9 rows · Claude Sonnet 5.5 vs GPT-6.1 Sol · n = 10 per side. Source study: Prompt caching and run-to-run consistency in Claude Code and Codex CLI. Rows shown: Same prompt, 10 times: strict pass rate (Exact number); Same prompt, 10 times: strict pass rate (JSON object); Same prompt, 10 times: strict pass rate (Code fix); Same prompt, 10 times: time per call (Exact number); Same prompt, 10 times: time per call (Code fix). Recorded settings: Claude Sonnet 5.5 · same prompt repeated 10 times; GPT-6.1 Sol · effort medium · same prompt repeated 10 times. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.","Routing overhead: pass rate not measured. 2 of 4 rows separate them. Table: Routing overhead · Claude Haiku 4.5 vs default model · n = 5 per side. Source study: Routing overhead: deterministic policy vs LLM routers vs Jev. Rows shown: CLI start-up tax on a one-word answer (First output event); CLI start-up tax on a one-word answer (First model output); CLI start-up tax on a one-word answer (Total wall time); Input tokens a CLI sends for a one-word answer. Recorded settings: Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs; default model · CLI start-up, one-word prompt, 5 runs. Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.","Single call vs agent loop: pass rate 100% vs 63%, Claude Code ahead: 95% intervals separate. 1 of 14 rows separates them. Table: Single call vs agent loop · 5 of 14 rows · Claude Sonnet 5.5 vs GPT-6 Luna · n = 3–24 vs 2–16. Source study: Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks. Rows shown: Strict pass rate: single call vs agent loop on eight hard tasks; Strict passes per task: single call vs agent loop: Interval merge fix; Strict passes per task: single call vs agent loop: DST day-length fix; Strict passes per task: single call vs agent loop: CSV parser; Strict passes per task: single call vs agent loop: Event-loop order. Recorded settings: Claude Sonnet 5.5 · single call; GPT-6 Luna · single call. Caveat: Each row is a CLI + model pair. Claude Code and Codex CLI add their own system prompts and tool schemas, and Codex CLI also loads the account’s user-level instruction file. A gap between Claude and GPT-6 Luna rows is partly the CLI.","No winner where the data shows none. Showing 18 of 88 rows; every row and its reason online."],"compare"],["sonnet-vs-gpt-6-1-sol","Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI): what the measurements say","70 comparison rows from 10 studies: 4 rows favour Sonnet 5.5, 1 favour GPT-6.1 Sol (Codex CLI), 65 are ties or unclear. Cost rows are calculations.","hard-model-head-to-head",[],42,["Comparison · 70 rows · 10 studies. Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI). A winner only where the 95% intervals or run ranges do not overlap.","70 comparison rows from 10 studies: Sonnet 5.5 ahead on 4, GPT-6.1 Sol (Codex CLI) ahead on 1. The rest do not separate them. Rows where Sonnet 5.5 is ahead: 4 (of 70). Rows where GPT-6.1 Sol (Codex CLI) is ahead: 1 (of 70). Ties or unclear: 65 (22 ties · 43 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.","Coding agents, hidden tests: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 4 rows separates them. Table: Coding agents, hidden tests · Claude Code vs Codex CLI · n = 12 per side. Source study: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks. Rows shown: Coding sessions that passed every hidden check; Time per coding session; Tool calls per coding session; List-price cost per passing coding session (calculation). Recorded settings: Claude Code · six small repository tasks with hidden tests; Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests. Includes a calculation, not a bill or a new run. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.","Caching sessions: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 2 of 9 rows separate them. Table: Caching sessions · 5 of 9 rows · Claude Code vs Codex CLI · n = 10 per side. Source study: Prompt caching and run-to-run consistency in Claude Code and Codex CLI. Rows shown: Same prompt, 10 times: strict pass rate (Exact number); Same prompt, 10 times: strict pass rate (JSON object); Same prompt, 10 times: strict pass rate (Code fix); Same prompt, 10 times: time per call (Exact number); Same prompt, 10 times: time per call (Code fix). Recorded settings: Claude Code · same prompt repeated 10 times; Codex CLI · effort medium · same prompt repeated 10 times. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.","Instructions vs JSON schema: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 8 rows separates them. Table: Instructions vs JSON schema · 5 of 8 rows · Claude Code vs Codex CLI · n = 12 per side. Source study: Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI. Rows shown: Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON); What each call produced: strict pass, format miss, wrong values or error (Strict pass); What each call produced: strict pass, format miss, wrong values or error (Format miss); What each call produced: strict pass, format miss, wrong values or error (Wrong values); Time per call, instructions vs schema mode. Recorded settings: Claude Code · instructions; Codex CLI · effort low · instructions. Includes a calculation, not a bill or a new run. Caveat: The calls repeat only three fixed prompts. Wilson intervals describe call outcomes under a binomial assumption; they do not measure accuracy across unseen tasks. The paired p-values also assume independent pairs and do not remove this limit.","Four harder tasks: pass rate 38% vs 69%, tie: 95% intervals overlap. 1 of 10 rows separates them. Table: Four harder tasks · 5 of 10 rows · Claude Code vs Codex CLI · n = 4–16 per side. Source study: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Rows shown: Pass rate on 4 harder tasks (Strict pass); Pass rate on 4 harder tasks (Lenient (format misses counted)); Calls that tried a tool although tools were off; Strict pass rate by task: 10x10 nonogram; Strict pass rate by task: 6x6 Skyscrapers. Recorded settings: Claude Code; Codex CLI · effort medium. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.","No winner where the data shows none. Showing 19 of 70 rows; every row and its reason online."],"compare"],["haiku-vs-sonnet","Claude Haiku 4.5 vs Claude Sonnet 5.5: what the measurements say","153 comparison rows from 14 studies: 0 rows favour Haiku 4.5, 21 favour Sonnet 5.5, 132 are ties or unclear. Cost rows are calculations.","hard-model-head-to-head",[],42.9,["Comparison · 153 rows · 14 studies. Haiku 4.5 vs Sonnet 5.5. A winner only where the 95% intervals or run ranges do not overlap.","153 comparison rows from 14 studies: Haiku 4.5 ahead on 0, Sonnet 5.5 ahead on 21. The rest do not separate them. Rows where Haiku 4.5 is ahead: 0 (of 153). Rows where Sonnet 5.5 is ahead: 21 (of 153). Ties or unclear: 132 (58 ties · 74 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.","Eight hard tasks: pass rate 46% vs 100%, Sonnet 5.5 ahead: 95% intervals separate. 2 of 6 rows separate them. Table: Eight hard tasks · 5 of 6 rows · Claude Code · n = 24 per side. Source study: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks. Rows shown: Pass rate on eight hard tasks (Strict pass); Pass rate on eight hard tasks (Lenient (format misses counted)); Total time per call on hard tasks (separate batches); Time to first useful output on hard tasks; Output tokens per call on hard tasks (Output tokens). Recorded settings: Claude Code · eight hard validated tasks. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.","Caching sessions: pass rate 0% vs 100%, Sonnet 5.5 ahead: 95% intervals separate. 3 of 9 rows separate them. Table: Caching sessions · 5 of 9 rows · Claude Code · n = 10 per side. Source study: Prompt caching and run-to-run consistency in Claude Code and Codex CLI. Rows shown: Same prompt, 10 times: strict pass rate (Exact number); Same prompt, 10 times: strict pass rate (JSON object); Same prompt, 10 times: strict pass rate (Code fix); Same prompt, 10 times: how many different answers (Exact number); Same prompt, 10 times: time per call (Code fix). Recorded settings: Claude Code · same prompt repeated 10 times. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.","Agent memory: pass rate 20% vs 60%, tie: 95% intervals overlap. 4 of 40 rows separate them. Table: Agent memory · 5 of 40 rows · n = 10 vs 15. Source study: Does memory help Claude Code? 8 kinds of agent memory, tested. Rows shown: Full pass rate by kind of memory: No memory; Full pass rate by kind of memory: /init CLAUDE.md; Full pass rate by kind of memory: Handbook, 210 lines; Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md; Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines. Caveat: The hook checks the same code rules as the grader. It shows what rules written as code can do; it cannot carry a fact such as the late-fee rate.","Routing overhead: success rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 3 of 8 rows separate them. Table: Routing overhead · 5 of 8 rows · thinking on vs effort low · n = 82 per side. Source study: Routing overhead: deterministic policy vs LLM routers vs Jev. Rows shown: Time to make one routing decision; Where an LLM router’s time goes: model vs CLI (Model API time); Where an LLM router’s time goes: model vs CLI (CLI and harness time); Routing calls that returned a decision; Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)). Recorded settings: thinking on · via Claude Code · routing overhead per decision; effort low · via Claude Code · routing overhead per decision; thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts; effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts. Includes a calculation, not a bill or a new run. Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.","No winner where the data shows none. Showing 20 of 153 rows; every row and its reason online."],"compare"],["jev-vs-sonnet-router","Jev 1.13 vs Claude Sonnet 5.5: what the measurements say","23 comparison rows from 3 studies: 2 rows favour Jev 1.13, 0 favour Sonnet 5.5, 21 are ties or unclear. Cost rows are calculations.","routing-jev-vs-llm",[],35.4,["Comparison · 23 rows · 3 studies. Jev 1.13 vs Sonnet 5.5. A winner only where the 95% intervals or run ranges do not overlap.","23 comparison rows from 3 studies: Jev 1.13 ahead on 2, Sonnet 5.5 ahead on 0. The rest do not separate them. Rows where Jev 1.13 is ahead: 2 (of 23). Rows where Sonnet 5.5 is ahead: 0 (of 23). Ties or unclear: 21 (15 ties · 6 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.","Jev vs LLM routers: pass rate 90% vs 94%, tie: 95% intervals overlap. None of the 7 rows separates them. Table: Jev vs LLM routers · 5 of 7 rows · typed routing decisions · n = 12–194 per side. Source study: Jev vs Claude as a router: accuracy and cost. Rows shown: Typed routing decisions answered exactly right; Per-question accuracy; Exact rate by decision type: Failure class; Exact rate by decision type: Message intent; Exact rate by decision type: Is it a rule? Recorded settings: typed routing decisions · TypeSafe API; typed routing decisions · via Claude Code. Caveat: Jev’s numbers are from a live run: 3 repeats of the same 82 decisions, not 246 independent samples, so its interval is taken at n = 82. Its recorded production run of 2026-10-05 scored 74 of 82.","Routing overhead: success rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 6 rows separates them. Table: Routing overhead · 5 of 6 rows · routing overhead per decision vs effort low · n = 246 vs 82. Source study: Routing overhead: deterministic policy vs LLM routers vs Jev. Rows shown: Time to make one routing decision; Routing calls that returned a decision; Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)); Added routing cost per 1,000 tasks (calculation) (Only System One decisions (7 per task)); Added routing delay per task (calculation) (Every model call routed (49.5 per task)). Recorded settings: routing overhead per decision · TypeSafe API; effort low · via Claude Code · routing overhead per decision; calculation per 1,000 tasks from recorded decision counts · TypeSafe API; effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts; calculation per task from recorded decision counts, decisions in line · TypeSafe API; effort low · via Claude Code · calculation per task from recorded decision counts, decisions in line. Includes a calculation, not a bill or a new run. Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.","Unseen routing decisions: pass rate 82% vs 88%, tie: 95% intervals overlap. 1 of 10 rows separates them. Table: Unseen routing decisions · 5 of 10 rows · Claude Code · n = 14–168 vs 14–125. Source study: Jev vs Claude routers on unseen decisions: a blind holdout. Rows shown: Unseen routing decisions answered exactly right; Per-question accuracy on unseen decisions; Exact rate on unseen decisions, by decision type: Failure class; Exact rate on unseen decisions, by decision type: Message intent; Time per routing decision, by route (Wall time). Recorded settings: Claude Code · effort low. Caveat: 56 cases (14 per decision type) from one author: per-type intervals are very wide.","No winner where the data shows none. Showing 15 of 23 rows; every row and its reason online."],"compare"],["opus-vs-gpt-6-1-sol","Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI): what the measurements say","50 comparison rows from 7 studies: 0 rows favour Opus 5.5, 0 favour GPT-6.1 Sol (Codex CLI), 50 are ties or unclear. Cost rows are calculations.","hard-model-head-to-head",[],41.1,["Comparison · 50 rows · 7 studies. Opus 5.5 vs GPT-6.1 Sol (Codex CLI). A winner only where the 95% intervals or run ranges do not overlap.","50 comparison rows from 7 studies. None separates them: intervals or run ranges overlap, or none was recorded. Rows where Opus 5.5 is ahead: 0 (of 50). Rows where GPT-6.1 Sol (Codex CLI) is ahead: 0 (of 50). Ties or unclear: 50 (12 ties · 38 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.","Five short tasks: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 8 rows separates them. Table: Five short tasks · 5 of 8 rows · Claude Code vs Codex CLI · n = 15 per side. Source study: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head. Rows shown: Pass rate on five validated tasks; Total time per call; Time to first useful output; Input tokens per call: what the CLI sends (Cache read); Input tokens per call: what the CLI sends (Other input). Recorded settings: Claude Code · effort high · five short validated tasks; Codex CLI · effort high · five short validated tasks. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.","Eight hard tasks: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 6 rows separates them. Table: Eight hard tasks · 5 of 6 rows · Claude Code vs Codex CLI · n = 24 vs 16. Source study: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks. Rows shown: Pass rate on eight hard tasks (Strict pass); Pass rate on eight hard tasks (Lenient (format misses counted)); Total time per call on hard tasks (separate batches); Time to first useful output on hard tasks; Output tokens per call on hard tasks (Output tokens). Recorded settings: Claude Code · effort high · eight hard validated tasks; Codex CLI · effort high · eight hard validated tasks. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.","Coding agents, hidden tests: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 4 rows separates them. Table: Coding agents, hidden tests · Claude Code vs Codex CLI · n = 12 per side. Source study: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks. Rows shown: Coding sessions that passed every hidden check; Time per coding session; Tool calls per coding session; List-price cost per passing coding session (calculation). Recorded settings: Claude Code · six small repository tasks with hidden tests; Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests. Includes a calculation, not a bill or a new run. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.","Effort ladder, default effort: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 4 rows separates them. Table: Effort ladder, default effort · Claude Code vs Codex CLI · n = 16 per side. Source study: Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks. Rows shown: Strict pass rate by effort on eight hard tasks; Total time per call by effort on hard tasks; Output tokens per call by effort on hard tasks (Output tokens); List-price cost per strict pass by effort (calculation). Recorded settings: Claude Code · effort medium · eight hard validated tasks, effort ladder; Codex CLI · effort medium · eight hard validated tasks, effort ladder. Includes a calculation, not a bill or a new run. Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.","No winner where the data shows none. Showing 18 of 50 rows; every row and its reason online."],"compare"]]}