{"$k":["i","story"],"$r":[[7,{"id":"hard-model-head-to-head","title":"Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on eight hard tasks","description":"139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Medians with ranges; costs are calculations.","studySlug":"hard-model-head-to-head","chartIds":["hard-h2h-pass-rate","hard-h2h-total-latency","hard-h2h-frontier"],"durationSeconds":44,"transcript":["Hard head-to-head · 152 scored calls · 7 configurations. Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol. Strict validators, tested against controls first. Every call kept, nothing retried.","139 of 152 calls passed strictly. 6 configurations passed every call; Haiku 4.5 passed 11/24. Calls that passed strictly: 91% (139/152) (n = 152, 95% CI 86–95%). Configurations that passed every call: 6 of 7 (4 at 24/24, 2 at 16/16) (n = 7). Codex CLI attempts blocked before any model call (not scored): 30 (of 182 attempts). Caveat: Claude Code and Codex CLI rows pair a CLI with a model, and each CLI adds its own system prompt and start-up time. A Claude-vs-GPT row compares the route + model pairs, not the models alone.","Perfect: 4 at 24/24 (86%–100%), 2 at 16/16 (81%–100%), 95% intervals. Haiku 4.5 at 11/24 (28%–65%). Chart: Strict pass rate · 95% Wilson intervals (n = 16–24 each). Caveat: 6 configurations passed every call, so the hard set still has a ceiling for them: a perfect 24/24 has a 95% interval of 86% to 100%; a perfect 16/16 has a 95% interval of 81% to 100%. Among them, only the latency, token and cost medians differ, and their per-call time ranges overlap.","Haiku 4.5 passed none of event-loop order, room schedule and SQL report query. 5 more replies were right but in the wrong format. Table: Every call, task by task · 3 or 2 calls per cell. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.","Median time per call: Sonnet 5.5 7.7 s, GPT-6.1 Sol 13.1 s and 18.1 s, Haiku 4.5 39.0 s. The ranges overlap: no speed ranking. Chart: Median total time per call · whiskers: fastest to slowest call (n = 16–24 each). Caveat: Strict format rules decide part of the result: a reply in a code fence fails. The lenient reading is shown next to it so the two can be told apart.","Quality vs list-price cost per strict pass, a calculation: Sonnet 5.5 at $0.0143 is the frontier alone. Chart: Quality vs cost frontier on hard tasks (n = 16–24 each). Calculation, not a run. Caveat: Default effort means the effort flag was not passed; the CLI chose. Haiku reported a median 4,556 reasoning tokens per call, which explains much of its extra time and output.","Open benchmarks: intervals, sources and every failure kept."]}],[14,{"id":"sonnet-vs-opus","title":"Claude Sonnet 5.5 vs Claude Opus 5.5: what the measurements say","description":"66 comparison rows from 10 studies: 0 rows favour Sonnet 5.5, 0 favour Opus 5.5, 66 are ties or unclear. Cost rows are calculations.","studySlug":"hard-model-head-to-head","placement":"compare","chartIds":[],"durationSeconds":101.1,"transcript":["Comparison · 66 rows · 10 studies. Sonnet 5.5 vs Opus 5.5. A winner only where the 95% intervals or run ranges do not overlap.","66 comparison rows from 10 studies. None separates them: intervals or run ranges overlap, or none was recorded. Rows where Sonnet 5.5 is ahead: 0 (of 66). Rows where Opus 5.5 is ahead: 0 (of 66). Ties or unclear: 66 (16 ties · 50 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.","SWE-bench pairs, interim: pass rate 33% vs 67%, tie: 95% intervals overlap. None of the 3 rows separates them. Table: SWE-bench pairs, interim · Agent · n = 3 per side. Caveat: Different platform builds: Opus ran on 4f6f4027; the Sonnet attempts ran on f0ac3a8a and 236c0d3f. Platform changes can move results by themselves, so a gap compares Sonnet on an older platform with Opus on the current one, not the models alone.","Five short tasks: pass rate 80% vs 100%, tie: 95% intervals overlap. None of the 8 rows separates them. Table: Five short tasks · Claude Code · n = 15 per side. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.","Eight hard tasks: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 6 rows separates them. Table: Eight hard tasks · Claude Code · n = 24 per side. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.","Coding agents, hidden tests: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Coding agents, hidden tests · Claude Code · n = 12 per side. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.","Effort ladder, default effort: pass rate 100% vs 100%, tie: 95% intervals overlap. None of the 4 rows separates them. Table: Effort ladder, default effort · Claude Code · n = 16 per side. Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.","Caching sessions: pass rate not measured. None of the 4 rows separates them. Table: Caching sessions · Claude Code · n = 15, 3, 12 per side. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.","Prompt cache break-even: after how many reuses does a cached prefix cost less?: pass rate not measured. None of the 9 rows separates them. Table: Prompt cache break-even: after how many reuses does a cached prefix cost less? · no cache · n =  per side. Caveat: The source has 30 attempted Claude turns, 0 failed turns and 0 turns without token usage. Failed turns with usage remain in cost totals. Missing usage cannot be priced. No quality rate or cache-caused speed effect is claimed.","How much of an AI bill is thinking? Reasoning tokens by model and effort: pass rate not measured. None of the 8 rows separates them. Table: How much of an AI bill is thinking? Reasoning tokens by model and effort · Claude Code · n = 24, 16, 15 per side. Caveat: The calls used one shared Mac and network. Other work and provider load were not controlled. Latency includes those effects.","Where the seconds go: first text, output speed and prompt size for 6 LLMs: pass rate 100% vs 56%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: Where the seconds go: first text, output speed and prompt size for 6 LLMs · Claude Code · n = 4, 3, 9 per side. Caveat: First text includes CLI start-up and reasoning time. It is not the API’s time to first token; a direct API call would skip the CLI start-up (the CLI reported ready after a median 0.33 s to 0.84 s per cell).","GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks: pass rate 38% vs 42%, tie: 95% intervals overlap. None of the 10 rows separates them. Table: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks · Claude Code · n =  per side. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.","No winner where the data shows none. Every row and its reason online."]}],[15,{"id":"claude-code-vs-codex-cli","title":"Claude Code vs Codex CLI: what the measurements say","description":"88 comparison rows from 12 studies: 7 rows favour Claude Code, 1 favour Codex CLI, 80 are ties or unclear. Cost rows are calculations.","studySlug":"hard-model-head-to-head","placement":"compare","chartIds":[],"durationSeconds":41.1,"transcript":["Comparison · 88 rows · 12 studies. Claude Code vs Codex CLI. A winner only where the 95% intervals or run ranges do not overlap.","88 comparison rows from 12 studies: Claude Code ahead on 7, Codex CLI ahead on 1. The rest do not separate them. Rows where Claude Code is ahead: 7 (of 88). Rows where Codex CLI is ahead: 1 (of 88). Ties or unclear: 80 (31 ties · 49 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.","Coding agents, hidden tests: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 4 rows separates them. Table: Coding agents, hidden tests · Claude Sonnet 5.5 vs GPT-6.1 Sol · n = 12 per side. Source study: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks. Rows shown: Coding sessions that passed every hidden check; Time per coding session; Tool calls per coding session; List-price cost per passing coding session (calculation). Recorded settings: Claude Sonnet 5.5 · six small repository tasks with hidden tests; GPT-6.1 Sol · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests. Includes a calculation, not a bill or a new run. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.","Caching sessions: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 2 of 9 rows separate them. Table: Caching sessions · 5 of 9 rows · Claude Sonnet 5.5 vs GPT-6.1 Sol · n = 10 per side. Source study: Prompt caching and run-to-run consistency in Claude Code and Codex CLI. Rows shown: Same prompt, 10 times: strict pass rate (Exact number); Same prompt, 10 times: strict pass rate (JSON object); Same prompt, 10 times: strict pass rate (Code fix); Same prompt, 10 times: time per call (Exact number); Same prompt, 10 times: time per call (Code fix). Recorded settings: Claude Sonnet 5.5 · same prompt repeated 10 times; GPT-6.1 Sol · effort medium · same prompt repeated 10 times. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.","Routing overhead: pass rate not measured. 2 of 4 rows separate them. Table: Routing overhead · Claude Haiku 4.5 vs default model · n = 5 per side. Source study: Routing overhead: deterministic policy vs LLM routers vs Jev. Rows shown: CLI start-up tax on a one-word answer (First output event); CLI start-up tax on a one-word answer (First model output); CLI start-up tax on a one-word answer (Total wall time); Input tokens a CLI sends for a one-word answer. Recorded settings: Claude Haiku 4.5 · CLI start-up, one-word prompt, 5 runs; default model · CLI start-up, one-word prompt, 5 runs. Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.","Single call vs agent loop: pass rate 100% vs 63%, Claude Code ahead: 95% intervals separate. 1 of 14 rows separates them. Table: Single call vs agent loop · 5 of 14 rows · Claude Sonnet 5.5 vs GPT-6 Luna · n = 3–24 vs 2–16. Source study: Single call vs agent loop: does letting the model run code help? Haiku 4.5, Sonnet 5.5 and GPT-6 Luna on 8 hard tasks. Rows shown: Strict pass rate: single call vs agent loop on eight hard tasks; Strict passes per task: single call vs agent loop: Interval merge fix; Strict passes per task: single call vs agent loop: DST day-length fix; Strict passes per task: single call vs agent loop: CSV parser; Strict passes per task: single call vs agent loop: Event-loop order. Recorded settings: Claude Sonnet 5.5 · single call; GPT-6 Luna · single call. Caveat: Each row is a CLI + model pair. Claude Code and Codex CLI add their own system prompts and tool schemas, and Codex CLI also loads the account’s user-level instruction file. A gap between Claude and GPT-6 Luna rows is partly the CLI.","No winner where the data shows none. Showing 18 of 88 rows; every row and its reason online."]}],[16,{"id":"sonnet-vs-gpt-6-1-sol","title":"Claude Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI): what the measurements say","description":"70 comparison rows from 10 studies: 4 rows favour Sonnet 5.5, 1 favour GPT-6.1 Sol (Codex CLI), 65 are ties or unclear. Cost rows are calculations.","studySlug":"hard-model-head-to-head","placement":"compare","chartIds":[],"durationSeconds":42,"transcript":["Comparison · 70 rows · 10 studies. Sonnet 5.5 vs GPT-6.1 Sol (Codex CLI). A winner only where the 95% intervals or run ranges do not overlap.","70 comparison rows from 10 studies: Sonnet 5.5 ahead on 4, GPT-6.1 Sol (Codex CLI) ahead on 1. The rest do not separate them. Rows where Sonnet 5.5 is ahead: 4 (of 70). Rows where GPT-6.1 Sol (Codex CLI) is ahead: 1 (of 70). Ties or unclear: 65 (22 ties · 43 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.","Coding agents, hidden tests: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 4 rows separates them. Table: Coding agents, hidden tests · Claude Code vs Codex CLI · n = 12 per side. Source study: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks. Rows shown: Coding sessions that passed every hidden check; Time per coding session; Tool calls per coding session; List-price cost per passing coding session (calculation). Recorded settings: Claude Code · six small repository tasks with hidden tests; Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests. Includes a calculation, not a bill or a new run. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.","Caching sessions: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 2 of 9 rows separate them. Table: Caching sessions · 5 of 9 rows · Claude Code vs Codex CLI · n = 10 per side. Source study: Prompt caching and run-to-run consistency in Claude Code and Codex CLI. Rows shown: Same prompt, 10 times: strict pass rate (Exact number); Same prompt, 10 times: strict pass rate (JSON object); Same prompt, 10 times: strict pass rate (Code fix); Same prompt, 10 times: time per call (Exact number); Same prompt, 10 times: time per call (Code fix). Recorded settings: Claude Code · same prompt repeated 10 times; Codex CLI · effort medium · same prompt repeated 10 times. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.","Instructions vs JSON schema: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 1 of 8 rows separates them. Table: Instructions vs JSON schema · 5 of 8 rows · Claude Code vs Codex CLI · n = 12 per side. Source study: Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI. Rows shown: Does a JSON schema raise the pass rate? Instructions vs schema mode (Strict pass: the whole reply is the right JSON); What each call produced: strict pass, format miss, wrong values or error (Strict pass); What each call produced: strict pass, format miss, wrong values or error (Format miss); What each call produced: strict pass, format miss, wrong values or error (Wrong values); Time per call, instructions vs schema mode. Recorded settings: Claude Code · instructions; Codex CLI · effort low · instructions. Includes a calculation, not a bill or a new run. Caveat: The calls repeat only three fixed prompts. Wilson intervals describe call outcomes under a binomial assumption; they do not measure accuracy across unseen tasks. The paired p-values also assume independent pairs and do not remove this limit.","Four harder tasks: pass rate 38% vs 69%, tie: 95% intervals overlap. 1 of 10 rows separates them. Table: Four harder tasks · 5 of 10 rows · Claude Code vs Codex CLI · n = 4–16 per side. Source study: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks. Rows shown: Pass rate on 4 harder tasks (Strict pass); Pass rate on 4 harder tasks (Lenient (format misses counted)); Calls that tried a tool although tools were off; Strict pass rate by task: 10x10 nonogram; Strict pass rate by task: 6x6 Skyscrapers. Recorded settings: Claude Code; Codex CLI · effort medium. Caveat: Opus and Haiku share Sonnet’s family and CLI, so tasks picked against Sonnet may be hard for them for the same reasons. GPT-6.1 Sol was not selected against, which can widen the gap between Sol and the Claude rows. A fixed set picked before any run would be fairer.","No winner where the data shows none. Showing 19 of 70 rows; every row and its reason online."]}],[17,{"id":"haiku-vs-sonnet","title":"Claude Haiku 4.5 vs Claude Sonnet 5.5: what the measurements say","description":"153 comparison rows from 14 studies: 0 rows favour Haiku 4.5, 21 favour Sonnet 5.5, 132 are ties or unclear. Cost rows are calculations.","studySlug":"hard-model-head-to-head","placement":"compare","chartIds":[],"durationSeconds":42.9,"transcript":["Comparison · 153 rows · 14 studies. Haiku 4.5 vs Sonnet 5.5. A winner only where the 95% intervals or run ranges do not overlap.","153 comparison rows from 14 studies: Haiku 4.5 ahead on 0, Sonnet 5.5 ahead on 21. The rest do not separate them. Rows where Haiku 4.5 is ahead: 0 (of 153). Rows where Sonnet 5.5 is ahead: 21 (of 153). Ties or unclear: 132 (58 ties · 74 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.","Eight hard tasks: pass rate 46% vs 100%, Sonnet 5.5 ahead: 95% intervals separate. 2 of 6 rows separate them. Table: Eight hard tasks · 5 of 6 rows · Claude Code · n = 24 per side. Source study: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks. Rows shown: Pass rate on eight hard tasks (Strict pass); Pass rate on eight hard tasks (Lenient (format misses counted)); Total time per call on hard tasks (separate batches); Time to first useful output on hard tasks; Output tokens per call on hard tasks (Output tokens). Recorded settings: Claude Code · eight hard validated tasks. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.","Caching sessions: pass rate 0% vs 100%, Sonnet 5.5 ahead: 95% intervals separate. 3 of 9 rows separate them. Table: Caching sessions · 5 of 9 rows · Claude Code · n = 10 per side. Source study: Prompt caching and run-to-run consistency in Claude Code and Codex CLI. Rows shown: Same prompt, 10 times: strict pass rate (Exact number); Same prompt, 10 times: strict pass rate (JSON object); Same prompt, 10 times: strict pass rate (Code fix); Same prompt, 10 times: how many different answers (Exact number); Same prompt, 10 times: time per call (Code fix). Recorded settings: Claude Code · same prompt repeated 10 times. Caveat: Cross-session reuse did not happen here. The cause was not tested: the CLI may add per-process context before the user message. Do not read it as a provider property.","Agent memory: pass rate 20% vs 60%, tie: 95% intervals overlap. 4 of 40 rows separate them. Table: Agent memory · 5 of 40 rows · n = 10 vs 15. Source study: Does memory help Claude Code? 8 kinds of agent memory, tested. Rows shown: Full pass rate by kind of memory: No memory; Full pass rate by kind of memory: /init CLAUDE.md; Full pass rate by kind of memory: Handbook, 210 lines; Team knowledge followed, Sonnet vs Haiku: /init CLAUDE.md; Team knowledge followed, Sonnet vs Haiku: Handbook, 210 lines. Caveat: The hook checks the same code rules as the grader. It shows what rules written as code can do; it cannot carry a fact such as the late-fee rate.","Routing overhead: success rate 100% vs 100%, both at the ceiling: this measure cannot separate them. 3 of 8 rows separate them. Table: Routing overhead · 5 of 8 rows · thinking on vs effort low · n = 82 per side. Source study: Routing overhead: deterministic policy vs LLM routers vs Jev. Rows shown: Time to make one routing decision; Where an LLM router’s time goes: model vs CLI (Model API time); Where an LLM router’s time goes: model vs CLI (CLI and harness time); Routing calls that returned a decision; Added routing cost per 1,000 tasks (calculation) (Every model call routed (49.5 per task)). Recorded settings: thinking on · via Claude Code · routing overhead per decision; effort low · via Claude Code · routing overhead per decision; thinking on · via Claude Code · calculation per 1,000 tasks from recorded decision counts; effort low · via Claude Code · calculation per 1,000 tasks from recorded decision counts. Includes a calculation, not a bill or a new run. Caveat: The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.","No winner where the data shows none. Showing 20 of 153 rows; every row and its reason online."]}],[19,{"id":"opus-vs-gpt-6-1-sol","title":"Claude Opus 5.5 vs GPT-6.1 Sol (Codex CLI): what the measurements say","description":"50 comparison rows from 7 studies: 0 rows favour Opus 5.5, 0 favour GPT-6.1 Sol (Codex CLI), 50 are ties or unclear. Cost rows are calculations.","studySlug":"hard-model-head-to-head","placement":"compare","chartIds":[],"durationSeconds":41.1,"transcript":["Comparison · 50 rows · 7 studies. Opus 5.5 vs GPT-6.1 Sol (Codex CLI). A winner only where the 95% intervals or run ranges do not overlap.","50 comparison rows from 7 studies. None separates them: intervals or run ranges overlap, or none was recorded. Rows where Opus 5.5 is ahead: 0 (of 50). Rows where GPT-6.1 Sol (Codex CLI) is ahead: 0 (of 50). Ties or unclear: 50 (12 ties · 38 unclear). Caveat: Calculation rows are derived from list prices and recorded counts; they are not bills or runs.","Five short tasks: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 8 rows separates them. Table: Five short tasks · 5 of 8 rows · Claude Code vs Codex CLI · n = 15 per side. Source study: Haiku vs Sonnet vs Opus vs Fable vs Codex: a timed head-to-head. Rows shown: Pass rate on five validated tasks; Total time per call; Time to first useful output; Input tokens per call: what the CLI sends (Cache read); Input tokens per call: what the CLI sends (Other input). Recorded settings: Claude Code · effort high · five short validated tasks; Codex CLI · effort high · five short validated tasks. Caveat: Few repetitions per cell (2 or 3 per task). Medians with ranges, not intervals.","Eight hard tasks: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 6 rows separates them. Table: Eight hard tasks · 5 of 6 rows · Claude Code vs Codex CLI · n = 24 vs 16. Source study: Haiku vs Sonnet vs Opus vs Fable vs GPT-6.1 Sol on 8 hard tasks. Rows shown: Pass rate on eight hard tasks (Strict pass); Pass rate on eight hard tasks (Lenient (format misses counted)); Total time per call on hard tasks (separate batches); Time to first useful output on hard tasks; Output tokens per call on hard tasks (Output tokens). Recorded settings: Claude Code · effort high · eight hard validated tasks; Codex CLI · effort high · eight hard validated tasks. Caveat: Claude Code: n = 24 per configuration (3 repetitions per task); Codex CLI: n = 16 per configuration (2 repetitions per task). Per-task cells have only 2 to 3 calls.","Coding agents, hidden tests: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 4 rows separates them. Table: Coding agents, hidden tests · Claude Code vs Codex CLI · n = 12 per side. Source study: Claude Code (Sonnet 5.5, Opus 5.5) vs Codex CLI on 6 hidden-test coding tasks. Rows shown: Coding sessions that passed every hidden check; Time per coding session; Tool calls per coding session; List-price cost per passing coding session (calculation). Recorded settings: Claude Code · six small repository tasks with hidden tests; Codex CLI · effort medium · tester’s AGENTS.md · six small repository tasks with hidden tests. Includes a calculation, not a bill or a new run. Caveat: Every agent passed every session, so the pass rates sit at the 100% ceiling. These tasks are too easy to separate the agents on quality; only time, tool use, tokens and diff size differ.","Effort ladder, default effort: pass rate 100% vs 100%, both at the ceiling: this measure cannot separate them. None of the 4 rows separates them. Table: Effort ladder, default effort · Claude Code vs Codex CLI · n = 16 per side. Source study: Does more effort buy quality? Sonnet, Opus and GPT-6.1 Sol on 8 hard tasks. Rows shown: Strict pass rate by effort on eight hard tasks; Total time per call by effort on hard tasks; Output tokens per call by effort on hard tasks (Output tokens); List-price cost per strict pass by effort (calculation). Recorded settings: Claude Code · effort medium · eight hard validated tasks, effort ladder; Codex CLI · effort medium · eight hard validated tasks, effort ladder. Includes a calculation, not a bill or a new run. Caveat: Reference cells ran in a different batch and hour than the new cells (provider load can differ), on the same host, CLI versions, cases and flags.","No winner where the data shows none. Showing 18 of 50 rows; every row and its reason online."]}]]}