Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
96 calls. Strict passes, schema vs instructions: Haiku 18/24 vs 0/24, Sonnet 12/12 vs 12/12, GPT-6.1 Sol 12/12 vs 12/12. With intervals.
Open data · October 7, 2026
Claude, Codex and routers, measured on real tasks: pass rates with intervals, latency, tokens and cost per passing answer. Every number links to its sample, its method and its raw data.
The first measured rate from each study, with its 95% interval on a 0–100% track. Open a study for the method and the caveats.
Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI
35% (17/48)
Format-miss rate with instructions only, all models
95% CI 23%–50% · n = 48
GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks
39% (22/56)
Counted calls that passed strictly (harder set)
95% CI 28%–52% · n = 56
Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts
46% (11/24)
Strict passes, Claude Haiku 4.5 · Claude Code, eight hard tasks
95% CI 28%–65% · n = 24
Where the seconds go: first text, output speed and prompt size for 6 LLMs
96% (23/24)
Part A replies that matched all 250 lines exactly (strict)
95% CI 80%–99% · n = 24
Each study answers one question. Newest first. The glyph is the study’s lead chart.
96 calls. Strict passes, schema vs instructions: Haiku 18/24 vs 0/24, Sonnet 12/12 vs 12/12, GPT-6.1 Sol 12/12 vs 12/12. With intervals.
30 calls: later Claude Code sessions showed near-full turn-1 cache reads with a fixed folder, but not with new folders in this sample. Codex CLI tested too.
56 counted calls on 4 harder tasks with strict validators: GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5. Pass rate, intervals, speed, cost.
Haiku 4.5 first, then retry or escalate to Sonnet 5.5? A calculation on 64 recorded calls across 8 hard tasks: cost and time per correct answer.
60 calls on 6 models: time to first text, tokens per second, and what a 1k, 16k or 64k prompt adds. Claude Code and Codex CLI.
36 graded sessions: Claude Code with Sonnet 5.5 and Opus 5.5, Codex CLI with GPT-6.1 Sol. All passed every hidden test; time, tool calls and diffs differ.
200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a long handbook, a Stop hook. What helped and what it cost.
176 calls on 8 hard tasks at low, medium, high and default effort: Sonnet, Opus and GPT-6.1 Sol. Pass rate with 95% intervals, time, tokens, cost per pass.
Claude Haiku 4.5 with extended thinking on and off: 82 routing decisions and 8 hard tasks. Accuracy with 95% intervals, time and cost.
152 calls on 8 hard tasks with strict validators. Pass rate with 95% intervals, format misses, speed, tokens and cost per pass.
Reasoning tokens in 378 recorded calls by model and effort: share of output, list-price cost per call and per pass, and time. Calculations.
Price per million tokens for 27 models across 52 providers, the spread between them and OpenRouter’s markup over first-party prices.
Typed routing decisions: Jev, a dedicated router model, against Claude Haiku and Sonnet. Accuracy with intervals, cost per 1,000 decisions, latency.
Jev 1.13, Claude Haiku 4.5 and Sonnet 5.5 on 56 routing decisions written blind and frozen first: exact rate with 95% intervals, tuned vs unseen.
Interim: 3 of 8 paired SWE-bench Verified instances graded. Opus 5.5 resolved 2, Sonnet 5.5 1 (McNemar p = 1.0). Cost 2.6×, a list-price calculation.
A calculation on list prices: reuses before a cached prompt prefix costs less, per Claude and GPT model, and cost per 1,000 sessions of 1 to 20 turns.
135 calls: cache hit share and list-price savings in multi-turn sessions, and pass rate and answer diversity over 10 repeats of the same prompt.
How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.
118 attempts on 8 hard tasks: one call vs an agent loop that runs code in a sandbox. Pass rate, time, tokens and cost.
Jev 1.13, Clef 27B, Clef-Flash and four small open decision models on 1,085 checkable decisions and in head-to-head games, Pong included.
Calculation: component times against assumed 300 ms, 800 ms and 1,500 ms budgets. No voice turn or audio was measured.
A calculation from measured runs: what 10,000 to 10 million routing decisions a day cost with rules, Jev and Claude, with median/p95 time scenarios.
Agent resolved 25 of 33 SWE-bench Verified instances (76%), inside the public panel's range on the same instances. Cost, time and calls.
A blind panel of Claude and GPT critics preferred Agent's change over the merged human change on 9 of 12 real tasks. Votes, scores, caveats.
194 timed runs: how long Claude Code, Codex CLI and the OpenAI API take to answer and to fix code, and how many hidden tokens a CLI adds.
One attempt per task, four platform builds, failures kept: how an AI worker did on real fastify/session, h3 and uvicorn issues, and what broke.
130 timed calls on five validated tasks: Claude Haiku, Sonnet, Opus and Fable against Codex GPT-6.1 Sol. Pass rate, speed, tokens, cost per pass.
Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.
The same numbers, arranged by pair, by model, by metric, by provider and by budget.
Every pair measured on at least two shared metrics, with the rows the data separates and the rows it does not.
One page per system with every value the studies recorded for it, its configuration and its interval.
No composite score. Pick a metric and see which systems the intervals can and cannot separate.
Every listed provider’s price for each model, and how far the prices spread for the same model.
Pick a job, then cost or speed. See measured configurations, sample sizes and what the data cannot separate.
Price a month of work from recorded token mixes at each vendor’s list price, with and without prompt caching.
Each rate shows n and a 95% Wilson interval. Timings show the median and the observed range.
When intervals overlap, we say the data does not separate the systems, even when our own product is in the row.
"What if every call ran on Opus" is a calculation from recorded tokens and list prices, never a run. The page says so.
Every study downloads as JSON and CSV. Failed and capped attempts stay in the data.
Read the task sets: pass rules, controls and recorded outcomes. Want everything at once? Download the full dataset (28 studies, schema agent-public-bench@1).

Seven dated AI updates from October 2–8, 2026, with practical limits for research, subagents, API credits, interfaces and decision models.

Anthropic's Frontier Academy plan points to a production skill gap: ownership, recovery, security and evidence matter after the demo.

Haiku 5.5 is positioned for focused coding subagents. Here is a practical lead-worker workflow with evidence, escalation and review.
Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.