AI week, October 2–8, 2026: research, workers and interfaces
Seven dated AI updates from October 2–8, 2026, with practical limits for research, subagents, API credits, interfaces and decision models.

Practical agent workflows, model release notes and reproducible benchmarks. Follow the original sources, measured results and limits behind each article.
Seven dated AI updates from October 2–8, 2026, with practical limits for research, subagents, API credits, interfaces and decision models.

Anthropic's Frontier Academy plan points to a production skill gap: ownership, recovery, security and evidence matter after the demo.
Haiku 5.5 is positioned for focused coding subagents. Here is a practical lead-worker workflow with evidence, escalation and review.
Eligible Claude Max plans may include monthly Claude Platform API credits. Learn the amounts, claim path, expiry and interactive-limit boundary.
A personal Claude-versus-Codex choice, with practical route comparisons that avoid universal benchmark claims.
A practical way to track AI allowances, local activity and reset dates when several agents run across a laptop and a remote Mac.
OpenAI's Intelligent UI points to answers with charts and controls. A Tiersel build shows why context, computation and evidence must stay separate.
Liquid AI's open-weight d1-3B targets structured decisions in one forward pass. Learn where a small decision model fits and what the latency means.
A real Mac Studio memory warning stopped parallel AI work. The dialog reported 625.49 GB for two apps; here is what the screenshots prove.
The OpenAI math repository lists 719 manuscripts and 372 result families. Its formalizations, withdrawals and history show how evidence changes.
136.5 ms for Jev, 0.82 s for a small-model API, 2.79 s to 3.79 s for Codex CLI: which steps fit a voice agent latency budget? A thought experiment.
Rules and Jev 1.13 fit every budget we assumed; a Claude router through a CLI fits none. 14 measured steps vs 300, 800 and 1,500 ms. A thought experiment.
12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.
AI coding cost per developer: a monthly budget formula using recorded tokens and list-price calculations. Replace our task volume and pass rate with yours.
Best LLM for JSON output? On one prompt, Sonnet and GPT-6.1 Sol passed 10/10 (95%: 72–100%); Haiku passed 1/10 (2–40%). CLI results, not JSON mode.
$0.087 per Claude Code coding session at API list prices, $0.0036 per short call, $2.81 per full agent attempt. Calculations on recorded tokens, not bills.
Claude Code Stop hook: 60/60 Sonnet code-rule checks (95% interval 94–100%). Median input was 1.6x no memory (calculation). Limits and raw data.
24 of 24: Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 each passed every hard task. Fable cost 3.3x Opus and 6.5x Sonnet per pass (list-price calculation).
Haiku 4.5 vs Sonnet 5.5 on 80 rows: Sonnet ahead on 14, Haiku on none, 31 ties. Hard tasks 11/24 vs 24/24, plus speed, memory and price.
60 timed calls: time to first text for Claude Haiku, Sonnet, Opus, Fable, GPT-6.1 Sol and Luna, their output speed, and what a 64k prompt adds.
Seven decision models played 1,620 games. Only Jev and Clef-Flash clearly beat random; speed alone won nothing.
Devin reports $0.60 per task. Our list-price calculation gives $3.71 per resolved SWE-bench task. Compare the units with a table and buyer checklist.
96 calls on 3 JSON tasks. Haiku 4.5 passed 0/24 with instructions and 18/24 with a JSON schema. Sonnet 5.5 and GPT-6.1 Sol passed 12/12 either way.
Claude Code showed near-full cache reads in later fixed-folder sessions (4/4; 95% interval 51–100%). New folders: 0/2 (0–66%). Exploratory, 30 calls.
Routing would save 2.8% ($3.01) on 2,362 recorded calls (a calculation). A Sonnet router on every call costs about $11.80, so the net is a loss.
5 of 13 failed calls on our hard set were right answers in the wrong format. How to grade LLM output strictly, test validators first and report both numbers.
GPT-6.1 Sol vs Claude Opus and Sonnet on four selected harder tasks. Pass intervals overlap for these models; only Sol vs Haiku separates on strict passes.
Of 70 comparison rows for GPT-6.1 Sol, Claude Sonnet 5.5 and Opus 5.5, only 2 have a winner (speed). All 15 pass-rate rows tie. Tokens, price and route differ.
Jev router latency over direct HTTPS: median 136.5 ms, p95 195.7 ms, range 100.9–297.3 ms, n = 246. Claude routers used a different CLI route.
Agent took a median 9.6 minutes per SWE-bench attempt (n = 33, range 1.6 to 54.1). Compare single calls, repairs and full tasks with limits.
Calculation: 100 runs can separate an observed 8-point gap at a perfect score. See Wilson intervals, paired tests and limits for LLM eval sample sizes.
Thinking tokens: 46% to 92% of output per median call, 9% to 80% of pooled list-price cost. Calculation over recorded Claude and GPT-6.1 Sol calls.
Set MAX_THINKING_TOKENS=0, then check both counters. Our Haiku 4.5 test covers routing and hard tasks, with timings, costs and uncertainty.
Haiku 4.5 first, one retry, then Sonnet 5.5? On 8 hard tasks, the calculation gives 4.3x the cost and 8.7x the time per correct answer versus Sonnet alone.
Haiku 4.5 lists at half the price of Sonnet 5.5, yet calculated cost per correct hard answer was 4.7x higher. Retry maths, two exceptions and every source.
The same 150 ms for every answer: Jev ranks 2 of 9 in Tron and Snake, 3 of 8 in Pong, 9 of 9 in Othello.
21 LLM API prices per million tokens, October 2026: Claude, GPT, Gemini and more. Blended prices span 100x. Price per token is not price per task.
44 of 1,196 AI model comparison rows show a gap under our overlap rules. Most gaps are timing rows. No effort row separates quality.
Opus 5.5 at low effort and Sonnet 5.5 at high effort both passed 16/16. Opus low cost 1.3x as much per pass (calculation). Sonnet at low effort cost least.
4.30 s p95 against a 2.60 s median for Claude Sonnet 5.5; 34.5 s against 12.5 s for Haiku 4.5. Measured LLM tail latency and what to do about it.
Cost calculations for 100,000 routing decisions a day: Jev $3.37, Sonnet $732.40, Haiku $892.40. Tested on 82 cases; not general classification.
7 levers that may cut an AI coding agent's bill, sized from our data: prompt cache 3.9x, Fable/Sonnet cost per pass 6.5x, and 5 more. List-price calculations.
Agent loop cost on 8 hard tasks: 1.3–8.7 times the median tokens and up to 2.1 times the price per strict pass (calculations), with no clear pass-rate gain.
What 10,000 to 10 million routing decisions a day cost with rules, Jev and Claude routers, how many run at once and how long they make requests wait.
4 models tied at the top of our hard coding set: Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol. Only Haiku 4.5 separated. A tier list built from intervals.
Haiku 4.5 spent a median 4,556 reasoning tokens per hard call, Sonnet 5.5 585, GPT-6.1 Sol 150. List price per 1,000 calls: $22.78, $5.85, $1.50 (calculation).
A 1-hour Anthropic cache write pays back after 2 reuses; a 5-minute write (assumed 1.25x) after 1. Break-even by model, cost per 1,000 sessions.
1.94 s was the lowest median on short calls (Fable 5.1). 7.75 s on hard calls (Sonnet 5.5). Haiku 4.5 took 4.43 s and 39.01 s with default thinking.
Which Claude model fits your job? Sonnet passed 24/24 hard calls (95% interval 86%–100%) at $0.01435 per pass (calculation). The top models hit a ceiling.
Only 5 to 7 of 82 routing cases split Claude and Jev. The exact McNemar test uses just those cases: p = 0.375 and p = 1. The method, with the arithmetic.
32 AI coding agent misses from our own runs, sorted into 8 classes: wrong answers, gates, caps, lost context and more. Counts, not rates.
A one-word Claude Code call took a median 2.5 s, with 1.7 s outside the model (n = 5). Where the rest goes: thinking, effort, route. Measured splits and fixes.
Calculation on 90 recorded calls: a vote over 10 Haiku replies returns 289, not 282. Repeated errors and format misses limit self-consistency voting.
6 of 6 Haiku 4.5 sessions invented a late-fee rate and reported done. Sonnet 5.5 flagged the gap in 9 of 9. Claimed vs verified, with small n.
Sixteen AI benchmark studies on one page: SWE-bench, Claude Code vs Codex CLI, agent memory, effort, caching, routing, decision models and provider prices.
Our leaderboard lists 46 models, CLIs, routers and providers with their best-supported facts, each with n and an interval. No single score. Here is why.
36 hidden-test sessions: Sonnet 5.5 and Opus 5.5 in Claude Code, GPT-6.1 Sol in Codex CLI. All passed. Time, tool calls, diffs and one confound differ.
Codex CLI sent a median 12,124 input tokens per call; Claude Code sent 2,130. What a coding CLI's hidden context costs, and how much time the route adds.
Sonnet 5.5 and Opus 5.5 tied on every quality test we ran, easy, hard and agentic. Opus cost 1.6x to 2.6x per unit of work. Where the gap comes from.
GPT-6.1 Sol in the Codex CLI passed 16/16 hard tasks at medium and high effort. Sonnet, Opus and Fable passed 24/24. What separates them: time and cost.
200 graded Claude Code sessions: no memory, /init, curated, raw notes, dreamed notes, a handbook and a Stop hook. Memory mattered where the repo was silent.
Claude Haiku 4.5, thinking off: 2.7x faster (a calculation from two medians), no routing accuracy gap shown. Hard tasks: 4/24 passed, thinking on 11/24.
Does tool use improve LLM accuracy? 8 hard tasks, 120 attempts (118 scored): single call vs agent loop for Haiku, Sonnet and GPT-6 Luna. No clear gain.
176 calls on 8 hard tasks: Sonnet, Opus and GPT-6.1 Sol at low, medium and high effort. Every cell passed 16/16. Higher effort cost more tokens and money.
Claude Code read 97% of later-turn input from the cache. At list price that halved a 5-turn session and cut an agent bill about 3.9x. Turn 1 costs more.
Estimate AI coding costs from recorded token mixes: an agent task at $2.64, a hard call at $0.014, a routing decision at $0.005. Calculations, limits stated.
Every row of our Jev, Haiku and Sonnet router comparisons: accuracy ties, Jev is far cheaper and quicker per call over its API, on a different route.
56 new routing decisions, written blind and frozen before any router call. Jev 1.13 vs Claude Haiku 4.5 vs Sonnet 5.5: accuracy, 95% intervals, cost, speed.
Seven decision models, 1,085 checkable decisions. The hosted model led; size bought less than you think.
OpenRouter charged the vendor's per-token price on 10 of 10 models we checked. The cost is a 5.5% credit fee. What we know, and what is still unmeasured.
Interim: 3 of 8 paired SWE-bench Verified issues graded. Opus 5.5 resolved 2, Sonnet 5.5 1; McNemar p = 1.0. Opus cost 2.6x, a calculation. Limits first.
3 prompts, 10 runs each, on Haiku, Sonnet and GPT-6.1 Sol. 7 of 9 cells passed 10/10. Haiku gave the same wrong number 10 times: consistent is not correct.
DeepSeek V4, gpt-oss-120b, Llama, Kimi K3 and GLM 5.3 across 52 providers. Prices differ up to 12.6x. Snapshot 2026-10-06, with the caveats.
A rule-based router decides in 1.42 µs for $0, Jev in 136.5 ms, Sonnet in 2.60 s. What each adds per 1,000 tasks, in delay and in dollars.
120 calls on 8 hard tasks with strict validators. Sonnet, Opus and Fable passed 24/24; Haiku 11/24. Speed, tokens and cost per pass. Codex results: see update.
194 timed runs. Codex CLI took 3.5x as long as the OpenAI API for a one-line answer and sent 19,551 input tokens instead of 17. What a coding CLI adds.
130 timed calls on five validated tasks. Pass rate, latency, hidden input tokens and cost per passing answer for Claude Code models and Codex CLI.
Blind Claude and GPT critics preferred an AI worker's change over the merged human change on 9 of 12 real tasks, and 6 of 12 on the first try. Caveats inside.
Is it the model or the harness? Our data shows the harness clearly moves speed, tokens and cost. Whether it moves accuracy, our samples cannot yet say.
82 typed routing decisions. Jev, Claude Haiku 4.5 and Sonnet 5.5 tie on accuracy, but Jev costs $0.0337 per 1,000 decisions against $5 to $8.92.
Agent resolved 25 of 33 SWE-bench Verified instances (76%). Eleven public models solved 21 to 28 of the same ones. Why that is a tie, not a win.
We repriced 162.9M recorded agent tokens at Haiku, Sonnet, Opus, Fable, Gemini Flash and GPT prices. A calculation, not a run, with clear limits.
A full agent pipeline spent a notional $3.71 per resolved SWE-bench instance. Where the money went, what caching saved, and the public panel range.
How we benchmark AI agents: declare the protocol first, keep every failure, show intervals, publish our own confounds and label interim results.
Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.