<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Agent benchmarks changelog</title>
    <link>https://agent.sasid.ai/changelog</link>
    <description>Every change to the public Agent benchmarks, by day: new and updated studies, posts, explainers, sources and dataset builds. Also as an RSS feed.</description>
    <language>en-us</language>
    <lastBuildDate>Sat, 10 Oct 2026 12:00:00 GMT</lastBuildDate>
    <atom:link href="https://agent.sasid.ai/changelog/rss.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>New post: AI week, October 2–8, 2026: research, workers and interfaces</title>
      <link>https://agent.sasid.ai/blog/ai-week-october-2-8-2026</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-ai-week-october-2-8-2026-2026-10-10</guid>
      <pubDate>Sat, 10 Oct 2026 12:00:00 GMT</pubDate>
      <description>Seven dated AI updates from October 2–8, 2026, with practical limits for research, subagents, API credits, interfaces and decision models.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Claude Frontier Academy and the hard part of production AI</title>
      <link>https://agent.sasid.ai/blog/claude-frontier-academy-production-ai</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-claude-frontier-academy-production-ai-2026-10-10</guid>
      <pubDate>Sat, 10 Oct 2026 12:00:00 GMT</pubDate>
      <description>Anthropic's Frontier Academy plan points to a production skill gap: ownership, recovery, security and evidence matter after the demo.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Claude Haiku 5.5 subagents: a lead-worker pattern worth testing</title>
      <link>https://agent.sasid.ai/blog/claude-haiku-5-5-subagents</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-claude-haiku-5-5-subagents-2026-10-10</guid>
      <pubDate>Sat, 10 Oct 2026 12:00:00 GMT</pubDate>
      <description>Haiku 5.5 is positioned for focused coding subagents. Here is a practical lead-worker workflow with evidence, escalation and review.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Claude Max included API credits: what the allowance covers</title>
      <link>https://agent.sasid.ai/blog/claude-max-included-api-credits</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-claude-max-included-api-credits-2026-10-10</guid>
      <pubDate>Sat, 10 Oct 2026 12:00:00 GMT</pubDate>
      <description>Eligible Claude Max plans may include monthly Claude Platform API credits. Learn the amounts, claim path, expiry and interactive-limit boundary.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Claude vs Codex for a personal AI workflow: choose the route that holds up</title>
      <link>https://agent.sasid.ai/blog/claude-vs-codex-personal-workflow</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-claude-vs-codex-personal-workflow-2026-10-10</guid>
      <pubDate>Sat, 10 Oct 2026 12:00:00 GMT</pubDate>
      <description>A personal Claude-versus-Codex choice, with practical route comparisons that avoid universal benchmark claims.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: How to measure AI subscriptions when your work spans accounts and machines</title>
      <link>https://agent.sasid.ai/blog/ai-subscription-usage-meter</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-ai-subscription-usage-meter-2026-10-10</guid>
      <pubDate>Sat, 10 Oct 2026 12:00:00 GMT</pubDate>
      <description>A practical way to track AI allowances, local activity and reset dates when several agents run across a laptop and a remote Mac.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Intelligent UI: when the question becomes the interface</title>
      <link>https://agent.sasid.ai/blog/intelligent-ui-houston-hackathon</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-intelligent-ui-houston-hackathon-2026-10-10</guid>
      <pubDate>Sat, 10 Oct 2026 12:00:00 GMT</pubDate>
      <description>OpenAI's Intelligent UI points to answers with charts and controls. A Tiersel build shows why context, computation and evidence must stay separate.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Liquid AI d1 decision models: small choices inside a larger workflow</title>
      <link>https://agent.sasid.ai/blog/liquid-ai-d1-decision-models</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-liquid-ai-d1-decision-models-2026-10-10</guid>
      <pubDate>Sat, 10 Oct 2026 12:00:00 GMT</pubDate>
      <description>Liquid AI's open-weight d1-3B targets structured decisions in one forward pass. Learn where a small decision model fits and what the latency means.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: My 96 GB Mac Studio still hit a memory warning</title>
      <link>https://agent.sasid.ai/blog/mac-studio-96gb-ai-memory</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-mac-studio-96gb-ai-memory-2026-10-10</guid>
      <pubDate>Sat, 10 Oct 2026 12:00:00 GMT</pubDate>
      <description>A real Mac Studio memory warning stopped parallel AI work. The dialog reported 625.49 GB for two apps; here is what the screenshots prove.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: OpenAI's math repository makes verification part of the release</title>
      <link>https://agent.sasid.ai/blog/openai-math-repository-verification</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-openai-math-repository-verification-2026-10-10</guid>
      <pubDate>Sat, 10 Oct 2026 12:00:00 GMT</pubDate>
      <description>The OpenAI math repository lists 719 manuscripts and 372 result families. Its formalizations, withdrawals and history show how evidence changes.</description>
      <category>New post</category>
    </item>
    <item>
      <title>Post updated: Claude Code vs Codex CLI vs the API: latency, time to first token and hidden prompts</title>
      <link>https://agent.sasid.ai/blog/claude-code-vs-codex-cli-vs-api-latency</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-update-claude-code-vs-codex-cli-vs-api-latency-2026-10-08</guid>
      <pubDate>Thu, 08 Oct 2026 12:00:00 GMT</pubDate>
      <description>194 timed runs. Codex CLI took 3.5x as long as the OpenAI API for a one-line answer and sent 19,551 input tokens instead of 17. What a coding CLI adds.</description>
      <category>Post updated</category>
    </item>
    <item>
      <title>Post updated: Claude Sonnet 5.5 vs Opus 5.5: when is Opus worth the price?</title>
      <link>https://agent.sasid.ai/blog/claude-sonnet-vs-opus-when-is-opus-worth-it</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-update-claude-sonnet-vs-opus-when-is-opus-worth-it-2026-10-08</guid>
      <pubDate>Thu, 08 Oct 2026 12:00:00 GMT</pubDate>
      <description>Sonnet 5.5 and Opus 5.5 tied on every quality test we ran, easy, hard and agentic. Opus cost 1.6x to 2.6x per unit of work. Where the gap comes from.</description>
      <category>Post updated</category>
    </item>
    <item>
      <title>Post updated: How much does prompt caching actually save? Measured in Claude Code and Codex</title>
      <link>https://agent.sasid.ai/blog/how-much-does-prompt-caching-save</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-update-how-much-does-prompt-caching-save-2026-10-08</guid>
      <pubDate>Thu, 08 Oct 2026 12:00:00 GMT</pubDate>
      <description>Claude Code read 97% of later-turn input from the cache. At list price that halved a 5-turn session and cut an agent bill about 3.9x. Turn 1 costs more.</description>
      <category>Post updated</category>
    </item>
    <item>
      <title>Dataset rebuilt: 28 studies, 206 charts, 37 sources</title>
      <link>https://agent.sasid.ai/benchmarks/dataset.json</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#dataset-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>The one file every page, story and download reads.</description>
      <category>Dataset rebuilt</category>
    </item>
    <item>
      <title>New study: Does a JSON schema stop format misses? Instructions vs schema mode in Claude Code and Codex CLI</title>
      <link>https://agent.sasid.ai/benchmarks/json-schema-vs-instructions</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#study-json-schema-vs-instructions-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Counted calls in this study (every one counted): 96 (72 Claude Code, 24 Codex CLI) · For the same extraction tasks, does enforcing a JSON schema through the CLI change the strict pass rate, the format misses and the wrong values, compared with asking for JSON in the prompt?</description>
      <category>New study</category>
    </item>
    <item>
      <title>New study: Does a new Claude Code session reuse the prompt cache of an earlier one?</title>
      <link>https://agent.sasid.ai/benchmarks/prompt-cache-across-sessions</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#study-prompt-cache-across-sessions-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Later sessions with at least 50% of turn-1 input cached, A: new folder each time: 0 of 2 (95% interval 0% to 66%) · In the earlier caching study, a new Claude Code session did not read the cache that an earlier session wrote. Does a fixed working folder change that, and does putting the ledger in the system prompt help?</description>
      <category>New study</category>
    </item>
    <item>
      <title>New study: GPT-6.1 Sol vs Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 on 4 harder tasks</title>
      <link>https://agent.sasid.ai/benchmarks/harder-tasks-head-to-head</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#study-harder-tasks-head-to-head-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Counted calls that passed strictly (harder set): 39% (22/56) · On a task set built so that Claude Sonnet 5.5 did not pass it every time, does pass rate separate GPT-6.1 Sol (Codex CLI) from Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 (Claude Code)?</description>
      <category>New study</category>
    </item>
    <item>
      <title>New study: Is Claude Haiku cheaper? Retry and escalate, calculated on real receipts</title>
      <link>https://agent.sasid.ai/benchmarks/haiku-retry-or-escalate</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#study-haiku-retry-or-escalate-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Strict passes, Claude Haiku 4.5 · Claude Code, eight hard tasks: 46% (11/24) · On the 8 hard tasks, what does one correct answer cost, and how long does it take, if you try Claude Haiku 4.5 first and retry or escalate, against Claude Sonnet 5.5 every time?</description>
      <category>New study</category>
    </item>
    <item>
      <title>New study: Where the seconds go: first text, output speed and prompt size for 6 LLMs</title>
      <link>https://agent.sasid.ai/benchmarks/llm-speed-anatomy</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#study-llm-speed-anatomy-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>The same 4,327-character reply: most tokens ÷ fewest tokens across models (calculation): 1.8x · n = 23 · For 6 models run through their own coding CLIs (Haiku, Sonnet, Opus, Fable, Sol (low) and Luna (low)), how long until the first text, how fast does text stream after it, and what does a longer prompt add?</description>
      <category>New study</category>
    </item>
    <item>
      <title>New post: A latency budget for voice agents: which LLM steps fit in one turn?</title>
      <link>https://agent.sasid.ai/blog/voice-agent-latency-budget-llm-steps</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-voice-agent-latency-budget-llm-steps-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>136.5 ms for Jev, 0.82 s for a small-model API, 2.79 s to 3.79 s for Codex CLI: which steps fit a voice agent latency budget? A thought experiment.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: A voice agent latency budget, with measured times: what fits in one turn?</title>
      <link>https://agent.sasid.ai/blog/voice-agent-latency-budget-what-fits-in-one-turn</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-voice-agent-latency-budget-what-fits-in-one-turn-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Rules and Jev 1.13 fit every budget we assumed; a Claude router through a CLI fits none. 14 measured steps vs 300, 800 and 1,500 ms. A thought experiment.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: AI coding agent best practices: 12 rules, each backed by a measurement</title>
      <link>https://agent.sasid.ai/blog/ai-coding-agent-best-practices-backed-by-data</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-ai-coding-agent-best-practices-backed-by-data-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>12 rules for running AI coding agents, each with one measured number: validation, model choice, effort, caching, memory, routing, CLIs and sample size.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: AI coding cost per developer: a formula built on recorded work</title>
      <link>https://agent.sasid.ai/blog/ai-coding-cost-per-developer-formula</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-ai-coding-cost-per-developer-formula-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>AI coding cost per developer: a monthly budget formula using recorded tokens and list-price calculations. Replace our task volume and pass rate with yours.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Best LLM for JSON output? Haiku, Sonnet and GPT-6.1 Sol, 10 runs each</title>
      <link>https://agent.sasid.ai/blog/best-llm-for-json-output-haiku-sonnet-gpt-6-1-sol</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-best-llm-for-json-output-haiku-sonnet-gpt-6-1-sol-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Best LLM for JSON output? On one prompt, Sonnet and GPT-6.1 Sol passed 10/10 (95%: 72–100%); Haiku passed 1/10 (2–40%). CLI results, not JSON mode.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Claude Code cost per task: a price ladder from one decision to one agent run</title>
      <link>https://agent.sasid.ai/blog/claude-code-cost-per-task-at-api-prices</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-claude-code-cost-per-task-at-api-prices-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>$0.087 per Claude Code coding session at API list prices, $0.0036 per short call, $2.81 per full agent attempt. Calculations on recorded tokens, not bills.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Claude Code Stop hook, tested: it enforced the rules, used 1.6x the median input and lacked a fact</title>
      <link>https://agent.sasid.ai/blog/claude-code-stop-hook-tested</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-claude-code-stop-hook-tested-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Claude Code Stop hook: 60/60 Sonnet code-rule checks (95% interval 94–100%). Median input was 1.6x no memory (calculation). Limits and raw data.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Claude Fable 5.1 vs Opus 5.5 vs Sonnet 5.5: speed, tokens and price tested</title>
      <link>https://agent.sasid.ai/blog/claude-fable-5-1-vs-opus-5-5-vs-sonnet-5-5</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-claude-fable-5-1-vs-opus-5-5-vs-sonnet-5-5-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>24 of 24: Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 each passed every hard task. Fable cost 3.3x Opus and 6.5x Sonnet per pass (list-price calculation).</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Claude Haiku 4.5 vs Sonnet 5.5: all 80 comparison rows, and where the small model loses</title>
      <link>https://agent.sasid.ai/blog/claude-haiku-4-5-vs-sonnet-5-5-every-measured-row</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-claude-haiku-4-5-vs-sonnet-5-5-every-measured-row-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Haiku 4.5 vs Sonnet 5.5 on 80 rows: Sonnet ahead on 14, Haiku on none, 31 ties. Hard tasks 11/24 vs 24/24, plus speed, memory and price.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Claude tokens per second and time to first token: six models timed, and what a longer prompt adds</title>
      <link>https://agent.sasid.ai/blog/llm-time-to-first-token-and-tokens-per-second</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-llm-time-to-first-token-and-tokens-per-second-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>60 timed calls: time to first text for Claude Haiku, Sonnet, Opus, Fable, GPT-6.1 Sol and Luna, their output speed, and what a 64k prompt adds.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Decision models play each other: Jev vs Clef at Connect Four, Nim and Pong</title>
      <link>https://agent.sasid.ai/blog/decision-models-play-connect-four-nim-pong</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-decision-models-play-connect-four-nim-pong-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Seven decision models played 1,620 games. Only Jev and Clef-Flash clearly beat random; speed alone won nothing.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Devin's $0.60 per task and our $3.71 per resolved task are different numbers</title>
      <link>https://agent.sasid.ai/blog/ai-agent-cost-per-task-claims-vs-receipts</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-ai-agent-cost-per-task-claims-vs-receipts-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Devin reports $0.60 per task. Our list-price calculation gives $3.71 per resolved SWE-bench task. Compare the units with a table and buyer checklist.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Does a JSON schema stop format misses? Haiku went from 0/24 to 18/24</title>
      <link>https://agent.sasid.ai/blog/json-schema-vs-prompt-instructions-format-misses</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-json-schema-vs-prompt-instructions-format-misses-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>96 calls on 3 JSON tasks. Haiku 4.5 passed 0/24 with instructions and 18/24 with a JSON schema. Sonnet 5.5 and GPT-6.1 Sol passed 12/12 either way.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Does Claude Code reuse the prompt cache across sessions? A fixed-folder test</title>
      <link>https://agent.sasid.ai/blog/does-claude-code-reuse-prompt-cache-across-sessions</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-does-claude-code-reuse-prompt-cache-across-sessions-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Claude Code showed near-full cache reads in later fixed-folder sessions (4/4; 95% interval 51–100%). New folders: 0/2 (0–66%). Exploratory, 30 calls.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Does LLM routing save money? The saving, the router and the net</title>
      <link>https://agent.sasid.ai/blog/does-llm-routing-save-money-net-of-router-cost</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-does-llm-routing-save-money-net-of-router-cost-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Routing would save 2.8% ($3.01) on 2,362 recorded calls (a calculation). A Sonnet router on every call costs about $11.80, so the net is a loss.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Format misses vs wrong answers: your LLM eval may be failing right answers</title>
      <link>https://agent.sasid.ai/blog/llm-eval-format-misses-vs-wrong-answers</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-llm-eval-format-misses-vs-wrong-answers-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>5 of 13 failed calls on our hard set were right answers in the wrong format. How to grade LLM output strictly, test validators first and report both numbers.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: GPT-6.1 Sol vs Claude Opus 5.5 and Sonnet 5.5 on harder tasks: only one strict pair separates</title>
      <link>https://agent.sasid.ai/blog/harder-coding-tasks-gpt-6-1-sol-vs-claude-sonnet-vs-opus</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-harder-coding-tasks-gpt-6-1-sol-vs-claude-sonnet-vs-opus-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>GPT-6.1 Sol vs Claude Opus and Sonnet on four selected harder tasks. Pass intervals overlap for these models; only Sol vs Haiku separates on strict passes.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: GPT-6.1 Sol vs Claude Sonnet 5.5 vs Opus 5.5: every row we measured</title>
      <link>https://agent.sasid.ai/blog/gpt-6-1-sol-vs-claude-sonnet-5-5-vs-opus-5-5</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-gpt-6-1-sol-vs-claude-sonnet-5-5-vs-opus-5-5-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Of 70 comparison rows for GPT-6.1 Sol, Claude Sonnet 5.5 and Opus 5.5, only 2 have a winner (speed). All 15 pass-rate rows tie. Tokens, price and route differ.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: How fast is Jev? 136.5 ms per routing decision, measured live</title>
      <link>https://agent.sasid.ai/blog/how-fast-is-jev-router-latency-measured</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-how-fast-is-jev-router-latency-measured-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Jev router latency over direct HTTPS: median 136.5 ms, p95 195.7 ms, range 100.9–297.3 ms, n = 246. Claude routers used a different CLI route.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: How long does an AI coding agent take per task? Minutes, calls and where the time goes</title>
      <link>https://agent.sasid.ai/blog/how-long-does-an-ai-coding-agent-take-per-task</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-how-long-does-an-ai-coding-agent-take-per-task-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Agent took a median 9.6 minutes per SWE-bench attempt (n = 33, range 1.6 to 54.1). Compare single calls, repairs and full tasks with limits.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: How many runs do you need to compare two AI models? A sample-size table</title>
      <link>https://agent.sasid.ai/blog/how-many-runs-to-compare-two-llms</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-how-many-runs-to-compare-two-llms-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Calculation: 100 runs can separate an observed 8-point gap at a perfect score. See Wilson intervals, paired tests and limits for LLM eval sample sizes.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: How much of your AI bill is thinking tokens? Claude and GPT-6.1 Sol, measured</title>
      <link>https://agent.sasid.ai/blog/how-much-of-your-ai-bill-is-thinking-tokens</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-how-much-of-your-ai-bill-is-thinking-tokens-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Thinking tokens: 46% to 92% of output per median call, 9% to 80% of pooled list-price cost. Calculation over recorded Claude and GPT-6.1 Sol calls.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: How to turn off extended thinking in Claude Code, and what we measured for Haiku 4.5</title>
      <link>https://agent.sasid.ai/blog/haiku-thinking-on-off</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-haiku-thinking-on-off-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Set MAX_THINKING_TOKENS=0, then check both counters. Our Haiku 4.5 test covers routing and hard tasks, with timings, costs and uncertainty.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Is Claude Haiku cheaper if you retry or escalate to Sonnet? A calculation on real receipts</title>
      <link>https://agent.sasid.ai/blog/is-claude-haiku-cheaper-retry-and-escalate</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-is-claude-haiku-cheaper-retry-and-escalate-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Haiku 4.5 first, one retry, then Sonnet 5.5? On 8 hard tasks, the calculation gives 4.3x the cost and 8.7x the time per correct answer versus Sonnet alone.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Is Claude Haiku cheaper than Sonnet? Cost per correct answer, with retries</title>
      <link>https://agent.sasid.ai/blog/is-claude-haiku-cheaper-than-sonnet-cost-per-correct-answer</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-is-claude-haiku-cheaper-than-sonnet-cost-per-correct-answer-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Haiku 4.5 lists at half the price of Sonnet 5.5, yet calculated cost per correct hard answer was 4.7x higher. Retry maths, two exceptions and every source.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Jev in fair mode: 2nd of 9 at Tron and Snake, last of 9 at Othello</title>
      <link>https://agent.sasid.ai/blog/decision-models-fair-mode-othello-tron-snake-pong</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-decision-models-fair-mode-othello-tron-snake-pong-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>The same 150 ms for every answer: Jev ranks 2 of 9 in Tron and Snake, 3 of 8 in Pong, 9 of 9 in Othello.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: LLM API pricing comparison, October 2026: Claude vs GPT vs Gemini per million tokens</title>
      <link>https://agent.sasid.ai/blog/llm-api-pricing-comparison-claude-gpt-gemini-october-2026</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-llm-api-pricing-comparison-claude-gpt-gemini-october-2026-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>21 LLM API prices per million tokens, October 2026: Claude, GPT, Gemini and more. Blended prices span 100x. Price per token is not price per task.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Most AI model comparisons are ties: 44 of 1,196 rows show a clear gap</title>
      <link>https://agent.sasid.ai/blog/most-ai-model-comparisons-are-ties</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-most-ai-model-comparisons-are-ties-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>44 of 1,196 AI model comparison rows show a gap under our overlap rules. Most gaps are timing rows. No effort row separates quality.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Opus at low effort or Sonnet at high effort? A bigger model that thinks less, tested</title>
      <link>https://agent.sasid.ai/blog/opus-low-effort-vs-sonnet-high-effort</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-opus-low-effort-vs-sonnet-high-effort-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Opus 5.5 at low effort and Sonnet 5.5 at high effort both passed 16/16. Opus low cost 1.3x as much per pass (calculation). Sonnet at low effort cost least.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Plan for p95, not the median: LLM tail latency in our runs</title>
      <link>https://agent.sasid.ai/blog/llm-tail-latency-p95-not-median</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-llm-tail-latency-p95-not-median-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>4.30 s p95 against a 2.60 s median for Claude Sonnet 5.5; 34.5 s against 12.5 s for Haiku 4.5. Measured LLM tail latency and what to do about it.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: The cheapest LLM for classification: 100,000 decisions a day, and why Haiku cost more than Sonnet</title>
      <link>https://agent.sasid.ai/blog/cheapest-llm-for-classification-at-scale</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-cheapest-llm-for-classification-at-scale-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Cost calculations for 100,000 routing decisions a day: Jev $3.37, Sonnet $732.40, Haiku $892.40. Tested on 82 cases; not general classification.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: The cheapest way to run an AI coding agent: 7 levers from measured runs</title>
      <link>https://agent.sasid.ai/blog/cheapest-way-to-run-an-ai-coding-agent</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-cheapest-way-to-run-an-ai-coding-agent-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>7 levers that may cut an AI coding agent's bill, sized from our data: prompt cache 3.9x, Fable/Sonnet cost per pass 6.5x, and 5 more. List-price calculations.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: What does an agent loop cost? 1.3 to 8.7 times the median tokens (calculation)</title>
      <link>https://agent.sasid.ai/blog/single-call-vs-agent-loop</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-single-call-vs-agent-loop-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Agent loop cost on 8 hard tasks: 1.3–8.7 times the median tokens and up to 2.1 times the price per strict pass (calculations), with no clear pass-rate gain.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: What does routing a million AI requests a day cost? Rules vs Jev vs Claude</title>
      <link>https://agent.sasid.ai/blog/what-does-routing-a-million-ai-requests-a-day-cost</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-what-does-routing-a-million-ai-requests-a-day-cost-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>What 10,000 to 10 million routing decisions a day cost with rules, Jev and Claude routers, how many run at once and how long they make requests wait.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: What is the best AI model for coding? Our data says four models tie</title>
      <link>https://agent.sasid.ai/blog/best-ai-model-for-coding-october-2026</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-best-ai-model-for-coding-october-2026-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>4 models tied at the top of our hard coding set: Sonnet 5.5, Opus 5.5, Fable 5.1 and GPT-6.1 Sol. Only Haiku 4.5 separated. A tier list built from intervals.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: What thinking costs: reasoning tokens per call for Claude and GPT-6.1 Sol</title>
      <link>https://agent.sasid.ai/blog/reasoning-tokens-cost-claude-gpt-6-1-sol</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-reasoning-tokens-cost-claude-gpt-6-1-sol-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Haiku 4.5 spent a median 4,556 reasoning tokens per hard call, Sonnet 5.5 585, GPT-6.1 Sol 150. List price per 1,000 calls: $22.78, $5.85, $1.50 (calculation).</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: When does prompt caching pay off? The Anthropic cache break-even, calculated</title>
      <link>https://agent.sasid.ai/blog/when-does-prompt-caching-pay-off</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-when-does-prompt-caching-pay-off-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>A 1-hour Anthropic cache write pays back after 2 reuses; a 5-minute write (assumed 1.25x) after 1. Break-even by model, cost per 1,000 sessions.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Which Claude model is fastest? It depends on the task, and on thinking</title>
      <link>https://agent.sasid.ai/blog/fastest-claude-model-haiku-sonnet-opus-fable-timed</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-fastest-claude-model-haiku-sonnet-opus-fable-timed-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>1.94 s was the lowest median on short calls (Fable 5.1). 7.75 s on hard calls (Sonnet 5.5). Haiku 4.5 took 4.43 s and 39.01 s with default thinking.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Which Claude model should you use? A task-by-task guide from our measurements</title>
      <link>https://agent.sasid.ai/blog/which-claude-model-should-i-use</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-which-claude-model-should-i-use-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Which Claude model fits your job? Sonnet passed 24/24 hard calls (95% interval 86%–100%) at $0.01435 per pass (calculation). The top models hit a ceiling.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Why a paired test: comparing two LLMs on the same cases with McNemar</title>
      <link>https://agent.sasid.ai/blog/mcnemar-test-compare-two-llms</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-mcnemar-test-compare-two-llms-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>Only 5 to 7 of 82 routing cases split Claude and Jev. The exact McNemar test uses just those cases: p = 0.375 and p = 1. The method, with the arithmetic.</description>
      <category>New post</category>
    </item>
    <item>
      <title>New post: Why AI coding agents fail on real pull requests: every failure from 12 attempts</title>
      <link>https://agent.sasid.ai/blog/why-ai-coding-agents-fail-on-real-pull-requests</link>
      <guid isPermaLink="false">https://agent.sasid.ai/changelog#post-why-ai-coding-agents-fail-on-real-pull-requests-2026-10-07</guid>
      <pubDate>Wed, 07 Oct 2026 12:00:00 GMT</pubDate>
      <description>32 AI coding agent misses from our own runs, sorted into 8 classes: wrong answers, gates, caps, lost context and more. Counts, not rates.</description>
      <category>New post</category>
    </item>
  </channel>
</rss>
