{"i":13,"study":{"slug":"cost-thought-experiments","title":"What if every call ran on Opus? Repricing real agent tokens","seoTitle":"Agent token costs repriced: Haiku vs Sonnet vs Opus vs Fable","description":"Thought experiments on real tokens: the same SWE-bench agent work priced at Haiku, Sonnet, Opus, Fable, Gemini Flash, GPT and Jev list prices.","question":"Agent recorded every token it used on 33 SWE-bench instances. What would the same tokens cost at other models’ list prices, and what did caching save?","answer":"Calculation, not a run. Agent used 162.9M input tokens (94.0% cache reads) and 1.8M output tokens across 33 attempts. At Sonnet 5.5 list prices that is $87.23 ($3.49 per resolved instance); the platform's own notional figure, which also counts compaction calls, is $92.64. The same tokens at Opus 5.5 prices cost $143.83, at Fable 5.1 $321.31 and at Haiku 4.5 $43.61. Without prompt caching the Sonnet bill would be $343.33. A different model would have used different tokens and resolved a different set, so these figures bound price sensitivity; they do not predict outcomes.","date":"2026-10-05","updated":"2026-10-05","tags":["thought-experiment","llm-pricing","prompt-caching","opus","sonnet","haiku","jev"],"caveats":["Every repriced figure is a calculation, not a run. Only the Sonnet 5.5 row matches the model that produced the tokens.","Different models use different numbers of calls, tokens and cache hits, and they resolve different instances. Use these figures for price sensitivity only.","Jev is a routing model. Pricing coding tokens at Jev rates shows a floor, not a feasible configuration.","The router-overhead chart uses assumed decision prompt sizes.","Recorded costs are list-price estimates for subscription calls; no invoice backs them."],"sourceIds":["calc-repricing","agent-swebench-c1","agent-swebench-c2","swebench-leaderboard","price-anthropic","price-google","price-openai","price-jev"],"stats":{"$k":["id","label","value","unit","display","n"],"$r":[["tokens-input","Input tokens recorded",162861253,"tokens","162.9M",33],["tokens-output","Output tokens recorded",1760652,"tokens","1.8M",33],["cache-read-share","Share of input served from cache",0.9401,"rate","94.0%",33],["sonnet-repriced","Recorded tokens at Sonnet 5.5 list price",87.23,"usd","$87.23",33],["opus-repriced","Same tokens at Opus 5.5 list price (calculation)",143.83,"usd","$143.83",33],["haiku-repriced","Same tokens at Haiku 4.5 list price (calculation)",43.61,"usd","$43.61",33],["no-cache-sonnet","Sonnet 5.5 without caching (calculation)",343.33,"usd","$343.33",33],["panel-cost-per-resolved-mean","Public panel mean cost per resolved instance (recorded)",0.569,"usd","$0.57",11]]},"charts":{"$k":["id","title","subtitle","kind","unit","yLabel","series","note","sourceIds"],"$r":[["repriced-cost-per-resolved","Thought experiment: the same tokens at other list prices","Cost per resolved SWE-bench instance if 162.9M input and 1.8M output tokens had been billed at each model's list price","bar","usd","USD per resolved instance",[{"name":"Repriced cost per resolved instance","points":{"$k":["label","value","highlight"],"$r":[["Claude Fable 5.1",12.852,false],["Claude Opus 5",8.723,false],["Claude Opus 5.5",5.753,false],["Claude Sonnet 5.5",3.489,true],["GPT-6.1 Sol",2.097,false],["Claude Haiku 4.5",1.745,false],["Gemini 3.x Flash",1.016,false],["Jev 1.13 (router)",0.042,false]]}}],"Calculation, not a run: tokens recorded by Agent on claude-sonnet-5-5 (33 attempts, 25 resolved) times list prices effective 2026-09-21. Another model would use a different number of tokens and resolve a different set. Jev is a routing model and cannot do this work; its bar is a price floor only.",["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic","price-google","price-openai","price-jev"]],["cost-per-resolved-agent-vs-panel","Recorded cost per resolved instance: Agent vs the public panel","Same 33 SWE-bench Verified instances; all attempts in the numerator","bar","usd","USD per resolved instance",[{"name":"Cost per resolved instance","points":{"$k":["label","value","n","highlight"],"$r":[["Agent (notional)",3.706,25,true],["Claude 4.5 Opus (high)",1.184,24,"\u0001"],["Claude 4.5 Sonnet (high)",0.913,25,"\u0001"],["Claude 4.6 Opus",0.875,23,"\u0001"],["GLM 5 (high)",0.667,26,"\u0001"],["DeepSeek V3.2 (high)",0.637,24,"\u0001"],["GPT 5.2 (high)",0.628,28,"\u0001"],["Claude 4.5 Haiku (high)",0.479,25,"\u0001"],["Gemini 3 Flash (high)",0.436,27,"\u0001"],["Kimi K2.5 (high)",0.256,23,"\u0001"],["MiniMax M2.5 (high)",0.107,23,"\u0001"],["GPT 5 mini",0.08,21,"\u0001"]]}}],"Recorded figures, not repricing. Panel costs are published API costs for a bash-only agent. Agent's figure is a list-price estimate of subscription calls and includes onboarding, planning, verification and review.",["agent-swebench-c1","agent-swebench-c2","swebench-leaderboard"]],["prompt-cache-savings","Thought experiment: what prompt caching saved","The same recorded tokens with and without cache pricing","grouped-bar","usd","USD for all attempts",[{"name":"With caching (as recorded)","points":{"$k":["label","value","highlight"],"$r":[["Claude Haiku 4.5",43.61,false],["Claude Sonnet 5.5",87.23,true],["Claude Opus 5.5",143.83,false],["Claude Fable 5.1",321.31,false]]}},{"name":"Without caching","points":{"$k":["label","value"],"$r":[["Claude Haiku 4.5",171.66],["Claude Sonnet 5.5",343.33],["Claude Opus 5.5",686.66],["Claude Fable 5.1",1716.65]]}}],"94.0% of recorded input tokens were cache reads. Calculation, not a run.",["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic"]],["token-cost-mix","Where the token dollars go","Recorded tokens at Sonnet 5.5 list price, by token kind","bar","usd","USD for all attempts",[{"name":"Cost","points":{"$k":["label","value"],"$r":[["Cache writes (1 h)",38.99],["Cache reads",30.62],["Output",17.61],["Uncached input",0.01]]}}],"List-price calculation on recorded tokens: 153.1M cache reads, 9.7M cache writes, 1.8M output, 3.2k uncached input. An agent loop re-reads its context on every call, so cache reads dominate the token count; by price the largest part is cache writes (1 h).",["calc-repricing","agent-swebench-c1","agent-swebench-c2","price-anthropic"]],["router-overhead-jev","Thought experiment: the price of a routing decision on every call","1,632 model calls; one Jev decision per call at an assumed prompt size","bar","usd","USD for all attempts",[{"name":"Jev decision cost","points":{"$k":["label","value"],"$r":[["1k-token decision prompt",0.0685],["2k-token decision prompt",0.1371],["5k-token decision prompt",0.3427]]}}],"Assumption-based calculation: the decision prompt sizes are assumptions, not measurements. For scale, the recorded work cost $92.64 (notional). A router only pays off if its choices save more than this.",["calc-repricing","price-jev","agent-swebench-c1","agent-swebench-c2"]]]}}}