{
  "schema": "agent-public-bench@1",
  "generatedAt": "2026-10-07T00:00:00.000Z",
  "url": "https://agent.sasid.ai/benchmarks/routing-at-scale",
  "study": {
    "slug": "routing-at-scale",
    "title": "What does routing a million AI requests a day cost? A calculation from measured runs",
    "seoTitle": "LLM router cost at scale: 1 million decisions a day",
    "description": "A calculation from measured runs: what 10,000 to 10 million routing decisions a day cost with rules, Jev and Claude, with median/p95 time scenarios.",
    "question": "At 10,000 to 10 million routing decisions a day, what does each router cost per day, what in-flight and waiting scenarios do median and p95 times give?",
    "answer": "This is a calculation, not a run. We multiplied measured numbers and made no model call. At 1 million routing decisions a day, the rule-based policy costs $0 and Jev costs $33.70. A Sonnet 5.5 router costs $7,324 and a Haiku 4.5 router costs $8,924. Cost samples: Jev n = 82, Sonnet n = 82, Haiku n = 82. Jev scales reported run cost. Claude uses list prices and reported tokens. These calculations fix the token mix and have no statistical cost bounds. At 10 million decisions a day, Jev costs $337.00, Sonnet $73,240 and Haiku $89,240. At 1 million a day, that is 11.57 decisions per second. At its median time, the Sonnet scenario gives about 30.1 decisions in flight at once (49.7 at its 95th percentile time). At its median time, the Haiku scenario gives about 146.7 (398.3). At its median time, the Jev scenario gives about 1.58 (2.27), timed live over direct HTTPS. If every decision took its median time, Sonnet would add 721.7 hours of waiting a day and Jev would add 37.9 hours. The policy median-time scenario adds 1.42 s. Say the router decides only the System One decisions (14.1% of model calls). The Sonnet cost at 1 million calls a day then falls from $7,324 to $1,036. Time samples: policy n = 20000, Sonnet n = 82, Haiku n = 82, Jev n = 246. These are median/p95 scenarios, not measured totals or confidence intervals. We did not measure rate limits, load or accuracy.",
    "date": "2026-10-06",
    "updated": "2026-10-06",
    "tags": [
      "thought-experiment",
      "calculation",
      "routing",
      "llm-router",
      "jev",
      "latency",
      "cost",
      "scale",
      "littles-law"
    ],
    "method": [
      "This is a calculation, not a run. We made no new model call and ran nothing at scale. Every input is a number from a recorded run in this dataset. The volumes (10,000, 100,000, 1 million, 10 million decisions a day) are parameters, not measurements.",
      "Jev cost per 1,000 is a calculation from the routing study’s provider-reported total cost and counted calls. Claude cost comes from its token totals and list prices, with one-hour cache writes. Jev costs $0.0337 per 1,000, a calculation from the reported production run cost (n = 82). \n\nSonnet 5.5 at effort low and Haiku 4.5 ran through the Claude Code CLI. Their cost is list price × the tokens the CLI reported (Sonnet n = 82, Haiku n = 82). The policy makes no model call, so it costs $0. The live Jev run gives $0.0337 per 1,000, a calculation from reported input tokens and the published price. The two figures agree at this precision.",
      "Claude median and p95 times come from the routing receipts; we checked them against the per-call rows. The overhead summary uses different quantiles for those calls, so we do not mix its median or p95 into this calculation. Run ranges and policy times come from the overhead results. The policy is the production routing decision, timed in process over 20,000 decisions (median 1.42 µs, 95th percentile 2.33 µs, maximum 2538.21 µs; minimum unavailable). \n\nSonnet 5.5 and Haiku 4.5 are wall time per call through the CLI, one call at a time, from the recorded routing runs (Sonnet n = 82, Haiku n = 82). Jev is the live run: 246 calls (82 typed decisions × 3 repetitions) over direct HTTPS from one Mac, one call at a time, as client wall time.",
      "Daily cost = decisions a day × cost per decision. Decisions per second = decisions a day ÷ 86,400. Decisions in flight = decisions per second × time per decision in seconds (median/p95 scenarios; Little’s law requires the mean). The median time gives the bar. The 95th percentile time gives the whisker. \n\nWaiting hours a day = decisions a day × time per decision ÷ 3,600, if each request waits for its decision. Cost per year = cost per day × 365.",
      "The scenarios use 1,000,000 model calls a day. In the first, a router decides every call. In the second, it decides only the System One decisions: 7 ÷ 49.5 = 14.1% of calls. That ratio uses the medians of 48 recorded bench runs. \n\nRouting was off in those runs, so each model call counts as one decision a router could make. In the third, the policy decides every call in process.",
      "Policy capacity is the measured 469,409 decisions per second in one process, set against the rate each volume needs. The measurement leaves out the database reads and the decision-record write of the production decision."
    ],
    "caveats": [
      "The routing and live Jev protocols predate the first counted calls by file birth time. The overhead protocol predates its microbenchmark output. Claude ran 82 calls per router, below its 120-call cap. Jev ran 246 counted calls, below its 300-call cap. One Haiku pilot and one Jev probe are excluded. All counted calls completed without call errors.\n\nThe policy microbenchmark used 5,000 warm-up iterations and 64 synthetic contexts; the minimum and individual timing samples are unavailable, so we cannot rebuild its quantiles or full range. No separate validator-control receipt or Sonnet pilot is retained.",
      "The source accuracy sets hit a ceiling in some purposes. These tuned cases do not establish equal quality on new traffic. The same 82 Jev cases repeat three times; 246 latency calls are not 246 distinct tasks.",
      "Nothing ran at scale. We did not measure rate limits, queueing or slowdown under load. The Claude routers ran one call at a time through a CLI, and the live Jev run also ran one call at a time. The calculation does not show that a provider accepts this many decisions a day, or this many at once.",
      "The routes differ. The policy runs in process. We timed Jev over direct HTTPS from one Mac on a home network, so its time includes the network round trip. The Claude routers ran through the Claude Code CLI. The CLI adds time (a median 975 ms for Sonnet and 1,701 ms for Haiku) and its own prompt to every call. \n\nWe did not time a direct API call, so the Claude rows show the CLI route and not the best a Claude router can do.",
      "The Claude costs are calculations on subscription calls, not invoices. Sonnet's receipts report one-hour cache writes. The original routing estimate used five-minute write rates and gave $5.00 per 1,000. This study uses the one-hour rate and gives $7.32.\n\nThe CLI estimate is $7.32 per 1,000. Mean output tokens: Haiku 1,419, Sonnet 107. Haiku used default thinking; this comparison does not isolate its effect.",
      "Little’s law uses the mean time per decision. The public CLI extracts contain the median and p95, but no mean. Neither value bounds the mean. These are assumed-time scenarios, not average concurrency, actual waiting totals or capacity guarantees. For Jev, the mean (142.1 ms) is 4.1% above the median (calculation). \n\nThe waiting total has the same limit. Each whisker runs from the median calculation to the 95th percentile calculation. It is not a confidence interval.",
      "A steady rate over 24 hours is an assumption. Real traffic has peaks. To size a peak, multiply the in-flight figures by the peak rate ÷ the average rate.",
      "We did not score accuracy here. The policy picks a model and effort from typed signals. Jev and the Claude routers answer typed System One decisions. The routing study scores the three model routers on 82 decisions. We tuned the case sets against Jev’s answers, so Jev has a home advantage. A cheap router that is wrong costs more in the work it misroutes, and this calculation does not price that.",
      "The System One share comes from 48 bench runs of one agent. The ratio of medians is 14.1%. Pooled over the runs, it is 17.0%, which raises the System One scenario cost by 20% (calculation). Your traffic may differ. The work cost per call is Agent’s notional list-price cost on Sonnet 5.5, not a bill.",
      "The policy’s 1.42 µs median leaves out database reads and the decision-record write, and it needs a server to run on. The calculation prices model calls only. It does not price servers, retries or engineering time.",
      "The samples are small. Each Claude router has 82 decisions (one sample per case). The Jev cost has 82, the live Jev time has 246 calls and the policy time has 20,000. A cost per decision is one estimate with no interval. The projection fixes that estimate and the token mix; it has no statistical cost bounds.",
      "We did not isolate the Mac that ran the timed calls. Other jobs, such as local model servers, may have run at the same time, so a wall time may include contention."
    ],
    "sourceIds": [
      "calc-routing-at-scale",
      "agent-routing-overhead",
      "agent-routing",
      "price-anthropic",
      "price-jev",
      "agent-jev-live"
    ],
    "stats": [
      {
        "id": "routing-at-scale-cost-per-1000-jev",
        "label": "Cost per 1,000 routing decisions, Jev (calculation)",
        "value": 0.0337,
        "unit": "usd",
        "display": "$0.0337",
        "n": 82,
        "note": "Calculation: provider-reported total cost ÷ counted calls × 1,000. This fixes the recorded token mix; no cost interval is available."
      },
      {
        "id": "routing-at-scale-cost-per-1000-sonnet",
        "label": "Cost per 1,000 routing decisions, Sonnet 5.5 router (calculation)",
        "value": 7.324,
        "unit": "usd",
        "display": "$7.32",
        "n": 82,
        "note": "A list-price calculation on the tokens the CLI reported. The calls ran on a subscription."
      },
      {
        "id": "routing-at-scale-cost-per-1000-haiku",
        "label": "Cost per 1,000 routing decisions, Haiku 4.5 router (calculation)",
        "value": 8.924,
        "unit": "usd",
        "display": "$8.92",
        "n": 82,
        "note": "A list-price calculation on the tokens the CLI reported. The calls ran on a subscription."
      },
      {
        "id": "routing-at-scale-system-one-share",
        "label": "System One decisions as a share of model calls (ratio of the two medians, calculation)",
        "value": 0.1414,
        "unit": "rate",
        "display": "14.1%",
        "n": 48,
        "note": "7 ÷ 49.5. A ratio of medians, not a sample proportion, so it carries no interval."
      },
      {
        "id": "routing-at-scale-rate-1m",
        "label": "Decisions per second at 1 million a day (calculation)",
        "value": 11.5741,
        "unit": "count",
        "display": "11.57 per second",
        "note": "A steady rate over 24 hours."
      },
      {
        "id": "routing-at-scale-cost-1m-policy",
        "label": "Daily cost at 1 million decisions, rule-based policy (calculation)",
        "value": 0,
        "unit": "usd",
        "display": "$0",
        "n": 20000,
        "note": "$0 a year at a flat volume."
      },
      {
        "id": "routing-at-scale-cost-1m-jev",
        "label": "Daily cost at 1 million decisions, Jev (calculation)",
        "value": 33.7,
        "unit": "usd",
        "display": "$33.70",
        "n": 82,
        "note": "$12,301 a year at a flat volume."
      },
      {
        "id": "routing-at-scale-cost-1m-sonnet",
        "label": "Daily cost at 1 million decisions, Sonnet 5.5 router (calculation)",
        "value": 7324,
        "unit": "usd",
        "display": "$7,324",
        "n": 82,
        "note": "$2,673,260 a year at a flat volume."
      },
      {
        "id": "routing-at-scale-cost-1m-haiku",
        "label": "Daily cost at 1 million decisions, Haiku 4.5 router (calculation)",
        "value": 8924,
        "unit": "usd",
        "display": "$8,924",
        "n": 82,
        "note": "$3,257,260 a year at a flat volume."
      },
      {
        "id": "routing-at-scale-flight-1m-jev",
        "label": "Decisions in flight at 1 million a day, Jev (calculation)",
        "value": 1.5799,
        "unit": "calls",
        "display": "1.58 (2.27 at the 95th percentile time)",
        "n": 246,
        "note": "Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."
      },
      {
        "id": "routing-at-scale-flight-1m-sonnet",
        "label": "Decisions in flight at 1 million a day, Sonnet 5.5 router (calculation)",
        "value": 30.069,
        "unit": "calls",
        "display": "30.1 (49.7 at the 95th percentile time)",
        "n": 82,
        "note": "Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."
      },
      {
        "id": "routing-at-scale-flight-1m-haiku",
        "label": "Decisions in flight at 1 million a day, Haiku 4.5 router (calculation)",
        "value": 146.69,
        "unit": "calls",
        "display": "146.7 (398.3 at the 95th percentile time)",
        "n": 82,
        "note": "Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."
      },
      {
        "id": "routing-at-scale-hours-1m-policy",
        "label": "Waiting hours a day at 1 million decisions, rule-based policy (calculation)",
        "value": 0.00039444,
        "unit": "count",
        "display": "1.42 s in total",
        "n": 20000,
        "note": "Decisions × assumed median or p95 time. Actual total waiting requires the mean; these are not confidence bounds."
      },
      {
        "id": "routing-at-scale-hours-1m-jev",
        "label": "Waiting hours a day at 1 million decisions, Jev (calculation)",
        "value": 37.917,
        "unit": "count",
        "display": "37.9 hours (54.4 at the 95th percentile time)",
        "n": 246,
        "note": "Decisions × assumed median or p95 time. Actual total waiting requires the mean; these are not confidence bounds."
      },
      {
        "id": "routing-at-scale-hours-1m-sonnet",
        "label": "Waiting hours a day at 1 million decisions, Sonnet 5.5 router (calculation)",
        "value": 721.67,
        "unit": "count",
        "display": "721.7 hours (1,194 at the 95th percentile time)",
        "n": 82,
        "note": "Decisions × assumed median or p95 time. Actual total waiting requires the mean; these are not confidence bounds."
      }
    ],
    "charts": [
      {
        "id": "routing-at-scale-daily-cost",
        "title": "Daily cost of routing at 10,000 to 10 million decisions a day (calculation)",
        "subtitle": "Decisions a day × cost per decision, USD per day, every decision routed. The rule-based policy costs $0 and is not on the log axis",
        "kind": "grouped-bar",
        "unit": "usd",
        "yLabel": "USD per day",
        "series": [
          {
            "name": "Jev 1.13 (TypeSafe)",
            "points": [
              {
                "label": "10,000 a day",
                "value": 0.337,
                "n": 82
              },
              {
                "label": "100,000 a day",
                "value": 3.37,
                "n": 82
              },
              {
                "label": "1 million a day",
                "value": 33.7,
                "n": 82
              },
              {
                "label": "10 million a day",
                "value": 337,
                "n": 82
              }
            ]
          },
          {
            "name": "Claude Sonnet 5.5 (low) · Claude Code",
            "points": [
              {
                "label": "10,000 a day",
                "value": 73.24,
                "n": 82
              },
              {
                "label": "100,000 a day",
                "value": 732.4,
                "n": 82
              },
              {
                "label": "1 million a day",
                "value": 7324,
                "n": 82
              },
              {
                "label": "10 million a day",
                "value": 73240,
                "n": 82
              }
            ]
          },
          {
            "name": "Claude Haiku 4.5 · Claude Code",
            "points": [
              {
                "label": "10,000 a day",
                "value": 89.24,
                "n": 82
              },
              {
                "label": "100,000 a day",
                "value": 892.4,
                "n": 82
              },
              {
                "label": "1 million a day",
                "value": 8924,
                "n": 82
              },
              {
                "label": "10 million a day",
                "value": 89240,
                "n": 82
              }
            ]
          }
        ],
        "note": "This is a calculation, not a run. Daily cost = decisions a day × cost per decision. Cost per 1,000 decisions: Jev $0.0337 (calculation from provider-reported run cost, n = 82), Sonnet 5.5 $7.32, Haiku 4.5 $8.92. The Claude figures use one-hour cache-write rates and list price × the tokens the CLI reported (Sonnet n = 82, Haiku n = 82). \n\nThe rule-based policy makes no model call, so it costs $0 at every volume and has no bar: a log axis cannot show zero. The chart prices model calls only, not servers, retries or the work itself.",
        "sourceIds": [
          "calc-routing-at-scale",
          "agent-routing-overhead",
          "agent-routing",
          "price-anthropic",
          "price-jev"
        ]
      },
      {
        "id": "routing-at-scale-in-flight",
        "title": "In-flight scenarios at 10,000 to 10 million a day (calculation)",
        "subtitle": "Decisions per second × assumed time per decision; bar = median time, whisker = the same calculation at the 95th percentile. The in-process policy is in the table",
        "kind": "grouped-bar",
        "unit": "calls",
        "yLabel": "In-flight scenarios at median / p95 time",
        "whisker": "p50-p95",
        "series": [
          {
            "name": "Jev 1.13 (TypeSafe)",
            "points": [
              {
                "label": "10,000 a day",
                "value": 0.015799,
                "lo": 0.015799,
                "hi": 0.02265,
                "n": 246
              },
              {
                "label": "100,000 a day",
                "value": 0.15799,
                "lo": 0.15799,
                "hi": 0.2265,
                "n": 246
              },
              {
                "label": "1 million a day",
                "value": 1.5799,
                "lo": 1.5799,
                "hi": 2.265,
                "n": 246
              },
              {
                "label": "10 million a day",
                "value": 15.799,
                "lo": 15.799,
                "hi": 22.65,
                "n": 246
              }
            ]
          },
          {
            "name": "Claude Sonnet 5.5 (low) · Claude Code",
            "points": [
              {
                "label": "10,000 a day",
                "value": 0.30069,
                "lo": 0.30069,
                "hi": 0.49745,
                "n": 82
              },
              {
                "label": "100,000 a day",
                "value": 3.0069,
                "lo": 3.0069,
                "hi": 4.9745,
                "n": 82
              },
              {
                "label": "1 million a day",
                "value": 30.069,
                "lo": 30.069,
                "hi": 49.745,
                "n": 82
              },
              {
                "label": "10 million a day",
                "value": 300.69,
                "lo": 300.69,
                "hi": 497.45,
                "n": 82
              }
            ]
          },
          {
            "name": "Claude Haiku 4.5 · Claude Code",
            "points": [
              {
                "label": "10,000 a day",
                "value": 1.4669,
                "lo": 1.4669,
                "hi": 3.983,
                "n": 82
              },
              {
                "label": "100,000 a day",
                "value": 14.669,
                "lo": 14.669,
                "hi": 39.83,
                "n": 82
              },
              {
                "label": "1 million a day",
                "value": 146.69,
                "lo": 146.69,
                "hi": 398.3,
                "n": 82
              },
              {
                "label": "10 million a day",
                "value": 1466.9,
                "lo": 1466.9,
                "hi": 3983,
                "n": 82
              }
            ]
          }
        ],
        "note": "This is a calculation, not a run. In flight = decisions a day ÷ 86,400 × time per decision, at a steady rate over 24 hours. The bar uses the median time. The whisker uses the 95th percentile time and is not a confidence interval. These are assumed-time scenarios. Actual average concurrency requires the mean time. \n\nThe inputs table lists the times. Jev ran over direct HTTPS and the Claude routers ran through a CLI, so the routes differ. Traffic has peaks, so multiply by the peak rate ÷ the average rate to size a peak. \n\nThe rule-based policy has no row. It runs in process, so at 1 million decisions a day the median-time scenario gives 0.000016 decisions in flight, a small share of one decision and not a connection. Labels round small values. The grid table lists every value.",
        "sourceIds": [
          "calc-routing-at-scale",
          "agent-routing-overhead",
          "agent-routing",
          "agent-jev-live"
        ]
      },
      {
        "id": "routing-at-scale-waiting-hours",
        "title": "Waiting scenarios per day at 1 million decisions a day (calculation)",
        "subtitle": "Decisions × time per decision, in hours of waiting added across all requests; whisker = the same calculation at the 95th percentile",
        "kind": "bar",
        "unit": "count",
        "yLabel": "Waiting hours at assumed median / p95 time",
        "whisker": "p50-p95",
        "series": [
          {
            "name": "Hours of waiting per day",
            "points": [
              {
                "label": "Jev 1.13 (TypeSafe)",
                "value": 37.917,
                "lo": 37.917,
                "hi": 54.361,
                "n": 246
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code",
                "value": 721.67,
                "lo": 721.67,
                "hi": 1193.9,
                "n": 82
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 3520.6,
                "lo": 3520.6,
                "hi": 9559.2,
                "n": 82
              }
            ]
          }
        ],
        "note": "This is a calculation, not a run. Hours of waiting a day = decisions a day × time per decision ÷ 3,600, if each request waits for its decision. The bar uses the median time. The whisker uses the 95th percentile time and is not a confidence interval. Total waiting equals decisions × the mean time. The public extracts do not contain the mean for the CLI routers, so the true total is not known. \n\nThe rule-based policy has no bar: at this volume the median-time scenario gives 1.42 s in total. A decision outside the request’s critical path may add no waiting for that request.",
        "sourceIds": [
          "calc-routing-at-scale",
          "agent-routing-overhead",
          "agent-routing",
          "agent-jev-live"
        ]
      },
      {
        "id": "routing-at-scale-scenarios",
        "title": "Daily cost at 1 million model calls a day: route every call or only System One decisions (calculation)",
        "subtitle": "Every call is one decision; System One decisions are 7 of 49.5 model calls per task in the recorded runs",
        "kind": "grouped-bar",
        "unit": "usd",
        "yLabel": "USD per day",
        "series": [
          {
            "name": "Every model call routed",
            "points": [
              {
                "label": "Deterministic routing policy",
                "value": 0,
                "n": 20000
              },
              {
                "label": "Jev 1.13 (TypeSafe)",
                "value": 33.7,
                "n": 82
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code",
                "value": 7324,
                "n": 82
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 8924,
                "n": 82
              }
            ]
          },
          {
            "name": "Only System One decisions routed (14.1% of calls)",
            "points": [
              {
                "label": "Deterministic routing policy",
                "value": 0,
                "n": 20000
              },
              {
                "label": "Jev 1.13 (TypeSafe)",
                "value": 4.7657,
                "n": 82
              },
              {
                "label": "Claude Sonnet 5.5 (low) · Claude Code",
                "value": 1035.7,
                "n": 82
              },
              {
                "label": "Claude Haiku 4.5 · Claude Code",
                "value": 1262,
                "n": 82
              }
            ]
          }
        ],
        "note": "This is a calculation, not a run. It assumes 1,000,000 model calls a day. Route every call: 1,000,000 decisions. Route only System One decisions: 141,414 decisions, which is 7 ÷ 49.5 = 14.1% of calls. \n\nSystem One decisions are the typed choices an agent asks a router to make. The 7 and the 49.5 are medians over 48 recorded bench runs. \n\nRouting was off in those runs, so each model call counts as one decision a router could make. Pooled over the runs, System One decisions are 17.0% of calls (419 of 2,467). The policy is the third scenario: it decides every call in process at no model cost.",
        "sourceIds": [
          "calc-routing-at-scale",
          "agent-routing-overhead",
          "agent-routing",
          "price-anthropic",
          "price-jev"
        ]
      }
    ],
    "tables": [
      {
        "id": "routing-at-scale-inputs",
        "title": "Recorded inputs and cost calculations (per decision)",
        "columns": [
          {
            "key": "router",
            "label": "Router",
            "unit": "text"
          },
          {
            "key": "costPer1000",
            "label": "Cost per 1,000 decisions",
            "unit": "usd"
          },
          {
            "key": "costBasis",
            "label": "Cost basis",
            "unit": "text"
          },
          {
            "key": "costN",
            "label": "n (cost)",
            "unit": "count"
          },
          {
            "key": "p50",
            "label": "Median time per decision",
            "unit": "ms"
          },
          {
            "key": "p95",
            "label": "95th percentile",
            "unit": "ms"
          },
          {
            "key": "min",
            "label": "Fastest call (range, not an interval)",
            "unit": "ms"
          },
          {
            "key": "max",
            "label": "Slowest call (range, not an interval)",
            "unit": "ms"
          },
          {
            "key": "latN",
            "label": "n (time)",
            "unit": "count"
          },
          {
            "key": "latBasis",
            "label": "Time basis",
            "unit": "text"
          }
        ],
        "rows": [
          {
            "router": "Deterministic routing policy",
            "costPer1000": 0,
            "costBasis": "no model call",
            "costN": 20000,
            "p50": 0.00142,
            "p95": 0.00233,
            "min": null,
            "max": 2.53821,
            "latN": 20000,
            "latBasis": "timed in process (microbenchmark, no database reads)"
          },
          {
            "router": "Jev 1.13 (TypeSafe)",
            "costPer1000": 0.0337,
            "costBasis": "provider-reported run cost ÷ calls × 1,000 (calculation)",
            "costN": 82,
            "p50": 136.5,
            "p95": 195.7,
            "min": 100.9,
            "max": 297.3,
            "latN": 246,
            "latBasis": "live run over direct HTTPS, one call at a time (client wall time)"
          },
          {
            "router": "Claude Sonnet 5.5 (low) · Claude Code",
            "costPer1000": 7.324,
            "costBasis": "list price × CLI-reported tokens (calculation)",
            "costN": 82,
            "p50": 2598,
            "p95": 4298,
            "min": 1993,
            "max": 5583,
            "latN": 82,
            "latBasis": "recorded routing study, through the Claude Code CLI, one call at a time"
          },
          {
            "router": "Claude Haiku 4.5 · Claude Code",
            "costPer1000": 8.924,
            "costBasis": "list price × CLI-reported tokens (calculation)",
            "costN": 82,
            "p50": 12674,
            "p95": 34413,
            "min": 5857,
            "max": 51278,
            "latN": 82,
            "latBasis": "recorded routing study, through the Claude Code CLI, one call at a time (default thinking on)"
          }
        ]
      },
      {
        "id": "routing-at-scale-formulas",
        "title": "The formulas",
        "columns": [
          {
            "key": "quantity",
            "label": "Quantity",
            "unit": "text"
          },
          {
            "key": "formula",
            "label": "Formula",
            "unit": "text"
          },
          {
            "key": "inputs",
            "label": "Inputs",
            "unit": "text"
          }
        ],
        "rows": [
          {
            "quantity": "Cost per day",
            "formula": "decisions a day × cost per 1,000 decisions ÷ 1,000",
            "inputs": "Routing receipts: Jev reported total cost ÷ calls × 1,000 (calculation); Claude tokensPerDecision × list prices, with one-hour cache writes"
          },
          {
            "quantity": "Decisions per second",
            "formula": "decisions a day ÷ 86,400",
            "inputs": "A steady rate over 24 hours"
          },
          {
            "quantity": "Decisions in flight",
            "formula": "decisions per second × assumed median or p95 time per decision in seconds",
            "inputs": "Routing-overhead results: policy latencyUs p50 and p95; Claude routing receipts latencyMs wallP50 and wallP95; live Jev run: latencyMs.calls median and p95"
          },
          {
            "quantity": "Waiting hours a day",
            "formula": "decisions a day × time per decision in seconds ÷ 3,600",
            "inputs": "The same times, if each request waits for its decision"
          },
          {
            "quantity": "System One scenario",
            "formula": "decisions = model calls a day × (System One decisions per task ÷ model calls per task) = × 7 ÷ 49.5",
            "inputs": "Routing-overhead results: perTask.systemOneDecisions.median and perTask.modelCalls.median"
          },
          {
            "quantity": "Policy share of one process",
            "formula": "decisions per second ÷ measured decisions per second",
            "inputs": "Routing-overhead results: policy decisionsPerSecond"
          },
          {
            "quantity": "Cost per year",
            "formula": "cost per day × 365",
            "inputs": "A flat daily volume"
          }
        ]
      },
      {
        "id": "routing-at-scale-context",
        "title": "Counts, inputs and scale checks behind the calculation",
        "columns": [
          {
            "key": "quantity",
            "label": "Quantity",
            "unit": "text"
          },
          {
            "key": "value",
            "label": "Value",
            "unit": "text"
          },
          {
            "key": "n",
            "label": "n",
            "unit": "count"
          },
          {
            "key": "note",
            "label": "Note",
            "unit": "text"
          }
        ],
        "rows": [
          {
            "quantity": "Cost per 1,000 routing decisions, rule-based policy",
            "value": "$0",
            "n": 20000,
            "note": "The policy makes no model call."
          },
          {
            "quantity": "Median time per routing decision, rule-based policy",
            "value": "1.42 µs (p95 2.33 µs); maximum 2538.21 µs (minimum unavailable)",
            "n": 20000,
            "note": "timed in process (microbenchmark, no database reads). The 95th percentile is a point on the time distribution. It is not a confidence interval. The run range, where shown, is not an interval. The minimum and individual timing samples are unavailable; we cannot rebuild the quantiles or full range."
          },
          {
            "quantity": "Median time per routing decision, Jev",
            "value": "136.5 ms (p95 195.7 ms); range 100.9 ms to 297.3 ms",
            "n": 246,
            "note": "live run over direct HTTPS, one call at a time (client wall time). The 95th percentile is a point on the time distribution. It is not a confidence interval. The run range, where shown, is not an interval."
          },
          {
            "quantity": "Median time per routing decision, Sonnet 5.5 router",
            "value": "2.60 s (p95 4.30 s); range 1.99 s to 5.58 s",
            "n": 82,
            "note": "recorded routing study, through the Claude Code CLI, one call at a time. The 95th percentile is a point on the time distribution. It is not a confidence interval. The run range, where shown, is not an interval."
          },
          {
            "quantity": "Median time per routing decision, Haiku 4.5 router",
            "value": "12.67 s (p95 34.41 s); range 5.86 s to 51.28 s",
            "n": 82,
            "note": "recorded routing study, through the Claude Code CLI, one call at a time (default thinking on). The 95th percentile is a point on the time distribution. It is not a confidence interval. The run range, where shown, is not an interval."
          },
          {
            "quantity": "Model calls per task, median of the recorded runs",
            "value": "49.5",
            "n": 48,
            "note": "Each model call counts as one decision a router could make. Recorded range: 13 to 73 calls; not an interval."
          },
          {
            "quantity": "System One decisions per task, median of the recorded runs",
            "value": "7",
            "n": 48,
            "note": "Recorded range: 2 to 24 decisions; not an interval."
          },
          {
            "quantity": "System One decisions as a share of model calls, pooled over the runs (calculation)",
            "value": "17.0% (419 of 2,467)",
            "n": 48,
            "note": "Total System One decisions ÷ total model calls over the runs with at least one call. It differs from the ratio of medians. The scenario chart uses the ratio of medians, as the routing-overhead study does."
          },
          {
            "quantity": "Recorded work cost per model call, median task cost ÷ median calls (calculation)",
            "value": "$0.0611",
            "n": 48,
            "note": "Notional list-price cost of Agent’s recorded bench work on Sonnet 5.5; the runs used a subscription."
          },
          {
            "quantity": "Decisions in flight at 1 million a day, rule-based policy (calculation)",
            "value": "0.000016 (0.000027 at the 95th percentile time)",
            "n": 20000,
            "note": "Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."
          },
          {
            "quantity": "Decisions per second at 10 million a day (calculation)",
            "value": "115.7 per second",
            "n": null,
            "note": "A steady rate over 24 hours."
          },
          {
            "quantity": "Daily cost at 10 million decisions, rule-based policy (calculation)",
            "value": "$0",
            "n": 20000,
            "note": "$0 a year at a flat volume."
          },
          {
            "quantity": "Daily cost at 10 million decisions, Jev (calculation)",
            "value": "$337.00",
            "n": 82,
            "note": "$123,005 a year at a flat volume."
          },
          {
            "quantity": "Daily cost at 10 million decisions, Sonnet 5.5 router (calculation)",
            "value": "$73,240",
            "n": 82,
            "note": "$26,732,600 a year at a flat volume."
          },
          {
            "quantity": "Daily cost at 10 million decisions, Haiku 4.5 router (calculation)",
            "value": "$89,240",
            "n": 82,
            "note": "$32,572,600 a year at a flat volume."
          },
          {
            "quantity": "Decisions in flight at 10 million a day, rule-based policy (calculation)",
            "value": "0.00016 (0.00027 at the 95th percentile time)",
            "n": 20000,
            "note": "Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."
          },
          {
            "quantity": "Decisions in flight at 10 million a day, Jev (calculation)",
            "value": "15.8 (22.7 at the 95th percentile time)",
            "n": 246,
            "note": "Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."
          },
          {
            "quantity": "Decisions in flight at 10 million a day, Sonnet 5.5 router (calculation)",
            "value": "300.7 (497.5 at the 95th percentile time)",
            "n": 82,
            "note": "Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."
          },
          {
            "quantity": "Decisions in flight at 10 million a day, Haiku 4.5 router (calculation)",
            "value": "1,467 (3,983 at the 95th percentile time)",
            "n": 82,
            "note": "Decisions per second × assumed median or p95 time. These are scenarios, not actual mean concurrency or confidence bounds."
          },
          {
            "quantity": "Waiting hours a day at 1 million decisions, Haiku 4.5 router (calculation)",
            "value": "3,521 hours (9,559 at the 95th percentile time)",
            "n": 82,
            "note": "Decisions × assumed median or p95 time. Actual total waiting requires the mean; these are not confidence bounds."
          },
          {
            "quantity": "Work cost of 1 million model calls a day at the recorded cost per call (calculation)",
            "value": "$61,131",
            "n": 48,
            "note": "A notional list-price cost; your tasks may differ."
          },
          {
            "quantity": "Router cost as a share of the work cost per call, Sonnet 5.5 router (calculation)",
            "value": "12.0%",
            "n": 82,
            "note": "Cost per decision ÷ recorded work cost per model call, if the router decides every call."
          },
          {
            "quantity": "Router cost as a share of the work cost per call, Haiku 4.5 router (calculation)",
            "value": "14.6%",
            "n": 82,
            "note": "Cost per decision ÷ recorded work cost per model call, if the router decides every call."
          },
          {
            "quantity": "Router cost as a share of the work cost per call, Jev (calculation)",
            "value": "0.06%",
            "n": 82,
            "note": "Cost per decision ÷ recorded work cost per model call, if the router decides every call."
          },
          {
            "quantity": "Rule-based policy: measured decisions per second in one batch",
            "value": "469,409 per second",
            "n": 200000,
            "note": "One process, one 200,000-iteration throughput batch of the routing-overhead study. Database reads and the decision-record write are not included."
          },
          {
            "quantity": "Share of that measured policy rate used at 10 million decisions a day (calculation)",
            "value": "0.025%",
            "n": null,
            "note": "Decisions per second at that volume ÷ measured decisions per second."
          }
        ]
      },
      {
        "id": "routing-at-scale-grid",
        "title": "Every scenario, volume and router (calculation)",
        "columns": [
          {
            "key": "scenario",
            "label": "Scenario",
            "unit": "text"
          },
          {
            "key": "calls",
            "label": "Model calls a day",
            "unit": "count"
          },
          {
            "key": "router",
            "label": "Router",
            "unit": "text"
          },
          {
            "key": "decisions",
            "label": "Decisions routed a day",
            "unit": "count"
          },
          {
            "key": "perSecond",
            "label": "Decisions per second (24-hour average)",
            "unit": "count"
          },
          {
            "key": "costDay",
            "label": "Cost per day",
            "unit": "usd"
          },
          {
            "key": "costYear",
            "label": "Cost per year (× 365)",
            "unit": "usd"
          },
          {
            "key": "flightP50",
            "label": "In flight at median time",
            "unit": "calls"
          },
          {
            "key": "flightP95",
            "label": "In flight at 95th percentile time",
            "unit": "calls"
          },
          {
            "key": "hoursP50",
            "label": "Waiting hours a day at median time",
            "unit": "count"
          },
          {
            "key": "hoursP95",
            "label": "Waiting hours a day at 95th percentile time",
            "unit": "count"
          }
        ],
        "rows": [
          {
            "scenario": "Every model call routed",
            "calls": 10000,
            "router": "Deterministic routing policy",
            "decisions": 10000,
            "perSecond": 0.11574,
            "costDay": 0,
            "costYear": 0,
            "flightP50": 1.6435e-7,
            "flightP95": 2.6968e-7,
            "hoursP50": 0.0000039444,
            "hoursP95": 0.0000064722
          },
          {
            "scenario": "Every model call routed",
            "calls": 10000,
            "router": "Jev 1.13 (TypeSafe)",
            "decisions": 10000,
            "perSecond": 0.11574,
            "costDay": 0.337,
            "costYear": 123.01,
            "flightP50": 0.015799,
            "flightP95": 0.02265,
            "hoursP50": 0.37917,
            "hoursP95": 0.54361
          },
          {
            "scenario": "Every model call routed",
            "calls": 10000,
            "router": "Claude Sonnet 5.5 (low) · Claude Code",
            "decisions": 10000,
            "perSecond": 0.11574,
            "costDay": 73.24,
            "costYear": 26733,
            "flightP50": 0.30069,
            "flightP95": 0.49745,
            "hoursP50": 7.2167,
            "hoursP95": 11.939
          },
          {
            "scenario": "Every model call routed",
            "calls": 10000,
            "router": "Claude Haiku 4.5 · Claude Code",
            "decisions": 10000,
            "perSecond": 0.11574,
            "costDay": 89.24,
            "costYear": 32573,
            "flightP50": 1.4669,
            "flightP95": 3.983,
            "hoursP50": 35.206,
            "hoursP95": 95.592
          },
          {
            "scenario": "Every model call routed",
            "calls": 100000,
            "router": "Deterministic routing policy",
            "decisions": 100000,
            "perSecond": 1.1574,
            "costDay": 0,
            "costYear": 0,
            "flightP50": 0.0000016435,
            "flightP95": 0.0000026968,
            "hoursP50": 0.000039444,
            "hoursP95": 0.000064722
          },
          {
            "scenario": "Every model call routed",
            "calls": 100000,
            "router": "Jev 1.13 (TypeSafe)",
            "decisions": 100000,
            "perSecond": 1.1574,
            "costDay": 3.37,
            "costYear": 1230,
            "flightP50": 0.15799,
            "flightP95": 0.2265,
            "hoursP50": 3.7917,
            "hoursP95": 5.4361
          },
          {
            "scenario": "Every model call routed",
            "calls": 100000,
            "router": "Claude Sonnet 5.5 (low) · Claude Code",
            "decisions": 100000,
            "perSecond": 1.1574,
            "costDay": 732.4,
            "costYear": 267330,
            "flightP50": 3.0069,
            "flightP95": 4.9745,
            "hoursP50": 72.167,
            "hoursP95": 119.39
          },
          {
            "scenario": "Every model call routed",
            "calls": 100000,
            "router": "Claude Haiku 4.5 · Claude Code",
            "decisions": 100000,
            "perSecond": 1.1574,
            "costDay": 892.4,
            "costYear": 325730,
            "flightP50": 14.669,
            "flightP95": 39.83,
            "hoursP50": 352.06,
            "hoursP95": 955.92
          },
          {
            "scenario": "Every model call routed",
            "calls": 1000000,
            "router": "Deterministic routing policy",
            "decisions": 1000000,
            "perSecond": 11.574,
            "costDay": 0,
            "costYear": 0,
            "flightP50": 0.000016435,
            "flightP95": 0.000026968,
            "hoursP50": 0.00039444,
            "hoursP95": 0.00064722
          },
          {
            "scenario": "Every model call routed",
            "calls": 1000000,
            "router": "Jev 1.13 (TypeSafe)",
            "decisions": 1000000,
            "perSecond": 11.574,
            "costDay": 33.7,
            "costYear": 12301,
            "flightP50": 1.5799,
            "flightP95": 2.265,
            "hoursP50": 37.917,
            "hoursP95": 54.361
          },
          {
            "scenario": "Every model call routed",
            "calls": 1000000,
            "router": "Claude Sonnet 5.5 (low) · Claude Code",
            "decisions": 1000000,
            "perSecond": 11.574,
            "costDay": 7324,
            "costYear": 2673300,
            "flightP50": 30.069,
            "flightP95": 49.745,
            "hoursP50": 721.67,
            "hoursP95": 1193.9
          },
          {
            "scenario": "Every model call routed",
            "calls": 1000000,
            "router": "Claude Haiku 4.5 · Claude Code",
            "decisions": 1000000,
            "perSecond": 11.574,
            "costDay": 8924,
            "costYear": 3257300,
            "flightP50": 146.69,
            "flightP95": 398.3,
            "hoursP50": 3520.6,
            "hoursP95": 9559.2
          },
          {
            "scenario": "Every model call routed",
            "calls": 10000000,
            "router": "Deterministic routing policy",
            "decisions": 10000000,
            "perSecond": 115.74,
            "costDay": 0,
            "costYear": 0,
            "flightP50": 0.00016435,
            "flightP95": 0.00026968,
            "hoursP50": 0.0039444,
            "hoursP95": 0.0064722
          },
          {
            "scenario": "Every model call routed",
            "calls": 10000000,
            "router": "Jev 1.13 (TypeSafe)",
            "decisions": 10000000,
            "perSecond": 115.74,
            "costDay": 337,
            "costYear": 123010,
            "flightP50": 15.799,
            "flightP95": 22.65,
            "hoursP50": 379.17,
            "hoursP95": 543.61
          },
          {
            "scenario": "Every model call routed",
            "calls": 10000000,
            "router": "Claude Sonnet 5.5 (low) · Claude Code",
            "decisions": 10000000,
            "perSecond": 115.74,
            "costDay": 73240,
            "costYear": 26733000,
            "flightP50": 300.69,
            "flightP95": 497.45,
            "hoursP50": 7216.7,
            "hoursP95": 11939
          },
          {
            "scenario": "Every model call routed",
            "calls": 10000000,
            "router": "Claude Haiku 4.5 · Claude Code",
            "decisions": 10000000,
            "perSecond": 115.74,
            "costDay": 89240,
            "costYear": 32573000,
            "flightP50": 1466.9,
            "flightP95": 3983,
            "hoursP50": 35206,
            "hoursP95": 95592
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 10000,
            "router": "Deterministic routing policy",
            "decisions": 1414,
            "perSecond": 0.016366,
            "costDay": 0,
            "costYear": 0,
            "flightP50": 2.3239e-8,
            "flightP95": 3.8132e-8,
            "hoursP50": 5.5774e-7,
            "hoursP95": 9.1517e-7
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 10000,
            "router": "Jev 1.13 (TypeSafe)",
            "decisions": 1414,
            "perSecond": 0.016366,
            "costDay": 0.047652,
            "costYear": 17.393,
            "flightP50": 0.0022339,
            "flightP95": 0.0032028,
            "hoursP50": 0.053614,
            "hoursP95": 0.076867
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 10000,
            "router": "Claude Sonnet 5.5 (low) · Claude Code",
            "decisions": 1414,
            "perSecond": 0.016366,
            "costDay": 10.356,
            "costYear": 3780,
            "flightP50": 0.042518,
            "flightP95": 0.07034,
            "hoursP50": 1.0204,
            "hoursP95": 1.6882
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 10000,
            "router": "Claude Haiku 4.5 · Claude Code",
            "decisions": 1414,
            "perSecond": 0.016366,
            "costDay": 12.619,
            "costYear": 4605.8,
            "flightP50": 0.20742,
            "flightP95": 0.56319,
            "hoursP50": 4.9781,
            "hoursP95": 13.517
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 100000,
            "router": "Deterministic routing policy",
            "decisions": 14141,
            "perSecond": 0.16367,
            "costDay": 0,
            "costYear": 0,
            "flightP50": 2.3241e-7,
            "flightP95": 3.8135e-7,
            "hoursP50": 0.0000055778,
            "hoursP95": 0.0000091524
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 100000,
            "router": "Jev 1.13 (TypeSafe)",
            "decisions": 14141,
            "perSecond": 0.16367,
            "costDay": 0.47655,
            "costYear": 173.94,
            "flightP50": 0.022341,
            "flightP95": 0.03203,
            "hoursP50": 0.53618,
            "hoursP95": 0.76872
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 100000,
            "router": "Claude Sonnet 5.5 (low) · Claude Code",
            "decisions": 14141,
            "perSecond": 0.16367,
            "costDay": 103.57,
            "costYear": 37803,
            "flightP50": 0.42521,
            "flightP95": 0.70345,
            "hoursP50": 10.205,
            "hoursP95": 16.883
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 100000,
            "router": "Claude Haiku 4.5 · Claude Code",
            "decisions": 14141,
            "perSecond": 0.16367,
            "costDay": 126.19,
            "costYear": 46061,
            "flightP50": 2.0743,
            "flightP95": 5.6323,
            "hoursP50": 49.784,
            "hoursP95": 135.18
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 1000000,
            "router": "Deterministic routing policy",
            "decisions": 141414,
            "perSecond": 1.6367,
            "costDay": 0,
            "costYear": 0,
            "flightP50": 0.0000023242,
            "flightP95": 0.0000038136,
            "hoursP50": 0.00005578,
            "hoursP95": 0.000091526
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 1000000,
            "router": "Jev 1.13 (TypeSafe)",
            "decisions": 141414,
            "perSecond": 1.6367,
            "costDay": 4.7657,
            "costYear": 1739.5,
            "flightP50": 0.22341,
            "flightP95": 0.32031,
            "hoursP50": 5.3619,
            "hoursP95": 7.6874
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 1000000,
            "router": "Claude Sonnet 5.5 (low) · Claude Code",
            "decisions": 141414,
            "perSecond": 1.6367,
            "costDay": 1035.7,
            "costYear": 378040,
            "flightP50": 4.2522,
            "flightP95": 7.0347,
            "hoursP50": 102.05,
            "hoursP95": 168.83
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 1000000,
            "router": "Claude Haiku 4.5 · Claude Code",
            "decisions": 141414,
            "perSecond": 1.6367,
            "costDay": 1262,
            "costYear": 460620,
            "flightP50": 20.744,
            "flightP95": 56.325,
            "hoursP50": 497.86,
            "hoursP95": 1351.8
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 10000000,
            "router": "Deterministic routing policy",
            "decisions": 1414141,
            "perSecond": 16.367,
            "costDay": 0,
            "costYear": 0,
            "flightP50": 0.000023242,
            "flightP95": 0.000038136,
            "hoursP50": 0.0005578,
            "hoursP95": 0.00091526
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 10000000,
            "router": "Jev 1.13 (TypeSafe)",
            "decisions": 1414141,
            "perSecond": 16.367,
            "costDay": 47.657,
            "costYear": 17395,
            "flightP50": 2.2341,
            "flightP95": 3.2031,
            "hoursP50": 53.62,
            "hoursP95": 76.874
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 10000000,
            "router": "Claude Sonnet 5.5 (low) · Claude Code",
            "decisions": 1414141,
            "perSecond": 16.367,
            "costDay": 10357,
            "costYear": 3780400,
            "flightP50": 42.522,
            "flightP95": 70.347,
            "hoursP50": 1020.5,
            "hoursP95": 1688.3
          },
          {
            "scenario": "Only System One decisions routed (14.1% of calls)",
            "calls": 10000000,
            "router": "Claude Haiku 4.5 · Claude Code",
            "decisions": 1414141,
            "perSecond": 16.367,
            "costDay": 12620,
            "costYear": 4606200,
            "flightP50": 207.44,
            "flightP95": 563.25,
            "hoursP50": 4978.6,
            "hoursP95": 13518
          }
        ]
      }
    ],
    "related": [
      "routing-overhead",
      "routing-jev-vs-llm"
    ]
  },
  "sources": [
    {
      "id": "agent-routing",
      "title": "Routing runs: Jev router vs LLM routing",
      "kind": "run",
      "date": "2026-10-05",
      "note": "Routing decisions recorded per case and arm.",
      "data": [
        "/benchmarks/raw/routing/receipts.json"
      ]
    },
    {
      "id": "agent-jev-live",
      "title": "Jev live run: 246 timed calls on the 82 routing decisions",
      "kind": "run",
      "date": "2026-10-06",
      "note": "Jev 1.13 called over HTTPS, 3 repeats of the same 82 typed decisions, one call at a time, from one Mac over a home network: client wall time, with the network inside it. The API reports no server time. Cost per 1,000 decisions is a calculation from the reported input tokens and the published price. The case sets were revised against Jev answers, so Jev has a home advantage.",
      "data": [
        "/benchmarks/raw/jev-live/summary.json",
        "/benchmarks/raw/jev-live/calls.json"
      ]
    },
    {
      "id": "price-anthropic",
      "title": "Anthropic list prices (Claude models)",
      "kind": "price-list",
      "date": "2026-09-21",
      "url": "https://platform.claude.com/docs/en/about-claude/pricing",
      "note": "Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price."
    },
    {
      "id": "price-jev",
      "title": "Jev 1.13 list price",
      "kind": "price-list",
      "date": "2026-09-23",
      "url": "https://docs.typesafe.ai/models",
      "note": "Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free."
    },
    {
      "id": "agent-routing-overhead",
      "title": "Routing overhead runs: policy microbenchmark and CLI start-up",
      "kind": "run",
      "date": "2026-10-06",
      "note": "In-process timing of the deterministic routing policy (20,000 timed decisions), CLI start-up with a one-word prompt (5 runs per CLI), and decision counts read from recorded bench runs. LLM router timings are reused from the routing runs.",
      "data": [
        "/benchmarks/raw/routing-overhead/results.json"
      ]
    },
    {
      "id": "calc-routing-at-scale",
      "title": "Routing at scale calculation",
      "kind": "calculation",
      "date": "2026-10-07",
      "note": "Recorded routing costs and times scaled to assumed daily volumes. No load test ran; median and p95 scenarios are not measured mean concurrency."
    }
  ]
}
