{
  "schema": "agent-public-bench@1",
  "generatedAt": "2026-10-07T00:00:00.000Z",
  "url": "https://agent.sasid.ai/benchmarks/routing-overhead",
  "study": {
    "slug": "routing-overhead",
    "title": "Routing overhead: deterministic policy vs LLM routers vs Jev",
    "seoTitle": "Routing overhead: rules vs LLM routers vs Jev",
    "description": "How much delay and cost a router adds per decision: an in-process policy, Claude routers through a CLI, and Jev. Plus CLI start-up tax and per-task totals.",
    "question": "What delay and what cost does each kind of router add before the real work of a call starts?",
    "answer": "The deterministic routing policy decided in a median 1.42 µs (p95 2.33 µs, 20,000 decisions, $0). The fastest LLM router, Claude Sonnet 5.5 through the Claude Code CLI, took a median 2.60 s per decision (p95 4.30 s, n = 82), about 1.8 million times longer; 973 ms of that median was CLI time, not model time. Haiku 4.5 with its default thinking took 12.54 s (p95 34.48 s). Jev 1.13, called directly over HTTPS from the same Mac, took a median 137 ms per decision (p95 196 ms, n = 246 calls, client wall time with the network inside it) and cost $0.0337 per 1,000 decisions (a calculation from its reported input tokens). That is a different route from the CLI routers, so it is not a model-against-model comparison. As a calculation over 48 recorded tasks (median 49.5 model calls each), routing every call would add $1.67 per 1,000 tasks and up to 7 s of waiting per task with Jev, $247.30 with Sonnet (8.2% of the work cost) and up to 129 s of waiting per task with Sonnet; routing only the 7 System One decisions cuts Sonnet to $34.97 and 18 s. CLI start-up alone, for a one-word answer: Claude Code (Haiku 4.5) took 2.53 s and sent 6,761 input tokens; Codex CLI took 6.00 s and sent 17,051 input tokens, 13,184 of them read from the cache (5 runs each, different models).",
    "date": "2026-10-06",
    "updated": "2026-10-06",
    "tags": [
      "routing",
      "latency",
      "overhead",
      "jev",
      "llm-router",
      "cli",
      "calculation"
    ],
    "method": [
      "Protocol declared before any measurement.",
      "Deterministic policy: the platform’s production routing decision (plan, model ladder and effort) on its default policy, timed in process on one Apple M3 Ultra Mac with Node 25: 5,000 warm-up calls, then 20,000 timed decisions over 64 synthetic routing contexts, plus a 200,000-call batch for throughput. Database reads and the decision record write are out of scope; the recorded rule-based System One decisions (millisecond resolution, record write included) are shown as a stat.",
      "LLM routers: the recorded routing runs of the routing study (Claude Sonnet 5.5 (effort low, via Claude Code), 82 calls; Claude Haiku 4.5 (thinking on, via Claude Code), 82 calls), one call at a time. Wall time per call, the API time the CLI reported, and the difference (CLI and harness time). p50 and p95 recomputed from the per-call log.",
      "Jev: a live run on 2026-10-06. The same 82 typed decisions as the routing study, 3 repeats, 246 counted calls sent one at a time over HTTPS to the TypeSafe API from the same Mac, on a home network. Time per call is client wall time from before the request to after the body was read, so the network is inside it; the API reports no server time. Cost per decision is a calculation: the mean input tokens its API reported × the published price. The recorded production run (82 decisions) has the same cost as a provider-reported figure.",
      "CLI start-up: Claude Code · Claude Haiku 4.5 5 runs, Codex CLI (default model) 5 runs, a one-word prompt, one call at a time. Times to the first output event, the first model output and exit.",
      "Per 1,000 tasks (a calculation): decisions per task from 48 recorded bench runs read only (2 empty runs excluded) × cost and median time per decision. Decisions are assumed to wait in line, so the delay is an upper bound."
    ],
    "caveats": [
      "The policy is timed in process and the LLM routers through a CLI: this compares the two ways of routing as deployed, not two models on equal footing. A direct API call would skip the CLI time (shown separately).",
      "The routing study reports its latency medians from its own summary; this study recomputes them from the per-call log, so the medians can differ by a few tens of milliseconds.",
      "Haiku 4.5 ran with the CLI’s default thinking, which makes it slower than Sonnet 5.5 at effort low here.",
      "OpenRouter’s Auto Router and cheaper hosted inference were not timed: no key in the environment. A local router server was not running, so it was not timed either.",
      "CLI timings come from one Mac with 5 runs per CLI; a range is not a confidence interval. Codex CLI reports no API time, so its CLI time cannot be separated from model time.",
      "Per-task numbers are calculations on runs where routing was off; the work cost is the recorded list-price estimate for Sonnet 5.5.",
      "Jev was timed over a direct HTTPS call; the Claude routers ran through the CLI. These are different routes, so the gap is what a caller waits per decision, not model compute time. A caller closer to the API would see less than this Mac on a home network did.",
      "Jev’s timing is one 35-second window on 2026-10-06, 246 calls from one machine; server load at that time is unknown. The 246 calls are 3 repeats of 82 requests, so they are not independent draws.",
      "Corrected 2026-10-07: an earlier version of this page counted the cached tokens twice (30,235). The CLI reports 17,051 input tokens including 13,184 read from the cache."
    ],
    "sourceIds": [
      "agent-routing-overhead",
      "calc-routing-overhead",
      "agent-routing",
      "price-anthropic",
      "price-jev",
      "agent-jev-live"
    ],
    "stats": [
      {
        "id": "router-overhead-policy-p50",
        "label": "Deterministic routing policy: median decision time",
        "value": 0.00142,
        "unit": "ms",
        "display": "1.42 µs (p95 2.33 µs, p99 3.04 µs)",
        "n": 20000,
        "note": "Timer resolution 0.041 µs; 469,409 decisions per second in a 200,000-call batch."
      },
      {
        "id": "router-overhead-speedup",
        "label": "Median LLM router call ÷ median policy decision",
        "value": 1828873,
        "unit": "ratio",
        "display": "about 1.8 million times",
        "n": 82
      },
      {
        "id": "router-overhead-system-one-rule-arm",
        "label": "Recorded rule-based System One decision, record write included",
        "value": 1,
        "unit": "ms",
        "display": "1 ms median, 2 ms p95 (millisecond resolution)",
        "n": 419
      },
      {
        "id": "router-overhead-decisions-per-task",
        "label": "Model calls per task (each one a routing decision)",
        "value": 49.5,
        "unit": "calls",
        "display": "49.5 median (13 to 73)",
        "n": 48
      },
      {
        "id": "router-overhead-jev-p50",
        "label": "Jev 1.13 (TypeSafe): median decision time over the API",
        "value": 136.5,
        "unit": "ms",
        "display": "137 ms (p95 196 ms)",
        "n": 246,
        "note": "Client wall time from one Mac over a home network, 246 calls in a 35-second window. The API reports no server time."
      },
      {
        "id": "router-overhead-sonnet-vs-jev",
        "label": "Median Claude Sonnet 5.5 call through the CLI ÷ median Jev call over the API",
        "value": 19,
        "unit": "ratio",
        "display": "about 19 times",
        "n": 82,
        "note": "A calculation across two routes (CLI vs a direct API call from one Mac), not a model-against-model comparison."
      },
      {
        "id": "router-overhead-haiku-vs-jev",
        "label": "Median Claude Haiku 4.5 call through the CLI ÷ median Jev call over the API",
        "value": 91.9,
        "unit": "ratio",
        "display": "about 92 times",
        "n": 82,
        "note": "A calculation across two routes (CLI vs a direct API call from one Mac), not a model-against-model comparison."
      },
      {
        "id": "cli-startup-claude-harness-ms",
        "label": "Claude Code time outside the model on a one-word answer",
        "value": 1690,
        "unit": "ms",
        "display": "1,690 ms median (1,533 to 1,811)",
        "n": 5
      }
    ],
    "charts": [
      {
        "id": "router-overhead-decision-latency",
        "title": "Time to make one routing decision",
        "subtitle": "Median; whiskers = median to 95th percentile",
        "kind": "dot-range",
        "unit": "ms",
        "yLabel": "Time per decision",
        "whisker": "p50-p95",
        "series": [
          {
            "name": "Decision time",
            "points": [
              {
                "label": "Deterministic routing policy (Agent, in process)",
                "value": 0.00142,
                "lo": 0.00142,
                "hi": 0.00233,
                "n": 20000,
                "highlight": true
              },
              {
                "label": "Jev 1.13 (TypeSafe)",
                "value": 136.5,
                "lo": 136.5,
                "hi": 195.7,
                "n": 246
              },
              {
                "label": "Claude Sonnet 5.5 (effort low, via Claude Code)",
                "value": 2597,
                "lo": 2597,
                "hi": 4298,
                "n": 82
              },
              {
                "label": "Claude Haiku 4.5 (thinking on, via Claude Code)",
                "value": 12543,
                "lo": 12543,
                "hi": 34481,
                "n": 82
              }
            ]
          }
        ],
        "note": "The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.",
        "sourceIds": [
          "agent-routing-overhead",
          "agent-routing",
          "agent-jev-live"
        ]
      },
      {
        "id": "router-overhead-cli-vs-model-time",
        "title": "Where an LLM router’s time goes: model vs CLI",
        "subtitle": "Median per call; whiskers = median to 95th percentile",
        "kind": "dot-range",
        "unit": "ms",
        "yLabel": "Time per call",
        "whisker": "p50-p95",
        "series": [
          {
            "name": "Model API time",
            "points": [
              {
                "label": "Claude Sonnet 5.5 (effort low, via Claude Code)",
                "value": 1596,
                "lo": 1596,
                "hi": 2583,
                "n": 82
              },
              {
                "label": "Claude Haiku 4.5 (thinking on, via Claude Code)",
                "value": 10508,
                "lo": 10508,
                "hi": 32132,
                "n": 82
              }
            ]
          },
          {
            "name": "CLI and harness time",
            "points": [
              {
                "label": "Claude Sonnet 5.5 (effort low, via Claude Code)",
                "value": 973,
                "lo": 973,
                "hi": 1277,
                "n": 82
              },
              {
                "label": "Claude Haiku 4.5 (thinking on, via Claude Code)",
                "value": 1698,
                "lo": 1698,
                "hi": 2677,
                "n": 82
              }
            ]
          }
        ],
        "note": "Model API time is the API duration the CLI reports; CLI and harness time is wall time minus that, per call. Medians of the parts do not add up to the median of the whole. Jev is not split: its API reports no server time. The whisker is the median to the 95th percentile, not a confidence interval.",
        "sourceIds": [
          "agent-routing-overhead",
          "agent-routing"
        ]
      },
      {
        "id": "router-overhead-completed",
        "title": "Routing calls that returned a decision",
        "subtitle": "Completed calls ÷ calls; whiskers = 95% Wilson interval",
        "kind": "dot-range",
        "unit": "rate",
        "yLabel": "Completed",
        "whisker": "ci95",
        "series": [
          {
            "name": "Completed",
            "points": [
              {
                "label": "Deterministic routing policy (Agent, in process)",
                "value": 1,
                "lo": 0.9998,
                "hi": 1,
                "n": 20000,
                "highlight": true
              },
              {
                "label": "Jev 1.13 (TypeSafe)",
                "value": 1,
                "lo": 0.9846,
                "hi": 1,
                "n": 246
              },
              {
                "label": "Claude Sonnet 5.5 (effort low, via Claude Code)",
                "value": 1,
                "lo": 0.9552,
                "hi": 1,
                "n": 82
              },
              {
                "label": "Claude Haiku 4.5 (thinking on, via Claude Code)",
                "value": 1,
                "lo": 0.9552,
                "hi": 1,
                "n": 82
              }
            ]
          }
        ],
        "note": "A completed call returned a decision, right or wrong (accuracy is in the routing study). Whiskers are 95% Wilson intervals.",
        "sourceIds": [
          "agent-routing-overhead",
          "agent-routing",
          "agent-jev-live"
        ]
      },
      {
        "id": "router-overhead-cost-reported",
        "title": "Cost per 1,000 routing decisions: no model call vs provider-reported",
        "subtitle": "USD per 1,000 decisions",
        "kind": "bar",
        "unit": "usd",
        "yLabel": "USD per 1,000 decisions",
        "series": [
          {
            "name": "Cost per 1,000 decisions",
            "points": [
              {
                "label": "Deterministic routing policy (Agent, in process)",
                "value": 0,
                "n": 20000,
                "highlight": true
              },
              {
                "label": "Jev 1.13 (TypeSafe)",
                "value": 0.0337,
                "n": 82,
                "highlight": false
              }
            ]
          }
        ],
        "note": "The policy makes no model call, so it costs nothing per decision. Jev’s figure is the cost its provider reported for the recorded production run (82 decisions). The input tokens of the live run give the same figure at the published price. The Claude routers are in the next chart: their cost is derived from list prices.",
        "sourceIds": [
          "agent-routing-overhead",
          "agent-routing",
          "price-jev"
        ]
      },
      {
        "id": "router-overhead-cost-list-price",
        "title": "Cost per 1,000 routing decisions for the model routers (calculation)",
        "subtitle": "List price × the tokens each route reported, USD per 1,000 decisions",
        "kind": "bar",
        "unit": "usd",
        "yLabel": "USD per 1,000 decisions",
        "series": [
          {
            "name": "Cost per 1,000 decisions (list price)",
            "points": [
              {
                "label": "Jev 1.13 (TypeSafe)",
                "value": 0.0337,
                "n": 246
              },
              {
                "label": "Claude Sonnet 5.5 (effort low, via Claude Code)",
                "value": 4.996,
                "n": 82
              },
              {
                "label": "Claude Haiku 4.5 (thinking on, via Claude Code)",
                "value": 8.924,
                "n": 82
              }
            ]
          }
        ],
        "note": "A calculation, not a bill: the Claude calls ran on a subscription. Jev: the 803 input tokens per decision its API reported × the published $0.042 per million input tokens (output tokens are free). Same tokens and prices as the routing study.",
        "sourceIds": [
          "agent-routing-overhead",
          "calc-routing-overhead",
          "agent-routing",
          "price-anthropic",
          "price-jev",
          "agent-jev-live"
        ]
      },
      {
        "id": "router-overhead-cost-per-1000-tasks",
        "title": "Added routing cost per 1,000 tasks (calculation)",
        "subtitle": "Decisions per task from recorded runs × cost per decision",
        "kind": "grouped-bar",
        "unit": "usd",
        "yLabel": "USD per 1,000 tasks",
        "series": [
          {
            "name": "Every model call routed (49.5 per task)",
            "points": [
              {
                "label": "Deterministic routing policy (Agent, in process)",
                "value": 0
              },
              {
                "label": "Jev 1.13 (TypeSafe)",
                "value": 1.67
              },
              {
                "label": "Claude Sonnet 5.5 (effort low, via Claude Code)",
                "value": 247.3
              },
              {
                "label": "Claude Haiku 4.5 (thinking on, via Claude Code)",
                "value": 441.74
              }
            ]
          },
          {
            "name": "Only System One decisions (7 per task)",
            "points": [
              {
                "label": "Deterministic routing policy (Agent, in process)",
                "value": 0
              },
              {
                "label": "Jev 1.13 (TypeSafe)",
                "value": 0.24
              },
              {
                "label": "Claude Sonnet 5.5 (effort low, via Claude Code)",
                "value": 34.97
              },
              {
                "label": "Claude Haiku 4.5 (thinking on, via Claude Code)",
                "value": 62.47
              }
            ]
          }
        ],
        "note": "A calculation. Decisions per task: the median of 48 recorded bench runs (routing was off in them, so every model call counts as one decision a router would make). Median recorded work cost per task: $3.03. Claude router costs are list-price calculations; Jev’s is a list-price calculation too (its recorded run’s provider-reported cost is the same).",
        "sourceIds": [
          "agent-routing-overhead",
          "calc-routing-overhead",
          "agent-routing",
          "price-anthropic",
          "price-jev",
          "agent-jev-live"
        ]
      },
      {
        "id": "router-overhead-delay-per-task",
        "title": "Added routing delay per task (calculation)",
        "subtitle": "Decisions per task × median decision time, if every decision waits in line",
        "kind": "grouped-bar",
        "unit": "seconds",
        "yLabel": "Seconds per task",
        "series": [
          {
            "name": "Every model call routed (49.5 per task)",
            "points": [
              {
                "label": "Deterministic routing policy (Agent, in process)",
                "value": 0.0000703
              },
              {
                "label": "Jev 1.13 (TypeSafe)",
                "value": 6.7568
              },
              {
                "label": "Claude Sonnet 5.5 (effort low, via Claude Code)",
                "value": 128.5515
              },
              {
                "label": "Claude Haiku 4.5 (thinking on, via Claude Code)",
                "value": 620.8785
              }
            ]
          },
          {
            "name": "Only System One decisions (7 per task)",
            "points": [
              {
                "label": "Deterministic routing policy (Agent, in process)",
                "value": 0.0000099
              },
              {
                "label": "Jev 1.13 (TypeSafe)",
                "value": 0.9555
              },
              {
                "label": "Claude Sonnet 5.5 (effort low, via Claude Code)",
                "value": 18.179
              },
              {
                "label": "Claude Haiku 4.5 (thinking on, via Claude Code)",
                "value": 87.801
              }
            ]
          }
        ],
        "note": "A calculation and an upper bound: it assumes each decision waits for the one before. Median recorded task wall time: 10.3 min. Jev’s delay uses its live median over the API from one Mac (network included); the Claude routers’ includes the CLI.",
        "sourceIds": [
          "agent-routing-overhead",
          "calc-routing-overhead",
          "agent-routing",
          "price-anthropic",
          "price-jev",
          "agent-jev-live"
        ]
      },
      {
        "id": "cli-startup-tax",
        "title": "CLI start-up tax on a one-word answer",
        "subtitle": "Median of 5 runs; whiskers = fastest and slowest run",
        "kind": "dot-range",
        "unit": "ms",
        "yLabel": "Time to output",
        "whisker": "minmax",
        "series": [
          {
            "name": "First output event",
            "points": [
              {
                "label": "Claude Code · Claude Haiku 4.5",
                "value": 563,
                "lo": 519,
                "hi": 726,
                "n": 5
              },
              {
                "label": "Codex CLI (default model)",
                "value": 489,
                "lo": 354,
                "hi": 1304,
                "n": 5
              }
            ]
          },
          {
            "name": "First model output",
            "points": [
              {
                "label": "Claude Code · Claude Haiku 4.5",
                "value": 1461,
                "lo": 1206,
                "hi": 2308,
                "n": 5
              },
              {
                "label": "Codex CLI (default model)",
                "value": 5059,
                "lo": 4391,
                "hi": 5478,
                "n": 5
              }
            ]
          },
          {
            "name": "Total wall time",
            "points": [
              {
                "label": "Claude Code · Claude Haiku 4.5",
                "value": 2529,
                "lo": 2273,
                "hi": 3382,
                "n": 5
              },
              {
                "label": "Codex CLI (default model)",
                "value": 5999,
                "lo": 5367,
                "hi": 6506,
                "n": 5
              }
            ]
          }
        ],
        "note": "Prompt: reply with one word. Claude Code · Claude Haiku 4.5: 5/5 runs completed; Codex CLI (default model): 5/5 runs completed. Isolated flags (no tools, no MCP servers, no session) for Claude Code; read-only sandbox and a fresh folder for Codex. The two CLIs ran different models, so CLI and model are not separated. A range, not a confidence interval.",
        "sourceIds": [
          "agent-routing-overhead"
        ]
      },
      {
        "id": "cli-startup-input-tokens",
        "title": "Input tokens a CLI sends for a one-word answer",
        "subtitle": "Per call, mostly the CLI’s own system prompt and tool definitions",
        "kind": "bar",
        "unit": "tokens",
        "yLabel": "Input tokens per call",
        "series": [
          {
            "name": "Input tokens per call",
            "points": [
              {
                "label": "Claude Code · Claude Haiku 4.5",
                "value": 6761,
                "n": 5
              },
              {
                "label": "Codex CLI (default model)",
                "value": 17051,
                "n": 5
              }
            ]
          }
        ],
        "note": "Claude Code sums its disjoint input, cache-read and cache-write fields. Codex CLI reports 17,051 input tokens, 13,184 of them read from the cache. The prompt itself is a few tokens.",
        "sourceIds": [
          "agent-routing-overhead"
        ]
      }
    ],
    "tables": [
      {
        "id": "router-overhead-status",
        "title": "What was measured, recorded, calculated or not measured",
        "columns": [
          {
            "key": "router",
            "label": "Router",
            "unit": "text"
          },
          {
            "key": "kind",
            "label": "Kind",
            "unit": "text"
          },
          {
            "key": "latency",
            "label": "Latency",
            "unit": "text"
          },
          {
            "key": "cost",
            "label": "Cost",
            "unit": "text"
          }
        ],
        "rows": [
          {
            "router": "Deterministic routing policy (Agent, in process)",
            "kind": "in-process rules",
            "latency": "measured 2026-10-06: 1.42 µs median",
            "cost": "$0 (no model call)"
          },
          {
            "router": "Jev 1.13 (TypeSafe)",
            "kind": "hosted decision model",
            "latency": "measured 2026-10-06: 137 ms median, p95 196 ms (246 calls over the API from one Mac; client wall time, no server time)",
            "cost": "$0.0337 per 1,000 (list-price calculation from input tokens; the recorded run’s own cost figure agrees)"
          },
          {
            "router": "Claude Sonnet 5.5 (effort low, via Claude Code)",
            "kind": "LLM router via agent CLI",
            "latency": "recorded: 2.60 s median (82 calls)",
            "cost": "$4.996 per 1,000 (list-price calculation)"
          },
          {
            "router": "Claude Haiku 4.5 (thinking on, via Claude Code)",
            "kind": "LLM router via agent CLI",
            "latency": "recorded: 12.54 s median (82 calls)",
            "cost": "$8.924 per 1,000 (list-price calculation)"
          },
          {
            "router": "OpenRouter Auto Router / cheaper hosted inference",
            "kind": "hosted LLM router / gateway",
            "latency": "Not measured: no key in the environment. The harness is ready and runs when a key is set.",
            "cost": "Not measured: no key in the environment. The harness is ready and runs when a key is set."
          },
          {
            "router": "Clef / Clef-Flash (local)",
            "kind": "local router model",
            "latency": "Not measured: no local server running.",
            "cost": "Not measured: no local server running."
          }
        ]
      },
      {
        "id": "router-overhead-jev-live",
        "title": "Jev live run: time per call and cost",
        "columns": [
          {
            "key": "measure",
            "label": "Measure",
            "unit": "text"
          },
          {
            "key": "value",
            "label": "Value",
            "unit": "text"
          }
        ],
        "rows": [
          {
            "measure": "Counted calls",
            "value": "246 (3 repeats of the same 82 decisions, one at a time), 0 failed"
          },
          {
            "measure": "Median time per call",
            "value": "136.5 ms"
          },
          {
            "measure": "90th percentile",
            "value": "171.0 ms"
          },
          {
            "measure": "95th percentile",
            "value": "195.7 ms"
          },
          {
            "measure": "Fastest and slowest call",
            "value": "100.9 ms and 297.3 ms"
          },
          {
            "measure": "Cold first call (new process, fresh connection)",
            "value": "224.7 ms"
          },
          {
            "measure": "Median per repeat",
            "value": "130.3 ms, 141.7 ms, 137.1 ms (the first repeat without the cold call)"
          },
          {
            "measure": "Input and output tokens per decision (mean)",
            "value": "803 and 148 (output tokens are free at the published price)"
          },
          {
            "measure": "Cost per 1,000 decisions (calculation)",
            "value": "$0.0337 = 803 input tokens × 1,000 × $0.042 per million"
          },
          {
            "measure": "Server-side time",
            "value": "not available: the API sends no timing header or field"
          }
        ]
      }
    ],
    "related": [
      "routing-jev-vs-llm",
      "cli-model-latency-tokens",
      "inference-provider-index"
    ]
  },
  "sources": [
    {
      "id": "agent-routing",
      "title": "Routing runs: Jev router vs LLM routing",
      "kind": "run",
      "date": "2026-10-05",
      "note": "Routing decisions recorded per case and arm.",
      "data": [
        "/benchmarks/raw/routing/receipts.json"
      ]
    },
    {
      "id": "agent-jev-live",
      "title": "Jev live run: 246 timed calls on the 82 routing decisions",
      "kind": "run",
      "date": "2026-10-06",
      "note": "Jev 1.13 called over HTTPS, 3 repeats of the same 82 typed decisions, one call at a time, from one Mac over a home network: client wall time, with the network inside it. The API reports no server time. Cost per 1,000 decisions is a calculation from the reported input tokens and the published price. The case sets were revised against Jev answers, so Jev has a home advantage.",
      "data": [
        "/benchmarks/raw/jev-live/summary.json",
        "/benchmarks/raw/jev-live/calls.json"
      ]
    },
    {
      "id": "price-anthropic",
      "title": "Anthropic list prices (Claude models)",
      "kind": "price-list",
      "date": "2026-09-21",
      "url": "https://platform.claude.com/docs/en/about-claude/pricing",
      "note": "Prices as listed by the vendor on 2026-09-21 and recorded in the product price table. Cache reads at the listed rate, one-hour cache writes at twice the input price."
    },
    {
      "id": "price-jev",
      "title": "Jev 1.13 list price",
      "kind": "price-list",
      "date": "2026-09-23",
      "url": "https://docs.typesafe.ai/models",
      "note": "Prices as listed by the vendor on 2026-09-23: input tokens only, output tokens free."
    },
    {
      "id": "agent-routing-overhead",
      "title": "Routing overhead runs: policy microbenchmark and CLI start-up",
      "kind": "run",
      "date": "2026-10-06",
      "note": "In-process timing of the deterministic routing policy (20,000 timed decisions), CLI start-up with a one-word prompt (5 runs per CLI), and decision counts read from recorded bench runs. LLM router timings are reused from the routing runs.",
      "data": [
        "/benchmarks/raw/routing-overhead/results.json"
      ]
    },
    {
      "id": "calc-routing-overhead",
      "title": "Routing overhead per 1,000 tasks (calculation)",
      "kind": "calculation",
      "date": "2026-10-06",
      "note": "Decisions per task from recorded bench runs multiplied by the cost and the median time per decision. Decisions are assumed to wait in line, so the delay is an upper bound. A calculation, not a run.",
      "data": [
        "/benchmarks/raw/routing-overhead/results.json"
      ]
    }
  ]
}
