- Home
- Why Agent
Evidence with n, intervals and caveats
Keep your work. Change the AI underneath.
Agent is an AI worker that belongs to your organization: memory, rules and receipts that stay when the model changes. We publish the numbers behind our model, routing and effort choices, and we show where those numbers stop.
Every figure on this page links to its chart. How we measure · One real task, as a receipt
Motion reduced: press Replay zoom to animate
1st, 2nd …: place by median, given only to a lane whose run range overlaps no other lane’s. A lane marked ~ overlaps another lane’s range, so it gets no place.
about 96 thousand×Jev 1.13 (TypeSafe) takes about 96 thousand times as long as Deterministic routing policy (Agent, in process). Calculation: ratio of the two medians (137 ms ÷ 1.42 µs). Log scale: each gridline is 10 times the one before.
| Item | Decision time | Median to p95 | n |
|---|---|---|---|
| Deterministic routing policy (Agent, in process) | 1.42 µs | 1.42 µs–2.33 µs | 20000 |
| Jev 1.13 (TypeSafe) | 137 ms | 137 ms–196 ms | 246 |
| Claude Sonnet 5.5 (effort low, via Claude Code) | 2.6 s | 2.6 s–4.3 s | 82 |
| Claude Haiku 4.5 (thinking on, via Claude Code) | 12.54 s | 12.54 s–34.48 s | 82 |
4 rows. Slowest Claude Haiku 4.5 (thinking on, via Claude Code) 12.54 s (median to p95 12.54 s–34.48 s, n 82). Fastest Deterministic routing policy (Agent, in process) 1.42 µs (median to p95 1.42 µs–2.33 µs, n 20000). Not all run ranges overlap.
NotesLines: median to p95 (not an interval)n 82–20000 per row
Median; whiskers = median to 95th percentile
The policy is the pure in-process decision (median 1.42 µs, p95 2.33 µs), timed over 20,000 decisions after 5,000 warm-up calls. The Claude routers are wall time per call through the Claude Code CLI, one call at a time, from the recorded routing runs. Jev is wall time of a direct HTTPS call from the same Mac over a home network (246 calls in a 35-second window; its API reports no server time), so the network is inside it. A different route from the CLI, so the chart shows what a caller waits, not model compute time. The whisker is the median to the 95th percentile, not a confidence interval.
Sources: Routing overhead runs: policy microbenchmark and CLI start-up, Routing runs: Jev router vs LLM routing, Jev live run: 246 timed calls on the 82 routing decisions
Routing
Recorded routing accuracy and list-price cost
Jev 1.13 and Claude Sonnet 5.5 answered the same 82 typed routing cases. Every row of this comparison
150xlower
cost per 1,000 routing decisions for Jev 1.13: $0.034 vs $5.00. List-price calculation
Measured comparison
Jev 1.13 vs Claude Sonnet 5.5: accuracy and cost
One row per metric, each on its own axis. The shaded band is where the two 95% intervals overlap: a side is ahead only when they do not.
- Jev 1.13
- Claude Sonnet 5.5
- 95% interval
- where the two overlap
- hollow: list-price calculation
Jev vs Claude as a router: accuracy and cost
- Typed routing decisions answered exactly right90%n 8294% (77/82)n 82TieTyped routing decisions answered exactly right: Jev 1.13 90% (n 82, 95% interval 82%–95%); Claude Sonnet 5.5 94% (77/82) (n 82, 95% interval 87%–97%). Tie.
- Per-question accuracy95% (184/194)n 19497% (189/194)n 194TiePer-question accuracy: Jev 1.13 95% (184/194) (n 194, 95% interval 91%–97%); Claude Sonnet 5.5 97% (189/194) (n 194, 95% interval 94%–99%). Tie.
- Cost per 1,000 routing decisionsCalculation$0.034n 246$5.00n 82UnclearCost per 1,000 routing decisions, calculation: Jev 1.13 $0.034 (n 246); Claude Sonnet 5.5 $5.00 (n 82). Unclear.
| Metric | Jev 1.13 | Claude Sonnet 5.5 | n | Outcome | Basis |
|---|---|---|---|---|---|
| Typed routing decisions answered exactly right | 90% | 94% (77/82) | 82 | Tie | The 95% intervals overlap (Jev 1.13 82% to 95%; Claude Sonnet 5.5 87% to 97%), so this sample cannot separate them. |
| Per-question accuracy | 95% (184/194) | 97% (189/194) | 194 | Tie | The 95% intervals overlap (Jev 1.13 91% to 97%; Claude Sonnet 5.5 94% to 99%), so this sample cannot separate them. |
| Cost per 1,000 routing decisions | $0.034 | $5.00 | 246 and 82 | Unclear | No interval or range was recorded for either side, so the gap ($0.034 vs $5.00, 148x) is not tested against run-to-run variation. |
3 rows: Typed routing decisions answered exactly right, 90% vs 94% (77/82) (n = 82), tie; Per-question accuracy, 95% (184/194) vs 97% (189/194) (n = 194), tie; Cost per 1,000 routing decisions, $0.034 vs $5.00, unclear. 2 of 3 rows are ties: the 95% intervals overlap.
The cost row is a list-price calculation with no interval, so the gap is stated, not tested. Press a row (or Enter) for its basis.
Source: Every row of this comparison
The evidence
What we measured, and what it means
Six results from the public studies. Each one keeps its sample size, its interval and its caveat.
Blind review
75% (9/12)
95% CI 47%–91% · n = 12
Critics who could not see the labels preferred the AI change over the merged human change.
Caveat Small sample and a wide interval. The first scored attempt was 50% (6/12); later attempts had the earlier attempts' lessons.
See the chartSWE-bench Verified
76% (25/33)
95% CI 59%–87% · n = 33
On the same 33 instances, Agent resolved as many as the public frontier panel, within the intervals.
Caveat Public panel mean 74.1%. The intervals overlap: this is "on par", not "ahead".
See the chartPrompt cachingCalculation
50%
($0.1350 vs $0.2698)
Reading the stable prefix from the cache halved the list-price cost of the recorded turns.
Caveat A list-price calculation on recorded tokens, not a bill: the calls ran on a subscription.
See the chartReasoning effort
100% (176/176)
95% CI 98%–100% · n = 176
On the hard task set, every call at every effort level passed strictly: more effort could not buy quality here.
Caveat A ceiling: this task set cannot separate the efforts. Harder tasks may.
See the chartRule-based routing
1.42µs
p95 2.33 µs, p99 3.04 µs
n = 20000
A deterministic routing decision takes microseconds; an LLM router call takes seconds.
Caveat Timed in process over 20,000 decisions. A rule covers only the decisions it describes.
See the chartCLI context tax
19,551input tokens (median)
The API sends 17 tokens for the same request.
n = 15
A coding CLI wraps a one-line request in its own system prompt and tools before your words reach the model.
Caveat Part of that input is served from the cache, at a lower price.
See the chart
Limits
What the evidence does not show
Agent is not the cheapest way to resolve a SWE-bench task
Agent's notional model cost per resolved instance was $3.71; the public panel's recorded mean was $0.57. Agent runs a full pipeline with many calls per attempt.
Check itMost comparisons have no winner
1,507 of 1,571 comparison rows are a tie or unclear. We say so on every page instead of ranking by a gap the data cannot support.
Check itNo customer savings are measured yet
The cost figures here are list-price calculations on our own recorded runs. Savings on your work stay unproven until your runs show them.
Some routes are not measured, and one speed is a route gap
OpenRouter's Auto Router, cheaper hosted inference and a local router model are "not measured" until a run exists; the pages never show a number for them. Jev's speed is measured (a median 137 ms per call over its API), but from one Mac on a home network and as a direct call, while the Claude routers ran through a CLI. That gap is between routes, not only between models.
Check it
The product
What Agent does with this
Keep the work when the model changes
Start a task on Claude and finish it on Codex. The task, its history and its corrections stay with your organization, not with one vendor.
Correct it once
Agent keeps your standards and past corrections as memory and rules, and uses them on the next task, whichever model runs it.
Your subscriptions, visible costs
Workers run on your own Claude and Codex subscriptions, with budgets, effort controls and usage you can inspect.
Receipts, not promises
The past-PR trial shows the independent checks that ran, the time and the cost side by side, the same way these pages show n and intervals.
Start with a task you already know
Replay a past pull request. Compare it with yours.
The reference stays sealed while Agent works; then you see the independent checks, the time and the cost side by side.
- Your own Claude and Codex subscriptions
- Memory and rules that outlive the model
- Receipts: checks, time and cost for every task