Routing

Recorded routing accuracy and list-price cost

Jev 1.13 and Claude Sonnet 5.5 answered the same 82 typed routing cases. Every row of this comparison

Measured comparison

Jev 1.13 vs Claude Sonnet 5.5: accuracy and cost

One row per metric, each on its own axis. The shaded band is where the two 95% intervals overlap: a side is ahead only when they do not.

MeasuredCost: calculation
  • Jev 1.13
  • Claude Sonnet 5.5
  • 95% interval
  • where the two overlap
  • hollow: list-price calculation

Jev vs Claude as a router: accuracy and cost

2 ties · 1 unclear
  • Typed routing decisions answered exactly right: Jev 1.13 90% (n 82, 95% interval 82%–95%); Claude Sonnet 5.5 94% (77/82) (n 82, 95% interval 87%–97%). Tie.
  • Per-question accuracy: Jev 1.13 95% (184/194) (n 194, 95% interval 91%–97%); Claude Sonnet 5.5 97% (189/194) (n 194, 95% interval 94%–99%). Tie.
  • Cost per 1,000 routing decisions, calculation: Jev 1.13 $0.034 (n 246); Claude Sonnet 5.5 $5.00 (n 82). Unclear.

3 rows: Typed routing decisions answered exactly right, 90% vs 94% (77/82) (n = 82), tie; Per-question accuracy, 95% (184/194) vs 97% (189/194) (n = 194), tie; Cost per 1,000 routing decisions, $0.034 vs $5.00, unclear. 2 of 3 rows are ties: the 95% intervals overlap.

The cost row is a list-price calculation with no interval, so the gap is stated, not tested. Press a row (or Enter) for its basis.

Source: Every row of this comparison

The evidence

What we measured, and what it means

Six results from the public studies. Each one keeps its sample size, its interval and its caveat.

Limits

What the evidence does not show

The product

What Agent does with this