Free tools · Sources included

AI benchmark tools

Choose a model from recorded results, or estimate a cost with your own inputs. Each tool keeps measured evidence separate from a calculation.

Calculation

LLM routing cost and delay simulator

Which recorded router fits my timing cap and cost basis?

You enter: Enter decisions per day, working days and a recorded median or p95 timing cap.

  • Qualified cost per day and month
  • Serial waiting scenarios from observed timings
  • Conditional choice with unknowns kept visible

Recorded direct-API and CLI routes differ. This is not a quality forecast, queue model or SLA.

Recorded evidence

Which AI model should I use?

Which measured configuration fits my task?

You enter: Pick a recorded task, execution route and cost or speed preference.

  • Recorded strict pass rates
  • Measured latency and qualified cost
  • Ties and missing results kept visible

A choice from the recorded sample, not a prediction for your work.

Calculation

AI cost calculator

What would my workload cost at API list prices?

You enter: Choose a workload or enter token counts, tasks per month and a price route.

  • Cost per task and per month
  • Same workload priced on each model
  • Token and cache cost breakdown

A list-price estimate, not a subscription bill or vendor invoice.

Calculation

Prompt caching calculator

When could a repeated prompt pay back its cache write?

You enter: Enter a repeated prefix, new tokens, turns and sessions.

  • Break-even turn
  • Cost with and without the cache
  • Daily and monthly list-price estimates

The estimate assumes the stated cache reuse. Recorded sessions remain separate.

Calculation

AI cost of delay calculator

How much time could waiting on a model add?

You enter: Choose a recorded task; enter tasks per person per day, people and working days. Set whether someone waits and an optional hourly cost.

  • Estimated waiting hours
  • A scenario value for that time
  • Retry and routing assumptions

Scaled benchmark timing, not measured team productivity or realized savings.

Calculation

LLM eval sample size calculator

How many runs would separate two fixed pass rates?

You enter: Enter two pass rates and an optional cost per run.

  • First equal sample size with non-overlapping Wilson intervals
  • Interval charts and tables
  • Total runs and optional scenario cost

An independent fixed-rate calculation, not statistical power or a prediction for your tasks.

Calculation

Cost per correct answer calculator

What could independent retries add to answer cost?

You enter: Choose two recorded hard-task configurations and a requested number of correct answers.

  • Expected attempts at the recorded strict pass rate
  • Qualified list-price cost and sensitivity range
  • Recorded median time scaled as an illustrative proxy

Fixed independent retries are an assumption; unknown costs remain unknown. This is not a vendor invoice.

Recorded evidence

AI model efficiency explorer

How do two recorded metrics compare within one study?

You enter: Choose one study and two source-qualified metrics.

  • Exact entity/configuration coordinates
  • Both axes retain recorded intervals, ranges and units
  • Missing and ambiguous joins remain visible

Aggregate study values, not paired raw runs or a rank. No cross-study plot.

Read the result with its limits

Recorded results describe the tested tasks, routes and dates. Calculator outputs depend on your inputs and the stated prices or timing. They do not prove a bill, customer savings or future performance.

Start with our benchmark methodology, or inspect the source studies.

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.