- Learn
29 explainers · updated October 8, 2026
The terms behind the benchmarks.
Short explainers on routing, caching, reasoning effort, benchmark statistics and inference cost. Each one starts with a definition and shows the idea on charts from our own measured studies, with sample sizes and intervals.
All explainers
Agent harness
What is an agent harness? The code around the model, measured
Agent harness is the code that wraps a language model and turns it into an agent. The model answers; the harness runs the loop, offers tools, builds prompts, keeps memory, manages context and checks the result. Two systems can share one model and still differ in time and tokens, so a benchmark scores a system, not a model.
Benchmark saturation
Benchmark saturation: when every model scores 100%
Benchmark saturation means most or all models score near the top of a test, so scores stop telling them apart. This is the ceiling effect. A perfect score means every observed call passed. It does not prove equal ability. Our sets hit ceilings while time, tokens and calculated list-price cost still differed.
Claude Code hook
Claude Code hooks, explained with a measured Stop hook
A Claude Code hook runs a command at an event, such as a tool call or session stop. Our Stop hook blocked finishing and sent standard-error feedback with exit code 2.
CLAUDE.md
What is CLAUDE.md? Agent memory files, measured
CLAUDE.md is a Markdown file in your project that Claude Code reads into its context at the start of every session. It holds facts the code cannot show, such as a team decision or the current test command, and it helped most there in our tests. Codex CLI reads AGENTS.md for the same job, but we measured Claude Code only.
CLI context tax
Tokens per call and the CLI context tax
Tokens per call is the number of input and output tokens one model request uses, and the CLI context tax is the part of the input that a coding CLI such as Claude Code or the Codex CLI adds on its own: its system prompt, tool definitions and environment context, sent with every request before your prompt. You pay for these tokens (or spend subscription quota on them), and the model must read them, even when your question is one line.
Context engineering
What is context engineering? What to put in the context, measured
Context engineering chooses what enters a model's context window on each call: instructions, memory, tools, retrieved files and history. It covers more than wording. In our tests, curated memory improved team-knowledge checks. Haiku used a stale command less often with consolidated notes than raw notes.
Cost per correct answer
Cost per correct answer: the LLM price that counts failures
Cost per correct answer is the total cost of every call, failed calls included, divided by the number of attempts that pass a check. An attempt can use one call or a full agent session. A lower token price can cost more per pass when a model writes more tokens or fails more often.
Eval sample size
How many runs does an LLM eval need? Sample size, with real intervals
Eval sample size counts the units behind a rate: tasks, attempts or checks. Choose independent cases for the precision you need. A narrow 95% interval does not guarantee power to detect a gap between systems.
Inference gateway
Inference gateways and OpenRouter, explained
An inference gateway is a service that puts many model providers behind one API: you send a request for a model to the gateway, and it forwards the request to one of the providers that serve that model, then returns the answer and one bill. OpenRouter is the best-known public example. A gateway gives you one key, one schema and fallback between providers; it can add a fee and a network hop, and it routes between providers of a model, not between models for a task.
Inference provider
What is an inference provider? Bedrock, Vertex, Azure and first-party prices
Inference provider is a company that sells API access to an AI model: the vendor, a cloud, a specialist host or a gateway. In our 2026-10-06 snapshot, Bedrock listed Claude Sonnet 5.5 at the same price as the Anthropic API: $2 input and $10 output per million tokens. For seven open-weight models, standard-tier price spreads ranged from 1.7x to 12.6x (a calculation). The largest spread covered 15 providers.
LLM as judge
LLM as judge: how blind review works, and where it fails
LLM as judge (or LLM-as-a-judge) is an evaluation method in which one or more language models grade or compare outputs, such as answers, essays or code changes, in place of human reviewers or fixed tests. In a blind review, the judge does not know which output came from which source, and in a pairwise setup it sees both outputs in both orders, so that labels and position cannot decide the verdict. It scales to tasks that have no exact answer, but its verdicts carry the judge's own biases.
LLM nondeterminism
Why the same prompt gives different answers: LLM nondeterminism, measured
LLM nondeterminism means a model can answer the same prompt differently. General causes include how the model samples each next token and how batching on shared servers can change tiny numeric details. In 90 repeated calls, the wording changed in 5 of 9 prompt-and-model pairs and the answer in 3 of 9.
LLM router
What is an LLM router?
An LLM router is the component that decides, before a request runs, which language model should answer it, at which reasoning effort and with how much context. The router can be a set of rules in code, a small model trained to make the decision, or a general LLM that you ask to choose. A good router sends easy work to cheap, fast models and hard work to strong ones, and adds less delay and cost than it saves.
McNemar test
The McNemar test: compare two models on the same cases
McNemar test, also called McNemar's test, compares two models on the same cases, such as two LLMs or two classifiers. It reads cases where exactly one model was right; the others cannot separate them. The exact version is a binomial test on those disagreements with a probability of 0.5.
Memory consolidation
Memory consolidation (dreaming) for AI agents, measured
Memory consolidation, also called dreaming, is a pass in which a model rewrites saved agent notes into one shorter file. It aims to keep current facts and drop stale notes and one-off events.
Open-weight model
Open-weight models: why the same model has many prices
Open-weight models are AI models whose trained weights are public, so anyone who accepts the license can download and run them. Many companies host the same open model. Our snapshot shows a price spread up to 12.6x (calculation, 15 standard-tier providers for DeepSeek V4 Flash 0423). All 15 closed models with 2 or more providers had one standard-tier price across those providers. This applies to our snapshot, not every closed model.
p95 latency
p95 latency, explained: median, tail and range for LLM calls
p95 latency marks the time at or below which about 95% of calls fall. About 1 call in 20 is slower. Small samples and tied times can change that share. It shows slow calls that the median (p50) hides. In 82 recorded routing calls, Claude Sonnet 5.5 through Claude Code had a median of 2.60 s and a p95 of 4.30 s (range 1.993 to 5.583 s).
Pareto frontier
The Pareto frontier: reading LLM cost against quality
Pareto frontier is the set of options that no rival matches or beats on every axis and beats on at least one. For large language models (LLMs), the axes are cost and quality, and sometimes speed. The frontier is your short list.
pass@k
pass@k explained: pass@1, pass@k and pass^k with real runs
pass@k is the chance that at least one of k tries at a task passes its check. pass@1 is the chance that one try passes, and pass^k (with a caret) is the chance that all k tries pass. So pass@k rewards a model that is sometimes right, and pass^k rewards one that is right every time.
Price per million tokens
LLM pricing per million tokens, explained with real token counts
Price per million tokens is an LLM vendor’s dollar rate per 1,000,000 tokens. Input, output and cached tokens have separate rates. Multiply each token count by its rate, divide by 1,000,000, then add the results. Our stored Sonnet 5.5 rates per million tokens are $2 for input, $10 for output and $0.20 for cache reads.
Prompt caching
Prompt caching, explained with measured sessions
Prompt caching is a feature of model APIs that stores the processed form of a prompt prefix, so that a later request that starts with the same prefix reads it from the cache instead of processing it again. A cache read is billed at a fraction of the normal input price, and the first request that writes the prefix may cost more than normal input. Caching saves money when the same long context, such as a system prompt, tool definitions or a document, is sent many times.
Quantization
Quantization and cheap inference, explained
Quantization is the practice of storing and running a model's weights (and sometimes its activations and cache) in fewer bits, such as 8-bit (fp8) or 4-bit (fp4) numbers instead of 16-bit (bf16 or fp16). A quantized model needs less memory and less compute per token, so a provider can serve it faster or cheaper on the same hardware. The trade is accuracy: a lower precision can change the model's outputs, by a little or a lot, depending on the model, the method and the task.
Reading an AI benchmark honestly
How to read AI benchmarks honestly
Reading an AI benchmark honestly means checking what was measured, on how many cases, under which configuration and with what uncertainty before you accept a ranking. A trustworthy result states its sample size, gives an interval for every rate and a range for every time, counts every failed attempt, separates measured runs from calculations, and names the setup (model, effort, route, harness) behind each number. A table that hides any of these can make a tie look like a win.
Reasoning effort
What is reasoning effort?
Reasoning effort is a setting on a reasoning model that controls how much internal work, usually in the form of hidden thinking tokens, the model does before it gives its answer. Typical levels are low, medium and high. A higher effort can help on hard problems, but it costs more output tokens, more time and more money per call, so the useful question is whether a higher effort changes the result on your tasks.
Reasoning tokens
Reasoning tokens, explained: the hidden output you pay for
Reasoning tokens are the tokens a reasoning model writes to think before it writes its visible answer. Vendors generally bill them as output tokens, and a tool may report their count without showing their text. A short answer can hide a long chain of thought that you pay for.
Strict grading
Strict grading and format misses: when a right answer fails
Strict grading requires the whole reply to pass the declared check. A right answer inside a code fence can fail. Lenient grading can count it after extraction. On our hard set, 139/152 passed strictly (95% interval 85.9% to 94.9%). With format misses included, 144/152 passed (90.0% to 97.3%). Here, "correct" means it passed the task validator, not a proof for all possible inputs.
SWE-bench Verified
SWE-bench Verified, explained
SWE-bench Verified is a benchmark of 500 real GitHub issues from popular Python repositories, each checked by people to have a clear problem statement and a fair test. A coding agent gets the repository and the issue text, writes a patch, and the patch counts as resolved only when the project's hidden tests that the real fix made pass now pass, and the tests that passed before still pass. The score is the share of instances resolved.
Time to first token
Time to first token (TTFT), explained with CLI and API timings
Time to first token (TTFT) is the wait from sending a model request to receiving its first reply token. Start-up, processing, queues and the network can add delay. We measure first useful output, a different clock from raw TTFT.
Wilson confidence interval
Wilson confidence intervals for AI benchmarks
A Wilson confidence interval (also called the Wilson score interval) is a range around a measured pass rate that shows which true pass rates are consistent with the result, given the number of trials. For a benchmark result of k passes in n attempts, the 95% Wilson interval is the range that would contain the true rate in about 95% of repeated experiments. Unlike the simple "plus or minus" interval, it stays inside 0% to 100% and stays honest at small n and at rates near 0% or 100%.