The vocabulary on these boards comes from three places: the field, the sources, and this site. Which is which matters, because only the third kind is ours to define.
SOTA
State of the art. On this site it is not a compliment but a position: a model is SOTA when it sits above the executioner’s cross: it scores higher than the model that kills the most others, whatever it costs.
Benchmark
One published measurement, named as its publisher names it. A benchmark belongs to whoever ran it; the same name run by two organizations is two benchmarks here, because the harness, the prompts and the grading are part of the number.
Capability
A dimension several benchmarks speak to — coding, agentic work, cost. A capability page puts the contributing benchmarks side by side without combining them, so the question it answers is “where do the sources disagree about this”, not “what is the score”.
Board
A question a visitor arrives with — what should I use for this kind of work — answered by every source that speaks to it. A board adds sections, which are the sub-questions a domain actually has: models against harnesses, graded coding against coding judged by preference.
Elo
A rating from pairwise comparisons: people are shown two answers and pick one. An Elo has no zero. It is an arbitrary anchor, not “no skill”, which is why an Elo is never drawn as a bar growing from the left edge on this site — the distance between two rows is the reading, the distance from the edge is not.
Kill line and the executioner
The executioner is the model that kills the most others, where to kill is to be at least as cheap and at least as capable. The kill line is its cross, and it splits the board into SOTA, low-cost and killed. Both axes must come from one source; borrowing one source’s prices for another’s scores would draw a line nobody measured. The kill line walks through it in full.
The fainter line joining the models nothing beats on both axes at once is the frontier. It runs point to point, so a score read off a slope falls between two models rather than on one you can buy.
Variant
A run configuration, as distinct from a model. Reasoning effort, a harness, a
context window, a provider, a size, a release stage, a dated build — all of
these are settings on a model, not separate models. Which words count as
settings is a closed, hand-checked list, because Max is a setting on one
developer’s model and part of the name on another’s.
Checkpoint
A dated build — 20260813. A checkpoint is a rung, not a model: it never gets
an id of its own.
Line, or family
The tier a model belongs to, declared rather than parsed from its name. Claude Opus is a line and holds Opus 3 through Opus 5; GPT 5.6 carries a version and
is therefore not a line but a group over the models inside it. Deriving this
from names does not work — the longest shared prefix of two ids splits one line
by version, and an id coined before a rename, such as claude-4-opus, shares
no prefix with the line it belongs to.
Blended price
A single price per million tokens, mixing input and output at a fixed ratio, published by the source that publishes it. It is a comparison rate, not a quote: nothing on this site is an offer, and a real bill depends on your own traffic shape.
Unmatched
An upstream row whose name matches no model in the registry yet. It is shown, flagged, and kept on its source’s own board. It is absent from cross-source views for one reason only: it cannot be joined to anything until somebody confirms what it is.
Legacy
No longer listed as current by the source that published it. That is the source’s word, not the developer’s: a legacy model may still be served. Kept on the board and flagged, read at the best configuration still listed as current, and absent from the price chart and the kill line. The flag is per configuration: a model can have a legacy high-effort rung and a current one.
Estimated
Marked by the source as projected rather than measured. It is carried through and labeled, including on the kill line’s most load-bearing point.