OnlySOTA does not evaluate models. Every number on this site was published by somebody else, and every one of them is shown beside the name of whoever published it. What this site adds is the join: the same model, under whatever name each source spells it, placed next to itself.
That sounds smaller than it is. A leaderboard is easy to read and hard to compare, because the hard part is never the ranking — it is knowing whether the row you are looking at on one board is the same thing as the row you are looking at on another.
Where the numbers come from
Every source is fetched on a schedule and stored as a committed snapshot. Fetching and building are separate jobs: the site is rebuilt from files that are already on disk, never from a live call to anyone’s API. Two consequences follow, and both are deliberate. A deploy cannot be broken by an upstream outage. And the date beside a figure is the date it was fetched, not the date the page was rebuilt — so a page that has not changed says so honestly.
Every source’s own board is reproduced at /leaderboards, close to the shape
its authors publish it in, with a provenance bar naming the publisher, the
license where one is stated, and how often they refresh.
What normalizing means here
Three things, and nothing beyond them.
Names are resolved to models. claude-opus-4-5, Claude Opus 4.5 and
anthropic/claude-opus-4.5 are one model. The match is exact against a
hand-curated alias list — there is no fuzzy fallback, because a fuzzy match that
is wrong is invisible, and an aggregator that quietly merges two different
models is worse than one that admits it cannot tell.
Run configuration is separated from identity. Claude Opus 5 (Max Effort)
is one model and one setting, not a second model. The same goes for a size, a
release stage, and a dated build: eleven scales of one release are one model at
eleven rungs.
A name that matches nothing is still shown. It is flagged as unmatched and left on its source’s board rather than dropped. Silent dropping is the failure mode that rots an aggregator invisibly — the board stays plausible while it quietly stops being complete.
What is never done: there is no blended score
Nothing on this site averages, weights or otherwise fuses scores across sources into a single ranking. Not on the home page, not in the structured data, not anywhere.
This is the one refusal worth stating plainly, because it is the feature every aggregator is asked for. Two boards averaged read as one board that agrees — and the agreement is manufactured. Benchmarks measure different things under different conditions: a graded examination and a count of which answer people preferred do not become more true by being added together. Where two sources disagree, the disagreement is the reading, and both are right about what they measured.
So panels sit beside each other, each keeping its own order, and the columns never merge.
The kill line
The one figure here that is derived rather than reproduced, and it is derived from a single source’s own two axes — capability and price — never from two sources borrowed from each other.
The executioner is the model that kills the most others, where killing means being at least as cheap and at least as capable. Its cross splits the board in three: everything that scores higher is SOTA, everything cheaper and lower-scoring is low-cost, and everything that costs at least as much and scores no higher is killed. It changes hands on a repricing, not only on a release, which is why it is recomputed on every refresh.
A legacy model is placed against that cross but never draws it: the line is drawn among the models the source lists as current.
What the New badge means
Nothing upstream publishes an arrival date, and this site does not invent one. A New badge means exactly one thing: a row turned up on a board we were already watching, within the last seven days.
A source’s first snapshot is a baseline and badges nothing — otherwise the day a board is added, every model on it looks new. A snapshot that brings a fifth of a board at once is treated as a different population rather than as news. And the ledger is keyed per source by the model each row resolves to, so neither a source respelling a name nor this site learning a new alias counts as an arrival.
What the Legacy flag means
One source hides more than half its rows behind a “current” filter. Those rows are kept and flagged rather than hidden, because “what should I use” wants them gone and “how good did this get” does not.
The flag is the source’s, not the developer’s. A legacy model has left the source’s current list, which usually means something newer replaced it; its developer may well still serve it. Claude Opus 5 is legacy on that board and can still be called.
A legacy model keeps its place on a board, reads the best configuration the source still lists as current, and is off the price chart and the kill line, which compare only what the source is still tracking.
What is not here
No re-run evaluations. No blended score. No editorial ranking of the sources against each other. No number without the name of whoever published it.