Families, each fielding its best model
Each line, such as Claude Opus or Gemini Flash, is represented by its strongest model, row intact.
- SOTA
- 13models
- Scores above the executioner
- No. 1 · AA
- Claude Opus 5.5
- 57.6 · $8.00
- Executioner · price
- Claude Haiku 5.5
- $0.20
- Cheapest SOTA
- MiMo V2.6 Pro
- 46.3 · $0.54
Ranking · AA 28–58 · ARENA 1416–1534Every model that outscores the executioner, with the executioner itself as the last row for reference. Each column is one board’s own score, side by side and never combined; a dash means that board does not list the model. Click a column head to sort by it. The rank is that board’s own, and hovering shows every board’s. A model marked Legacy is no longer listed as current by its source, and its row shows the best setting still listed.
- 1scores from Claude Opus 5.557.61507
- 2scores from GPT 6 Astra52.71475
- 352.61525
- 448.11494
- 5scores from Grok 4.746.41443
- 646.31480
- 745.41483
- 844.81478
- 943.71447
- 1043.61488
- 11NewExecutioner43.4—
Kill line · score × price The executioner is the best value on the board: every model that costs more and scores no higher is killed.Each dot is a model: further right costs more, higher up scores more. Where the dashed lines cross stands the executioner. Above it is SOTA, the top tier; below and to the left is cheaper and weaker, the low-cost picks.65 of 102 models are killed. The faint line joins the models nothing beats on both price and score. Legacy models are left off the chart. Select a model to open its details below.How the kill line is drawn →Who held it before →
The executioner is the best value on the board: every model that costs more and scores no higher is killed. How the kill line is drawn →Who held it before →
Kill lineFrontier: no model beats these on price and score at onceSOTALow-costKilledHollow dot or faint cross — estimated by the source
Claude Opus 5.5
Scores on each board Each rank is the board’s own: the rank it publishes, or the model’s place in the board’s own order where it publishes none. The bar shows how much of that board the model ranks ahead of. Scores from different boards cannot be compared.
- Artificial Analysis LLM Leaderboard57.6No. 1 of 375
- Arena — Agent0.143No. 1 of 52
- Artificial Analysis Coding Agents0.660No. 2 of 24
- Epoch AI Benchmarking Hub91.2%No. 3 of 82
- Terminal-Bench 4.064.8%No. 1 of 35
- Arena — Code1813No. 1 of 142
- Arena — Text1507No. 2 of 414
- Arena — Vision1293No. 11 of 165
Scores at each effort setting The same model at each effort setting. A bar runs from zero to that board’s best score, so read each board on its own and never across boards.
Artificial Analysis LLM Leaderboardbest at max
- low42.3
- medium51.2
- high53.6
- xhigh56.0
- max57.6
Tested at one setting only
- Arena — Agent · agenthigh0.143
- Artificial Analysis Coding Agents · Claude Codemax0.660
- Epoch AI Benchmarking Hubmax91.2%
- Terminal-Bench 4.0 · Claude Codemax64.8%
- Arena — Codemax1813
- Arena — Texthigh1507
- Arena — Visionhigh1293