OnlySOTA

SOTA AI models ranked: LLM leaderboards and the kill line

The public AI leaderboards, gathered daily: who leads, and who charges more for no better score.

As of 2026-10-10: Claude Opus 5.5 leads Artificial Analysis Intelligence Index (57.6). Gemini 4 Argon leads Arena — Text · Overall (style control) (1525). The kill line’s executioner is Claude Haiku 5.5, at a blended $0.20 per million tokens.

SOTA
13models
Scores above the executioner
No. 1 · AA
Claude Opus 5.5
57.6 · $8.00
Executioner · price
Claude Haiku 5.5
$0.20
Cheapest SOTA
MiMo V2.6 Pro
46.3 · $0.54
Ranking · AA 28–58 · ARENA 1416–1534
  1. 157.61507
  2. 256.01476
  3. 353.41501
  4. 452.71475
  5. 552.61525
  6. 651.81484
  7. 7Legacy50.81490
  8. 8Legacy49.61504
  9. 948.11494
  10. 10Legacy47.61456
  11. 11Legacy47.01485
  12. 1246.41443
  13. 1346.31480
  14. 1445.41483
  15. 1544.81478
  16. 16Legacy44.31454
  17. 1743.71447
  18. 1843.61488
  19. 19NewExecutioner43.4—
All 375 models on AA →
Kill line · score × price

The executioner is the best value on the board: every model that costs more and scores no higher is killed. How the kill line is drawn →Who held it before →

Artificial Analysis Intelligence Index · starts at 02040600$0.05$0.25$1.00$5.00$20.0057.6$8.00Claude Opus 5.5ExecutionerClaude Haiku 5.543.4 · $0.20SOTALow-costKilledBlended Price, USD per 1M tokens (3:1 input:output) · log10
Kill lineFrontier: no model beats these on price and score at onceSOTALow-costKilledHollow dot or faint cross — estimated by the source

Claude Opus 5.5

anthropic/claude-opus-5-5 · closed · $8.00 · SOTA

Open full page →
Scores on each board
  • Artificial Analysis LLM Leaderboard57.6No. 1 of 375
  • Arena — Agent0.143No. 1 of 52
  • Artificial Analysis Coding Agents0.660No. 2 of 24
  • Epoch AI Benchmarking Hub91.2%No. 3 of 82
  • Terminal-Bench 4.064.8%No. 1 of 35
  • Arena — Code1813No. 1 of 142
  • Arena — Text1507No. 2 of 414
  • Arena — Vision1293No. 11 of 165
Scores at each effort setting

Artificial Analysis LLM Leaderboardbest at max

  1. low42.3
  2. medium51.2
  3. high53.6
  4. xhigh56.0
  5. max57.6

Tested at one setting only

  • Arena — Agent · agenthigh0.143
  • Artificial Analysis Coding Agents · Claude Codemax0.660
  • Epoch AI Benchmarking Hubmax91.2%
  • Terminal-Bench 4.0 · Claude Codemax64.8%
  • Arena — Codemax1813
  • Arena — Texthigh1507
  • Arena — Visionhigh1293