OnlySOTA

Coding

Writing and changing code, split by who does the judging: a grader running tests and a person picking the better of two answers measure different things, and often disagree.

As of 2026-10-10: Claude Opus 5.5 leads Artificial Analysis LLM Leaderboard · SciCode (66.9%). Claude Opus 5.5 leads Arena — Code · Overall (1813). The kill line’s executioner is Claude Haiku 5.5, at a blended $0.20 per million tokens.

SOTA
14models
Scores above the executioner
No. 1 · AA
Claude Opus 5.5
66.9% · $8.00
Executioner · price
Claude Haiku 5.5
$0.20
Cheapest SOTA
MiMo V2.6 Pro
60.9% · $0.54
Ranking · AA 42–67% · A-CODE 1442–1828
  1. 166.9%1813
  2. 263.1%1744
  3. 361.8%1678
  4. 4Legacy61.0%1625
  5. 561.0%1774
  6. 660.9%1629
  7. 7Legacy59.8%1592
  8. 859.7%1657
  9. 959.5%1654
  10. 1059.0%1622
  11. 1158.9%1564
  12. 12Legacy58.8%1542
  13. 1358.7%1447
  14. 14Legacy57.8%1618
  15. 1557.8%1639
  16. 16Legacy57.6%1688
  17. 17Legacy57.4%1533
  18. 1856.6%1583
  19. 1956.5%1786
  20. 20Legacy56.5%1617
  21. 21Legacy56.4%1691
  22. 22Legacy56.1%1513
  23. 2355.8%1755
  24. 24NewExecutioner55.0%1587
All 125 models on AA →
Kill line · score × price

The executioner is the best value on the board: every model that costs more and scores no higher is killed. How the kill line is drawn →Who held it before →

SciCode · starts at 020%40%60%0$0.05$0.25$1.00$5.00$20.0066.9%$8.00Claude Opus 5.5ExecutionerClaude Haiku 5.555.0% · $0.20SOTALow-costKilledBlended Price, USD per 1M tokens (3:1 input:output) · log10
Kill lineFrontier: no model beats these on price and score at onceSOTALow-costKilledHollow dot or faint cross — estimated by the source

Claude Opus 5.5

anthropic/claude-opus-5-5 · closed · $8.00 · SOTA

Open full page →
Scores on each board
  • Artificial Analysis LLM Leaderboard57.6No. 1 of 375
  • Arena — Agent0.143No. 1 of 52
  • Artificial Analysis Coding Agents0.660No. 2 of 24
  • Epoch AI Benchmarking Hub91.2%No. 3 of 82
  • Terminal-Bench 4.064.8%No. 1 of 35
  • Arena — Code1813No. 1 of 142
  • Arena — Text1507No. 2 of 414
  • Arena — Vision1293No. 11 of 165
Scores at each effort setting

Artificial Analysis LLM Leaderboardbest at max

  1. low42.3
  2. medium51.2
  3. high53.6
  4. xhigh56.0
  5. max57.6

Tested at one setting only

  • Arena — Agent · agenthigh0.143
  • Artificial Analysis Coding Agents · Claude Codemax0.660
  • Epoch AI Benchmarking Hubmax91.2%
  • Terminal-Bench 4.0 · Claude Codemax64.8%
  • Arena — Codemax1813
  • Arena — Texthigh1507
  • Arena — Visionhigh1293

Where this board’s numbers come from

Graded

Scored against tests or a reference, with no human in the loop and no agent loop either — these measure the model's output directly.

Web development

Human preference between two finished web applications. Arena's Code Arena is web work end to end — every category it publishes is a web application category — so the board is named for what it measures rather than passed off as general coding.