Coding
Writing and changing code, split by who does the judging: a grader running tests and a person picking the better of two answers measure different things, and often disagree.
As of 2026-10-10: Claude Opus 5.5 leads Artificial Analysis LLM Leaderboard · SciCode (66.9%). Claude Opus 5.5 leads Arena — Code · Overall (1813). The kill line’s executioner is Claude Haiku 5.5, at a blended $0.20 per million tokens.
- SOTA
- 14models
- Scores above the executioner
- No. 1 · AA
- Claude Opus 5.5
- 66.9% · $8.00
- Executioner · price
- Claude Haiku 5.5
- $0.20
- Cheapest SOTA
- MiMo V2.6 Pro
- 60.9% · $0.54
- 166.9%1813
- 263.1%1744
- 361.8%1678
- 4Legacy61.0%1625
- 561.0%1774
- 660.9%1629
- 7Legacy59.8%1592
- 859.7%1657
- 959.5%1654
- 1059.0%1622
- 1158.9%1564
- 12Legacy58.8%1542
- 1358.7%1447
- 14Legacy57.8%1618
- 1557.8%1639
- 16Legacy57.6%1688
- 17Legacy57.4%1533
- 1856.6%1583
- 1956.5%1786
- 20Legacy56.5%1617
- 21Legacy56.4%1691
- 22Legacy56.1%1513
- 2355.8%1755
- 24NewExecutioner55.0%1587
The executioner is the best value on the board: every model that costs more and scores no higher is killed. How the kill line is drawn →Who held it before →
Claude Opus 5.5
- Artificial Analysis LLM Leaderboard57.6No. 1 of 375
- Arena — Agent0.143No. 1 of 52
- Artificial Analysis Coding Agents0.660No. 2 of 24
- Epoch AI Benchmarking Hub91.2%No. 3 of 82
- Terminal-Bench 4.064.8%No. 1 of 35
- Arena — Code1813No. 1 of 142
- Arena — Text1507No. 2 of 414
- Arena — Vision1293No. 11 of 165
Artificial Analysis LLM Leaderboardbest at max
- low42.3
- medium51.2
- high53.6
- xhigh56.0
- max57.6
Tested at one setting only
- Arena — Agent · agenthigh0.143
- Artificial Analysis Coding Agents · Claude Codemax0.660
- Epoch AI Benchmarking Hubmax91.2%
- Terminal-Bench 4.0 · Claude Codemax64.8%
- Arena — Codemax1813
- Arena — Texthigh1507
- Arena — Visionhigh1293
Where this board’s numbers come from Each panel is one board’s ranking as published, never averaged. When two disagree, they are usually measuring different things.
Graded
Scored against tests or a reference, with no human in the loop and no agent loop either — these measure the model's output directly.
Web development
Human preference between two finished web applications. Arena's Code Arena is web work end to end — every category it publishes is a web application category — so the board is named for what it measures rather than passed off as general coding.