Agent
Not one answer but a task seen through, with tools. Two questions live here, and they are not the same one: which model to send in, and which harness to drive it with.
As of 2026-10-10: Claude Sonnet 5.5 leads Artificial Analysis LLM Leaderboard · Terminal-Bench 4.0 (63.6%). Claude Sonnet 5.5 leads Artificial Analysis Coding Agent Index (0.684). Claude Opus 5.5 leads Arena — Agent · Overall (0.143). The kill line’s executioner is Claude Haiku 5.5, at a blended $0.20 per million tokens.
- SOTA
- 13models
- Scores above the executioner
- No. 1 · AA
- Claude Sonnet 5.5
- 63.6% · $4.00
- Executioner · price
- Claude Haiku 5.5
- $0.20
- Cheapest SOTA
- Ling 3.1 Flash
- 33.3% · $0.45
- 163.6%0.6840.120
- 259.6%0.6600.143
- 359.6%0.6160.131
- 457.1%0.6380.093
- 556.1%0.6290.117
- 655.1%0.6220.127
- 7Legacy49.0%0.5970.081
- 8Legacy43.9%0.5670.099
- 9Legacy42.4%—0.089
- 1041.9%0.5360.021
- 11Legacy39.9%0.5460.062
- 1238.9%0.4330.023
- 1335.4%—0.009
- 1434.8%—0.035
- 15New33.3%——
- 1633.3%0.5430.040
- 1733.3%0.5120.000
- 18NewExecutioner32.8%0.414—
The executioner is the best value on the board: every model that costs more and scores no higher is killed. How the kill line is drawn →Who held it before →
Claude Sonnet 5.5
- Artificial Analysis LLM Leaderboard56.0No. 2 of 375
- Arena — Agent0.120No. 4 of 52
- Artificial Analysis Coding Agents0.684No. 1 of 24
- Epoch AI Benchmarking Hub88.8%No. 7 of 82
- Terminal-Bench 4.061.8%No. 2 of 35
- Arena — Code1774No. 3 of 142
- Arena — Text1476No. 32 of 414
- Arena — Vision1266No. 39 of 165
Artificial Analysis LLM Leaderboardbest at max
- low35.9
- medium40.8
- high46.8
- xhigh51.9
- max56.0
Artificial Analysis Coding Agents · Claude Codebest at max
- low0.421
- medium0.459
- high0.550
- xhigh0.629
- max0.684
Arena — Codebest at xhigh
- high1715
- xhigh1774
Tested at one setting only
- Arena — Agent · agentmax0.120
- Epoch AI Benchmarking Hubmax88.8%
- Terminal-Bench 4.0 · Claude Codemax61.8%
- Arena — Textxhigh1476
- Arena — Visionxhigh1266
Where this board’s numbers come from Each panel is one board’s ranking as published, never averaged. When two disagree, they are usually measuring different things.
Models
The model itself, whatever it was driven by. Artificial Analysis grades completed tasks; Arena observes real sessions and scores tool reliability, task completion and steerability over its own traffic.
Harnesses
The same rows read the other way round: an agent is a program, and the model is what it drives. HarnessTax runs every model under every harness, so a gap between two of its rows is the harness's alone, but it ran once and stopped. Artificial Analysis covers more harnesses, mostly on their own vendor's models, and Terminal-Bench's own board keeps taking runs.