OnlySOTA

Agent

Not one answer but a task seen through, with tools. Two questions live here, and they are not the same one: which model to send in, and which harness to drive it with.

As of 2026-10-10: Claude Sonnet 5.5 leads Artificial Analysis LLM Leaderboard · Terminal-Bench 4.0 (63.6%). Claude Sonnet 5.5 leads Artificial Analysis Coding Agent Index (0.684). Claude Opus 5.5 leads Arena — Agent · Overall (0.143). The kill line’s executioner is Claude Haiku 5.5, at a blended $0.20 per million tokens.

SOTA
13models
Scores above the executioner
No. 1 · AA
Claude Sonnet 5.5
63.6% · $4.00
Executioner · price
Claude Haiku 5.5
$0.20
Cheapest SOTA
Ling 3.1 Flash
33.3% · $0.45
Ranking · AA 1–64% · AA-CODE 0.28–0.68 · A-AGENT -0.02–0.17
  1. 163.6%0.6840.120
  2. 259.6%0.6600.143
  3. 359.6%0.6160.131
  4. 457.1%0.6380.093
  5. 556.1%0.6290.117
  6. 655.1%0.6220.127
  7. 7Legacy49.0%0.5970.081
  8. 8Legacy43.9%0.5670.099
  9. 9Legacy42.4%—0.089
  10. 1041.9%0.5360.021
  11. 11Legacy39.9%0.5460.062
  12. 1238.9%0.4330.023
  13. 1335.4%—0.009
  14. 1434.8%—0.035
  15. 15New33.3%——
  16. 1633.3%0.5430.040
  17. 1733.3%0.5120.000
  18. 18NewExecutioner32.8%0.414—
All 119 models on AA →
Kill line · score × price

The executioner is the best value on the board: every model that costs more and scores no higher is killed. How the kill line is drawn →Who held it before →

Terminal-Bench 4.0 · starts at 020%40%60%0$0.05$0.25$1.00$20.0063.6%$4.00Claude Sonnet 5.5ExecutionerClaude Haiku 5.532.8% · $0.20SOTALow-costKilledBlended Price, USD per 1M tokens (3:1 input:output) · log10
Kill lineFrontier: no model beats these on price and score at onceSOTALow-costKilledHollow dot or faint cross — estimated by the source

Claude Sonnet 5.5

anthropic/claude-sonnet-5-5 · closed · $4.00 · SOTA

Open full page →
Scores on each board
  • Artificial Analysis LLM Leaderboard56.0No. 2 of 375
  • Arena — Agent0.120No. 4 of 52
  • Artificial Analysis Coding Agents0.684No. 1 of 24
  • Epoch AI Benchmarking Hub88.8%No. 7 of 82
  • Terminal-Bench 4.061.8%No. 2 of 35
  • Arena — Code1774No. 3 of 142
  • Arena — Text1476No. 32 of 414
  • Arena — Vision1266No. 39 of 165
Scores at each effort setting

Artificial Analysis LLM Leaderboardbest at max

  1. low35.9
  2. medium40.8
  3. high46.8
  4. xhigh51.9
  5. max56.0

Artificial Analysis Coding Agents · Claude Codebest at max

  1. low0.421
  2. medium0.459
  3. high0.550
  4. xhigh0.629
  5. max0.684

Arena — Codebest at xhigh

  1. high1715
  2. xhigh1774

Tested at one setting only

  • Arena — Agent · agentmax0.120
  • Epoch AI Benchmarking Hubmax88.8%
  • Terminal-Bench 4.0 · Claude Codemax61.8%
  • Arena — Textxhigh1476
  • Arena — Visionxhigh1266

Where this board’s numbers come from

Models

The model itself, whatever it was driven by. Artificial Analysis grades completed tasks; Arena observes real sessions and scores tool reliability, task completion and steerability over its own traffic.

Harnesses

The same rows read the other way round: an agent is a program, and the model is what it drives. HarnessTax runs every model under every harness, so a gap between two of its rows is the harness's alone, but it ran once and stopped. Artificial Analysis covers more harnesses, mostly on their own vendor's models, and Terminal-Bench's own board keeps taking runs.