Agentic work
Multi-step tasks drawn from real occupations rather than from software engineering — spreadsheets, documents, customer-facing procedures — run through a tool-using loop. Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged. A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.
| Model | In top 10 | Artificial Analysis LLM Leaderboard GDPval-AA | Arena — Agent Overall |
|---|---|---|---|
| | 2/2 | #1 68.3% | #1 0.143 |
| | 2/2 | #2 67.0% | #4 0.120 |
| | 2/2 | #3 62.8% | #3 0.127 |
| | 2/2 | #4 61.2% | #9 0.081 |
| | 1/2 | #23 52.1% | #2 0.131 |
| | 1/2 | #5 60.6% | #20 0.028 |
| | 1/2 | #21 53.8% | #5 0.117 |
| | 1/2 | #6 59.2% | #18 0.035 |
| | 1/2 | #25 50.4% | #6 0.099 |
| | 1/2 | #7 59.0% | #14 0.040 |
| | 1/2 | #13 56.3% | #7 0.093 |
| | 1/2 | #8 58.6% | #22 0.023 |
| | 1/2 | #16 55.6% | #8 0.089 |
| | 1/2 | #9 57.5% | #23 0.021 |
| | 1/2 | #10 57.3% | #28 0.011 |
| | 1/2 | #29 47.8% | #10 0.071 |