Agentic coding
Resolving real repository and terminal tasks through a tool-using loop. Every score here measures a model together with the harness that drives it, so one model can hold two rows, one per harness, and neither is a duplicate. Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged. A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.
| Model | In top 10 | Artificial Analysis LLM Leaderboard Terminal-Bench 4.0 | Artificial Analysis Coding Agents Artificial Analysis Coding Agent Index | Artificial Analysis Coding Agents Terminal-Bench v4 | HarnessTax — SWE-bench Lite SWE-bench Lite | HarnessTax — Terminal-Bench 2.0 Terminal-Bench 2.0 | Terminal-Bench 4.0 Terminal-Bench 4.0 |
|---|---|---|---|---|---|---|---|
| | 4/6 | #1 63.6% | #1 0.684 | #1 66.2% | — | — | #2 61.8% |
| | 4/6 | #2 59.6% | #2 0.660 | #2 63.1% | — | — | #1 64.8% |
| | 4/6 | #9 42.4% | — | — | #1 97.8% | #3 75.6% | #8 44.5% |
| | 4/6 | #11 39.9% | #10 0.546 | #10 37.4% | #3 77.8% | #1 83.3% | #11 37.3% |
| | 4/6 | #3 59.6% | #6 0.616 | #5 55.6% | — | — | #3 58.2% |
| | 4/6 | #6 55.1% | #5 0.622 | #3 57.6% | — | — | #5 57.9% |
| | 4/6 | #5 56.1% | #4 0.629 | #6 54.5% | — | — | #4 58.2% |
| | 4/6 | #7 49.0% | #7 0.597 | #7 54.5% | — | — | #6 53.9% |
| | 4/6 | #8 43.9% | #8 0.567 | #8 43.4% | — | — | #7 49.4% |
| | 3/6 | #4 57.1% | #3 0.638 | #4 56.1% | — | — | — |
| | 3/6 | #10 41.9% | #12 0.536 | #9 39.9% | — | — | #9 41.8% |
| | 2/6 | #25 21.7% | — | — | #2 88.9% | #5 72.2% | #14 23.6% |
| | 2/6 | #37 11.6% | #17 0.432 | #19 14.6% | #7 55.6% | #2 76.7% | #18 17.3% |
| | 2/6 | #34 12.6% | #13 0.519 | #15 21.2% | #4 76.7% | #4 73.3% | — |
| | 2/6 | #45 3.0% | — | — | #5 68.9% | #6 65.6% | — |
| | 2/6 | #75 0.0% | — | — | #6 60.0% | #7 47.8% | — |
| | 2/6 | #22 25.8% | #9 0.563 | #11 33.3% | — | — | #10 37.6% |