Reasoning and knowledge
Hard questions with verifiable answers, where the difficulty is in the reasoning rather than in recall alone. Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged. A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.
| Model | In top 10 | Artificial Analysis LLM Leaderboard Humanity's Last Exam | Artificial Analysis LLM Leaderboard CritPt | Epoch AI Benchmarking Hub FrontierMath Tiers 1–3 | Epoch AI Benchmarking Hub FrontierMath Tier 4 |
|---|---|---|---|---|---|
| | 4/4 | #1 61.4% | #2 31.7% | #3 91.2% | #3 95.0% |
| | 4/4 | #7 54.7% | #3 31.7% | #1 93.7% | #2 97.6% |
| | 4/4 | #8 52.9% | #4 31.7% | #2 93.7% | #1 100.0% |
| | 4/4 | #9 49.5% | #1 32.3% | #6 89.1% | #7 82.9% |
| | 4/4 | #2 59.1% | #6 31.1% | #4 90.2% | #6 87.8% |
| | 4/4 | #5 55.0% | #5 31.4% | #7 88.8% | #8 80.5% |
| | 3/4 | #4 55.5% | #12 28.6% | #9 87.0% | #4 90.2% |
| | 3/4 | #13 47.9% | #7 30.9% | #5 89.8% | #5 90.0% |
| | 3/4 | — | #8 30.6% | #8 87.7% | #9 78.0% |
| | 2/4 | #6 54.9% | #11 29.1% | #11 85.6% | #10 73.2% |
| | 2/4 | #28 42.9% | #10 30.0% | #10 86.0% | #12 70.7% |
| | 1/4 | #3 57.1% | #14 27.1% | — | — |
| | 1/4 | — | #9 30.0% | #13 82.5% | #14 58.5% |
| | 1/4 | #10 49.4% | #15 26.6% | — | — |