Factuality
Whether the model states things that are not so, and whether it declines when it does not know. Measured both as accuracy and as the rate of confident wrong answers, which move independently. Ordered by how many of these benchmarks put a model in their top 10, then by its best placing. Each column keeps its own ranking; nothing is averaged. A dash means the board does not list that model — not a score of zero. Entries not yet matched to a model are left out here; they still appear on their own board.
| Model | In top 10 | Artificial Analysis LLM Leaderboard AA-Omniscience Index | Artificial Analysis LLM Leaderboard AA-Omniscience Non-Hallucination Rate | Epoch AI Benchmarking Hub SimpleQA Verified |
|---|---|---|---|---|
| | 2/3 | #1 46.4 | #96 41.4% | #4 72.2% |
| | 2/3 | #2 43.7 | #68 55.2% | #1 75.6% |
| | 2/3 | #6 41.5 | #77 50.6% | #2 73.9% |
| | 2/3 | #3 43.5 | #116 34.4% | #5 70.8% |
| | 2/3 | #10 31.9 | #81 49.1% | #3 73.5% |
| | 2/3 | #4 43.3 | #110 36.4% | #6 70.7% |
| | 2/3 | #5 42.4 | #5 84.9% | — |
| | 1/3 | #73 -0.800 | #1 99.1% | — |
| | 1/3 | #81 -4.367 | #2 88.3% | — |
| | 1/3 | #54 3.767 | #3 87.0% | — |
| | 1/3 | #79 -4.017 | #4 85.8% | — |
| | 1/3 | #101 -10.9 | #6 84.0% | — |
| | 1/3 | #7 37.1 | #98 40.5% | #16 59.9% |
| | 1/3 | #12 29.6 | #86 48.1% | #7 69.7% |
| | 1/3 | #27 18.0 | #7 83.1% | #54 33.2% |
| | 1/3 | #8 32.3 | #72 53.0% | #33 46.5% |
| | 1/3 | #22 22.0 | #245 10.6% | #8 69.7% |
| | 1/3 | #31 14.8 | #8 82.6% | #58 30.2% |
| | 1/3 | #9 32.0 | #29 70.7% | #18 56.0% |
| | 1/3 | #18 26.5 | #113 35.5% | #9 69.2% |
| | 1/3 | #90 -7.950 | #9 81.9% | — |
| | 1/3 | #42 10.1 | #267 7.6% | #10 66.8% |
| | 1/3 | #59 1.350 | #10 81.6% | — |