OnlySOTA

Speed

How fast output arrives — sustained throughput and time to the first token, which trade off against each other and against cost. Measured on a provider's serving of the model, so it describes an endpoint rather than the weights.

Model In top 10 Artificial Analysis LLM Leaderboard Output Speed Artificial Analysis LLM Leaderboard Time To First Token
Celeris 1 Celeris 2/2 #1 1523/s #6 0.56s
Nemotron 3.5 Lightning NVIDIA 2/2 #10 307/s #7 0.57s
Gemini 2.5 Flash Lite Google 1/2 #12 289/s #1 0.32s
Mercury 2.5 Inception 1/2 #2 818/s #135 3.21s
Gemini 2.5 Flash Google 1/2 #28 208/s #2 0.46s
Mercury 2 Inception 1/2 #3 456/s #148 6.51s
Claude Haiku 4.5 Anthropic 1/2 #95 90/s #3 0.47s
Gemini 3.5 Flash Lite Google 1/2 #4 365/s #151 8.96s
Command A+ Cohere 1/2 #35 175/s #4 0.48s
Trinity Large Thinking Arcee AI 1/2 #5 339/s #60 1.21s
Granite 4.2 IBM 1/2 #25 215/s #5 0.53s
Ling 3.0 Flash InclusionAI 1/2 #6 334/s #116 2.50s
Muse Spark 1.2 Meta 1/2 #7 325/s #154 12.05s
Gemini 3.1 Flash Lite Google 1/2 #8 323/s #146 5.54s
Ministral 3 Mistral 1/2 #23 222/s #8 0.59s
Ling 3.0 Flash Fin InclusionAI 1/2 #9 321/s #80 1.70s
Grok 4.20 SpaceXAI 1/2 #84 103/s #9 0.59s
North Mini Code Cohere 1/2 #145 49/s #10 0.60s