AI Model Rankings — 11 leaderboards combined

RankProviderOverall scoreTop modelModels
186.7Claude Opus 5.527
285.2GPT-6 Astra54
377.7Muse Spark 1.323
474.4Gemini 3.8 Flash33
572.7Grok 4.713
670.1MiMo-V2.6-Pro4
769.8Kimi K35
868.1Qwen3.8 Max46
967.8DeepSeek V4.1 Flash22
1066.0GLM-5.310
1153.3—2
1242.0Inkling2
1340.3MiniMax-M35
1432.7—5
1529.6Nemotron 3 Ultra 550B A55B9
1617.6Mistral 3.526
1716.7Phi-46
1815.8—3
1913.3Mercury 2.51
2012.9Command R3

A provider’s result on each site is that of its best model there.

FAQ

What is ARBITER?

An overall ranking of AI models and providers that statistically combines 11 authoritative AI leaderboards.

How often is it updated?

Every hour. New models on any leaderboard are added automatically.

How is the overall score calculated?

Each site’s main metric is standardized, then a factor model corrects for differences between sites and estimates each model’s shared ability.

Are models rated on fewer leaderboards at a disadvantage?

No. Missing ratings are never penalized. Uncertainty is shown as a 90% interval instead.

Which AI model is the best right now?

The #1 model in the overall ranking. The best choice depends on your use case, so also check the domain rankings.