Performance vs price
Top left = cheap and strong. Click a point for details.
| Rank | Provider | Overall score | Top model | Models |
|---|---|---|---|---|
| 1 | 86.7 | Claude Opus 5.5 | 27 | |
| 2 | 85.2 | GPT-6 Astra | 54 | |
| 3 | 77.7 | Muse Spark 1.3 | 23 | |
| 4 | 74.4 | Gemini 3.8 Flash | 33 | |
| 5 | 72.7 | Grok 4.7 | 13 | |
| 6 | 70.1 | MiMo-V2.6-Pro | 4 | |
| 7 | 69.8 | Kimi K3 | 5 | |
| 8 | 68.1 | Qwen3.8 Max | 46 | |
| 9 | 67.8 | DeepSeek V4.1 Flash | 22 | |
| 10 | 66.0 | GLM-5.3 | 10 | |
| 11 | 53.3 | — | 2 | |
| 12 | 42.0 | Inkling | 2 | |
| 13 | 40.3 | MiniMax-M3 | 5 | |
| 14 | 32.7 | — | 5 | |
| 15 | 29.6 | Nemotron 3 Ultra 550B A55B | 9 | |
| 16 | 17.6 | Mistral 3.5 | 26 | |
| 17 | 16.7 | Phi-4 | 6 | |
| 18 | 15.8 | — | 3 | |
| 19 | 13.3 | Mercury 2.5 | 1 | |
| 20 | 12.9 | Command R | 3 |
A provider’s result on each site is that of its best model there.
An overall ranking of AI models and providers that statistically combines 11 authoritative AI leaderboards.
Every hour. New models on any leaderboard are added automatically.
Each site’s main metric is standardized, then a factor model corrects for differences between sites and estimates each model’s shared ability.
No. Missing ratings are never penalized. Uncertainty is shown as a 90% interval instead.
The #1 model in the overall ranking. The best choice depends on your use case, so also check the domain rankings.