AI Model Rankings — 11 leaderboards combined

FAQ

What is ARBITER?

An overall ranking of AI models and providers that statistically combines 11 authoritative AI leaderboards.

How often is it updated?

Every hour. New models on any leaderboard are added automatically.

How is the overall score calculated?

Each site’s main metric is standardized, then a factor model corrects for differences between sites and estimates each model’s shared ability.

Are models rated on fewer leaderboards at a disadvantage?

No. Missing ratings are never penalized. Uncertainty is shown as a 90% interval instead.

Which AI model is the best right now?

The #1 model in the overall ranking. The best choice depends on your use case, so also check the domain rankings.