AI Model Rankings — 11 leaderboards combined

How the ranking works

  1. Collect — Official data from the 11 leaderboards below is fetched every hour. Scores marked as estimates are not used.
  2. Match — Different spellings of the same model are merged. For variants such as reasoning effort, each site’s best result is used.
  3. Standardize — Each site’s main metric is standardized within that site, keeping score gaps rather than just ranks.
  4. Equate — A factor model, bridged by models listed on several sites, corrects for each site’s difficulty and model mix. Noisier sites get less weight.
  5. No penalty for missing data — Models are scored only from the sites that rated them. Uncertainty is shown as a 90% interval. Models rated on 2+ sites are ranked.
  6. Arena — Arena (formerly LMArena) is used for domain rankings only, not for the overall score. New models are added automatically.

FAQ

What is ARBITER?

An overall ranking of AI models and providers that statistically combines 11 authoritative AI leaderboards.

How often is it updated?

Every hour. New models on any leaderboard are added automatically.

How is the overall score calculated?

Each site’s main metric is standardized, then a factor model corrects for differences between sites and estimates each model’s shared ability.

Are models rated on fewer leaderboards at a disadvantage?

No. Missing ratings are never penalized. Uncertainty is shown as a 90% interval instead.

Which AI model is the best right now?

The #1 model in the overall ranking. The best choice depends on your use case, so also check the domain rankings.