AI Model Rankings — 11 leaderboards combined

RankModelOverall scoreArtificial AnalysisLiveBenchEpoch AIARC-AGIScale SEALVals AITerminal-BenchEQ-BenchSimpleBenchBenchLM
192.5#1#2—#2—#1——#1—
292.1#3#4#1#1#1#4#1—#4#1
391.7#2#1#2#5#2#2#2—#2#2
490.6#5#3#3#6—#5#4#2—#5
5
Anthropic
90.2#4#9#4#4—#3#3#1#8#4
689.0#6#5#15——#6——#7—
788.1#8#6#6#3—#7#7#9#23#3
886.9——#5#9————#10—
9
OpenAI
86.5#7#11———#8——#15—
1085.2——#9#13——————
1185.2#10————#16————
1284.7#22#6#28——#20——#21—
1384.6#9#19———#11#6———
14
OpenAI
84.4#25#8#8#8—#22—#4#19#11
15
Moonshot AI
83.3#15#11#12#24—#21—#3#33#8
1683.2#19#31#13#6#4#9#11—#5#7
1782.8#23#13#11#9—#17#15———
18
Alibaba (Qwen)
82.4#11#14#26——#33———#9
1982.3#13#15#18#20—#18#10—#13#12
20
OpenAI
82.2—#15#14#16———#7—#12
2182.0#21#15#27——#23——#14—
2182.0#16#18#7#12—#13#9#11#53#6
23
Z.ai (Zhipu)
82.0#12#28#22——#24#5—#22—
2481.4#17#26#10#17—#10#8#6#23#10
2580.1—#25#19#15#9#25—#5#26—
26
Alibaba (Qwen)
78.7#20#26————————
2777.8#26#30#20——#13#13#10#34#14
2877.3#39#22#29#14#3#41—#21#9—
2976.2#29#19#24#23—#32—#18#51—
3076.1#33#33#33——#28—#8—#15
3175.5#24#31#34#32—#34#13—#18#16
3274.8—#37#25#19#10——#13#20#18
3374.4——#39—#6————#24
3474.0#31#38#31#22—#29——#31—
3574.0#36#35#30#17—#30—#24#11#19
3673.7#32#40#32#24—#26———#17
3773.7#27#40#21#28—#12#12#19#55—
38
Alibaba (Qwen)
73.5#33#33———#36————
3873.5#27#46—#29—#19————
4070.3——#37#37#7———#12#25
4169.9—#35——#12—————
42
Z.ai (Zhipu)
69.7#17#48#41——#37————
4369.5——#23#30————#41—
4469.2——#16——————#23
4568.8#38#44#38#24—#35—#16—#28
46
Z.ai (Zhipu)
68.5#33#42#43#38—#30—#14#38#21
47
Alibaba (Qwen)
66.0#40#43#35——#38—#23—#20
48
Alibaba (Qwen)
65.8——#53—————#25—
49
OpenAI
65.4——#36#31————#58#22
50
Alibaba (Qwen)
63.4——#52—————#36—

FAQ

What is ARBITER?

An overall ranking of AI models and providers that statistically combines 11 authoritative AI leaderboards.

How often is it updated?

Every hour. New models on any leaderboard are added automatically.

How is the overall score calculated?

Each site’s main metric is standardized, then a factor model corrects for differences between sites and estimates each model’s shared ability.

Are models rated on fewer leaderboards at a disadvantage?

No. Missing ratings are never penalized. Uncertainty is shown as a 90% interval instead.

Which AI model is the best right now?

The #1 model in the overall ranking. The best choice depends on your use case, so also check the domain rankings.