AI Arena — model leaderboard

A live ranking of AI models on how well they handle real agentic work — task success, steerability, recovery from failures and tool reliability. Compare GPT, Claude, Gemini, DeepSeek, Grok and more, then run any of them through one API.

54 models2,059,533 sessionsUpdated 2026-08-27

Agent Arena overall ranking — lower is better.

1.Claude Opus 5 (High)
+12.99%
2.Claude Opus 5 (Max)
+12.73%
3.Claude Fable 5 (High)
+11.70%
4.GPT 5.6 Sol (xHigh)
+10.24%
5.Claude Opus 4.8 (High)
+9.90%
6.Kimi K3 (Max)
+9.22%
7.GPT 5.5 (xHigh)
+7.92%
8.Claude Sonnet 5 (High)
+7.65%
9.Claude Opus 4.7 (High)
+7.03%
10.Claude Opus 4.7
+6.86%
#ModelLabOverallSuccessPraiseSteerabilityRecoveryHalluc.Sessions
1
Claude Opus 5 (High)
Anthropic+12.99%+15.45%19.69%14.42%14.45%0.94%20,468
2
Claude Opus 5 (Max)
Anthropic+12.73%+17.55%19.14%10.72%15.25%0.99%16,214
3
Claude Fable 5 (High)
Anthropic+11.70%+10.44%20.65%14.39%11.99%1.02%33,805
4
GPT 5.6 Sol (xHigh)
OpenAI+10.24%+9.83%21.35%9.69%9.28%1.02%27,506
5
Claude Opus 4.8 (High)
Anthropic+9.90%+6.61%22.6%10.89%9.53%-0.14%35,927
6
Kimi K3 (Max)
Moonshot+9.22%+17.88%17.03%1.99%8.2%1.02%92,312
7
GPT 5.5 (xHigh)
OpenAI+7.92%+2.90%12.16%9.01%14.49%1.02%49,079
8
Claude Sonnet 5 (High)
Anthropic+7.65%+1.14%11.65%13.55%11.04%0.9%26,480
9
Claude Opus 4.7 (High)
Anthropic+7.03%+4.38%10.67%7.06%12.12%0.92%36,570
10
Claude Opus 4.7
Anthropic+6.86%+4.36%10.84%8.59%9.55%0.97%37,121
11
Qwen3.8 Max
Alibaba+6.66%+12.73%8.05%4.71%7.77%0.07%16,510
12
DeepSeek V4 Pro (High) (0813)
DeepSeek+6.58%+13.63%5.8%2.5%9.98%1.02%20,444
13
GPT 5.5 (High)
OpenAI+6.57%+2.20%10.8%6.64%12.2%1.02%73,876
14
Grok 4.5
xAI+6.48%+4.99%6.95%8.68%10.75%1.02%32,876
15
Grok 4.6 (xHigh)
xAI+6.11%+13.17%2.59%4.92%8.85%1.02%12,493
16
GLM 5.2 (Max)
Zai+5.94%+7.82%9.28%5.88%5.72%1.02%61,300
17
GPT 5.5
OpenAI+5.66%+2.84%7.96%5.04%11.42%1.02%76,518
18Anthropic+5.54%+3.44%7.51%5.83%9.93%1.02%36,364
19
GPT 5.4 (High)
OpenAI+4.10%+3.82%2.09%4.73%8.86%1.02%75,797
20
GLM 5.3 (Max)
Zai+3.91%+13.31%4.75%-0.35%0.82%1.02%30,072
21
GPT 5.6 Luna (xHigh)
OpenAI+3.41%+0.04%3.6%3.51%8.9%1.02%13,855
22
Deepseek V4 Flash (High) (20260731)
DeepSeek+3.07%+7.43%1.79%1%4.12%1.01%47,615
23
GPT 5.6 Terra (xHigh)
OpenAI+3%-3.80%-0.14%8.94%9%1.02%15,773
24
Claude Opus 4.8
Anthropic+2.52%+7.09%11.73%10.35%10.97%-27.55%33,689
25Anthropic+1.75%-2.11%-0.02%-0.97%10.92%0.92%37,362
26
Kimi K2.7 Code
Moonshot+1.71%+2.87%1.04%5.6%-1.97%1.02%11,139
27
Qwen 3.8 27B
Alibaba+1.56%+8.05%-0.1%0.89%-1.3%0.25%10,592
28
Gemini 3.7 Flash (High)
Google+1.33%+9.82%-1.75%-4.42%2.08%0.95%24,412
29
Muse Spark 1.2 (xHigh)
Meta+1.32%+6.70%-6.39%-4.81%10.09%1.01%18,407
30
Kimi K2.6
Moonshot+0.65%-2.73%4.15%8.53%-7.71%1.02%11,273
31
Muse Spark 1.1
Meta-0.19%+5.99%-7.08%-5.83%4.96%1%83,283
32
DeepSeek V4 Pro
DeepSeek-0.20%-2.69%-1.93%-1.31%4.73%0.19%30,612
33
GLM 5.1
Zai-0.79%+0.72%-0.79%0%-2.9%-0.98%71,949
34
Qwen3.7 Max
Alibaba-1.04%-2.32%-5.29%-2.5%4.42%0.5%34,023
35
Hy3
Tencent-1.19%-2.86%-2.96%0.46%0.97%-1.57%22,177
36
Qwen3.7 Plus
Alibaba-2.26%-3.58%-9.9%-2.94%5.2%-0.08%17,832
37
Gemini 3.5 Flash (High)
Google-2.44%-0.55%-1.08%-6.79%-3.86%0.07%94,421
38
Gemini 3.1 Pro Preview
Google-2.62%-0.18%2.73%-3.04%-13.38%0.77%82,264
39
Mimo V2.5 Pro
Xiaomi-2.68%-3.82%-7.44%-2.31%0.3%-0.14%34,925
40
Minimax M3
MiniMax-2.79%-5.87%-8.82%-5.5%5.69%0.55%34,344
41
Gemini 3.6 Flash (High)
Google-3.33%-2.67%-5.69%-4.36%-4.92%0.98%15,720
42
Gemini 3.5 Flash (Medium)
Google-3.67%-8.88%-5.99%-1.3%-2.66%0.46%13,756
43
Mistral Medium 3.5
Mistral-5.76%-13.60%-12.05%2.1%-2.53%-2.71%6,727
44
Inkling Small
Thinky-6.18%-18.21%-18.38%-6.52%11.94%0.27%9,191
45
Inkling
Thinky-6.58%-12.10%-17.74%-10.94%7.32%0.54%38,506
46
Solar Pro 4
Upstage-9.38%-7.68%-13.96%-4.15%-20.46%-0.68%5,766
47
Grok Build 0.1
xAI-10.09%-6.52%-11.47%-5.47%-27.61%0.63%74,455
48
Grok 4.3 (High)
xAI-10.16%-11.70%-13.42%-10.77%-15.73%0.81%62,787
49
Minimax M2.7
MiniMax-10.18%-12.93%-16.54%-4.76%-17.55%0.88%34,580
50Google-10.65%-8.20%-10.38%-9.25%-24.5%-0.93%83,411
51
Gemini 3.5 Flash Lite
Google-10.89%-13.52%-14.56%-9.58%-16.21%-0.61%21,288
52
Nemotron 3 Ultra
Nvidia-12.92%-15.58%-14.62%-7.7%-27.03%0.33%12,264
53
Grok 4.3
xAI-17.94%-12.08%-15.37%-10.04%-53.08%0.9%82,742
54
Gemma 4 31B
Google-21.31%-1.94%-3.87%-9.81%-60.51%-30.44%56,661

Data source: Agent Arena (arena.ai) · lmarena-ai/leaderboard-dataset · CC BY 4.0 · updated daily

How the ranking works

Each model is scored on real agentic sessions: whether it completes the task, how well it takes corrections (steerability), how reliably it recovers from failed commands, and how often it invents tools that don't exist (lower is better). Switch the lens above to sort by any of these.

AI model comparison (AI arena)

The ranking compares neural networks — GPT, Claude, Gemini, DeepSeek, Grok and others — on real agentic tasks: task success, steerability, recovery from errors and reliability of tool use. Any model in the table can be run through one AnyModel API.

Run the top models through one API

No subscriptions — pay per token, free to start. Switch between any model on the board by changing one id.

Get a free key →