AI Arena — model leaderboard
A live ranking of AI models on how well they handle real agentic work — task success, steerability, recovery from failures and tool reliability. Compare GPT, Claude, Gemini, DeepSeek, Grok and more, then run any of them through one API.
Agent Arena overall ranking — lower is better.
| # | Model | Lab | Overall | Success | Praise | Steerability | Recovery | Halluc. | Sessions |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (High) | Anthropic | +12.99% | +15.45% | 19.69% | 14.42% | 14.45% | 0.94% | 20,468 |
| 2 | Claude Opus 5 (Max) | Anthropic | +12.73% | +17.55% | 19.14% | 10.72% | 15.25% | 0.99% | 16,214 |
| 3 | Claude Fable 5 (High) | Anthropic | +11.70% | +10.44% | 20.65% | 14.39% | 11.99% | 1.02% | 33,805 |
| 4 | GPT 5.6 Sol (xHigh) | OpenAI | +10.24% | +9.83% | 21.35% | 9.69% | 9.28% | 1.02% | 27,506 |
| 5 | Claude Opus 4.8 (High) | Anthropic | +9.90% | +6.61% | 22.6% | 10.89% | 9.53% | -0.14% | 35,927 |
| 6 | Kimi K3 (Max) | Moonshot | +9.22% | +17.88% | 17.03% | 1.99% | 8.2% | 1.02% | 92,312 |
| 7 | GPT 5.5 (xHigh) | OpenAI | +7.92% | +2.90% | 12.16% | 9.01% | 14.49% | 1.02% | 49,079 |
| 8 | Claude Sonnet 5 (High) | Anthropic | +7.65% | +1.14% | 11.65% | 13.55% | 11.04% | 0.9% | 26,480 |
| 9 | Claude Opus 4.7 (High) | Anthropic | +7.03% | +4.38% | 10.67% | 7.06% | 12.12% | 0.92% | 36,570 |
| 10 | Claude Opus 4.7 | Anthropic | +6.86% | +4.36% | 10.84% | 8.59% | 9.55% | 0.97% | 37,121 |
| 11 | Qwen3.8 Max | Alibaba | +6.66% | +12.73% | 8.05% | 4.71% | 7.77% | 0.07% | 16,510 |
| 12 | DeepSeek V4 Pro (High) (0813) | DeepSeek | +6.58% | +13.63% | 5.8% | 2.5% | 9.98% | 1.02% | 20,444 |
| 13 | GPT 5.5 (High) | OpenAI | +6.57% | +2.20% | 10.8% | 6.64% | 12.2% | 1.02% | 73,876 |
| 14 | Grok 4.5 | xAI | +6.48% | +4.99% | 6.95% | 8.68% | 10.75% | 1.02% | 32,876 |
| 15 | Grok 4.6 (xHigh) | xAI | +6.11% | +13.17% | 2.59% | 4.92% | 8.85% | 1.02% | 12,493 |
| 16 | GLM 5.2 (Max) | Zai | +5.94% | +7.82% | 9.28% | 5.88% | 5.72% | 1.02% | 61,300 |
| 17 | GPT 5.5 | OpenAI | +5.66% | +2.84% | 7.96% | 5.04% | 11.42% | 1.02% | 76,518 |
| 18 | Anthropic | +5.54% | +3.44% | 7.51% | 5.83% | 9.93% | 1.02% | 36,364 | |
| 19 | GPT 5.4 (High) | OpenAI | +4.10% | +3.82% | 2.09% | 4.73% | 8.86% | 1.02% | 75,797 |
| 20 | GLM 5.3 (Max) | Zai | +3.91% | +13.31% | 4.75% | -0.35% | 0.82% | 1.02% | 30,072 |
| 21 | GPT 5.6 Luna (xHigh) | OpenAI | +3.41% | +0.04% | 3.6% | 3.51% | 8.9% | 1.02% | 13,855 |
| 22 | Deepseek V4 Flash (High) (20260731) | DeepSeek | +3.07% | +7.43% | 1.79% | 1% | 4.12% | 1.01% | 47,615 |
| 23 | GPT 5.6 Terra (xHigh) | OpenAI | +3% | -3.80% | -0.14% | 8.94% | 9% | 1.02% | 15,773 |
| 24 | Claude Opus 4.8 | Anthropic | +2.52% | +7.09% | 11.73% | 10.35% | 10.97% | -27.55% | 33,689 |
| 25 | Anthropic | +1.75% | -2.11% | -0.02% | -0.97% | 10.92% | 0.92% | 37,362 | |
| 26 | Kimi K2.7 Code | Moonshot | +1.71% | +2.87% | 1.04% | 5.6% | -1.97% | 1.02% | 11,139 |
| 27 | Qwen 3.8 27B | Alibaba | +1.56% | +8.05% | -0.1% | 0.89% | -1.3% | 0.25% | 10,592 |
| 28 | Gemini 3.7 Flash (High) | +1.33% | +9.82% | -1.75% | -4.42% | 2.08% | 0.95% | 24,412 | |
| 29 | Muse Spark 1.2 (xHigh) | Meta | +1.32% | +6.70% | -6.39% | -4.81% | 10.09% | 1.01% | 18,407 |
| 30 | Kimi K2.6 | Moonshot | +0.65% | -2.73% | 4.15% | 8.53% | -7.71% | 1.02% | 11,273 |
| 31 | Muse Spark 1.1 | Meta | -0.19% | +5.99% | -7.08% | -5.83% | 4.96% | 1% | 83,283 |
| 32 | DeepSeek V4 Pro | DeepSeek | -0.20% | -2.69% | -1.93% | -1.31% | 4.73% | 0.19% | 30,612 |
| 33 | GLM 5.1 | Zai | -0.79% | +0.72% | -0.79% | 0% | -2.9% | -0.98% | 71,949 |
| 34 | Qwen3.7 Max | Alibaba | -1.04% | -2.32% | -5.29% | -2.5% | 4.42% | 0.5% | 34,023 |
| 35 | Hy3 | Tencent | -1.19% | -2.86% | -2.96% | 0.46% | 0.97% | -1.57% | 22,177 |
| 36 | Qwen3.7 Plus | Alibaba | -2.26% | -3.58% | -9.9% | -2.94% | 5.2% | -0.08% | 17,832 |
| 37 | Gemini 3.5 Flash (High) | -2.44% | -0.55% | -1.08% | -6.79% | -3.86% | 0.07% | 94,421 | |
| 38 | Gemini 3.1 Pro Preview | -2.62% | -0.18% | 2.73% | -3.04% | -13.38% | 0.77% | 82,264 | |
| 39 | Mimo V2.5 Pro | Xiaomi | -2.68% | -3.82% | -7.44% | -2.31% | 0.3% | -0.14% | 34,925 |
| 40 | Minimax M3 | MiniMax | -2.79% | -5.87% | -8.82% | -5.5% | 5.69% | 0.55% | 34,344 |
| 41 | Gemini 3.6 Flash (High) | -3.33% | -2.67% | -5.69% | -4.36% | -4.92% | 0.98% | 15,720 | |
| 42 | Gemini 3.5 Flash (Medium) | -3.67% | -8.88% | -5.99% | -1.3% | -2.66% | 0.46% | 13,756 | |
| 43 | Mistral Medium 3.5 | Mistral | -5.76% | -13.60% | -12.05% | 2.1% | -2.53% | -2.71% | 6,727 |
| 44 | Inkling Small | Thinky | -6.18% | -18.21% | -18.38% | -6.52% | 11.94% | 0.27% | 9,191 |
| 45 | Inkling | Thinky | -6.58% | -12.10% | -17.74% | -10.94% | 7.32% | 0.54% | 38,506 |
| 46 | Solar Pro 4 | Upstage | -9.38% | -7.68% | -13.96% | -4.15% | -20.46% | -0.68% | 5,766 |
| 47 | Grok Build 0.1 | xAI | -10.09% | -6.52% | -11.47% | -5.47% | -27.61% | 0.63% | 74,455 |
| 48 | Grok 4.3 (High) | xAI | -10.16% | -11.70% | -13.42% | -10.77% | -15.73% | 0.81% | 62,787 |
| 49 | Minimax M2.7 | MiniMax | -10.18% | -12.93% | -16.54% | -4.76% | -17.55% | 0.88% | 34,580 |
| 50 | -10.65% | -8.20% | -10.38% | -9.25% | -24.5% | -0.93% | 83,411 | ||
| 51 | Gemini 3.5 Flash Lite | -10.89% | -13.52% | -14.56% | -9.58% | -16.21% | -0.61% | 21,288 | |
| 52 | Nemotron 3 Ultra | Nvidia | -12.92% | -15.58% | -14.62% | -7.7% | -27.03% | 0.33% | 12,264 |
| 53 | Grok 4.3 | xAI | -17.94% | -12.08% | -15.37% | -10.04% | -53.08% | 0.9% | 82,742 |
| 54 | Gemma 4 31B | -21.31% | -1.94% | -3.87% | -9.81% | -60.51% | -30.44% | 56,661 |
Data source: Agent Arena (arena.ai) · lmarena-ai/leaderboard-dataset · CC BY 4.0 · updated daily
How the ranking works
Each model is scored on real agentic sessions: whether it completes the task, how well it takes corrections (steerability), how reliably it recovers from failed commands, and how often it invents tools that don't exist (lower is better). Switch the lens above to sort by any of these.
AI model comparison (AI arena)
The ranking compares neural networks — GPT, Claude, Gemini, DeepSeek, Grok and others — on real agentic tasks: task success, steerability, recovery from errors and reliability of tool use. Any model in the table can be run through one AnyModel API.
Run the top models through one API
No subscriptions — pay per token, free to start. Switch between any model on the board by changing one id.
AnyModel