0.34
0.25–0.43
Benchmarks / Arena (formerly LMArena)
Reported by Arena (formerly LMArena)
Arena (formerly LMArena)
Estimated effect on users praising rather than complaining.
- Results dated
- 28 Sep 2026
- Models
- 46
- Unit
- IPS effect estimate
Full results
| # | Model | Arena Agent: praise over complaint · 95% range IPS effect estimate, higher is better |
|---|---|---|
| 1 | GPT-6 AstraOpenAI | |
| 1 | Claude Fable 5.1Anthropic |
0.32 0.25–0.40 |
| 1 | Claude Opus 5.5Anthropic |
0.19 0.12–0.27 |
| 3 | GPT-5.6 SolOpenAI |
0.19 0.14–0.23 |
| 3 | Claude Opus 5Anthropic |
0.18 0.13–0.23 |
| 3 | Claude Opus 5Anthropic |
0.17 0.11–0.22 |
| 3 | Claude Fable 5Anthropic |
0.17 0.12–0.21 |
| 3 | Claude Opus 4.8Anthropic |
0.15 0.11–0.20 |
| 3 | GPT-6 SolOpenAI |
0.11 0.02–0.19 |
| 3 | Claude Sonnet 5Anthropic |
0.09 0.04–0.15 |
| 6 | GLM 5.2Z.ai |
0.10 0.07–0.12 |
| 8 | GPT-5.5OpenAI |
0.08 0.05–0.11 |
| 8 | Hy4 previewTencent |
0.07 0.04–0.11 |
| 9 | Kimi K3Moonshot AI |
0.07 0.06–0.09 |
| 9 | GLM 5.3Z.ai |
0.06 0.03–0.08 |
| 9 | Gemini 3.8 FlashGoogle |
0.05 0.02–0.08 |
| 9 | GPT-5.5OpenAI |
0.05 0.02–0.08 |
| 9 | Grok 4.7xAI |
0.04 -0.01–0.10 |
| 9 | DeepSeek V4 Pro 0813DeepSeek |
0.03 -0.03–0.09 |
| 10 | Muse Spark 1.3Meta |
0.04 0.02–0.06 |
| 10 | Qwen3.8 Max (0902)Alibaba |
0.04 0.01–0.06 |
| 10 | GPT-5.6 TerraOpenAI |
0.03 -0.01–0.06 |
| 11 | DeepSeek V4.1 FlashDeepSeek |
0.03 0.01–0.05 |
| 15 | Grok 4.6xAI |
-0.00 -0.03–0.03 |
| 15 | GPT-6 LunaOpenAI |
-0.03 -0.09–0.03 |
| 17 | Gemini 3.1 Pro PreviewGoogle |
-0.01 -0.03–0.02 |
| 17 | GPT-5.4OpenAI |
-0.01 -0.04–0.02 |
| 17 | DeepSeek V4 Pro 0423DeepSeek |
-0.02 -0.05–0.02 |
| 20 | GLM 5.3 FlashZ.ai |
-0.00 -0.02–0.01 |
| 21 | Qwen3.8 27BAlibaba |
-0.03 -0.05–-0.01 |
| 21 | Grok 4.5xAI |
-0.03 -0.06–0.00 |
| 21 | Qwen3.8 Flash Next (arena agent)Alibaba |
-0.03 -0.05–-0.01 |
| 23 | Gemini 3.7 FlashGoogle |
-0.03 -0.05–-0.01 |
| 23 | GPT-5.6 LunaOpenAI |
-0.04 -0.06–-0.02 |
| 28 | Hy3Tencent |
-0.07 -0.11–-0.04 |
| 32 | Qwen3.7 MaxAlibaba |
-0.08 -0.11–-0.06 |
| 32 | Gemini 3.6 FlashGoogle |
-0.09 -0.11–-0.06 |
| 34 | MiMo-V2.5-ProXiaomi |
-0.11 -0.13–-0.08 |
| 35 | Muse Spark 1.1Meta |
-0.11 -0.12–-0.10 |
| 35 | Muse Spark 1.2Meta |
-0.12 -0.14–-0.10 |
| 35 | Qwen3.7 PlusAlibaba |
-0.13 -0.16–-0.09 |
| 35 | MiniMax M3MiniMax |
-0.13 -0.16–-0.11 |
| 38 | Mistral Medium 3.5Mistral |
-0.15 -0.19–-0.11 |
| 41 | Solar Pro 4Upstage |
-0.20 -0.25–-0.15 |
| 43 | Inkling SmallThinkingmachines |
-0.21 -0.24–-0.17 |
| 44 | InklingThinkingmachines |
-0.22 -0.25–-0.20 |
Models share a rank when their ranges overlap. Results as published by Arena (formerly LMArena); we do not re-run them.
What it measures
Estimated effect on users praising rather than complaining.
What it does not measure
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
Source
Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.