Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Estimated effect on users praising rather than complaining.

Results dated
28 Sep 2026
Models
46
Unit
IPS effect estimate
Licence
Creative Commons Attribution 4.0 International

Full results

Arena Agent: praise over complaint, IPS effect estimate, higher is better
#ModelArena Agent: praise over complaint · 95% range
IPS effect estimate, higher is better
1 GPT-6 AstraOpenAI
0.34
0.25–0.43
1 Claude Fable 5.1Anthropic
0.32
0.25–0.40
1 Claude Opus 5.5Anthropic
0.19
0.12–0.27
3 GPT-5.6 SolOpenAI
0.19
0.14–0.23
3 Claude Opus 5Anthropic
0.18
0.13–0.23
3 Claude Opus 5Anthropic
0.17
0.11–0.22
3 Claude Fable 5Anthropic
0.17
0.12–0.21
3 Claude Opus 4.8Anthropic
0.15
0.11–0.20
3 GPT-6 SolOpenAI
0.11
0.02–0.19
3 Claude Sonnet 5Anthropic
0.09
0.04–0.15
6 GLM 5.2Z.ai
0.10
0.07–0.12
8 GPT-5.5OpenAI
0.08
0.05–0.11
8 Hy4 previewTencent
0.07
0.04–0.11
9 Kimi K3Moonshot AI
0.07
0.06–0.09
9 GLM 5.3Z.ai
0.06
0.03–0.08
9 Gemini 3.8 FlashGoogle
0.05
0.02–0.08
9 GPT-5.5OpenAI
0.05
0.02–0.08
9 Grok 4.7xAI
0.04
-0.01–0.10
9 DeepSeek V4 Pro 0813DeepSeek
0.03
-0.03–0.09
10 Muse Spark 1.3Meta
0.04
0.02–0.06
10 Qwen3.8 Max (0902)Alibaba
0.04
0.01–0.06
10 GPT-5.6 TerraOpenAI
0.03
-0.01–0.06
11 DeepSeek V4.1 FlashDeepSeek
0.03
0.01–0.05
15 Grok 4.6xAI
-0.00
-0.03–0.03
15 GPT-6 LunaOpenAI
-0.03
-0.09–0.03
17 Gemini 3.1 Pro PreviewGoogle
-0.01
-0.03–0.02
17 GPT-5.4OpenAI
-0.01
-0.04–0.02
17 DeepSeek V4 Pro 0423DeepSeek
-0.02
-0.05–0.02
20 GLM 5.3 FlashZ.ai
-0.00
-0.02–0.01
21 Qwen3.8 27BAlibaba
-0.03
-0.05–-0.01
21 Grok 4.5xAI
-0.03
-0.06–0.00
21 Qwen3.8 Flash Next (arena agent)Alibaba
-0.03
-0.05–-0.01
23 Gemini 3.7 FlashGoogle
-0.03
-0.05–-0.01
23 GPT-5.6 LunaOpenAI
-0.04
-0.06–-0.02
28 Hy3Tencent
-0.07
-0.11–-0.04
32 Qwen3.7 MaxAlibaba
-0.08
-0.11–-0.06
32 Gemini 3.6 FlashGoogle
-0.09
-0.11–-0.06
34 MiMo-V2.5-ProXiaomi
-0.11
-0.13–-0.08
35 Muse Spark 1.1Meta
-0.11
-0.12–-0.10
35 Muse Spark 1.2Meta
-0.12
-0.14–-0.10
35 Qwen3.7 PlusAlibaba
-0.13
-0.16–-0.09
35 MiniMax M3MiniMax
-0.13
-0.16–-0.11
38 Mistral Medium 3.5Mistral
-0.15
-0.19–-0.11
41 Solar Pro 4Upstage
-0.20
-0.25–-0.15
43 Inkling SmallThinkingmachines
-0.21
-0.24–-0.17
44 InklingThinkingmachines
-0.22
-0.25–-0.20

Models share a rank when their ranges overlap. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Estimated effect on users praising rather than complaining.

What it does not measure

Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.