Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Estimated effect on avoiding calls to tools that do not exist.

Results dated
28 Sep 2026
Models
46
Unit
IPS effect estimate
Licence
Creative Commons Attribution 4.0 International

Full results

Arena Agent: tool grounding, IPS effect estimate, higher is better
#ModelArena Agent: tool grounding · 95% range
IPS effect estimate, higher is better
1 Grok 4.6xAI
0.00
0.00–0.00
1 GPT-5.4OpenAI
0.00
0.00–0.00
1 GPT-6 AstraOpenAI
0.00
0.00–0.00
1 Grok 4.5xAI
0.00
0.00–0.00
1 GPT-5.5OpenAI
0.00
0.00–0.00
1 GPT-6 LunaOpenAI
0.00
0.00–0.00
1 GPT-5.6 SolOpenAI
0.00
0.00–0.00
1 Grok 4.7xAI
0.00
0.00–0.00
1 GPT-5.6 TerraOpenAI
0.00
0.00–0.00
1 GPT-6 SolOpenAI
0.00
0.00–0.00
1 DeepSeek V4 Pro 0813DeepSeek
0.00
0.00–0.00
1 GPT-5.6 LunaOpenAI
0.00
0.00–0.00
1 GLM 5.3 FlashZ.ai
0.00
0.00–0.00
1 GLM 5.3Z.ai
0.00
0.00–0.00
1 Kimi K3Moonshot AI
0.00
0.00–0.00
1 Claude Fable 5Anthropic
0.00
0.00–0.00
1 GPT-5.5OpenAI
0.00
0.00–0.00
1 GLM 5.2Z.ai
0.00
0.00–0.00
1 Muse Spark 1.3Meta
0.00
0.00–0.00
1 Claude Fable 5.1Anthropic
0.00
0.00–0.00
1 Claude Opus 5Anthropic
0.00
0.00–0.00
1 Muse Spark 1.1Meta
0.00
0.00–0.00
1 Muse Spark 1.2Meta
0.00
0.00–0.00
1 Claude Opus 5Anthropic
0.00
0.00–0.00
1 Gemini 3.6 FlashGoogle
0.00
0.00–0.00
1 Claude Opus 5.5Anthropic
0.00
0.00–0.00
1 Hy4 previewTencent
0.00
-0.00–0.00
1 Gemini 3.1 Pro PreviewGoogle
-0.00
-0.01–0.00
24 Gemini 3.8 FlashGoogle
0.00
0.00–0.00
24 Claude Sonnet 5Anthropic
0.00
0.00–0.00
25 DeepSeek V4.1 FlashDeepSeek
0.00
0.00–0.00
27 Claude Opus 4.8Anthropic
-0.00
-0.00–0.00
27 Gemini 3.7 FlashGoogle
-0.00
-0.00–0.00
29 InklingThinkingmachines
-0.00
-0.00–0.00
29 MiniMax M3MiniMax
-0.00
-0.00–0.00
30 DeepSeek V4 Pro 0423DeepSeek
-0.00
-0.00–-0.00
30 Qwen3.7 MaxAlibaba
-0.00
-0.00–-0.00
30 Inkling SmallThinkingmachines
-0.00
-0.01–-0.00
31 Qwen3.8 27BAlibaba
-0.00
-0.00–-0.00
38 MiMo-V2.5-ProXiaomi
-0.01
-0.01–-0.01
38 Qwen3.7 PlusAlibaba
-0.01
-0.01–-0.01
39 Qwen3.8 Flash Next (arena agent)Alibaba
-0.01
-0.01–-0.01
39 Qwen3.8 Max (0902)Alibaba
-0.01
-0.01–-0.01
40 Solar Pro 4Upstage
-0.02
-0.03–-0.01
44 Mistral Medium 3.5Mistral
-0.02
-0.03–-0.02
44 Hy3Tencent
-0.03
-0.04–-0.02

Models share a rank when their ranges overlap. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Estimated effect on avoiding calls to tools that do not exist.

What it does not measure

Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.