1,509
1,496–1,521
Benchmarks / Arena (formerly LMArena)
Reported by Arena (formerly LMArena)
Arena (formerly LMArena)
Head-to-head human preference across all text prompts.
- Results dated
- 25 Sep 2026
- Models
- 409
- Unit
- Arena rating
Full results
| # | Model | Arena Text: overall · 95% range Arena rating, higher is better |
|---|---|---|
| 1 | Claude Opus 5.5 (high)Anthropic | |
| 1 | Claude Opus 4.6 (high)Anthropic |
1,505 1,502–1,509 |
| 1 | Claude Fable 5 (high)Anthropic |
1,504 1,499–1,508 |
| 1 | Claude Opus 4.7 (high)Anthropic |
1,502 1,498–1,506 |
| 1 | Claude Fable 5.1 (max)Anthropic |
1,501 1,495–1,508 |
| 1 | Muse Spark 1.2Meta |
1,496 1,486–1,506 |
| 2 | Claude Opus 4.6Anthropic |
1,498 1,494–1,501 |
| 2 | Muse Spark 1.3 (max)Meta |
1,494 1,488–1,501 |
| 3 | Claude Opus 4.7Anthropic |
1,495 1,491–1,498 |
| 4 | Gemini 3.8 Flash (high)Google |
1,492 1,487–1,498 |
| 5 | Claude Opus 5 (high)Anthropic |
1,491 1,487–1,495 |
| 5 | Muse Spark 1.1Meta |
1,491 1,486–1,495 |
| 5 | muse-sparkMeta |
1,489 1,483–1,495 |
| 7 | Claude Opus 5 (max)Anthropic |
1,488 1,483–1,493 |
| 7 | Gemini 3.7 Flash (high)Google |
1,488 1,483–1,493 |
| 7 | Kimi K3 (max)Moonshot AI |
1,488 1,483–1,493 |
| 8 | Gemini 3.1 Pro PreviewGoogle |
1,487 1,484–1,490 |
| 8 | gemini-3-proGoogle |
1,485 1,482–1,489 |
| 8 | GPT-5.6 Sol (xhigh)OpenAI |
1,483 1,479–1,488 |
| 8 | MiMo-V2.6-ProXiaomi |
1,480 1,470–1,489 |
| 11 | Gemini 3.6 Flash (high)Google |
1,482 1,477–1,486 |
| 13 | GPT-5.5 (high)OpenAI |
1,481 1,478–1,485 |
| 13 | Claude Opus 4.8 (high)Anthropic |
1,480 1,476–1,484 |
| 13 | GLM 5.3 (max)Z.ai |
1,480 1,474–1,485 |
| 13 | Qwen3.8 Max (0902)Alibaba |
1,479 1,474–1,485 |
| 13 | GPT-6 Astra (max)OpenAI |
1,478 1,470–1,485 |
| 13 | DeepSeek V4.1 Flash (max)DeepSeek |
1,477 1,469–1,484 |
| 13 | qwen3.7-max-previewAlibaba |
1,475 1,465–1,485 |
| 19 | Gemini 3.5 Flash (high)Google |
1,477 1,473–1,481 |
| 19 | GPT-5.5OpenAI |
1,477 1,473–1,480 |
| 19 | gpt-5.2-chat-latest-20260210OpenAI |
1,476 1,472–1,480 |
| 19 | GLM 5.2 (max)Z.ai |
1,476 1,471–1,480 |
| 19 | GPT-5.4 (high)OpenAI |
1,475 1,471–1,479 |
| 19 | grok-4.20-beta1xAI |
1,475 1,470–1,479 |
| 19 | GLM 5.3 FlashZ.ai |
1,474 1,469–1,480 |
| 20 | Gemini 3.5 Flash (medium)Google |
1,474 1,470–1,478 |
| 20 | gpt-5.5-instantOpenAI |
1,473 1,468–1,478 |
| 21 | Claude Opus 4.8Anthropic |
1,474 1,470–1,478 |
| 21 | Claude Opus 4.5Anthropic |
1,474 1,470–1,477 |
| 22 | gemini-3-flashGoogle |
1,473 1,468–1,477 |
| 23 | Claude Sonnet 4.6Anthropic |
1,472 1,469–1,476 |
| 23 | grok-4.20-beta-0309-reasoningxAI |
1,471 1,468–1,475 |
| 24 | grok-4.20-multi-agent-beta-0309xAI |
1,470 1,467–1,474 |
| 27 | Claude Opus 4.5Anthropic |
1,470 1,467–1,473 |
| 27 | ernie-5.1Baidu |
1,468 1,463–1,472 |
| 30 | MiMo-V2.5-ProXiaomi |
1,467 1,464–1,471 |
| 30 | DeepSeek V4 Pro 0813DeepSeek |
1,464 1,458–1,471 |
| 34 | Grok 4.5xAI |
1,465 1,461–1,470 |
| 36 | GPT-5.6 Terra (xhigh)OpenAI |
1,465 1,461–1,469 |
| 36 | qwen3.5-max-previewAlibaba |
1,465 1,460–1,470 |
| 37 | GLM 5.1Z.ai |
1,465 1,461–1,469 |
| 38 | GPT-5.4OpenAI |
1,465 1,461–1,469 |
| 40 | grok-4.1-thinkingxAI |
1,465 1,462–1,468 |
| 40 | Qwen3.6 Max PreviewAlibaba |
1,460 1,452–1,468 |
| 44 | Claude Sonnet 5 (high)Anthropic |
1,462 1,457–1,466 |
| 44 | Kimi K2.6Moonshot AI |
1,461 1,456–1,465 |
| 44 | GPT-6 Sol (max)OpenAI |
1,457 1,449–1,466 |
| 46 | Hy3Tencent |
1,457 1,450–1,464 |
| 46 | MiMo-V2.6-FlashXiaomi |
1,454 1,444–1,463 |
| 47 | grok-4.1xAI |
1,459 1,456–1,462 |
| 48 | DeepSeek V4 Pro 0423DeepSeek |
1,458 1,454–1,461 |
| 48 | GLM 5Z.ai |
1,457 1,453–1,461 |
| 49 | gemini-3-flash (minimal reasoning)Google |
1,458 1,455–1,461 |
| 51 | Gemini 3.5 Flash LiteGoogle |
1,456 1,451–1,461 |
| 52 | dola-seed-2.0-proBytedance |
1,456 1,453–1,460 |
| 52 | Qwen3.7 PlusAlibaba |
1,456 1,451–1,460 |
| 52 | Gemma 4 31BGoogle |
1,453 1,445–1,460 |
| 53 | Claude Sonnet 4.5Anthropic |
1,457 1,454–1,459 |
| 53 | GPT-5.1 (high)OpenAI |
1,456 1,452–1,459 |
| 53 | Claude Sonnet 4.5Anthropic |
1,455 1,452–1,458 |
| 53 | deepseek-v4-pro-preview (high reasoning)DeepSeek |
1,455 1,451–1,459 |
| 53 | GPT-5.6 Luna (xhigh)OpenAI |
1,454 1,449–1,458 |
| 53 | Grok 4.6 (high)xAI |
1,453 1,448–1,459 |
| 57 | ernie-5.0-preview-1203Baidu |
1,449 1,442–1,455 |
| 58 | gpt-5.3-chat-latestOpenAI |
1,450 1,446–1,454 |
| 60 | Kimi K2.5Moonshot AI |
1,450 1,447–1,454 |
| 60 | Claude Opus 4.1Anthropic |
1,450 1,446–1,453 |
| 62 | mimo-v2-proXiaomi |
1,448 1,443–1,453 |
| 67 | Claude Opus 4.1Anthropic |
1,448 1,445–1,451 |
| 67 | GPT-5.4 Mini (high)OpenAI |
1,447 1,443–1,451 |
| 67 | GPT-6 Luna (max)OpenAI |
1,442 1,434–1,451 |
| 68 | ernie-5.0-0110Baidu |
1,447 1,443–1,450 |
| 68 | gpt-4.5-preview-2025-02-27OpenAI |
1,445 1,439–1,450 |
| 71 | Grok 4.7 (xhigh)xAI |
1,439 1,430–1,449 |
| 72 | Gemini 2.5 ProGoogle |
1,446 1,443–1,448 |
| 72 | Qwen3.6 PlusAlibaba |
1,443 1,439–1,448 |
| 72 | GLM 4.7Z.ai |
1,441 1,435–1,448 |
| 73 | InklingThinkingmachines |
1,442 1,438–1,447 |
| 75 | chatgpt-4o-latest-20250326OpenAI |
1,443 1,440–1,446 |
| 76 | Grok 4.3xAI |
1,442 1,438–1,445 |
| 76 | Gemma 4 26B A4BGoogle |
1,438 1,430–1,445 |
| 77 | Qwen3.5 397B A17BAlibaba |
1,442 1,439–1,445 |
| 77 | MiniMax M3MiniMax |
1,440 1,436–1,444 |
| 79 | Qwen3.8 27BAlibaba |
1,438 1,432–1,443 |
| 82 | deepseek-v4-flash-preview (high reasoning)DeepSeek |
1,439 1,434–1,443 |
| 83 | GPT-5.1OpenAI |
1,438 1,435–1,442 |
| 83 | GPT-5.2 (high)OpenAI |
1,437 1,434–1,441 |
| 83 | longcat-flash-chat-2602-expMeituan |
1,437 1,432–1,441 |
| 84 | DeepSeek V4 Flash 0423DeepSeek |
1,436 1,432–1,440 |
| 84 | GLM 5V TurboZ.ai |
1,433 1,427–1,440 |
| 85 | qwen3-max-previewAlibaba |
1,435 1,430–1,439 |
| 85 | GPT-5 (high)OpenAI |
1,435 1,430–1,439 |
| 86 | GPT-5.2OpenAI |
1,436 1,432–1,439 |
| 88 | MiMo-V2.5Xiaomi |
1,434 1,429–1,438 |
| 89 | mimo-v2-omniXiaomi |
1,431 1,426–1,437 |
| 89 | Kimi K2.5Moonshot AI |
1,430 1,424–1,437 |
| 90 | Gemini 3.1 Flash Lite PreviewGoogle |
1,433 1,429–1,436 |
| 90 | amazon-nova-experimental-chat-26-02-10Amazon |
1,427 1,417–1,436 |
| 91 | o3OpenAI |
1,432 1,428–1,435 |
| 93 | muse-glimmerMeta |
1,424 1,415–1,434 |
| 95 | grok-4-1-fast-reasoningxAI |
1,430 1,427–1,433 |
| 95 | kimi-k2-thinking-turboMoonshot AI |
1,430 1,427–1,433 |
| 95 | Mistral Medium 3.5Mistral |
1,426 1,420–1,433 |
| 96 | nvidia-nemotron-3-ultra-550b-a55b-nvfp4NVIDIA |
1,425 1,419–1,432 |
| 99 | gpt-5-chatOpenAI |
1,427 1,422–1,431 |
| 99 | claude-opus-4 (thinking-16k reasoning)Anthropic |
1,426 1,422–1,430 |
| 99 | DeepSeek V3.2 ExpDeepSeek |
1,425 1,418–1,432 |
| 103 | Qwen3 MaxAlibaba |
1,423 1,417–1,430 |
| 105 | DeepSeek V3.2DeepSeek |
1,425 1,421–1,428 |
| 105 | GLM 4.6Z.ai |
1,424 1,420–1,428 |
| 105 | DeepSeek V3.2 ExpDeepSeek |
1,422 1,416–1,428 |
| 106 | R1 0528DeepSeek |
1,422 1,416–1,427 |
| 106 | ernie-5.0-preview-1022Baidu |
1,418 1,410–1,427 |
| 106 | DeepSeek V3.1 TerminusDeepSeek |
1,417 1,407–1,427 |
| 109 | DeepSeek V3.2DeepSeek |
1,423 1,419–1,426 |
| 109 | grok-4-fast-chatxAI |
1,419 1,411–1,427 |
| 110 | Qwen3 235B A22B Instruct 2507Alibaba |
1,422 1,420–1,425 |
| 110 | Kimi K2 0905Moonshot AI |
1,419 1,412–1,425 |
| 110 | DeepSeek V3.1 TerminusDeepSeek |
1,415 1,405–1,424 |
| 111 | Kimi K2 0711Moonshot AI |
1,418 1,413–1,423 |
| 111 | DeepSeek V3.1DeepSeek |
1,417 1,411–1,423 |
| 111 | DeepSeek V3.1DeepSeek |
1,416 1,409–1,423 |
| 111 | amazon-nova-experimental-chat-26-01-10Amazon |
1,413 1,403–1,423 |
| 114 | Qwen3.5-122B-A10BAlibaba |
1,416 1,412–1,420 |
| 115 | Qwen3 VL 235B A22B InstructAlibaba |
1,413 1,407–1,420 |
| 118 | GPT-4.1OpenAI |
1,415 1,411–1,419 |
| 119 | claude-opus-4Anthropic |
1,414 1,410–1,418 |
| 120 | MiniMax M2.7MiniMax |
1,415 1,411–1,418 |
| 121 | Hy3 previewTencent |
1,409 1,402–1,417 |
| 122 | Claude Haiku 4.5Anthropic |
1,414 1,411–1,416 |
| 122 | Mistral Large 3 2512Mistral |
1,413 1,410–1,416 |
| 122 | GLM 4.5Z.ai |
1,411 1,406–1,416 |
| 123 | grok-3-preview-02-24xAI |
1,411 1,407–1,416 |
| 124 | grok-4-0709xAI |
1,411 1,407–1,415 |
| 126 | Qwen3.5-27BAlibaba |
1,409 1,404–1,413 |
| 128 | Gemini 2.5 FlashGoogle |
1,409 1,407–1,412 |
| 133 | Mistral Medium 3.1Mistral |
1,408 1,405–1,411 |
| 133 | Inkling SmallThinkingmachines |
1,405 1,400–1,410 |
| 134 | grok-4-fast-reasoningxAI |
1,405 1,400–1,410 |
| 137 | gemini-2.5-flash-preview-09-2025Google |
1,403 1,399–1,407 |
| 137 | longcat-flash-chatMeituan |
1,401 1,395–1,408 |
| 137 | hunyuan-vision-1.5-thinkingTencent |
1,395 1,382–1,407 |
| 139 | Qwen3 235B A22BAlibaba |
1,403 1,398–1,407 |
| 142 | o1OpenAI |
1,402 1,398–1,407 |
| 142 | Qwen3 235B A22B Thinking 2507Alibaba |
1,400 1,393–1,407 |
| 143 | Claude Sonnet 4Anthropic |
1,402 1,397–1,406 |
| 145 | GPT-5.4 Nano (high)OpenAI |
1,401 1,397–1,405 |
| 146 | Qwen3 Next 80B A3B InstructAlibaba |
1,399 1,394–1,404 |
| 147 | R1DeepSeek |
1,398 1,393–1,403 |
| 147 | amazon-nova-experimental-chat-12-10Amazon |
1,394 1,384–1,403 |
| 148 | Qwen3 VL 235B A22B ThinkingAlibaba |
1,395 1,388–1,402 |
| 149 | Qwen3.5-FlashAlibaba |
1,396 1,393–1,400 |
| 150 | DeepSeek V3 0324DeepSeek |
1,396 1,392–1,400 |
| 151 | Qwen3.5-35B-A3BAlibaba |
1,394 1,390–1,398 |
| 155 | Step 3.5 FlashStepfun |
1,393 1,390–1,397 |
| 155 | mimo-v2-flash (no reasoning)Xiaomi |
1,392 1,388–1,395 |
| 155 | o4 MiniOpenAI |
1,391 1,387–1,395 |
| 155 | Claude Sonnet 4Anthropic |
1,391 1,386–1,395 |
| 155 | hunyuan-t1-20250711Tencent |
1,387 1,378–1,396 |
| 156 | MiniMax M2.5MiniMax |
1,391 1,387–1,395 |
| 156 | GPT-5 Mini (high)OpenAI |
1,390 1,385–1,394 |
| 157 | o1-previewOpenAI |
1,389 1,384–1,394 |
| 159 | mimo-v2-flash (thinking reasoning)Xiaomi |
1,387 1,381–1,393 |
| 160 | claude-3-7-sonnet-20250219 (thinking-32k reasoning)Anthropic |
1,388 1,384–1,393 |
| 160 | Mistral Medium 3Mistral |
1,387 1,383–1,392 |
| 160 | Qwen3 Coder 480B A35BAlibaba |
1,387 1,382–1,392 |
| 161 | Solar Pro 4Upstage |
1,385 1,379–1,392 |
| 162 | minimax-m2.1-previewMiniMax |
1,385 1,379–1,390 |
| 162 | GLM 4.6VZ.ai |
1,379 1,367–1,390 |
| 163 | hunyuan-turbos-20250416Tencent |
1,383 1,376–1,389 |
| 165 | GPT-4.1 MiniOpenAI |
1,383 1,379–1,387 |
| 165 | Qwen3 30B A3B Instruct 2507Alibaba |
1,383 1,378–1,388 |
| 172 | gemini-2.5-flash-lite-preview-09-2025 (no reasoning)Google |
1,379 1,376–1,383 |
| 172 | trinity-large-previewArcee Ai |
1,379 1,374–1,383 |
| 176 | Qwen3 235B A22BAlibaba |
1,375 1,371–1,380 |
| 178 | gemini-2.5-flash-lite-preview-06-17 (thinking reasoning)Google |
1,375 1,370–1,379 |
| 180 | qwen2.5-maxAlibaba |
1,374 1,370–1,378 |
| 181 | claude-3-5-sonnet-20241022Anthropic |
1,374 1,371–1,377 |
| 181 | GLM 4.5 AirZ.ai |
1,373 1,369–1,378 |
| 181 | claude-3-7-sonnet-20250219Anthropic |
1,373 1,369–1,377 |
| 183 | Qwen3 Next 80B A3B ThinkingAlibaba |
1,369 1,363–1,375 |
| 184 | Trinity Large ThinkingArcee Ai |
1,368 1,363–1,372 |
| 185 | GLM 4.7 FlashZ.ai |
1,365 1,359–1,371 |
| 188 | Gemma 3 27BGoogle |
1,365 1,362–1,369 |
| 188 | grok-3-mini (high reasoning)xAI |
1,364 1,358–1,369 |
| 189 | o3 Mini HighOpenAI |
1,364 1,358–1,369 |
| 190 | amazon-nova-experimental-chat-11-10Amazon |
1,364 1,360–1,369 |
| 190 | MiniMax M1MiniMax |
1,364 1,360–1,368 |
| 190 | Nemotron 3 SuperNVIDIA |
1,360 1,353–1,368 |
| 191 | gemini-2.0-flash-001Google |
1,360 1,356–1,364 |
| 191 | intellect-3Primeintellect |
1,356 1,348–1,364 |
| 192 | DeepSeek V3DeepSeek |
1,358 1,354–1,363 |
| 193 | grok-3-mini-betaxAI |
1,358 1,353–1,363 |
| 193 | Mistral Small 3.2 24BMistral |
1,357 1,351–1,362 |
| 194 | GLM 4.5VZ.ai |
1,352 1,344–1,361 |
| 194 | hunyuan-turbos-20250226Tencent |
1,349 1,337–1,361 |
| 196 | llama-3.1-nemotron-ultra-253b-v1NVIDIA |
1,348 1,336–1,359 |
| 199 | Command ACohere |
1,354 1,350–1,357 |
| 199 | gemini-2.0-flash-lite-preview-02-05Google |
1,354 1,349–1,358 |
| 199 | step-3Stepfun |
1,349 1,342–1,357 |
| 199 | amazon-nova-experimental-chat-10-09Amazon |
1,347 1,336–1,358 |
| 199 | Qwen3 32BAlibaba |
1,347 1,337–1,356 |
| 200 | gpt-oss-120bOpenAI |
1,352 1,347–1,356 |
| 200 | gemini-1.5-pro-002Google |
1,351 1,348–1,355 |
| 200 | amazon-nova-experimental-chat-10-20Amazon |
1,348 1,342–1,354 |
| 200 | Qwen-PlusAlibaba |
1,346 1,338–1,354 |
| 201 | Mercury 2Inception |
1,343 1,333–1,354 |
| 202 | nvidia-nemotron-3.5-lightning-30b-a3b-nvfp4NVIDIA |
1,347 1,340–1,353 |
| 202 | nvidia-llama-3.3-nemotron-super-49b-v1.5NVIDIA |
1,343 1,333–1,353 |
| 203 | o3 MiniOpenAI |
1,348 1,345–1,352 |
| 203 | granite-4.2-30bIBM |
1,341 1,331–1,352 |
| 203 | hunyuan-turbo-0110Tencent |
1,341 1,329–1,352 |
| 204 | ling-flash-2.0Ant Group |
1,344 1,336–1,351 |
| 204 | MiniMax M2MiniMax |
1,343 1,336–1,351 |
| 204 | glm-4-plus-0111Z.ai |
1,343 1,334–1,351 |
| 204 | Gemma 3 12BGoogle |
1,342 1,332–1,351 |
| 205 | GPT-4o (2024-05-13)OpenAI |
1,346 1,343–1,350 |
| 209 | claude-3-5-sonnet-20240620Anthropic |
1,343 1,340–1,347 |
| 209 | GPT-5 Nano (high)OpenAI |
1,338 1,331–1,345 |
| 211 | molmo-2-8bAi2 |
1,322 1,301–1,343 |
| 212 | step-2-16k-exp-202412Stepfun |
1,334 1,325–1,342 |
| 214 | o1-miniOpenAI |
1,337 1,333–1,341 |
| 214 | gemini-advanced-0514Google |
1,336 1,331–1,341 |
| 214 | Nova 2 LiteAmazon |
1,335 1,329–1,341 |
| 215 | qwq-32bAlibaba |
1,336 1,331–1,340 |
| 215 | GPT-4o (2024-08-06)OpenAI |
1,336 1,331–1,340 |
| 215 | llama-3.3-nemotron-49b-super-v1NVIDIA |
1,328 1,316–1,340 |
| 216 | grok-2-2024-08-13xAI |
1,336 1,332–1,339 |
| 216 | llama-3.1-405b-instruct-bf16Meta |
1,335 1,332–1,339 |
| 219 | llama-3.1-405b-instruct-fp8Meta |
1,334 1,330–1,337 |
| 220 | hunyuan-large-2025-02-10Tencent |
1,326 1,317–1,336 |
| 223 | olmo-3.1-32b-instructAi2 |
1,329 1,323–1,335 |
| 224 | yi-lightning01 Ai |
1,328 1,323–1,333 |
| 229 | deepseek-v2.5-1210DeepSeek |
1,323 1,315–1,332 |
| 232 | Llama 4 MaverickMeta |
1,327 1,323–1,331 |
| 232 | Qwen3 30B A3BAlibaba |
1,327 1,322–1,331 |
| 235 | GPT-4.1 NanoOpenAI |
1,322 1,315–1,330 |
| 236 | ring-flash-2.0Ant Group |
1,322 1,315–1,330 |
| 238 | claude-3-5-haiku-20241022Anthropic |
1,325 1,321–1,328 |
| 238 | GPT-4 TurboOpenAI |
1,324 1,321–1,328 |
| 238 | gemini-1.5-pro-001Google |
1,324 1,320–1,328 |
| 238 | claude-3-opus-20240229Anthropic |
1,322 1,319–1,325 |
| 238 | Llama 4 ScoutMeta |
1,322 1,317–1,326 |
| 238 | step-1o-turbo-202506Stepfun |
1,320 1,313–1,327 |
| 239 | glm-4-plusZ.ai |
1,319 1,314–1,324 |
| 239 | qwen-max-0919Alibaba |
1,318 1,312–1,324 |
| 239 | gpt-oss-20bOpenAI |
1,318 1,311–1,324 |
| 242 | gemma-3n-e4b-itGoogle |
1,317 1,312–1,323 |
| 244 | Llama 3.3 70B InstructMeta |
1,318 1,314–1,321 |
| 244 | GPT-4o-mini (2024-07-18)OpenAI |
1,318 1,314–1,321 |
| 244 | qwen2.5-plus-1127Alibaba |
1,315 1,308–1,321 |
| 244 | hunyuan-standard-2025-02-10Tencent |
1,311 1,302–1,321 |
| 247 | Mistral Large 2407Mistral |
1,314 1,310–1,318 |
| 247 | athene-v2-chatNexusflow |
1,314 1,310–1,319 |
| 247 | Nemotron 3 Nano 30B A3BNVIDIA |
1,313 1,308–1,319 |
| 247 | gpt-4-0125-previewOpenAI |
1,313 1,309–1,317 |
| 247 | mercuryInception |
1,305 1,291–1,319 |
| 248 | gpt-4-1106-previewOpenAI |
1,313 1,309–1,317 |
| 253 | olmo-3-32b-thinkAi2 |
1,306 1,298–1,315 |
| 256 | gemini-1.5-flash-002Google |
1,309 1,305–1,313 |
| 256 | granite-4.1-8bIBM |
1,304 1,294–1,314 |
| 257 | Gemma 3 4BGoogle |
1,303 1,294–1,313 |
| 258 | athene-70b-0725Nexusflow |
1,307 1,301–1,312 |
| 259 | grok-2-mini-2024-08-13xAI |
1,308 1,305–1,312 |
| 259 | deepseek-v2.5DeepSeek |
1,307 1,303–1,312 |
| 260 | magistral-medium-2506Mistral |
1,305 1,298–1,311 |
| 261 | mistral-large-2411Mistral |
1,306 1,301–1,310 |
| 266 | Mistral Small 3.1 24BMistral |
1,303 1,299–1,308 |
| 266 | Qwen2.5 72B InstructAlibaba |
1,303 1,299–1,307 |
| 266 | llama-3.1-nemotron-70b-instructNVIDIA |
1,299 1,291–1,307 |
| 268 | hunyuan-large-visionTencent |
1,294 1,285–1,303 |
| 268 | granite-4.2-3bIBM |
1,292 1,280–1,303 |
| 273 | Granite 4.2 8BIBM |
1,289 1,278–1,300 |
| 277 | Llama 3.1 70B InstructMeta |
1,293 1,290–1,297 |
| 277 | Nova Pro 1.0Amazon |
1,290 1,286–1,295 |
| 277 | jamba-1.5-largeAI21 Labs |
1,289 1,282–1,297 |
| 277 | reka-core-20240904Rekaai |
1,288 1,281–1,295 |
| 277 | ibm-granite-h-smallIBM |
1,287 1,278–1,295 |
| 277 | llama-3.1-nemotron-51b-instructNVIDIA |
1,287 1,277–1,297 |
| 277 | llama-3.1-tulu-3-70bAi2 |
1,286 1,276–1,297 |
| 279 | Gemma 2 27BGoogle |
1,289 1,286–1,293 |
| 279 | gpt-4-0314OpenAI |
1,288 1,283–1,293 |
| 279 | olmo-3.1-32b-thinkAi2 |
1,287 1,279–1,294 |
| 280 | gemini-1.5-flash-001Google |
1,287 1,282–1,291 |
| 282 | gemma-2-9b-it-simpoPrinceton Nlp |
1,280 1,274–1,287 |
| 284 | claude-3-sonnet-20240229Anthropic |
1,281 1,277–1,285 |
| 286 | nemotron-4-340b-instructNVIDIA |
1,277 1,272–1,282 |
| 286 | Command R+ (08-2024)Cohere |
1,276 1,270–1,283 |
| 289 | Mistral Small 3Mistral |
1,274 1,268–1,280 |
| 289 | glm-4-0520Z.ai |
1,273 1,266–1,280 |
| 290 | llama-3-70b-instructMeta |
1,277 1,273–1,280 |
| 290 | gpt-4-0613OpenAI |
1,276 1,272–1,280 |
| 290 | reka-flash-20240904Rekaai |
1,272 1,265–1,279 |
| 291 | Qwen2.5 Coder 32B InstructAlibaba |
1,271 1,262–1,279 |
| 299 | c4ai-aya-expanse-32bCohere |
1,267 1,262–1,272 |
| 300 | gemma-2-9b-itGoogle |
1,267 1,263–1,271 |
| 300 | deepseek-coder-v2DeepSeek |
1,265 1,259–1,272 |
| 302 | qwen2-72b-instructAlibaba |
1,262 1,257–1,267 |
| 303 | command-r-plusCohere |
1,262 1,257–1,266 |
| 303 | Nova Lite 1.0Amazon |
1,260 1,255–1,266 |
| 304 | claude-3-haiku-20240307Anthropic |
1,262 1,258–1,265 |
| 305 | gemini-1.5-flash-8b-001Google |
1,259 1,254–1,263 |
| 306 | olmo-2-0325-32b-instructAi2 |
1,252 1,241–1,262 |
| 307 | Phi 4Microsoft |
1,256 1,252–1,261 |
| 310 | Command R (08-2024)Cohere |
1,250 1,244–1,257 |
| 314 | mistral-large-2402Mistral |
1,243 1,238–1,247 |
| 314 | Nova Micro 1.0Amazon |
1,241 1,236–1,246 |
| 314 | jamba-1.5-miniAI21 Labs |
1,240 1,233–1,247 |
| 314 | ministral-8b-2410Mistral |
1,238 1,229–1,247 |
| 314 | gemini-pro-dev-apiGoogle |
1,237 1,230–1,244 |
| 314 | hunyuan-standard-256kTencent |
1,233 1,222–1,245 |
| 315 | reka-flash-21b-20240226-onlineRekaai |
1,234 1,226–1,241 |
| 316 | qwen1.5-110b-chatAlibaba |
1,234 1,229–1,240 |
| 316 | qwen1.5-72b-chatAlibaba |
1,234 1,228–1,239 |
| 318 | Mixtral 8x22B InstructMistral |
1,230 1,225–1,234 |
| 318 | reka-flash-21b-20240226Rekaai |
1,227 1,221–1,233 |
| 318 | gemini-proGoogle |
1,224 1,212–1,235 |
| 319 | command-rCohere |
1,227 1,222–1,232 |
| 319 | gpt-3.5-turbo-0125OpenAI |
1,226 1,221–1,230 |
| 319 | c4ai-aya-expanse-8bCohere |
1,223 1,216–1,230 |
| 319 | llama-3.1-tulu-3-8bAi2 |
1,220 1,210–1,231 |
| 323 | llama-3-8b-instructMeta |
1,224 1,220–1,227 |
| 323 | mistral-mediumMistral |
1,223 1,217–1,228 |
| 325 | zephyr-orpo-141b-A35b-v0.1Huggingface |
1,213 1,202–1,224 |
| 330 | yi-1.5-34b-chat01 Ai |
1,213 1,208–1,218 |
| 330 | granite-3.1-8b-instructIBM |
1,208 1,197–1,220 |
| 332 | Llama 3.1 8B InstructMeta |
1,211 1,207–1,215 |
| 332 | gpt-3.5-turbo-1106OpenAI |
1,204 1,195–1,213 |
| 333 | qwen1.5-32b-chatAlibaba |
1,204 1,198–1,210 |
| 336 | gemma-2-2b-itGoogle |
1,200 1,196–1,204 |
| 336 | phi-3-medium-4k-instructMicrosoft |
1,198 1,193–1,203 |
| 337 | mixtral-8x7b-instruct-v0.1Mistral |
1,197 1,193–1,201 |
| 337 | dbrx-instruct-previewDatabricks |
1,196 1,189–1,202 |
| 337 | qwen1.5-14b-chatAlibaba |
1,191 1,184–1,198 |
| 337 | internlm2_5-20b-chatShanghai Ai Lab |
1,191 1,184–1,198 |
| 339 | deepseek-llm-67b-chatDeepSeek |
1,185 1,173–1,196 |
| 341 | wizardlm-70bMicrosoft |
1,185 1,175–1,194 |
| 343 | yi-34b-chat01 Ai |
1,184 1,177–1,191 |
| 343 | granite-3.0-8b-instructIBM |
1,183 1,174–1,192 |
| 343 | openchat-3.5Openchat |
1,183 1,173–1,193 |
| 343 | openchat-3.5-0106Openchat |
1,183 1,175–1,191 |
| 343 | granite-3.1-2b-instructIBM |
1,179 1,168–1,190 |
| 344 | gemma-1.1-7b-itGoogle |
1,183 1,177–1,189 |
| 344 | snowflake-arctic-instructSnowflake |
1,180 1,174–1,186 |
| 344 | tulu-2-dpo-70bAi2 |
1,178 1,168–1,187 |
| 344 | openhermes-2.5-mistral-7bTeknium |
1,176 1,166–1,186 |
| 346 | vicuna-33bLmsys |
1,173 1,167–1,179 |
| 346 | phi-3-small-8k-instructMicrosoft |
1,171 1,165–1,177 |
| 346 | starling-lm-7b-betaNexusflow |
1,171 1,164–1,178 |
| 348 | llama-2-70b-chatMeta |
1,171 1,165–1,176 |
| 348 | starling-lm-7b-alphaBerkeley |
1,167 1,159–1,175 |
| 348 | nous-hermes-2-mixtral-8x7b-dpoNousresearch |
1,164 1,152–1,176 |
| 350 | Llama 3.2 3B InstructMeta |
1,167 1,159–1,174 |
| 356 | llama2-70b-steerlm-chatNVIDIA |
1,155 1,142–1,167 |
| 356 | dolphin-2.2.1-mistral-7bCognitivecomputations |
1,152 1,137–1,168 |
| 357 | qwq-32b-previewAlibaba |
1,154 1,143–1,166 |
| 358 | solar-10.7b-instruct-v1.0Upstage |
1,152 1,139–1,166 |
| 359 | falcon-180b-chatTII |
1,148 1,131–1,165 |
| 360 | granite-3.0-2b-instructIBM |
1,157 1,148–1,165 |
| 361 | mpt-30b-chatMosaicml |
1,151 1,139–1,163 |
| 363 | wizardlm-13bMicrosoft |
1,150 1,140–1,159 |
| 363 | mistral-7b-instruct-v0.2Mistral |
1,149 1,143–1,156 |
| 363 | qwen1.5-7b-chatAlibaba |
1,144 1,134–1,154 |
| 364 | phi-3-mini-4k-instruct-june-2024Microsoft |
1,143 1,137–1,150 |
| 364 | qwen-14b-chatAlibaba |
1,139 1,128–1,150 |
| 364 | palm-2Google |
1,139 1,130–1,149 |
| 365 | vicuna-13bLmsys |
1,142 1,135–1,148 |
| 365 | llama-2-13b-chatMeta |
1,141 1,135–1,148 |
| 365 | gemma-7b-itGoogle |
1,138 1,128–1,147 |
| 365 | codellama-34b-instructMeta |
1,137 1,128–1,146 |
| 367 | zephyr-7b-alphaHuggingface |
1,127 1,111–1,142 |
| 369 | zephyr-7b-betaHuggingface |
1,131 1,122–1,140 |
| 369 | guanaco-33bTimdettmers |
1,127 1,115–1,140 |
| 371 | phi-3-mini-128k-instructMicrosoft |
1,130 1,122–1,137 |
| 371 | codellama-70b-instructMeta |
1,119 1,101–1,137 |
| 375 | phi-3-mini-4k-instructMicrosoft |
1,128 1,122–1,134 |
| 376 | stripedhyena-nous-7bTogether |
1,122 1,111–1,133 |
| 378 | smollm2-1.7b-instructHuggingface |
1,115 1,100–1,129 |
| 381 | gemma-1.1-2b-itGoogle |
1,117 1,109–1,124 |
| 381 | vicuna-7bLmsys |
1,115 1,106–1,124 |
| 384 | Llama 3.2 1B InstructMeta |
1,111 1,103–1,119 |
| 384 | mistral-7b-instructMistral |
1,110 1,101–1,120 |
| 385 | llama-2-7b-chatMeta |
1,108 1,101–1,115 |
| 389 | gemma-2b-itGoogle |
1,094 1,082–1,105 |
| 393 | qwen1.5-4b-chatAlibaba |
1,091 1,082–1,100 |
| 394 | olmo-7b-instructAi2 |
1,074 1,062–1,085 |
| 394 | gpt4all-13b-snoozyNomic |
1,068 1,053–1,083 |
| 396 | koala-13bBerkeley |
1,071 1,061–1,081 |
| 396 | alpaca-13bStanford |
1,070 1,059–1,082 |
| 396 | mpt-7b-chatMosaicml |
1,063 1,052–1,075 |
| 396 | chatglm3-6bZ.ai |
1,057 1,045–1,068 |
| 399 | RWKV-4-Raven-14BRwkv |
1,042 1,031–1,054 |
| 402 | chatglm2-6bZ.ai |
1,025 1,011–1,038 |
| 402 | oasst-pythia-12bOpenassistant |
1,024 1,013–1,034 |
| 405 | chatglm-6bZ.ai |
996 983–1,008 |
| 405 | fastchat-t5-3bLmsys |
993 980–1,005 |
| 405 | dolly-v2-12bDatabricks |
982 969–996 |
| 405 | llama-13bMeta |
975 959–991 |
| 408 | stablelm-tuned-alpha-7bStabilityai |
953 941–966 |
Models share a rank when their ranges overlap. Results as published by Arena (formerly LMArena); we do not re-run them.
What it measures
Head-to-head human preference across all text prompts.
What it does not measure
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Source
Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.