Benchmarks / Arena (formerly LMArena)

Reported by Arena (formerly LMArena)

Arena (formerly LMArena)

Human preference on prompts that set explicit instructions.

Results dated
25 Sep 2026
Models
409
Unit
Arena rating
Licence
Creative Commons Attribution 4.0 International

Full results

Arena Text: instruction following, Arena rating, higher is better
#ModelArena Text: instruction following · 95% range
Arena rating, higher is better
1 Claude Opus 5.5 (high)Anthropic
1,516
1,496–1,537
1 Claude Opus 4.6 (high)Anthropic
1,513
1,508–1,519
1 Claude Fable 5 (high)Anthropic
1,509
1,503–1,515
1 Claude Fable 5.1 (max)Anthropic
1,497
1,486–1,509
2 Claude Opus 4.7 (high)Anthropic
1,503
1,497–1,508
2 Claude Opus 4.6Anthropic
1,500
1,495–1,505
2 Claude Opus 5 (high)Anthropic
1,497
1,491–1,503
3 Claude Opus 4.7Anthropic
1,494
1,488–1,499
3 Claude Opus 5 (max)Anthropic
1,493
1,485–1,500
4 Claude Opus 4.8 (high)Anthropic
1,491
1,486–1,496
4 Muse Spark 1.3 (max)Meta
1,486
1,476–1,496
5 GPT-5.6 Sol (xhigh)OpenAI
1,489
1,483–1,496
5 Gemini 3.8 Flash (high)Google
1,488
1,480–1,496
5 Gemini 3.7 Flash (high)Google
1,487
1,479–1,495
6 Kimi K3 (max)Moonshot AI
1,486
1,479–1,493
6 Muse Spark 1.2Meta
1,475
1,458–1,492
7 Claude Opus 4.5Anthropic
1,483
1,477–1,490
7 GLM 5.3 (max)Z.ai
1,480
1,471–1,488
7 DeepSeek V4.1 Flash (max)DeepSeek
1,478
1,466–1,490
8 MiMo-V2.6-ProXiaomi
1,471
1,455–1,487
9 GPT-6 Astra (max)OpenAI
1,473
1,461–1,486
11 Gemini 3.1 Pro PreviewGoogle
1,479
1,475–1,484
11 GPT-5.5 (high)OpenAI
1,478
1,473–1,483
11 Claude Opus 4.8Anthropic
1,478
1,472–1,483
12 GLM 5.3 FlashZ.ai
1,474
1,466–1,482
12 qwen3.7-max-previewAlibaba
1,466
1,449–1,482
13 GPT-5.5OpenAI
1,475
1,470–1,480
13 Claude Sonnet 4.6Anthropic
1,475
1,469–1,480
13 Claude Opus 4.5Anthropic
1,475
1,470–1,480
13 Muse Spark 1.1Meta
1,473
1,467–1,479
13 Gemini 3.6 Flash (high)Google
1,473
1,466–1,479
13 Qwen3.8 Max (0902)Alibaba
1,472
1,465–1,480
15 gemini-3-proGoogle
1,473
1,466–1,479
15 GPT-5.4 (high)OpenAI
1,471
1,466–1,477
16 GLM 5.2 (max)Z.ai
1,470
1,464–1,476
18 MiMo-V2.5-ProXiaomi
1,469
1,464–1,474
18 qwen3.5-max-previewAlibaba
1,467
1,459–1,474
18 Claude Sonnet 5 (high)Anthropic
1,467
1,461–1,473
18 muse-sparkMeta
1,465
1,456–1,474
20 DeepSeek V4 Pro 0813DeepSeek
1,461
1,451–1,471
21 Gemini 3.5 Flash (high)Google
1,465
1,459–1,470
21 Gemini 3.5 Flash (medium)Google
1,464
1,458–1,470
24 Claude Sonnet 4.5Anthropic
1,464
1,459–1,468
24 Grok 4.5xAI
1,462
1,456–1,468
24 GPT-5.6 Terra (xhigh)OpenAI
1,462
1,455–1,468
24 Gemma 4 31BGoogle
1,453
1,439–1,467
28 Claude Sonnet 4.5Anthropic
1,462
1,457–1,466
29 GLM 5.1Z.ai
1,460
1,455–1,466
29 GPT-5.4OpenAI
1,460
1,455–1,466
31 gemini-3-flashGoogle
1,458
1,451–1,465
31 MiMo-V2.6-FlashXiaomi
1,449
1,434–1,465
32 Claude Opus 4.1Anthropic
1,458
1,452–1,464
33 gpt-5.2-chat-latest-20260210OpenAI
1,457
1,450–1,463
33 gpt-5.5-instantOpenAI
1,456
1,449–1,463
33 Kimi K2.6Moonshot AI
1,456
1,449–1,462
33 Qwen3.6 Max PreviewAlibaba
1,449
1,435–1,463
33 GPT-6 Sol (max)OpenAI
1,447
1,432–1,461
37 Claude Opus 4.1Anthropic
1,454
1,450–1,459
38 ernie-5.1Baidu
1,452
1,445–1,459
38 Grok 4.6 (high)xAI
1,451
1,443–1,458
40 DeepSeek V4 Pro 0423DeepSeek
1,452
1,447–1,458
40 GPT-6 Luna (max)OpenAI
1,443
1,429–1,457
41 GPT-5.1 (high)OpenAI
1,450
1,444–1,457
41 grok-4.20-beta1xAI
1,450
1,443–1,457
42 Hy3Tencent
1,446
1,435–1,456
47 deepseek-v4-pro-preview (high reasoning)DeepSeek
1,448
1,442–1,454
47 Qwen3.7 PlusAlibaba
1,447
1,441–1,453
47 GLM 5Z.ai
1,446
1,439–1,452
47 Grok 4.7 (xhigh)xAI
1,437
1,421–1,453
48 grok-4.20-beta-0309-reasoningxAI
1,446
1,441–1,452
48 GPT-5.6 Luna (xhigh)OpenAI
1,446
1,439–1,452
48 mimo-v2-proXiaomi
1,445
1,437–1,452
48 claude-opus-4 (thinking-16k reasoning)Anthropic
1,444
1,437–1,451
48 Gemma 4 26B A4BGoogle
1,437
1,424–1,451
51 Gemini 3.5 Flash LiteGoogle
1,444
1,437–1,450
54 grok-4.20-multi-agent-beta-0309xAI
1,444
1,438–1,449
55 gemini-3-flash (minimal reasoning)Google
1,443
1,438–1,448
57 Kimi K2.5Moonshot AI
1,439
1,434–1,444
57 gpt-4.5-preview-2025-02-27OpenAI
1,437
1,428–1,445
58 Kimi K2.5Moonshot AI
1,432
1,420–1,444
60 gpt-5.3-chat-latestOpenAI
1,436
1,429–1,442
61 Gemini 2.5 ProGoogle
1,438
1,434–1,442
61 Qwen3.6 PlusAlibaba
1,436
1,430–1,442
63 dola-seed-2.0-proBytedance
1,435
1,430–1,440
63 MiniMax M3MiniMax
1,435
1,429–1,440
63 DeepSeek V3.1 TerminusDeepSeek
1,420
1,400–1,440
65 GPT-5.4 Mini (high)OpenAI
1,434
1,428–1,439
66 Qwen3.5 397B A17BAlibaba
1,434
1,429–1,439
66 deepseek-v4-flash-preview (high reasoning)DeepSeek
1,433
1,428–1,439
71 grok-4.1xAI
1,431
1,426–1,436
71 MiMo-V2.5Xiaomi
1,430
1,424–1,437
71 GLM 4.7Z.ai
1,425
1,415–1,436
73 grok-4.1-thinkingxAI
1,429
1,424–1,434
73 mimo-v2-omniXiaomi
1,427
1,419–1,435
75 DeepSeek V4 Flash 0423DeepSeek
1,428
1,422–1,434
76 GPT-5.2 (high)OpenAI
1,427
1,421–1,433
76 InklingThinkingmachines
1,427
1,420–1,433
76 GPT-5.1OpenAI
1,426
1,420–1,432
76 GLM 5V TurboZ.ai
1,423
1,413–1,433
76 amazon-nova-experimental-chat-26-02-10Amazon
1,414
1,397–1,432
77 chatgpt-4o-latest-20250326OpenAI
1,427
1,422–1,431
77 ernie-5.0-0110Baidu
1,425
1,419–1,431
77 qwen3-max-previewAlibaba
1,424
1,417–1,432
77 Qwen3.8 27BAlibaba
1,423
1,415–1,432
77 Mistral Medium 3.5Mistral
1,421
1,411–1,431
77 ernie-5.0-preview-1203Baidu
1,419
1,408–1,431
83 GPT-5.2OpenAI
1,424
1,419–1,429
85 muse-glimmerMeta
1,412
1,396–1,428
86 DeepSeek V3.2DeepSeek
1,420
1,414–1,426
86 DeepSeek V3.1DeepSeek
1,416
1,405–1,427
86 DeepSeek V3.2 ExpDeepSeek
1,416
1,404–1,427
87 DeepSeek V3.2DeepSeek
1,420
1,414–1,425
87 DeepSeek V3.2 ExpDeepSeek
1,414
1,404–1,424
87 Qwen3 MaxAlibaba
1,413
1,402–1,425
90 kimi-k2-thinking-turboMoonshot AI
1,418
1,413–1,423
90 longcat-flash-chat-2602-expMeituan
1,416
1,409–1,423
92 Claude Sonnet 4Anthropic
1,415
1,408–1,422
92 claude-opus-4Anthropic
1,415
1,409–1,422
92 gpt-5-chatOpenAI
1,415
1,408–1,422
94 Qwen3 VL 235B A22B InstructAlibaba
1,409
1,398–1,421
96 Grok 4.3xAI
1,415
1,410–1,420
99 Claude Haiku 4.5Anthropic
1,415
1,411–1,419
100 Qwen3 235B A22B Instruct 2507Alibaba
1,414
1,410–1,419
100 GLM 4.6Z.ai
1,412
1,406–1,419
100 claude-3-7-sonnet-20250219 (thinking-32k reasoning)Anthropic
1,412
1,405–1,418
100 GPT-5 (high)OpenAI
1,410
1,403–1,417
100 hunyuan-vision-1.5-thinkingTencent
1,396
1,374–1,419
101 amazon-nova-experimental-chat-26-01-10Amazon
1,398
1,380–1,416
103 nvidia-nemotron-3-ultra-550b-a55b-nvfp4NVIDIA
1,404
1,394–1,414
104 o1OpenAI
1,407
1,401–1,414
107 Gemini 3.1 Flash Lite PreviewGoogle
1,407
1,402–1,413
107 MiniMax M2.7MiniMax
1,407
1,402–1,412
107 Qwen3.5-122B-A10BAlibaba
1,406
1,399–1,412
107 GLM 4.5Z.ai
1,405
1,397–1,413
107 DeepSeek V3.1DeepSeek
1,402
1,392–1,412
107 ernie-5.0-preview-1022Baidu
1,398
1,382–1,413
107 DeepSeek V3.1 TerminusDeepSeek
1,394
1,377–1,412
110 grok-3-preview-02-24xAI
1,403
1,397–1,410
110 grok-4-fast-chatxAI
1,395
1,381–1,410
111 GPT-4.1OpenAI
1,403
1,398–1,409
111 Hy3 previewTencent
1,397
1,385–1,410
115 o3OpenAI
1,403
1,397–1,408
116 grok-4-1-fast-reasoningxAI
1,401
1,396–1,407
116 Qwen3.5-27BAlibaba
1,400
1,394–1,407
116 gemini-2.5-flash-preview-09-2025Google
1,400
1,393–1,406
119 Mistral Large 3 2512Mistral
1,400
1,396–1,405
119 R1DeepSeek
1,397
1,390–1,405
121 Gemini 2.5 FlashGoogle
1,400
1,396–1,403
121 grok-4-0709xAI
1,397
1,391–1,403
122 Claude Sonnet 4Anthropic
1,396
1,389–1,402
125 Kimi K2 0905Moonshot AI
1,390
1,379–1,401
126 Inkling SmallThinkingmachines
1,393
1,386–1,400
126 R1 0528DeepSeek
1,391
1,381–1,400
127 Mistral Medium 3.1Mistral
1,395
1,391–1,399
128 longcat-flash-chatMeituan
1,388
1,377–1,399
130 Qwen3 235B A22B Thinking 2507Alibaba
1,385
1,372–1,397
131 grok-4-fast-reasoningxAI
1,389
1,380–1,397
136 amazon-nova-experimental-chat-12-10Amazon
1,377
1,359–1,396
138 Qwen3 VL 235B A22B ThinkingAlibaba
1,383
1,370–1,395
140 minimax-m2.1-previewMiniMax
1,385
1,376–1,394
141 Qwen3.5-35B-A3BAlibaba
1,386
1,380–1,393
142 GPT-5.4 Nano (high)OpenAI
1,387
1,381–1,392
142 claude-3-7-sonnet-20250219Anthropic
1,385
1,380–1,391
142 Qwen3 Coder 480B A35BAlibaba
1,384
1,376–1,392
142 hunyuan-t1-20250711Tencent
1,374
1,356–1,391
145 Step 3.5 FlashStepfun
1,384
1,379–1,389
146 mimo-v2-flash (no reasoning)Xiaomi
1,382
1,376–1,388
146 Qwen3.5-FlashAlibaba
1,381
1,376–1,387
146 MiniMax M2.5MiniMax
1,381
1,375–1,387
146 o1-previewOpenAI
1,381
1,374–1,388
146 Kimi K2 0711Moonshot AI
1,381
1,373–1,389
146 Qwen3 235B A22BAlibaba
1,381
1,374–1,388
146 Solar Pro 4Upstage
1,379
1,370–1,388
146 mimo-v2-flash (thinking reasoning)Xiaomi
1,375
1,365–1,386
146 GLM 4.6VZ.ai
1,367
1,346–1,388
148 DeepSeek V3 0324DeepSeek
1,378
1,372–1,384
148 Qwen3 Next 80B A3B InstructAlibaba
1,376
1,369–1,384
152 GPT-5 Mini (high)OpenAI
1,373
1,365–1,381
156 GPT-4.1 MiniOpenAI
1,373
1,366–1,379
158 claude-3-5-sonnet-20241022Anthropic
1,374
1,370–1,378
161 trinity-large-previewArcee Ai
1,370
1,363–1,376
165 o4 MiniOpenAI
1,369
1,362–1,375
165 gemini-2.5-flash-lite-preview-06-17 (thinking reasoning)Google
1,368
1,361–1,375
165 Qwen3 30B A3B Instruct 2507Alibaba
1,366
1,359–1,374
167 o3 Mini HighOpenAI
1,366
1,359–1,374
168 Mistral Medium 3Mistral
1,366
1,359–1,373
169 llama-3.1-nemotron-ultra-253b-v1NVIDIA
1,351
1,330–1,373
174 gemini-2.5-flash-lite-preview-09-2025 (no reasoning)Google
1,364
1,358–1,369
174 grok-3-mini (high reasoning)xAI
1,360
1,351–1,369
174 Qwen3 Next 80B A3B ThinkingAlibaba
1,359
1,349–1,369
175 GLM 4.5 AirZ.ai
1,361
1,354–1,368
175 amazon-nova-experimental-chat-11-10Amazon
1,361
1,354–1,368
175 hunyuan-turbos-20250226Tencent
1,351
1,333–1,368
177 hunyuan-turbos-20250416Tencent
1,353
1,341–1,365
178 Trinity Large ThinkingArcee Ai
1,357
1,351–1,364
178 Qwen3 235B A22BAlibaba
1,357
1,349–1,365
179 qwen2.5-maxAlibaba
1,357
1,351–1,363
181 grok-3-mini-betaxAI
1,353
1,344–1,361
183 step-3Stepfun
1,344
1,330–1,359
185 GLM 4.7 FlashZ.ai
1,348
1,338–1,358
186 granite-4.2-30bIBM
1,339
1,321–1,357
187 amazon-nova-experimental-chat-10-20Amazon
1,345
1,335–1,356
187 Nemotron 3 SuperNVIDIA
1,343
1,330–1,355
187 GLM 4.5VZ.ai
1,338
1,323–1,354
189 gemini-2.0-flash-001Google
1,347
1,342–1,353
189 MiniMax M1MiniMax
1,346
1,340–1,353
189 nvidia-nemotron-3.5-lightning-30b-a3b-nvfp4NVIDIA
1,343
1,333–1,353
192 DeepSeek V3DeepSeek
1,343
1,336–1,350
192 Qwen3 32BAlibaba
1,331
1,312–1,350
194 o3 MiniOpenAI
1,344
1,339–1,349
194 Gemma 3 27BGoogle
1,343
1,337–1,349
194 Command ACohere
1,342
1,336–1,347
194 Mistral Small 3.2 24BMistral
1,338
1,329–1,347
194 MiniMax M2MiniMax
1,334
1,321–1,347
194 intellect-3Primeintellect
1,332
1,317–1,348
194 Mercury 2Inception
1,327
1,307–1,346
195 gemini-1.5-pro-002Google
1,340
1,335–1,345
195 llama-3.3-nemotron-49b-super-v1NVIDIA
1,325
1,306–1,345
196 claude-3-5-sonnet-20240620Anthropic
1,338
1,333–1,343
196 Qwen-PlusAlibaba
1,332
1,320–1,344
196 molmo-2-8bAi2
1,304
1,267–1,342
198 GPT-5 Nano (high)OpenAI
1,327
1,314–1,340
198 nvidia-llama-3.3-nemotron-super-49b-v1.5NVIDIA
1,321
1,301–1,341
201 gemini-2.0-flash-lite-preview-02-05Google
1,331
1,325–1,338
201 Nova 2 LiteAmazon
1,327
1,317–1,338
202 o1-miniOpenAI
1,332
1,326–1,337
202 Gemma 3 12BGoogle
1,321
1,305–1,337
204 amazon-nova-experimental-chat-10-09Amazon
1,315
1,293–1,336
209 GPT-4o (2024-05-13)OpenAI
1,326
1,321–1,331
209 gpt-oss-120bOpenAI
1,323
1,316–1,330
209 olmo-3.1-32b-instructAi2
1,320
1,310–1,331
209 hunyuan-large-2025-02-10Tencent
1,316
1,300–1,332
209 hunyuan-turbo-0110Tencent
1,314
1,295–1,332
210 ring-flash-2.0Ant Group
1,316
1,303–1,330
211 qwq-32bAlibaba
1,323
1,316–1,330
212 ling-flash-2.0Ant Group
1,315
1,302–1,329
213 step-2-16k-exp-202412Stepfun
1,316
1,303–1,328
213 glm-4-plus-0111Z.ai
1,315
1,302–1,327
214 deepseek-v2.5-1210DeepSeek
1,314
1,303–1,326
215 GPT-4o (2024-08-06)OpenAI
1,319
1,313–1,325
215 gemini-advanced-0514Google
1,317
1,311–1,324
216 claude-3-5-haiku-20241022Anthropic
1,316
1,311–1,321
219 llama-3.1-405b-instruct-fp8Meta
1,315
1,310–1,320
219 llama-3.1-405b-instruct-bf16Meta
1,314
1,309–1,320
219 Llama 4 MaverickMeta
1,314
1,307–1,320
219 step-1o-turbo-202506Stepfun
1,307
1,294–1,320
220 claude-3-opus-20240229Anthropic
1,314
1,310–1,319
220 grok-2-2024-08-13xAI
1,312
1,307–1,317
220 gemini-1.5-pro-001Google
1,312
1,306–1,317
220 Qwen3 30B A3BAlibaba
1,310
1,302–1,318
224 yi-lightning01 Ai
1,308
1,301–1,315
224 olmo-3-32b-thinkAi2
1,298
1,283–1,314
225 GPT-4.1 NanoOpenAI
1,300
1,288–1,313
228 magistral-medium-2506Mistral
1,300
1,290–1,311
232 qwen-max-0919Alibaba
1,302
1,294–1,309
233 GPT-4 TurboOpenAI
1,303
1,298–1,309
233 athene-v2-chatNexusflow
1,302
1,296–1,308
233 glm-4-plusZ.ai
1,301
1,294–1,308
236 Llama 4 ScoutMeta
1,299
1,292–1,306
238 Mistral Large 2407Mistral
1,299
1,294–1,305
239 qwen2.5-plus-1127Alibaba
1,294
1,285–1,304
239 granite-4.1-8bIBM
1,287
1,270–1,305
239 hunyuan-large-visionTencent
1,287
1,270–1,303
239 granite-4.2-3bIBM
1,283
1,264–1,303
242 gpt-4-1106-previewOpenAI
1,297
1,291–1,302
247 mistral-large-2411Mistral
1,295
1,289–1,301
247 Mistral Small 3.1 24BMistral
1,294
1,287–1,301
248 GPT-4o-mini (2024-07-18)OpenAI
1,293
1,288–1,298
248 deepseek-v2.5DeepSeek
1,292
1,285–1,298
248 Nemotron 3 Nano 30B A3BNVIDIA
1,291
1,282–1,300
248 Granite 4.2 8BIBM
1,280
1,261–1,299
249 Llama 3.3 70B InstructMeta
1,293
1,288–1,298
249 Qwen2.5 72B InstructAlibaba
1,292
1,286–1,298
249 gpt-4-0125-previewOpenAI
1,290
1,285–1,296
249 gemini-1.5-flash-002Google
1,290
1,284–1,296
249 hunyuan-standard-2025-02-10Tencent
1,282
1,266–1,297
251 mercuryInception
1,269
1,243–1,294
254 gpt-oss-20bOpenAI
1,281
1,269–1,294
256 gpt-4-0314OpenAI
1,285
1,278–1,293
256 olmo-3.1-32b-thinkAi2
1,279
1,266–1,293
257 llama-3.1-nemotron-70b-instructNVIDIA
1,280
1,269–1,291
258 gemma-3n-e4b-itGoogle
1,282
1,273–1,290
259 grok-2-mini-2024-08-13xAI
1,284
1,278–1,289
262 llama-3.1-tulu-3-70bAi2
1,272
1,257–1,288
264 athene-70b-0725Nexusflow
1,279
1,271–1,286
269 gpt-4-0613OpenAI
1,277
1,271–1,283
269 Nova Pro 1.0Amazon
1,277
1,270–1,283
269 ibm-granite-h-smallIBM
1,268
1,252–1,283
269 Gemma 3 4BGoogle
1,267
1,251–1,284
272 Llama 3.1 70B InstructMeta
1,273
1,267–1,278
273 Gemma 2 27BGoogle
1,271
1,267–1,276
273 claude-3-sonnet-20240229Anthropic
1,268
1,263–1,274
273 jamba-1.5-largeAI21 Labs
1,267
1,256–1,277
273 Qwen2.5 Coder 32B InstructAlibaba
1,264
1,252–1,276
273 llama-3.1-nemotron-51b-instructNVIDIA
1,263
1,248–1,277
274 gemini-1.5-flash-001Google
1,265
1,259–1,271
274 reka-core-20240904Rekaai
1,262
1,252–1,272
281 nemotron-4-340b-instructNVIDIA
1,259
1,251–1,268
282 glm-4-0520Z.ai
1,256
1,246–1,267
286 llama-3-70b-instructMeta
1,259
1,253–1,264
286 Mistral Small 3Mistral
1,255
1,247–1,264
287 Command R+ (08-2024)Cohere
1,254
1,245–1,263
288 deepseek-coder-v2DeepSeek
1,253
1,244–1,262
288 gemma-2-9b-it-simpoPrinceton Nlp
1,252
1,242–1,262
289 reka-flash-20240904Rekaai
1,249
1,239–1,260
289 hunyuan-standard-256kTencent
1,243
1,227–1,260
292 c4ai-aya-expanse-32bCohere
1,249
1,243–1,256
295 Phi 4Microsoft
1,245
1,239–1,252
298 claude-3-haiku-20240307Anthropic
1,245
1,240–1,251
298 gemma-2-9b-itGoogle
1,245
1,240–1,251
298 Nova Lite 1.0Amazon
1,244
1,237–1,251
298 qwen2-72b-instructAlibaba
1,243
1,236–1,249
299 command-r-plusCohere
1,241
1,235–1,247
300 olmo-2-0325-32b-instructAi2
1,229
1,212–1,246
301 Command R (08-2024)Cohere
1,236
1,227–1,245
302 gemini-1.5-flash-8b-001Google
1,238
1,232–1,244
302 mistral-large-2402Mistral
1,238
1,231–1,244
310 gemini-proGoogle
1,220
1,203–1,237
317 qwen1.5-110b-chatAlibaba
1,218
1,210–1,226
317 gpt-3.5-turbo-0125OpenAI
1,216
1,209–1,222
317 Mixtral 8x22B InstructMistral
1,215
1,209–1,222
317 Nova Micro 1.0Amazon
1,215
1,208–1,222
317 qwen1.5-72b-chatAlibaba
1,212
1,205–1,219
317 ministral-8b-2410Mistral
1,211
1,198–1,224
317 mistral-mediumMistral
1,211
1,202–1,219
317 gemini-pro-dev-apiGoogle
1,210
1,200–1,221
317 llama-3.1-tulu-3-8bAi2
1,208
1,193–1,224
317 jamba-1.5-miniAI21 Labs
1,205
1,194–1,215
317 c4ai-aya-expanse-8bCohere
1,205
1,195–1,214
317 reka-flash-21b-20240226-onlineRekaai
1,202
1,192–1,212
319 gpt-3.5-turbo-1106OpenAI
1,198
1,186–1,210
319 granite-3.1-8b-instructIBM
1,193
1,176–1,210
320 zephyr-orpo-141b-A35b-v0.1Huggingface
1,194
1,178–1,209
322 command-rCohere
1,200
1,193–1,207
325 reka-flash-21b-20240226Rekaai
1,193
1,184–1,201
326 llama-3-8b-instructMeta
1,193
1,187–1,198
327 Llama 3.1 8B InstructMeta
1,191
1,185–1,197
327 yi-1.5-34b-chat01 Ai
1,188
1,180–1,196
327 dbrx-instruct-previewDatabricks
1,188
1,179–1,196
329 qwen1.5-32b-chatAlibaba
1,185
1,176–1,193
332 mixtral-8x7b-instruct-v0.1Mistral
1,181
1,174–1,187
332 internlm2_5-20b-chatShanghai Ai Lab
1,178
1,168–1,188
332 granite-3.1-2b-instructIBM
1,173
1,157–1,190
334 tulu-2-dpo-70bAi2
1,171
1,156–1,186
335 phi-3-medium-4k-instructMicrosoft
1,178
1,170–1,185
335 granite-3.0-8b-instructIBM
1,172
1,160–1,185
338 qwen1.5-14b-chatAlibaba
1,168
1,158–1,178
339 gemma-2-2b-itGoogle
1,171
1,165–1,177
339 wizardlm-70bMicrosoft
1,164
1,150–1,177
341 deepseek-llm-67b-chatDeepSeek
1,158
1,141–1,175
343 openchat-3.5Openchat
1,155
1,140–1,169
343 openhermes-2.5-mistral-7bTeknium
1,154
1,138–1,169
344 openchat-3.5-0106Openchat
1,156
1,145–1,166
345 gemma-1.1-7b-itGoogle
1,157
1,149–1,165
345 phi-3-small-8k-instructMicrosoft
1,154
1,146–1,163
345 snowflake-arctic-instructSnowflake
1,154
1,145–1,162
345 yi-34b-chat01 Ai
1,152
1,143–1,162
345 qwq-32b-previewAlibaba
1,146
1,130–1,161
345 falcon-180b-chatTII
1,135
1,106–1,163
347 Llama 3.2 3B InstructMeta
1,146
1,135–1,157
349 starling-lm-7b-betaNexusflow
1,146
1,136–1,156
349 mpt-30b-chatMosaicml
1,135
1,114–1,156
349 dolphin-2.2.1-mistral-7bCognitivecomputations
1,129
1,105–1,154
350 vicuna-33bLmsys
1,140
1,131–1,149
351 starling-lm-7b-alphaBerkeley
1,137
1,125–1,148
354 llama-2-70b-chatMeta
1,135
1,127–1,143
354 granite-3.0-2b-instructIBM
1,132
1,120–1,145
354 llama2-70b-steerlm-chatNVIDIA
1,125
1,106–1,143
357 wizardlm-13bMicrosoft
1,124
1,110–1,138
357 qwen1.5-7b-chatAlibaba
1,124
1,110–1,138
359 qwen-14b-chatAlibaba
1,119
1,103–1,135
360 phi-3-mini-4k-instruct-june-2024Microsoft
1,125
1,115–1,134
360 mistral-7b-instruct-v0.2Mistral
1,123
1,114–1,132
360 solar-10.7b-instruct-v1.0Upstage
1,116
1,098–1,135
361 palm-2Google
1,117
1,104–1,131
362 codellama-70b-instructMeta
1,098
1,068–1,127
363 vicuna-13bLmsys
1,117
1,108–1,127
363 smollm2-1.7b-instructHuggingface
1,104
1,083–1,125
364 phi-3-mini-4k-instructMicrosoft
1,113
1,105–1,122
364 nous-hermes-2-mixtral-8x7b-dpoNousresearch
1,108
1,091–1,124
364 zephyr-7b-alphaHuggingface
1,099
1,075–1,123
365 llama-2-13b-chatMeta
1,110
1,101–1,120
365 gemma-7b-itGoogle
1,105
1,092–1,118
365 codellama-34b-instructMeta
1,104
1,091–1,117
368 gemma-1.1-2b-itGoogle
1,101
1,091–1,112
368 phi-3-mini-128k-instructMicrosoft
1,100
1,090–1,111
372 stripedhyena-nous-7bTogether
1,091
1,076–1,106
377 zephyr-7b-betaHuggingface
1,089
1,076–1,102
378 mistral-7b-instructMistral
1,087
1,074–1,101
379 Llama 3.2 1B InstructMeta
1,086
1,074–1,097
382 vicuna-7bLmsys
1,077
1,063–1,091
384 gemma-2b-itGoogle
1,071
1,055–1,087
384 qwen1.5-4b-chatAlibaba
1,070
1,057–1,083
384 guanaco-33bTimdettmers
1,066
1,045–1,087
385 llama-2-7b-chatMeta
1,070
1,060–1,080
391 gpt4all-13b-snoozyNomic
1,038
1,013–1,063
395 chatglm3-6bZ.ai
1,037
1,019–1,054
395 olmo-7b-instructAi2
1,029
1,013–1,045
396 koala-13bBerkeley
1,027
1,011–1,042
396 alpaca-13bStanford
1,025
1,009–1,041
396 mpt-7b-chatMosaicml
1,009
991–1,028
401 oasst-pythia-12bOpenassistant
988
972–1,004
401 chatglm2-6bZ.ai
983
960–1,005
401 chatglm-6bZ.ai
978
960–996
402 RWKV-4-Raven-14BRwkv
970
953–987
402 fastchat-t5-3bLmsys
958
940–977
403 dolly-v2-12bDatabricks
940
918–961
406 llama-13bMeta
918
893–944
407 stablelm-tuned-alpha-7bStabilityai
910
890–930

Models share a rank when their ranges overlap. Results as published by Arena (formerly LMArena); we do not re-run them.

What it measures

Human preference on prompts that set explicit instructions.

What it does not measure

Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.

Contains data from the Arena Leaderboard Dataset by Arena, licensed under CC BY 4.0. Licence: Creative Commons Attribution 4.0 International.