Models / DeepSeek V3

DeepSeek

DeepSeek V3

11 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
DeepSeek
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

1,350
Business, management and finance · rank 183 of 402
Unit
Arena rating, higher is better
Range
1,338 to 1,362
Sample
2241 votes
Configuration
DeepSeek V3
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,349
Creative writing · rank 150 of 407
Unit
Arena rating, higher is better
Range
1,338 to 1,359
Sample
3623 votes
Configuration
DeepSeek V3
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,348
Expert prompts · rank 190 of 359
Unit
Arena rating, higher is better
Range
1,331 to 1,365
Sample
1236 votes
Configuration
DeepSeek V3
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,343
Instruction following · rank 192 of 409
Unit
Arena rating, higher is better
Range
1,336 to 1,350
Sample
8606 votes
Configuration
DeepSeek V3
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,358
Overall · rank 192 of 409
Unit
Arena rating, higher is better
Range
1,354 to 1,363
Sample
21770 votes
Configuration
DeepSeek V3
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,356
Writing, literature and language · rank 159 of 408
Unit
Arena rating, higher is better
Range
1,348 to 1,364
Sample
5913 votes
Configuration
DeepSeek V3
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by UGI Leaderboard
17.0%
Requested-length error · rank 140 of 370
Unit
% off the requested word count, lower is better
Configuration
DeepSeek V3
Measured
21 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.37
Style adherence · rank 84 of 370
Unit
score from 0 to 1, higher is better
Configuration
DeepSeek V3
Measured
21 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
43.2
Writing score · rank 188 of 370
Unit
score out of 100, higher is better
Configuration
DeepSeek V3
Measured
21 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
97.5%
Answer rate · rank 85 of 108
Unit
% of documents, higher is better
Configuration
DeepSeek V3
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
6.1%
Hallucination rate · rank 25 of 108
Unit
% of summaries, lower is better
Configuration
DeepSeek V3
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare DeepSeek V3 with