Models / Claude Opus 4.1

Anthropic

Claude Opus 4.1

23 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Anthropic
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

1,167
Overall · rank 23 of 34
Unit
Arena rating, higher is better
Range
1,162 to 1,172
Sample
76933 votes
Configuration
Claude Opus 4.1
Measured
24 Aug 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,435
Overall · rank 81 of 177
Unit
Arena rating, higher is better
Range
1,432 to 1,439
Sample
41433 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,433
Overall · rank 84 of 177
Unit
Arena rating, higher is better
Range
1,429 to 1,437
Sample
25840 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,447
Business, management and finance · rank 59 of 402
Unit
Arena rating, higher is better
Range
1,441 to 1,452
Sample
14034 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,447
Business, management and finance · rank 52 of 402
Unit
Arena rating, higher is better
Range
1,441 to 1,454
Sample
9072 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,442
Creative writing · rank 26 of 407
Unit
Arena rating, higher is better
Range
1,436 to 1,448
Sample
10814 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,445
Creative writing · rank 24 of 407
Unit
Arena rating, higher is better
Range
1,438 to 1,453
Sample
6823 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,466
Expert prompts · rank 56 of 359
Unit
Arena rating, higher is better
Range
1,457 to 1,475
Sample
4189 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,482
Expert prompts · rank 32 of 359
Unit
Arena rating, higher is better
Range
1,470 to 1,494
Sample
2510 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,454
Instruction following · rank 37 of 409
Unit
Arena rating, higher is better
Range
1,450 to 1,459
Sample
20685 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,458
Instruction following · rank 32 of 409
Unit
Arena rating, higher is better
Range
1,452 to 1,464
Sample
13129 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,448
Overall · rank 67 of 409
Unit
Arena rating, higher is better
Range
1,445 to 1,451
Sample
76940 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,450
Overall · rank 60 of 409
Unit
Arena rating, higher is better
Range
1,446 to 1,453
Sample
49332 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,444
Writing, literature and language · rank 39 of 408
Unit
Arena rating, higher is better
Range
1,439 to 1,449
Sample
17365 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,446
Writing, literature and language · rank 36 of 408
Unit
Arena rating, higher is better
Range
1,440 to 1,452
Sample
10992 votes
Configuration
Claude Opus 4.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by UGI Leaderboard
3.0%
Requested-length error · rank 18 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-4.1
Measured
30 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
9.0%
Requested-length error · rank 67 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-4.1
Measured
12 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.34
Style adherence · rank 201 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-4.1
Measured
30 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.38
Style adherence · rank 78 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-4.1
Measured
12 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
69.9
Writing score · rank 23 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-4.1
Measured
30 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
66.5
Writing score · rank 54 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-4.1
Measured
12 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
92.4%
Answer rate · rank 100 of 108
Unit
% of documents, higher is better
Configuration
Claude Opus 4.1
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
11.8%
Hallucination rate · rank 76 of 108
Unit
% of summaries, lower is better
Configuration
Claude Opus 4.1
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare Claude Opus 4.1 with