Models / o1

OpenAI

o1

12 published results from 2 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
OpenAI
Sources
2
Our benchmarks
0
Price
Not yet published

Reported by others

1,370
Business, management and finance · rank 164 of 402
Unit
Arena rating, higher is better
Range
1,358 to 1,382
Sample
2720 votes
Configuration
o1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,381
Creative writing · rank 101 of 407
Unit
Arena rating, higher is better
Range
1,372 to 1,390
Sample
4642 votes
Configuration
o1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,402
Expert prompts · rank 132 of 359
Unit
Arena rating, higher is better
Range
1,385 to 1,419
Sample
1330 votes
Configuration
o1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,407
Instruction following · rank 104 of 409
Unit
Arena rating, higher is better
Range
1,401 to 1,414
Sample
10246 votes
Configuration
o1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,402
Overall · rank 142 of 409
Unit
Arena rating, higher is better
Range
1,398 to 1,407
Sample
27807 votes
Configuration
o1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,398
Writing, literature and language · rank 94 of 408
Unit
Arena rating, higher is better
Range
1,391 to 1,405
Sample
7663 votes
Configuration
o1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by UGI Leaderboard
24.0%
Requested-length error · rank 217 of 370
Unit
% off the requested word count, lower is better
Configuration
o1 (high reasoning)
Measured
24 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
36.0%
Requested-length error · rank 281 of 370
Unit
% off the requested word count, lower is better
Configuration
o1 (low reasoning)
Measured
24 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.36
Style adherence · rank 127 of 370
Unit
score from 0 to 1, higher is better
Configuration
o1 (high reasoning)
Measured
24 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.35
Style adherence · rank 184 of 370
Unit
score from 0 to 1, higher is better
Configuration
o1 (low reasoning)
Measured
24 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
60.1
Writing score · rank 101 of 370
Unit
score out of 100, higher is better
Configuration
o1 (high reasoning)
Measured
24 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
61.0
Writing score · rank 94 of 370
Unit
score out of 100, higher is better
Configuration
o1 (low reasoning)
Measured
24 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare o1 with