1,323
Business, management and finance · rank 212 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,309 to 1,337
- Sample
- 1833 votes
- Configuration
- gpt-oss-20b
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,240
Creative writing · rank 266 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,221 to 1,258
- Sample
- 1223 votes
- Configuration
- gpt-oss-20b
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,314
Expert prompts · rank 204 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,287 to 1,342
- Sample
- 500 votes
- Configuration
- gpt-oss-20b
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,281
Instruction following · rank 254 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,269 to 1,294
- Sample
- 2532 votes
- Configuration
- gpt-oss-20b
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,318
Overall · rank 239 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,311 to 1,324
- Sample
- 10372 votes
- Configuration
- gpt-oss-20b
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,275
Writing, literature and language · rank 251 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,262 to 1,288
- Sample
- 2161 votes
- Configuration
- gpt-oss-20b
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
282.0%
Requested-length error · rank 370 of 370
- Unit
- % off the requested word count, lower is better
- Configuration
- gpt-oss-20b (high reasoning)
- Measured
- 11 Oct 2025
- Not shown
- Not other format limits such as character counts or bullet counts.
98.0%
Requested-length error · rank 347 of 370
- Unit
- % off the requested word count, lower is better
- Configuration
- gpt-oss-20b (low reasoning)
- Measured
- 6 Sep 2025
- Not shown
- Not other format limits such as character counts or bullet counts.
186.0%
Requested-length error · rank 366 of 370
- Unit
- % off the requested word count, lower is better
- Configuration
- gpt-oss-20b (medium reasoning)
- Measured
- 7 Sep 2025
- Not shown
- Not other format limits such as character counts or bullet counts.
0.27
Style adherence · rank 366 of 370
- Unit
- score from 0 to 1, higher is better
- Configuration
- gpt-oss-20b (high reasoning)
- Measured
- 11 Oct 2025
- Not shown
- Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
0.30
Style adherence · rank 335 of 370
- Unit
- score from 0 to 1, higher is better
- Configuration
- gpt-oss-20b (low reasoning)
- Measured
- 6 Sep 2025
- Not shown
- Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
0.31
Style adherence · rank 302 of 370
- Unit
- score from 0 to 1, higher is better
- Configuration
- gpt-oss-20b (medium reasoning)
- Measured
- 7 Sep 2025
- Not shown
- Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
10.9
Writing score · rank 362 of 370
- Unit
- score out of 100, higher is better
- Configuration
- gpt-oss-20b (high reasoning)
- Measured
- 11 Oct 2025
- Not shown
- Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
24.6
Writing score · rank 307 of 370
- Unit
- score out of 100, higher is better
- Configuration
- gpt-oss-20b (low reasoning)
- Measured
- 6 Sep 2025
- Not shown
- Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
24.5
Writing score · rank 308 of 370
- Unit
- score out of 100, higher is better
- Configuration
- gpt-oss-20b (medium reasoning)
- Measured
- 7 Sep 2025
- Not shown
- Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.