Models / gpt-oss-20b

OpenAI

gpt-oss-20b

15 published results from 2 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
OpenAI
Sources
2
Our benchmarks
0
Price
Not yet published

Reported by others

1,323
Business, management and finance · rank 212 of 402
Unit
Arena rating, higher is better
Range
1,309 to 1,337
Sample
1833 votes
Configuration
gpt-oss-20b
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,240
Creative writing · rank 266 of 407
Unit
Arena rating, higher is better
Range
1,221 to 1,258
Sample
1223 votes
Configuration
gpt-oss-20b
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,314
Expert prompts · rank 204 of 359
Unit
Arena rating, higher is better
Range
1,287 to 1,342
Sample
500 votes
Configuration
gpt-oss-20b
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,281
Instruction following · rank 254 of 409
Unit
Arena rating, higher is better
Range
1,269 to 1,294
Sample
2532 votes
Configuration
gpt-oss-20b
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,318
Overall · rank 239 of 409
Unit
Arena rating, higher is better
Range
1,311 to 1,324
Sample
10372 votes
Configuration
gpt-oss-20b
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,275
Writing, literature and language · rank 251 of 408
Unit
Arena rating, higher is better
Range
1,262 to 1,288
Sample
2161 votes
Configuration
gpt-oss-20b
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by UGI Leaderboard
282.0%
Requested-length error · rank 370 of 370
Unit
% off the requested word count, lower is better
Configuration
gpt-oss-20b (high reasoning)
Measured
11 Oct 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
98.0%
Requested-length error · rank 347 of 370
Unit
% off the requested word count, lower is better
Configuration
gpt-oss-20b (low reasoning)
Measured
6 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
186.0%
Requested-length error · rank 366 of 370
Unit
% off the requested word count, lower is better
Configuration
gpt-oss-20b (medium reasoning)
Measured
7 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.27
Style adherence · rank 366 of 370
Unit
score from 0 to 1, higher is better
Configuration
gpt-oss-20b (high reasoning)
Measured
11 Oct 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.30
Style adherence · rank 335 of 370
Unit
score from 0 to 1, higher is better
Configuration
gpt-oss-20b (low reasoning)
Measured
6 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.31
Style adherence · rank 302 of 370
Unit
score from 0 to 1, higher is better
Configuration
gpt-oss-20b (medium reasoning)
Measured
7 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
10.9
Writing score · rank 362 of 370
Unit
score out of 100, higher is better
Configuration
gpt-oss-20b (high reasoning)
Measured
11 Oct 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
24.6
Writing score · rank 307 of 370
Unit
score out of 100, higher is better
Configuration
gpt-oss-20b (low reasoning)
Measured
6 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
24.5
Writing score · rank 308 of 370
Unit
score out of 100, higher is better
Configuration
gpt-oss-20b (medium reasoning)
Measured
7 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare gpt-oss-20b with