Models / gpt-3.5-turbo-0125
OpenAIgpt-3.5-turbo-0125
6 published results from 1 source. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.
- Provider
- OpenAI
- Sources
- 1
- Our benchmarks
- 0
- Price
- Not yet published
Reported by others
1,198
Business, management and finance · rank 322 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,189 to 1,208
- Sample
- 6898 votes
- Configuration
- gpt-3.5-turbo-0125
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,191
Creative writing · rank 309 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,182 to 1,200
- Sample
- 9694 votes
- Configuration
- gpt-3.5-turbo-0125
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,198
Expert prompts · rank 306 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,185 to 1,210
- Sample
- 3207 votes
- Configuration
- gpt-3.5-turbo-0125
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,216
Instruction following · rank 317 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,209 to 1,222
- Sample
- 23523 votes
- Configuration
- gpt-3.5-turbo-0125
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,226
Overall · rank 319 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,221 to 1,230
- Sample
- 66207 votes
- Configuration
- gpt-3.5-turbo-0125
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,226
Writing, literature and language · rank 306 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,219 to 1,233
- Sample
- 16149 votes
- Configuration
- gpt-3.5-turbo-0125
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.