Models / gpt-4-0125-preview
OpenAIgpt-4-0125-preview
6 published results from 1 source. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.
- Provider
- OpenAI
- Sources
- 1
- Our benchmarks
- 0
- Price
- Not yet published
Reported by others
1,283
Business, management and finance · rank 260 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,274 to 1,292
- Sample
- 9697 votes
- Configuration
- gpt-4-0125-preview
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,280
Creative writing · rank 226 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,272 to 1,289
- Sample
- 13940 votes
- Configuration
- gpt-4-0125-preview
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,280
Expert prompts · rank 244 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,268 to 1,292
- Sample
- 4454 votes
- Configuration
- gpt-4-0125-preview
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,290
Instruction following · rank 249 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,285 to 1,296
- Sample
- 33252 votes
- Configuration
- gpt-4-0125-preview
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,313
Overall · rank 247 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,309 to 1,317
- Sample
- 93439 votes
- Configuration
- gpt-4-0125-preview
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,306
Writing, literature and language · rank 216 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,299 to 1,312
- Sample
- 23492 votes
- Configuration
- gpt-4-0125-preview
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.