Models / o3

OpenAI

o3

33 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
OpenAI
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

1,192
Overall · rank 8 of 34
Unit
Arena rating, higher is better
Range
1,185 to 1,199
Sample
20644 votes
Configuration
o3
Measured
24 Aug 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,405
Overall · rank 133 of 177
Unit
Arena rating, higher is better
Range
1,398 to 1,411
Sample
7512 votes
Configuration
o3
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,426
Business, management and finance · rank 89 of 402
Unit
Arena rating, higher is better
Range
1,420 to 1,433
Sample
9464 votes
Configuration
o3
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,383
Creative writing · rank 101 of 407
Unit
Arena rating, higher is better
Range
1,376 to 1,390
Sample
7564 votes
Configuration
o3
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,449
Expert prompts · rank 83 of 359
Unit
Arena rating, higher is better
Range
1,437 to 1,460
Sample
2983 votes
Configuration
o3
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,403
Instruction following · rank 115 of 409
Unit
Arena rating, higher is better
Range
1,397 to 1,408
Sample
15395 votes
Configuration
o3
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,432
Overall · rank 91 of 409
Unit
Arena rating, higher is better
Range
1,428 to 1,435
Sample
58587 votes
Configuration
o3
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,396
Writing, literature and language · rank 99 of 408
Unit
Arena rating, higher is better
Range
1,390 to 1,402
Sample
13145 votes
Configuration
o3
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
47.3%
Memory · rank 12 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
51.8%
Memory · rank 10 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
14.8%
Multi-turn tasks · rank 58 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
62.2%
Multi-turn tasks · rank 8 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
48.6%
Overall accuracy · rank 30 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
63.0%
Overall accuracy · rank 8 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
81.2%
Relevance detection · rank 41 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
93.8%
Relevance detection · rank 8 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
40.4%
Single-turn calls (curated) · rank 100 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
81.9%
Single-turn calls (curated) · rank 66 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
66.2%
Single-turn calls (user-contributed) · rank 76 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
73.2%
Single-turn calls (user-contributed) · rank 50 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
77.0%
Web search · rank 9 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
50.5%
Web search · rank 26 of 109
Unit
% correct, higher is better
Configuration
o3
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
Reported by UGI Leaderboard
18.0%
Requested-length error · rank 150 of 370
Unit
% off the requested word count, lower is better
Configuration
o3
Measured
24 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
29.0%
Requested-length error · rank 248 of 370
Unit
% off the requested word count, lower is better
Configuration
o3
Measured
24 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
18.0%
Requested-length error · rank 150 of 370
Unit
% off the requested word count, lower is better
Configuration
o3
Measured
24 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.30
Style adherence · rank 322 of 370
Unit
score from 0 to 1, higher is better
Configuration
o3
Measured
24 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.28
Style adherence · rank 355 of 370
Unit
score from 0 to 1, higher is better
Configuration
o3
Measured
24 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.29
Style adherence · rank 346 of 370
Unit
score from 0 to 1, higher is better
Configuration
o3
Measured
24 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
66.8
Writing score · rank 52 of 370
Unit
score out of 100, higher is better
Configuration
o3
Measured
24 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
65.8
Writing score · rank 57 of 370
Unit
score out of 100, higher is better
Configuration
o3
Measured
24 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
66.8
Writing score · rank 50 of 370
Unit
score out of 100, higher is better
Configuration
o3
Measured
24 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare o3 with