Benchmarks / SimpleQA Verified (Epoch AI)

Reported by SimpleQA Verified (Epoch AI)

SimpleQA Verified (Epoch AI)

Share of short factual questions answered correctly without search, as run by Epoch AI.

Results dated
10 Aug 2026 to 29 Sep 2026
Models
85
Unit
% of questions
Licence
Creative Commons Attribution 4.0 International

Full results

SimpleQA Verified: correct answers, % of questions, higher is better
#ModelCorrect answers
% of questions, higher is better
1 GPT-6 Astra (max)OpenAI
75.6%
2 GPT-6.1 Sol (max)OpenAI
73.9%
3 Gemini 3.1 Pro Preview (high reasoning)Google
73.5%
4 Claude Opus 5.5 (max reasoning)Anthropic
72.2%
5 Claude Fable 5.1 (max)Anthropic
70.8%
6 Claude Fable 5 (xhigh reasoning)Anthropic
70.7%
7 Gemini 3.8 Flash (high)Google
69.7%
7 GPT-5.6 Sol (max)OpenAI
69.7%
9 Gemini 3.7 Flash (high)Google
69.2%
10 Gemini 3 Flash Preview (high reasoning)Google
66.8%
11 Gemini 3.5 Flash (high)Google
66.2%
11 Gemini 3.6 Flash (high)Google
66.2%
13 GPT-5.5 (xhigh reasoning)OpenAI
63.0%
14 GPT-6 Sol (max)OpenAI
60.7%
15 Muse Spark 1.2Meta
60.3%
16 Claude Opus 5 (max)Anthropic
59.9%
17 Muse Spark 1.1Meta
57.8%
18 Qwen3.7 MaxAlibaba
55.8%
19 Claude Opus 4.8 (max reasoning)Anthropic
53.0%
20 DeepSeek V4 Pro 0813 (max)DeepSeek
52.9%
21 Qwen3.6 Max PreviewAlibaba
52.0%
22 Claude Opus 4.7 (xhigh)Anthropic
51.7%
23 Kimi K3 (max)Moonshot AI
50.6%
24 GPT-5 (high)OpenAI
50.1%
25 o3 (high reasoning)OpenAI
49.4%
26 Grok 4.6 (high)xAI
49.3%
27 Grok 4.6 (xhigh)xAI
48.9%
28 Qwen3 MaxAlibaba
48.7%
29 Grok 4.5 (high)xAI
48.3%
30 GPT-5.1 (high)OpenAI
48.0%
31 Qwen3.8 Max (0902) (xhigh)Alibaba
47.3%
32 Claude Opus 4.6 (max reasoning)Anthropic
47.0%
33 DeepSeek V4 Pro 0423 (max)DeepSeek
47.0%
34 Claude Sonnet 5.5 (max)Anthropic
46.5%
35 GPT-5.4 Pro (xhigh)OpenAI
46.3%
36 qwen3.8-max (xhigh reasoning)Alibaba
45.8%
37 Claude Opus 4.5 (thinking-32k reasoning)Anthropic
45.7%
38 GPT-5.4 (xhigh reasoning)OpenAI
45.1%
39 Qwen3.6 PlusAlibaba
44.1%
40 GPT-5.6 Terra (max)OpenAI
43.2%
41 GPT-6 Luna (max)OpenAI
41.4%
42 o1 (high reasoning)OpenAI
41.1%
43 GPT-5.6 Luna (max)OpenAI
41.0%
43 GLM 5.3 (max)Z.ai
41.0%
45 Qwen3 235B A22B Thinking 2507Alibaba
40.4%
46 Inkling (xhigh)Thinkingmachines
40.3%
47 GPT-5.2 (xhigh)OpenAI
37.1%
48 Kimi K2.7 CodeMoonshot AI
36.5%
49 Claude Sonnet 4.6 (high reasoning)Anthropic
35.5%
50 Kimi K2.6Moonshot AI
34.9%
51 Kimi K2.5Moonshot AI
34.3%
51 GPT-5.2 (high)OpenAI
34.3%
53 GLM 5.2 (max)Z.ai
34.2%
54 GLM 5.1Z.ai
34.0%
55 Claude Sonnet 5 (max reasoning)Anthropic
33.7%
56 DeepSeek V4 Flash 0731 (max)DeepSeek
33.6%
57 Grok 4.3 (high)xAI
33.2%
58 Claude Sonnet 5 (xhigh reasoning)Anthropic
32.9%
59 Claude Sonnet 4.6 (max reasoning)Anthropic
32.8%
59 GPT-5.2 (low reasoning)OpenAI
32.8%
61 GPT-5.2 (medium reasoning)OpenAI
32.7%
62 GLM 4.7Z.ai
32.2%
63 GPT-4.1OpenAI
31.1%
64 Claude Sonnet 4.5 (thinking-59k reasoning)Anthropic
30.7%
65 grok-4.20-0309-reasoningxAI
30.2%
66 GPT-5.4 Mini (high)OpenAI
29.4%
67 GPT-4o (2024-08-06)OpenAI
26.0%
68 Qwen3.5 Plus 2026-02-15Alibaba
25.4%
69 Claude Sonnet 4.5Anthropic
23.7%
70 GPT-5 Mini (high)OpenAI
21.6%
71 Qwen3.5-FlashAlibaba
20.3%
72 o4 Mini (low reasoning)OpenAI
19.6%
73 Inkling Small (xhigh)Thinkingmachines
19.1%
74 o4 Mini (high reasoning)OpenAI
18.8%
75 Qwen3.6 FlashAlibaba
15.9%
76 o3 Mini (high)OpenAI
15.3%
77 Claude Haiku 4.5Anthropic
13.2%
78 GPT-4.1 MiniOpenAI
12.7%
79 claude-3-opus-20240229Anthropic
12.6%
79 Claude Haiku 4.5 (thinking-32k reasoning)Anthropic
12.6%
81 GPT-5.4 Nano (high)OpenAI
11.7%
81 GPT-5 Nano (high)OpenAI
11.7%
83 Gemma 4 31BGoogle
10.4%
84 GPT-4o-mini (2024-07-18)OpenAI
8.3%
85 GPT-4.1 NanoOpenAI
6.0%

Results as published by SimpleQA Verified (Epoch AI); we do not re-run them.

What it measures

Share of short factual questions answered correctly without search, as run by Epoch AI.

What it does not measure

Not answers grounded in your documents; tests what the model remembers.

Epoch AI, Capabilities & benchmarking (CC BY 4.0). Licence: Creative Commons Attribution 4.0 International.