Benchmarks / SimpleQA Verified (Epoch AI)
Reported by SimpleQA Verified (Epoch AI)
SimpleQA Verified (Epoch AI)
Share of short factual questions answered correctly without search, as run by Epoch AI.
- Results dated
- 10 Aug 2026 to 29 Sep 2026
- Models
- 85
- Unit
- % of questions
Full results
| # | Model | Correct answers % of questions, higher is better |
|---|---|---|
| 1 | GPT-6 Astra (max)OpenAI |
75.6%
|
| 2 | GPT-6.1 Sol (max)OpenAI |
73.9%
|
| 3 | Gemini 3.1 Pro Preview (high reasoning)Google |
73.5%
|
| 4 | Claude Opus 5.5 (max reasoning)Anthropic |
72.2%
|
| 5 | Claude Fable 5.1 (max)Anthropic |
70.8%
|
| 6 | Claude Fable 5 (xhigh reasoning)Anthropic |
70.7%
|
| 7 | Gemini 3.8 Flash (high)Google |
69.7%
|
| 7 | GPT-5.6 Sol (max)OpenAI |
69.7%
|
| 9 | Gemini 3.7 Flash (high)Google |
69.2%
|
| 10 | Gemini 3 Flash Preview (high reasoning)Google |
66.8%
|
| 11 | Gemini 3.5 Flash (high)Google |
66.2%
|
| 11 | Gemini 3.6 Flash (high)Google |
66.2%
|
| 13 | GPT-5.5 (xhigh reasoning)OpenAI |
63.0%
|
| 14 | GPT-6 Sol (max)OpenAI |
60.7%
|
| 15 | Muse Spark 1.2Meta |
60.3%
|
| 16 | Claude Opus 5 (max)Anthropic |
59.9%
|
| 17 | Muse Spark 1.1Meta |
57.8%
|
| 18 | Qwen3.7 MaxAlibaba |
55.8%
|
| 19 | Claude Opus 4.8 (max reasoning)Anthropic |
53.0%
|
| 20 | DeepSeek V4 Pro 0813 (max)DeepSeek |
52.9%
|
| 21 | Qwen3.6 Max PreviewAlibaba |
52.0%
|
| 22 | Claude Opus 4.7 (xhigh)Anthropic |
51.7%
|
| 23 | Kimi K3 (max)Moonshot AI |
50.6%
|
| 24 | GPT-5 (high)OpenAI |
50.1%
|
| 25 | o3 (high reasoning)OpenAI |
49.4%
|
| 26 | Grok 4.6 (high)xAI |
49.3%
|
| 27 | Grok 4.6 (xhigh)xAI |
48.9%
|
| 28 | Qwen3 MaxAlibaba |
48.7%
|
| 29 | Grok 4.5 (high)xAI |
48.3%
|
| 30 | GPT-5.1 (high)OpenAI |
48.0%
|
| 31 | Qwen3.8 Max (0902) (xhigh)Alibaba |
47.3%
|
| 32 | Claude Opus 4.6 (max reasoning)Anthropic |
47.0%
|
| 33 | DeepSeek V4 Pro 0423 (max)DeepSeek |
47.0%
|
| 34 | Claude Sonnet 5.5 (max)Anthropic |
46.5%
|
| 35 | GPT-5.4 Pro (xhigh)OpenAI |
46.3%
|
| 36 | qwen3.8-max (xhigh reasoning)Alibaba |
45.8%
|
| 37 | Claude Opus 4.5 (thinking-32k reasoning)Anthropic |
45.7%
|
| 38 | GPT-5.4 (xhigh reasoning)OpenAI |
45.1%
|
| 39 | Qwen3.6 PlusAlibaba |
44.1%
|
| 40 | GPT-5.6 Terra (max)OpenAI |
43.2%
|
| 41 | GPT-6 Luna (max)OpenAI |
41.4%
|
| 42 | o1 (high reasoning)OpenAI |
41.1%
|
| 43 | GPT-5.6 Luna (max)OpenAI |
41.0%
|
| 43 | GLM 5.3 (max)Z.ai |
41.0%
|
| 45 | Qwen3 235B A22B Thinking 2507Alibaba |
40.4%
|
| 46 | Inkling (xhigh)Thinkingmachines |
40.3%
|
| 47 | GPT-5.2 (xhigh)OpenAI |
37.1%
|
| 48 | Kimi K2.7 CodeMoonshot AI |
36.5%
|
| 49 | Claude Sonnet 4.6 (high reasoning)Anthropic |
35.5%
|
| 50 | Kimi K2.6Moonshot AI |
34.9%
|
| 51 | Kimi K2.5Moonshot AI |
34.3%
|
| 51 | GPT-5.2 (high)OpenAI |
34.3%
|
| 53 | GLM 5.2 (max)Z.ai |
34.2%
|
| 54 | GLM 5.1Z.ai |
34.0%
|
| 55 | Claude Sonnet 5 (max reasoning)Anthropic |
33.7%
|
| 56 | DeepSeek V4 Flash 0731 (max)DeepSeek |
33.6%
|
| 57 | Grok 4.3 (high)xAI |
33.2%
|
| 58 | Claude Sonnet 5 (xhigh reasoning)Anthropic |
32.9%
|
| 59 | Claude Sonnet 4.6 (max reasoning)Anthropic |
32.8%
|
| 59 | GPT-5.2 (low reasoning)OpenAI |
32.8%
|
| 61 | GPT-5.2 (medium reasoning)OpenAI |
32.7%
|
| 62 | GLM 4.7Z.ai |
32.2%
|
| 63 | GPT-4.1OpenAI |
31.1%
|
| 64 | Claude Sonnet 4.5 (thinking-59k reasoning)Anthropic |
30.7%
|
| 65 | grok-4.20-0309-reasoningxAI |
30.2%
|
| 66 | GPT-5.4 Mini (high)OpenAI |
29.4%
|
| 67 | GPT-4o (2024-08-06)OpenAI |
26.0%
|
| 68 | Qwen3.5 Plus 2026-02-15Alibaba |
25.4%
|
| 69 | Claude Sonnet 4.5Anthropic |
23.7%
|
| 70 | GPT-5 Mini (high)OpenAI |
21.6%
|
| 71 | Qwen3.5-FlashAlibaba |
20.3%
|
| 72 | o4 Mini (low reasoning)OpenAI |
19.6%
|
| 73 | Inkling Small (xhigh)Thinkingmachines |
19.1%
|
| 74 | o4 Mini (high reasoning)OpenAI |
18.8%
|
| 75 | Qwen3.6 FlashAlibaba |
15.9%
|
| 76 | o3 Mini (high)OpenAI |
15.3%
|
| 77 | Claude Haiku 4.5Anthropic |
13.2%
|
| 78 | GPT-4.1 MiniOpenAI |
12.7%
|
| 79 | claude-3-opus-20240229Anthropic |
12.6%
|
| 79 | Claude Haiku 4.5 (thinking-32k reasoning)Anthropic |
12.6%
|
| 81 | GPT-5.4 Nano (high)OpenAI |
11.7%
|
| 81 | GPT-5 Nano (high)OpenAI |
11.7%
|
| 83 | Gemma 4 31BGoogle |
10.4%
|
| 84 | GPT-4o-mini (2024-07-18)OpenAI |
8.3%
|
| 85 | GPT-4.1 NanoOpenAI |
6.0%
|
Results as published by SimpleQA Verified (Epoch AI); we do not re-run them.
What it measures
Share of short factual questions answered correctly without search, as run by Epoch AI.
What it does not measure
Not answers grounded in your documents; tests what the model remembers.
Epoch AI, Capabilities & benchmarking (CC BY 4.0). Licence: Creative Commons Attribution 4.0 International.