Benchmarks / Vectara Hallucination Leaderboard

Reported by Vectara Hallucination Leaderboard

Vectara Hallucination Leaderboard

Share of document summaries that contain something the document does not support.

Results dated
22 Sep 2026
Models
108
Unit
% of summaries
Licence
Apache License 2.0

Full results

Vectara: hallucination rate, % of summaries, lower is better
#ModelHallucination rate
% of summaries, lower is better
1 finix_s1_32bAnt Group
1.8%
2 GPT-5.4 NanoOpenAI
3.1%
3 Gemini 2.5 Flash LiteGoogle
3.3%
4 Phi 4Microsoft
3.7%
5 llama-3.3-70b-instruct-turboMeta
4.1%
6 snowflake-arctic-instructSnowflake
4.3%
7 Gemma 3 12BGoogle
4.4%
8 mistral-large-2411Mistral
4.5%
9 Qwen3 8BAlibaba
4.8%
10 Nova 2 LiteAmazon
5.1%
10 Nova Pro 1.0Amazon
5.1%
10 Mistral Small 3Mistral
5.1%
13 Gemma 4 26B A4BGoogle
5.2%
13 granite-4.0-h-smallIBM
5.2%
15 jamba-mini-2AI21 Labs
5.3%
15 DeepSeek V3.2 ExpDeepSeek
5.3%
17 Qwen3 14BAlibaba
5.4%
18 Nova Micro 1.0Amazon
5.5%
18 DeepSeek V3.1DeepSeek
5.5%
18 GPT-5.4 MiniOpenAI
5.5%
21 GPT-4.1OpenAI
5.6%
22 qwen3-4bAlibaba
5.7%
23 grok-3xAI
5.8%
24 Qwen3 32BAlibaba
5.9%
25 Nova Lite 1.0Amazon
6.1%
25 DeepSeek V3DeepSeek
6.1%
27 DeepSeek V3.2DeepSeek
6.3%
28 Gemma 3 4BGoogle
6.4%
29 GPT-6 SolOpenAI
6.5%
30 trinity-large-previewArcee Ai
6.9%
30 Command R+ (08-2024)Cohere
6.9%
32 Gemini 2.5 ProGoogle
7.0%
32 GPT-5.4OpenAI
7.0%
34 ministral-3b-2410Mistral
7.3%
35 Gemma 3 27BGoogle
7.4%
35 Gemma 4 31BGoogle
7.4%
35 ministral-8b-2410Mistral
7.4%
38 Llama 4 ScoutMeta
7.7%
39 Gemini 2.5 FlashGoogle
7.8%
40 Gemini 3.1 Flash Lite PreviewGoogle
8.2%
40 llama-4-maverick-17b-128e-instruct-fp8Meta
8.2%
42 GPT-5.4 ProOpenAI
8.3%
43 GPT-5.2 (low reasoning)OpenAI
8.4%
44 DeepSeek V4 Pro 0423DeepSeek
8.6%
45 GPT-6 AstraOpenAI
8.7%
46 MiniMax M2.5MiniMax
9.1%
47 Qwen3 235B A22BAlibaba
9.3%
47 Qwen3 Next 80B A3B ThinkingAlibaba
9.3%
47 Command ACohere
9.3%
47 GPT-5.5OpenAI
9.3%
47 glm-4.5-air-fp8Z.ai
9.3%
47 GLM 4.7 FlashZ.ai
9.3%
53 c4ai-aya-expanse-8bCohere
9.5%
53 GLM 4.6Z.ai
9.5%
55 Nemotron 3 Nano 30B A3BNVIDIA
9.6%
55 GPT-4o (2024-08-06)OpenAI
9.6%
57 ai21-jamba-large-1.7AI21 Labs
9.7%
58 Claude Haiku 4.5Anthropic
9.8%
59 GLM 5Z.ai
10.1%
60 Claude Sonnet 4Anthropic
10.3%
61 Gemini 3.1 Pro PreviewGoogle
10.4%
62 Qwen3.5-35B-A3BAlibaba
10.5%
62 Qwen3.5-FlashAlibaba
10.5%
62 GPT-5 NanoOpenAI
10.5%
65 Claude Sonnet 4.6Anthropic
10.6%
65 granite-3.3-8b-instructIBM
10.6%
67 Qwen3.5 Plus 2026-02-15Alibaba
10.7%
68 Kimi K2.6Moonshot AI
10.8%
68 GPT-5.2 (high)OpenAI
10.8%
70 Claude Opus 4.5Anthropic
10.9%
70 c4ai-aya-expanse-32bCohere
10.9%
70 GPT-5.1 (low reasoning)OpenAI
10.9%
73 Qwen3.5-122B-A10BAlibaba
11.2%
74 R1DeepSeek
11.3%
75 GLM 4.7Z.ai
11.7%
76 Claude Opus 4.1Anthropic
11.8%
76 MiniMax M2.1MiniMax
11.8%
78 Claude Opus 4.7Anthropic
12.0%
78 claude-opus-4Anthropic
12.0%
78 Claude Sonnet 4.5Anthropic
12.0%
81 Qwen3.5-27BAlibaba
12.1%
81 GPT-5.1 (high)OpenAI
12.1%
83 Claude Opus 4.6Anthropic
12.2%
84 Mercury 2Inception
12.3%
85 GPT-5.6 SolOpenAI
12.4%
86 MiniMax M2.7MiniMax
12.9%
86 GPT-5 MiniOpenAI
12.9%
88 Gemini 3 Flash PreviewGoogle
13.5%
89 gemini-3-pro-previewGoogle
13.6%
90 Kimi K2.5Moonshot AI
14.2%
90 gpt-oss-120bOpenAI
14.2%
92 Mistral Large 3 2512Mistral
14.5%
93 ai21-jamba-mini-1.7AI21 Labs
14.7%
93 GPT-5 (minimal reasoning)OpenAI
14.7%
95 GPT-5 (high)OpenAI
15.1%
96 grok-4-1-fast-non-reasoningxAI
17.8%
97 Kimi K2 0905Moonshot AI
17.9%
98 o4 Mini HighOpenAI
18.6%
98 o4 Mini (low reasoning)OpenAI
18.6%
100 grok-4-1-fast-reasoningxAI
19.2%
101 Ministral 3 14B 2512Mistral
19.4%
102 grok-4-fast-non-reasoningxAI
19.7%
103 grok-4-fast-reasoningxAI
20.2%
104 Ministral 3 8B 2512Mistral
21.7%
105 Mistral Medium 3.1Mistral
22.7%
106 o3 ProOpenAI
23.3%
107 phi-4-mini-instructMicrosoft
23.5%
108 Ministral 3 3B 2512Mistral
24.2%

Results as published by Vectara Hallucination Leaderboard; we do not re-run them.

What it measures

Share of document summaries that contain something the document does not support.

What it does not measure

Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Vectara Hallucination Leaderboard by Vectara, Apache License 2.0.