1,368
Overall · rank 157 of 177
- Unit
- Arena rating, higher is better
- Range
- 1,359 to 1,378
- Sample
- 3199 votes
- Configuration
- Mercury 2
- Measured
- 25 Sep 2026
- Not shown
- Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,340
Business, management and finance · rank 182 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,316 to 1,364
- Sample
- 512 votes
- Configuration
- Mercury 2
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,292
Creative writing · rank 193 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,266 to 1,318
- Sample
- 548 votes
- Configuration
- Mercury 2
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,348
Expert prompts · rank 170 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,312 to 1,384
- Sample
- 253 votes
- Configuration
- Mercury 2
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,327
Instruction following · rank 194 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,307 to 1,346
- Sample
- 888 votes
- Configuration
- Mercury 2
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,343
Overall · rank 201 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,333 to 1,354
- Sample
- 3200 votes
- Configuration
- Mercury 2
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,307
Writing, literature and language · rank 204 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,285 to 1,329
- Sample
- 742 votes
- Configuration
- Mercury 2
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
100.0%
Answer rate · rank 1 of 108
- Unit
- % of documents, higher is better
- Configuration
- Mercury 2
- Measured
- 22 Sep 2026
- Not shown
- Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
12.3%
Hallucination rate · rank 84 of 108
- Unit
- % of summaries, lower is better
- Configuration
- Mercury 2
- Measured
- 22 Sep 2026
- Not shown
- Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.