Models / Gemini 2.5 Flash

Google

Gemini 2.5 Flash

25 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Google
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

1,437
Overall · rank 80 of 177
Unit
Arena rating, higher is better
Range
1,434 to 1,439
Sample
70399 votes
Configuration
Gemini 2.5 Flash
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,399
Business, management and finance · rank 130 of 402
Unit
Arena rating, higher is better
Range
1,395 to 1,404
Sample
22722 votes
Configuration
Gemini 2.5 Flash
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,394
Creative writing · rank 93 of 407
Unit
Arena rating, higher is better
Range
1,389 to 1,399
Sample
17561 votes
Configuration
Gemini 2.5 Flash
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,424
Expert prompts · rank 115 of 359
Unit
Arena rating, higher is better
Range
1,417 to 1,431
Sample
8345 votes
Configuration
Gemini 2.5 Flash
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,400
Instruction following · rank 121 of 409
Unit
Arena rating, higher is better
Range
1,396 to 1,403
Sample
34707 votes
Configuration
Gemini 2.5 Flash
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,409
Overall · rank 128 of 409
Unit
Arena rating, higher is better
Range
1,407 to 1,412
Sample
125044 votes
Configuration
Gemini 2.5 Flash
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,402
Writing, literature and language · rank 94 of 408
Unit
Arena rating, higher is better
Range
1,398 to 1,406
Sample
28309 votes
Configuration
Gemini 2.5 Flash
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
93.7%
Irrelevance detection · rank 6 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
91.1%
Irrelevance detection · rank 11 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
41.3%
Memory · rank 17 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
38.7%
Memory · rank 18 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
36.2%
Multi-turn tasks · rank 32 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
16.8%
Multi-turn tasks · rank 53 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
56.2%
Overall accuracy · rank 15 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
50.9%
Overall accuracy · rank 26 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
75.0%
Relevance detection · rank 62 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
62.5%
Relevance detection · rank 91 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
85.0%
Single-turn calls (curated) · rank 44 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
88.1%
Single-turn calls (curated) · rank 21 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
74.4%
Single-turn calls (user-contributed) · rank 45 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
78.2%
Single-turn calls (user-contributed) · rank 27 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
59.0%
Web search · rank 21 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
62.0%
Web search · rank 20 of 109
Unit
% correct, higher is better
Configuration
gemini-2.5-flash
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
99.0%
Answer rate · rank 68 of 108
Unit
% of documents, higher is better
Configuration
Gemini 2.5 Flash
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
7.8%
Hallucination rate · rank 39 of 108
Unit
% of summaries, lower is better
Configuration
Gemini 2.5 Flash
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare Gemini 2.5 Flash with