Benchmarks / Berkeley Function Calling Leaderboard (BFCL) V4
Reported by Berkeley Function Calling Leaderboard (BFCL) V4
Berkeley Function Calling Leaderboard (BFCL) V4
Declining to call a function when none of those offered fits.
- Results dated
- 16 Dec 2025
- Models
- 109
- Unit
- % correct
- Licence
- Apache License 2.0
Full results
| # | Model | BFCL: irrelevance detection % correct, higher is better |
|---|---|---|
| 1 | ministral-8b-2410 (bfcl fc)Mistral |
100.0%
|
| 1 | llama-3.1-nemotron-ultra-253b-v1 (bfcl fc)NVIDIA |
100.0%
|
| 3 | bitagent-bounty-8b (bfcl unspecified)Bittensor |
97.5%
|
| 4 | Claude Haiku 4.5 (bfcl prompt)Anthropic |
95.3%
|
| 5 | Claude Sonnet 4.5 (bfcl prompt)Anthropic |
95.0%
|
| 6 | Gemini 2.5 Flash (bfcl fc)Google |
93.7%
|
| 7 | Gemini 2.5 Flash Lite (bfcl prompt)Google |
93.3%
|
| 8 | DeepSeek V3.2 Exp (bfcl fc)DeepSeek |
93.2%
|
| 9 | Gemini 2.5 Flash Lite (bfcl fc)Google |
92.5%
|
| 10 | Mistral Medium 3 (bfcl fc)Mistral |
92.0%
|
| 11 | Gemini 2.5 Flash (bfcl prompt)Google |
91.1%
|
| 12 | GPT-5 Mini (bfcl fc)OpenAI |
91.0%
|
| 13 | toolace-2-8b (bfcl fc)Huawei Noah And Ustc |
90.8%
|
| 14 | Claude Opus 4.5 (bfcl prompt)Anthropic |
90.8%
|
| 15 | hammer2.1-7b (bfcl fc)Madeagents |
90.1%
|
| 16 | GPT-5 Nano (bfcl fc)OpenAI |
89.1%
|
| 17 | Mistral Small 3.2 24B (bfcl fc)Mistral |
87.9%
|
| 18 | Phi 4 (bfcl prompt)Microsoft |
87.5%
|
| 19 | Kimi K2 0711 (bfcl fc)Moonshot AI |
87.3%
|
| 20 | falcon3-1b-instruct (bfcl fc)TII |
87.3%
|
| 21 | GPT-5.2 (bfcl prompt)OpenAI |
87.3%
|
| 22 | Qwen3 14B (bfcl prompt)Alibaba |
87.2%
|
| 23 | o4 Mini (bfcl prompt)OpenAI |
87.2%
|
| 24 | coalm-8b (bfcl unspecified)Uiuc Oumi |
86.9%
|
| 25 | command-a-reasoning (bfcl fc)Cohere |
86.8%
|
| 26 | Claude Sonnet 4.5 (bfcl fc)Anthropic |
86.6%
|
| 27 | GPT-4.1 (bfcl fc)OpenAI |
86.5%
|
| 28 | o3 (bfcl fc)OpenAI |
86.1%
|
| 29 | hammer2.1-3b (bfcl fc)Madeagents |
86.1%
|
| 30 | coalm-70b (bfcl unspecified)Uiuc Oumi |
85.7%
|
| 31 | gemini-3-pro-preview (bfcl prompt)Google |
85.6%
|
| 32 | Claude Haiku 4.5 (bfcl fc)Anthropic |
85.1%
|
| 33 | GLM 4.6 (thinking reasoning, bfcl fc)Z.ai |
85.0%
|
| 34 | qwen3-4b-instruct-2507 (bfcl fc)Alibaba |
84.9%
|
| 35 | Claude Opus 4.5 (bfcl fc)Anthropic |
84.7%
|
| 36 | grok-4-0709 (bfcl prompt)xAI |
84.3%
|
| 37 | Command A (bfcl fc)Cohere |
84.2%
|
| 38 | GPT-4.1 (bfcl prompt)OpenAI |
84.0%
|
| 39 | o3 (bfcl prompt)OpenAI |
84.0%
|
| 40 | o4 Mini (bfcl fc)OpenAI |
83.9%
|
| 41 | GPT-4.1 Nano (bfcl prompt)OpenAI |
83.4%
|
| 42 | nanbeige4-3b-thinking-2511 (bfcl fc)Nanbeige |
83.1%
|
| 43 | qwen3-0.6b (bfcl prompt)Alibaba |
82.5%
|
| 44 | rzn-t (bfcl prompt)Phronetic Ai |
82.4%
|
| 45 | Qwen3 32B (bfcl prompt)Alibaba |
82.4%
|
| 46 | Qwen3 8B (bfcl prompt)Alibaba |
82.3%
|
| 47 | arch-agent-32b (bfcl unspecified)Katanemo |
82.2%
|
| 48 | Nova 2 Lite (bfcl fc)Amazon |
82.1%
|
| 49 | Qwen3 14B (bfcl fc)Alibaba |
81.9%
|
| 50 | Qwen3 235B A22B Instruct 2507 (bfcl fc)Alibaba |
81.7%
|
| 51 | GPT-4.1 Mini (bfcl fc)OpenAI |
81.7%
|
| 52 | Command R7B (12-2024) (bfcl fc)Cohere |
81.7%
|
| 53 | palmyra-x-004 (bfcl fc)Writer |
81.0%
|
| 54 | qwen3-0.6b (bfcl fc)Alibaba |
80.8%
|
| 55 | hammer2.1-0.5b (bfcl fc)Madeagents |
80.8%
|
| 56 | granite-3.2-8b-instruct (bfcl fc)IBM |
80.5%
|
| 57 | xlam-2-32b-fc-r (bfcl fc)Salesforce |
80.2%
|
| 58 | granite-3.1-8b-instruct (bfcl fc)IBM |
80.0%
|
| 59 | Qwen3 30B A3B Instruct 2507 (bfcl fc)Alibaba |
79.9%
|
| 60 | grok-4-1-fast-reasoning (bfcl fc)xAI |
79.4%
|
| 61 | GPT-5.2 (bfcl fc)OpenAI |
79.4%
|
| 62 | hammer2.1-1.5b (bfcl fc)Madeagents |
79.4%
|
| 63 | xlam-2-70b-fc-r (bfcl fc)Salesforce |
79.1%
|
| 64 | Qwen3 8B (bfcl fc)Alibaba |
79.1%
|
| 65 | Qwen3 235B A22B Instruct 2507 (bfcl prompt)Alibaba |
78.9%
|
| 66 | gemini-3-pro-preview (bfcl fc)Google |
77.8%
|
| 67 | qwen3-1.7b (bfcl fc)Alibaba |
76.5%
|
| 68 | Qwen3 32B (bfcl fc)Alibaba |
76.4%
|
| 69 | qwen3-4b-instruct-2507 (bfcl prompt)Alibaba |
75.9%
|
| 70 | grok-4-0709 (bfcl fc)xAI |
75.4%
|
| 71 | granite-20b-functioncalling (bfcl fc)IBM |
75.1%
|
| 72 | Qwen3 30B A3B Instruct 2507 (bfcl prompt)Alibaba |
74.8%
|
| 73 | arch-agent-1.5b (bfcl unspecified)Katanemo |
74.8%
|
| 74 | arch-agent-3b (bfcl unspecified)Katanemo |
74.7%
|
| 75 | Mistral Medium 3 (bfcl unspecified)Mistral |
74.5%
|
| 76 | nanbeige3.5-pro-thinking (bfcl fc)Nanbeige |
74.2%
|
| 77 | grok-4-1-fast-non-reasoning (bfcl fc)xAI |
74.1%
|
| 78 | GPT-4.1 Mini (bfcl prompt)OpenAI |
73.9%
|
| 79 | minicpm3-4b (bfcl prompt)Openbmb |
73.7%
|
| 80 | Gemma 3 27B (bfcl prompt)Google |
73.7%
|
| 81 | minicpm3-4b (bfcl fc)Openbmb |
72.8%
|
| 82 | Nova Micro 1.0 (bfcl fc)Amazon |
70.7%
|
| 83 | Gemma 3 12B (bfcl prompt)Google |
70.3%
|
| 84 | Nova Pro 1.0 (bfcl fc)Amazon |
70.1%
|
| 85 | mistral-large-2411 (bfcl fc)Mistral |
68.9%
|
| 86 | DeepSeek V3.2 Exp (thinking reasoning, bfcl prompt)DeepSeek |
67.0%
|
| 87 | GPT-4.1 Nano (bfcl fc)OpenAI |
66.0%
|
| 88 | Mistral Small 3.2 24B (bfcl prompt)Mistral |
65.7%
|
| 89 | xlam-2-1b-fc-r (bfcl fc)Salesforce |
64.5%
|
| 90 | xlam-2-3b-fc-r (bfcl fc)Salesforce |
63.5%
|
| 91 | xlam-2-8b-fc-r (bfcl fc)Salesforce |
63.3%
|
| 92 | Mistral Nemo (bfcl fc)Mistral |
61.8%
|
| 93 | granite-4.0-350m (bfcl fc)IBM |
60.8%
|
| 94 | llama-4-maverick-17b-128e-instruct-fp8 (bfcl fc)Meta |
56.0%
|
| 95 | GPT-5 Mini (bfcl prompt)OpenAI |
55.7%
|
| 96 | Gemma 3 4B (bfcl prompt)Google |
53.9%
|
| 97 | Llama 3.3 70B Instruct (bfcl fc)Meta |
53.5%
|
| 98 | Llama 3.2 3B Instruct (bfcl fc)Meta |
52.1%
|
| 99 | Llama 3.2 1B Instruct (bfcl fc)Meta |
51.6%
|
| 100 | GPT-5 Nano (bfcl prompt)OpenAI |
45.8%
|
| 101 | Llama 4 Scout (bfcl fc)Meta |
44.9%
|
| 102 | Llama 3.1 8B Instruct (bfcl prompt)Meta |
42.7%
|
| 103 | mistral-large-2411 (bfcl prompt)Mistral |
38.8%
|
| 104 | bielik-11b-v2.3-instruct (bfcl prompt)Speakleash And Ack Cyfronet Agh |
36.0%
|
| 105 | gemma-3-1b-it (bfcl prompt)Google |
33.2%
|
| 106 | falcon3-3b-instruct (bfcl fc)TII |
32.9%
|
| 107 | falcon3-10b-instruct (bfcl fc)TII |
32.1%
|
| 108 | falcon3-7b-instruct (bfcl fc)TII |
32.0%
|
| 109 | Mistral Nemo (bfcl prompt)Mistral |
6.3%
|
Results as published by Berkeley Function Calling Leaderboard (BFCL) V4; we do not re-run them.
What it measures
Declining to call a function when none of those offered fits.
What it does not measure
Not general refusal or safety behaviour.
Berkeley Function Calling Leaderboard (BFCL) V4 by the UC Berkeley Gorilla team, Apache License 2.0.