Benchmarks / Berkeley Function Calling Leaderboard (BFCL) V4

Reported by Berkeley Function Calling Leaderboard (BFCL) V4

Berkeley Function Calling Leaderboard (BFCL) V4

Declining to call a function when none of those offered fits.

Results dated
16 Dec 2025
Models
109
Unit
% correct
Licence
Apache License 2.0

Full results

BFCL: irrelevance detection, % correct, higher is better
#ModelBFCL: irrelevance detection
% correct, higher is better
1 ministral-8b-2410 (bfcl fc)Mistral
100.0%
1 llama-3.1-nemotron-ultra-253b-v1 (bfcl fc)NVIDIA
100.0%
3 bitagent-bounty-8b (bfcl unspecified)Bittensor
97.5%
4 Claude Haiku 4.5 (bfcl prompt)Anthropic
95.3%
5 Claude Sonnet 4.5 (bfcl prompt)Anthropic
95.0%
6 Gemini 2.5 Flash (bfcl fc)Google
93.7%
7 Gemini 2.5 Flash Lite (bfcl prompt)Google
93.3%
8 DeepSeek V3.2 Exp (bfcl fc)DeepSeek
93.2%
9 Gemini 2.5 Flash Lite (bfcl fc)Google
92.5%
10 Mistral Medium 3 (bfcl fc)Mistral
92.0%
11 Gemini 2.5 Flash (bfcl prompt)Google
91.1%
12 GPT-5 Mini (bfcl fc)OpenAI
91.0%
13 toolace-2-8b (bfcl fc)Huawei Noah And Ustc
90.8%
14 Claude Opus 4.5 (bfcl prompt)Anthropic
90.8%
15 hammer2.1-7b (bfcl fc)Madeagents
90.1%
16 GPT-5 Nano (bfcl fc)OpenAI
89.1%
17 Mistral Small 3.2 24B (bfcl fc)Mistral
87.9%
18 Phi 4 (bfcl prompt)Microsoft
87.5%
19 Kimi K2 0711 (bfcl fc)Moonshot AI
87.3%
20 falcon3-1b-instruct (bfcl fc)TII
87.3%
21 GPT-5.2 (bfcl prompt)OpenAI
87.3%
22 Qwen3 14B (bfcl prompt)Alibaba
87.2%
23 o4 Mini (bfcl prompt)OpenAI
87.2%
24 coalm-8b (bfcl unspecified)Uiuc Oumi
86.9%
25 command-a-reasoning (bfcl fc)Cohere
86.8%
26 Claude Sonnet 4.5 (bfcl fc)Anthropic
86.6%
27 GPT-4.1 (bfcl fc)OpenAI
86.5%
28 o3 (bfcl fc)OpenAI
86.1%
29 hammer2.1-3b (bfcl fc)Madeagents
86.1%
30 coalm-70b (bfcl unspecified)Uiuc Oumi
85.7%
31 gemini-3-pro-preview (bfcl prompt)Google
85.6%
32 Claude Haiku 4.5 (bfcl fc)Anthropic
85.1%
33 GLM 4.6 (thinking reasoning, bfcl fc)Z.ai
85.0%
34 qwen3-4b-instruct-2507 (bfcl fc)Alibaba
84.9%
35 Claude Opus 4.5 (bfcl fc)Anthropic
84.7%
36 grok-4-0709 (bfcl prompt)xAI
84.3%
37 Command A (bfcl fc)Cohere
84.2%
38 GPT-4.1 (bfcl prompt)OpenAI
84.0%
39 o3 (bfcl prompt)OpenAI
84.0%
40 o4 Mini (bfcl fc)OpenAI
83.9%
41 GPT-4.1 Nano (bfcl prompt)OpenAI
83.4%
42 nanbeige4-3b-thinking-2511 (bfcl fc)Nanbeige
83.1%
43 qwen3-0.6b (bfcl prompt)Alibaba
82.5%
44 rzn-t (bfcl prompt)Phronetic Ai
82.4%
45 Qwen3 32B (bfcl prompt)Alibaba
82.4%
46 Qwen3 8B (bfcl prompt)Alibaba
82.3%
47 arch-agent-32b (bfcl unspecified)Katanemo
82.2%
48 Nova 2 Lite (bfcl fc)Amazon
82.1%
49 Qwen3 14B (bfcl fc)Alibaba
81.9%
50 Qwen3 235B A22B Instruct 2507 (bfcl fc)Alibaba
81.7%
51 GPT-4.1 Mini (bfcl fc)OpenAI
81.7%
52 Command R7B (12-2024) (bfcl fc)Cohere
81.7%
53 palmyra-x-004 (bfcl fc)Writer
81.0%
54 qwen3-0.6b (bfcl fc)Alibaba
80.8%
55 hammer2.1-0.5b (bfcl fc)Madeagents
80.8%
56 granite-3.2-8b-instruct (bfcl fc)IBM
80.5%
57 xlam-2-32b-fc-r (bfcl fc)Salesforce
80.2%
58 granite-3.1-8b-instruct (bfcl fc)IBM
80.0%
59 Qwen3 30B A3B Instruct 2507 (bfcl fc)Alibaba
79.9%
60 grok-4-1-fast-reasoning (bfcl fc)xAI
79.4%
61 GPT-5.2 (bfcl fc)OpenAI
79.4%
62 hammer2.1-1.5b (bfcl fc)Madeagents
79.4%
63 xlam-2-70b-fc-r (bfcl fc)Salesforce
79.1%
64 Qwen3 8B (bfcl fc)Alibaba
79.1%
65 Qwen3 235B A22B Instruct 2507 (bfcl prompt)Alibaba
78.9%
66 gemini-3-pro-preview (bfcl fc)Google
77.8%
67 qwen3-1.7b (bfcl fc)Alibaba
76.5%
68 Qwen3 32B (bfcl fc)Alibaba
76.4%
69 qwen3-4b-instruct-2507 (bfcl prompt)Alibaba
75.9%
70 grok-4-0709 (bfcl fc)xAI
75.4%
71 granite-20b-functioncalling (bfcl fc)IBM
75.1%
72 Qwen3 30B A3B Instruct 2507 (bfcl prompt)Alibaba
74.8%
73 arch-agent-1.5b (bfcl unspecified)Katanemo
74.8%
74 arch-agent-3b (bfcl unspecified)Katanemo
74.7%
75 Mistral Medium 3 (bfcl unspecified)Mistral
74.5%
76 nanbeige3.5-pro-thinking (bfcl fc)Nanbeige
74.2%
77 grok-4-1-fast-non-reasoning (bfcl fc)xAI
74.1%
78 GPT-4.1 Mini (bfcl prompt)OpenAI
73.9%
79 minicpm3-4b (bfcl prompt)Openbmb
73.7%
80 Gemma 3 27B (bfcl prompt)Google
73.7%
81 minicpm3-4b (bfcl fc)Openbmb
72.8%
82 Nova Micro 1.0 (bfcl fc)Amazon
70.7%
83 Gemma 3 12B (bfcl prompt)Google
70.3%
84 Nova Pro 1.0 (bfcl fc)Amazon
70.1%
85 mistral-large-2411 (bfcl fc)Mistral
68.9%
86 DeepSeek V3.2 Exp (thinking reasoning, bfcl prompt)DeepSeek
67.0%
87 GPT-4.1 Nano (bfcl fc)OpenAI
66.0%
88 Mistral Small 3.2 24B (bfcl prompt)Mistral
65.7%
89 xlam-2-1b-fc-r (bfcl fc)Salesforce
64.5%
90 xlam-2-3b-fc-r (bfcl fc)Salesforce
63.5%
91 xlam-2-8b-fc-r (bfcl fc)Salesforce
63.3%
92 Mistral Nemo (bfcl fc)Mistral
61.8%
93 granite-4.0-350m (bfcl fc)IBM
60.8%
94 llama-4-maverick-17b-128e-instruct-fp8 (bfcl fc)Meta
56.0%
95 GPT-5 Mini (bfcl prompt)OpenAI
55.7%
96 Gemma 3 4B (bfcl prompt)Google
53.9%
97 Llama 3.3 70B Instruct (bfcl fc)Meta
53.5%
98 Llama 3.2 3B Instruct (bfcl fc)Meta
52.1%
99 Llama 3.2 1B Instruct (bfcl fc)Meta
51.6%
100 GPT-5 Nano (bfcl prompt)OpenAI
45.8%
101 Llama 4 Scout (bfcl fc)Meta
44.9%
102 Llama 3.1 8B Instruct (bfcl prompt)Meta
42.7%
103 mistral-large-2411 (bfcl prompt)Mistral
38.8%
104 bielik-11b-v2.3-instruct (bfcl prompt)Speakleash And Ack Cyfronet Agh
36.0%
105 gemma-3-1b-it (bfcl prompt)Google
33.2%
106 falcon3-3b-instruct (bfcl fc)TII
32.9%
107 falcon3-10b-instruct (bfcl fc)TII
32.1%
108 falcon3-7b-instruct (bfcl fc)TII
32.0%
109 Mistral Nemo (bfcl prompt)Mistral
6.3%

Results as published by Berkeley Function Calling Leaderboard (BFCL) V4; we do not re-run them.

What it measures

Declining to call a function when none of those offered fits.

What it does not measure

Not general refusal or safety behaviour.

Berkeley Function Calling Leaderboard (BFCL) V4 by the UC Berkeley Gorilla team, Apache License 2.0.