Benchmarks / Berkeley Function Calling Leaderboard (BFCL) V4

Reported by Berkeley Function Calling Leaderboard (BFCL) V4

Berkeley Function Calling Leaderboard (BFCL) V4

Making a call when a suitable function is available.

Results dated
16 Dec 2025
Models
109
Unit
% correct
Licence
Apache License 2.0

Full results

BFCL: relevance detection, % correct, higher is better
#ModelBFCL: relevance detection
% correct, higher is better
1 Gemma 3 4B (bfcl prompt)Google
100.0%
1 Llama 3.3 70B Instruct (bfcl fc)Meta
100.0%
1 llama-4-maverick-17b-128e-instruct-fp8 (bfcl fc)Meta
100.0%
1 Llama 4 Scout (bfcl fc)Meta
100.0%
1 nanbeige3.5-pro-thinking (bfcl fc)Nanbeige
100.0%
1 GPT-4.1 (bfcl prompt)OpenAI
100.0%
1 falcon3-7b-instruct (bfcl fc)TII
100.0%
8 Qwen3 235B A22B Instruct 2507 (bfcl prompt)Alibaba
93.8%
8 Qwen3 30B A3B Instruct 2507 (bfcl prompt)Alibaba
93.8%
8 Qwen3 32B (bfcl fc)Alibaba
93.8%
8 Qwen3 8B (bfcl fc)Alibaba
93.8%
8 Nova Pro 1.0 (bfcl fc)Amazon
93.8%
8 DeepSeek V3.2 Exp (thinking reasoning, bfcl prompt)DeepSeek
93.8%
8 Gemma 3 12B (bfcl prompt)Google
93.8%
8 Llama 3.1 8B Instruct (bfcl prompt)Meta
93.8%
8 mistral-large-2411 (bfcl fc)Mistral
93.8%
8 mistral-large-2411 (bfcl prompt)Mistral
93.8%
8 Mistral Nemo (bfcl prompt)Mistral
93.8%
8 Mistral Small 3.2 24B (bfcl prompt)Mistral
93.8%
8 GPT-4.1 Nano (bfcl fc)OpenAI
93.8%
8 GPT-5 Mini (bfcl prompt)OpenAI
93.8%
8 GPT-5 Nano (bfcl prompt)OpenAI
93.8%
8 o3 (bfcl prompt)OpenAI
93.8%
8 bielik-11b-v2.3-instruct (bfcl prompt)Speakleash And Ack Cyfronet Agh
93.8%
8 falcon3-10b-instruct (bfcl fc)TII
93.8%
8 coalm-70b (bfcl unspecified)Uiuc Oumi
93.8%
27 Qwen3 14B (bfcl fc)Alibaba
87.5%
27 Qwen3 235B A22B Instruct 2507 (bfcl fc)Alibaba
87.5%
27 qwen3-4b-instruct-2507 (bfcl fc)Alibaba
87.5%
27 qwen3-4b-instruct-2507 (bfcl prompt)Alibaba
87.5%
27 granite-20b-functioncalling (bfcl fc)IBM
87.5%
27 Llama 3.2 3B Instruct (bfcl fc)Meta
87.5%
27 Mistral Small 3.2 24B (bfcl fc)Mistral
87.5%
27 GPT-4.1 Mini (bfcl prompt)OpenAI
87.5%
27 GPT-4.1 (bfcl fc)OpenAI
87.5%
27 xlam-2-1b-fc-r (bfcl fc)Salesforce
87.5%
27 xlam-2-3b-fc-r (bfcl fc)Salesforce
87.5%
27 xlam-2-8b-fc-r (bfcl fc)Salesforce
87.5%
27 coalm-8b (bfcl unspecified)Uiuc Oumi
87.5%
27 grok-4-0709 (bfcl fc)xAI
87.5%
41 qwen3-1.7b (bfcl fc)Alibaba
81.2%
41 Qwen3 14B (bfcl prompt)Alibaba
81.2%
41 Qwen3 30B A3B Instruct 2507 (bfcl fc)Alibaba
81.2%
41 Qwen3 32B (bfcl prompt)Alibaba
81.2%
41 Nova Micro 1.0 (bfcl fc)Amazon
81.2%
41 Command A (bfcl fc)Cohere
81.2%
41 Gemma 3 27B (bfcl prompt)Google
81.2%
41 granite-4.0-350m (bfcl fc)IBM
81.2%
41 arch-agent-32b (bfcl unspecified)Katanemo
81.2%
41 Mistral Nemo (bfcl fc)Mistral
81.2%
41 GPT-4.1 Mini (bfcl fc)OpenAI
81.2%
41 o3 (bfcl fc)OpenAI
81.2%
41 o4 Mini (bfcl fc)OpenAI
81.2%
41 o4 Mini (bfcl prompt)OpenAI
81.2%
41 rzn-t (bfcl prompt)Phronetic Ai
81.2%
41 xlam-2-32b-fc-r (bfcl fc)Salesforce
81.2%
41 falcon3-3b-instruct (bfcl fc)TII
81.2%
41 palmyra-x-004 (bfcl fc)Writer
81.2%
41 grok-4-0709 (bfcl prompt)xAI
81.2%
41 grok-4-1-fast-non-reasoning (bfcl fc)xAI
81.2%
41 grok-4-1-fast-reasoning (bfcl fc)xAI
81.2%
62 qwen3-0.6b (bfcl fc)Alibaba
75.0%
62 qwen3-0.6b (bfcl prompt)Alibaba
75.0%
62 Qwen3 8B (bfcl prompt)Alibaba
75.0%
62 Nova 2 Lite (bfcl fc)Amazon
75.0%
62 Gemini 2.5 Flash (bfcl fc)Google
75.0%
62 gemini-3-pro-preview (bfcl fc)Google
75.0%
62 toolace-2-8b (bfcl fc)Huawei Noah And Ustc
75.0%
62 granite-3.2-8b-instruct (bfcl fc)IBM
75.0%
62 arch-agent-1.5b (bfcl unspecified)Katanemo
75.0%
62 hammer2.1-1.5b (bfcl fc)Madeagents
75.0%
62 Mistral Medium 3 (bfcl unspecified)Mistral
75.0%
62 Kimi K2 0711 (bfcl fc)Moonshot AI
75.0%
62 nanbeige4-3b-thinking-2511 (bfcl fc)Nanbeige
75.0%
62 GPT-5.2 (bfcl fc)OpenAI
75.0%
62 GPT-5.2 (bfcl prompt)OpenAI
75.0%
62 GPT-5 Nano (bfcl fc)OpenAI
75.0%
62 xlam-2-70b-fc-r (bfcl fc)Salesforce
75.0%
62 GLM 4.6 (thinking reasoning, bfcl fc)Z.ai
75.0%
80 Claude Opus 4.5 (bfcl prompt)Anthropic
68.8%
80 Claude Sonnet 4.5 (bfcl fc)Anthropic
68.8%
80 bitagent-bounty-8b (bfcl unspecified)Bittensor
68.8%
80 command-a-reasoning (bfcl fc)Cohere
68.8%
80 Command R7B (12-2024) (bfcl fc)Cohere
68.8%
80 gemini-3-pro-preview (bfcl prompt)Google
68.8%
80 granite-3.1-8b-instruct (bfcl fc)IBM
68.8%
80 arch-agent-3b (bfcl unspecified)Katanemo
68.8%
80 hammer2.1-0.5b (bfcl fc)Madeagents
68.8%
80 GPT-4.1 Nano (bfcl prompt)OpenAI
68.8%
80 minicpm3-4b (bfcl fc)Openbmb
68.8%
91 Claude Haiku 4.5 (bfcl fc)Anthropic
62.5%
91 Claude Opus 4.5 (bfcl fc)Anthropic
62.5%
91 Gemini 2.5 Flash (bfcl prompt)Google
62.5%
91 Mistral Medium 3 (bfcl fc)Mistral
62.5%
91 GPT-5 Mini (bfcl fc)OpenAI
62.5%
96 hammer2.1-3b (bfcl fc)Madeagents
56.2%
96 minicpm3-4b (bfcl prompt)Openbmb
56.2%
98 Gemini 2.5 Flash Lite (bfcl prompt)Google
50.0%
98 hammer2.1-7b (bfcl fc)Madeagents
50.0%
98 Phi 4 (bfcl prompt)Microsoft
50.0%
101 Gemini 2.5 Flash Lite (bfcl fc)Google
43.8%
101 Llama 3.2 1B Instruct (bfcl fc)Meta
43.8%
103 Claude Sonnet 4.5 (bfcl prompt)Anthropic
37.5%
103 DeepSeek V3.2 Exp (bfcl fc)DeepSeek
37.5%
103 gemma-3-1b-it (bfcl prompt)Google
37.5%
106 Claude Haiku 4.5 (bfcl prompt)Anthropic
31.2%
107 ministral-8b-2410 (bfcl fc)Mistral
0.0%
107 llama-3.1-nemotron-ultra-253b-v1 (bfcl fc)NVIDIA
0.0%
107 falcon3-1b-instruct (bfcl fc)TII
0.0%

Results as published by Berkeley Function Calling Leaderboard (BFCL) V4; we do not re-run them.

What it measures

Making a call when a suitable function is available.

What it does not measure

Not whether the call itself was right; only that one was attempted.

Berkeley Function Calling Leaderboard (BFCL) V4 by the UC Berkeley Gorilla team, Apache License 2.0.