Benchmarks / Berkeley Function Calling Leaderboard (BFCL) V4
Reported by Berkeley Function Calling Leaderboard (BFCL) V4
Berkeley Function Calling Leaderboard (BFCL) V4
Making a call when a suitable function is available.
- Results dated
- 16 Dec 2025
- Models
- 109
- Unit
- % correct
- Licence
- Apache License 2.0
Full results
| # | Model | BFCL: relevance detection % correct, higher is better |
|---|---|---|
| 1 | Gemma 3 4B (bfcl prompt)Google |
100.0%
|
| 1 | Llama 3.3 70B Instruct (bfcl fc)Meta |
100.0%
|
| 1 | llama-4-maverick-17b-128e-instruct-fp8 (bfcl fc)Meta |
100.0%
|
| 1 | Llama 4 Scout (bfcl fc)Meta |
100.0%
|
| 1 | nanbeige3.5-pro-thinking (bfcl fc)Nanbeige |
100.0%
|
| 1 | GPT-4.1 (bfcl prompt)OpenAI |
100.0%
|
| 1 | falcon3-7b-instruct (bfcl fc)TII |
100.0%
|
| 8 | Qwen3 235B A22B Instruct 2507 (bfcl prompt)Alibaba |
93.8%
|
| 8 | Qwen3 30B A3B Instruct 2507 (bfcl prompt)Alibaba |
93.8%
|
| 8 | Qwen3 32B (bfcl fc)Alibaba |
93.8%
|
| 8 | Qwen3 8B (bfcl fc)Alibaba |
93.8%
|
| 8 | Nova Pro 1.0 (bfcl fc)Amazon |
93.8%
|
| 8 | DeepSeek V3.2 Exp (thinking reasoning, bfcl prompt)DeepSeek |
93.8%
|
| 8 | Gemma 3 12B (bfcl prompt)Google |
93.8%
|
| 8 | Llama 3.1 8B Instruct (bfcl prompt)Meta |
93.8%
|
| 8 | mistral-large-2411 (bfcl fc)Mistral |
93.8%
|
| 8 | mistral-large-2411 (bfcl prompt)Mistral |
93.8%
|
| 8 | Mistral Nemo (bfcl prompt)Mistral |
93.8%
|
| 8 | Mistral Small 3.2 24B (bfcl prompt)Mistral |
93.8%
|
| 8 | GPT-4.1 Nano (bfcl fc)OpenAI |
93.8%
|
| 8 | GPT-5 Mini (bfcl prompt)OpenAI |
93.8%
|
| 8 | GPT-5 Nano (bfcl prompt)OpenAI |
93.8%
|
| 8 | o3 (bfcl prompt)OpenAI |
93.8%
|
| 8 | bielik-11b-v2.3-instruct (bfcl prompt)Speakleash And Ack Cyfronet Agh |
93.8%
|
| 8 | falcon3-10b-instruct (bfcl fc)TII |
93.8%
|
| 8 | coalm-70b (bfcl unspecified)Uiuc Oumi |
93.8%
|
| 27 | Qwen3 14B (bfcl fc)Alibaba |
87.5%
|
| 27 | Qwen3 235B A22B Instruct 2507 (bfcl fc)Alibaba |
87.5%
|
| 27 | qwen3-4b-instruct-2507 (bfcl fc)Alibaba |
87.5%
|
| 27 | qwen3-4b-instruct-2507 (bfcl prompt)Alibaba |
87.5%
|
| 27 | granite-20b-functioncalling (bfcl fc)IBM |
87.5%
|
| 27 | Llama 3.2 3B Instruct (bfcl fc)Meta |
87.5%
|
| 27 | Mistral Small 3.2 24B (bfcl fc)Mistral |
87.5%
|
| 27 | GPT-4.1 Mini (bfcl prompt)OpenAI |
87.5%
|
| 27 | GPT-4.1 (bfcl fc)OpenAI |
87.5%
|
| 27 | xlam-2-1b-fc-r (bfcl fc)Salesforce |
87.5%
|
| 27 | xlam-2-3b-fc-r (bfcl fc)Salesforce |
87.5%
|
| 27 | xlam-2-8b-fc-r (bfcl fc)Salesforce |
87.5%
|
| 27 | coalm-8b (bfcl unspecified)Uiuc Oumi |
87.5%
|
| 27 | grok-4-0709 (bfcl fc)xAI |
87.5%
|
| 41 | qwen3-1.7b (bfcl fc)Alibaba |
81.2%
|
| 41 | Qwen3 14B (bfcl prompt)Alibaba |
81.2%
|
| 41 | Qwen3 30B A3B Instruct 2507 (bfcl fc)Alibaba |
81.2%
|
| 41 | Qwen3 32B (bfcl prompt)Alibaba |
81.2%
|
| 41 | Nova Micro 1.0 (bfcl fc)Amazon |
81.2%
|
| 41 | Command A (bfcl fc)Cohere |
81.2%
|
| 41 | Gemma 3 27B (bfcl prompt)Google |
81.2%
|
| 41 | granite-4.0-350m (bfcl fc)IBM |
81.2%
|
| 41 | arch-agent-32b (bfcl unspecified)Katanemo |
81.2%
|
| 41 | Mistral Nemo (bfcl fc)Mistral |
81.2%
|
| 41 | GPT-4.1 Mini (bfcl fc)OpenAI |
81.2%
|
| 41 | o3 (bfcl fc)OpenAI |
81.2%
|
| 41 | o4 Mini (bfcl fc)OpenAI |
81.2%
|
| 41 | o4 Mini (bfcl prompt)OpenAI |
81.2%
|
| 41 | rzn-t (bfcl prompt)Phronetic Ai |
81.2%
|
| 41 | xlam-2-32b-fc-r (bfcl fc)Salesforce |
81.2%
|
| 41 | falcon3-3b-instruct (bfcl fc)TII |
81.2%
|
| 41 | palmyra-x-004 (bfcl fc)Writer |
81.2%
|
| 41 | grok-4-0709 (bfcl prompt)xAI |
81.2%
|
| 41 | grok-4-1-fast-non-reasoning (bfcl fc)xAI |
81.2%
|
| 41 | grok-4-1-fast-reasoning (bfcl fc)xAI |
81.2%
|
| 62 | qwen3-0.6b (bfcl fc)Alibaba |
75.0%
|
| 62 | qwen3-0.6b (bfcl prompt)Alibaba |
75.0%
|
| 62 | Qwen3 8B (bfcl prompt)Alibaba |
75.0%
|
| 62 | Nova 2 Lite (bfcl fc)Amazon |
75.0%
|
| 62 | Gemini 2.5 Flash (bfcl fc)Google |
75.0%
|
| 62 | gemini-3-pro-preview (bfcl fc)Google |
75.0%
|
| 62 | toolace-2-8b (bfcl fc)Huawei Noah And Ustc |
75.0%
|
| 62 | granite-3.2-8b-instruct (bfcl fc)IBM |
75.0%
|
| 62 | arch-agent-1.5b (bfcl unspecified)Katanemo |
75.0%
|
| 62 | hammer2.1-1.5b (bfcl fc)Madeagents |
75.0%
|
| 62 | Mistral Medium 3 (bfcl unspecified)Mistral |
75.0%
|
| 62 | Kimi K2 0711 (bfcl fc)Moonshot AI |
75.0%
|
| 62 | nanbeige4-3b-thinking-2511 (bfcl fc)Nanbeige |
75.0%
|
| 62 | GPT-5.2 (bfcl fc)OpenAI |
75.0%
|
| 62 | GPT-5.2 (bfcl prompt)OpenAI |
75.0%
|
| 62 | GPT-5 Nano (bfcl fc)OpenAI |
75.0%
|
| 62 | xlam-2-70b-fc-r (bfcl fc)Salesforce |
75.0%
|
| 62 | GLM 4.6 (thinking reasoning, bfcl fc)Z.ai |
75.0%
|
| 80 | Claude Opus 4.5 (bfcl prompt)Anthropic |
68.8%
|
| 80 | Claude Sonnet 4.5 (bfcl fc)Anthropic |
68.8%
|
| 80 | bitagent-bounty-8b (bfcl unspecified)Bittensor |
68.8%
|
| 80 | command-a-reasoning (bfcl fc)Cohere |
68.8%
|
| 80 | Command R7B (12-2024) (bfcl fc)Cohere |
68.8%
|
| 80 | gemini-3-pro-preview (bfcl prompt)Google |
68.8%
|
| 80 | granite-3.1-8b-instruct (bfcl fc)IBM |
68.8%
|
| 80 | arch-agent-3b (bfcl unspecified)Katanemo |
68.8%
|
| 80 | hammer2.1-0.5b (bfcl fc)Madeagents |
68.8%
|
| 80 | GPT-4.1 Nano (bfcl prompt)OpenAI |
68.8%
|
| 80 | minicpm3-4b (bfcl fc)Openbmb |
68.8%
|
| 91 | Claude Haiku 4.5 (bfcl fc)Anthropic |
62.5%
|
| 91 | Claude Opus 4.5 (bfcl fc)Anthropic |
62.5%
|
| 91 | Gemini 2.5 Flash (bfcl prompt)Google |
62.5%
|
| 91 | Mistral Medium 3 (bfcl fc)Mistral |
62.5%
|
| 91 | GPT-5 Mini (bfcl fc)OpenAI |
62.5%
|
| 96 | hammer2.1-3b (bfcl fc)Madeagents |
56.2%
|
| 96 | minicpm3-4b (bfcl prompt)Openbmb |
56.2%
|
| 98 | Gemini 2.5 Flash Lite (bfcl prompt)Google |
50.0%
|
| 98 | hammer2.1-7b (bfcl fc)Madeagents |
50.0%
|
| 98 | Phi 4 (bfcl prompt)Microsoft |
50.0%
|
| 101 | Gemini 2.5 Flash Lite (bfcl fc)Google |
43.8%
|
| 101 | Llama 3.2 1B Instruct (bfcl fc)Meta |
43.8%
|
| 103 | Claude Sonnet 4.5 (bfcl prompt)Anthropic |
37.5%
|
| 103 | DeepSeek V3.2 Exp (bfcl fc)DeepSeek |
37.5%
|
| 103 | gemma-3-1b-it (bfcl prompt)Google |
37.5%
|
| 106 | Claude Haiku 4.5 (bfcl prompt)Anthropic |
31.2%
|
| 107 | ministral-8b-2410 (bfcl fc)Mistral |
0.0%
|
| 107 | llama-3.1-nemotron-ultra-253b-v1 (bfcl fc)NVIDIA |
0.0%
|
| 107 | falcon3-1b-instruct (bfcl fc)TII |
0.0%
|
Results as published by Berkeley Function Calling Leaderboard (BFCL) V4; we do not re-run them.
What it measures
Making a call when a suitable function is available.
What it does not measure
Not whether the call itself was right; only that one was attempted.
Berkeley Function Calling Leaderboard (BFCL) V4 by the UC Berkeley Gorilla team, Apache License 2.0.