Benchmarks / Berkeley Function Calling Leaderboard (BFCL) V4

Reported by Berkeley Function Calling Leaderboard (BFCL) V4

Berkeley Function Calling Leaderboard (BFCL) V4

The same task on functions and questions contributed by real users.

Results dated
16 Dec 2025
Models
109
Unit
% correct
Licence
Apache License 2.0
BFCL: single-turn calls (user-contributed), % correct, higher is better
#ModelBFCL: single-turn calls (user-contributed)
% correct, higher is better
1 bitagent-bounty-8bBittensor
93.1%
2 gemini-3-pro-previewGoogle
83.1%
3 Qwen3 32BAlibaba · qwen3-32b
82.0%
3 Qwen3 32BAlibaba · qwen3-32b
82.0%
5 mistral-large-2411Mistral
81.9%
6 gemini-3-pro-previewGoogle
81.7%
7 Claude Sonnet 4.5Anthropic · claude-sonnet-4.5
81.1%
8 GLM 4.6Z.ai · glm-4.6
80.9%
9 Nova 2 LiteAmazon · nova-2-lite-v1
80.8%
10 arch-agent-32bKatanemo
80.7%
11 Qwen3 8BAlibaba · qwen3-8b
80.5%
12 Qwen3 8BAlibaba · qwen3-8b
80.1%
13 Qwen3 14BAlibaba · qwen3-14b
80.0%
14 Claude Opus 4.5Anthropic · claude-opus-4.5
79.8%
15 nanbeige4-3b-thinking-2511Nanbeige
79.4%
16 Qwen3 14BAlibaba · qwen3-14b
79.3%
17 Mistral Small 3.2 24BMistral · mistral-small-3.2-24b-instruct
79.0%
18 GPT-4.1OpenAI · gpt-4.1
78.9%
19 Qwen3 235B A22B Instruct 2507Alibaba · qwen3-235b-a22b-2507
78.7%
19 Claude Haiku 4.5Anthropic · claude-haiku-4.5
78.7%
19 Kimi K2 0711Moonshot AI · kimi-k2
78.7%
22 command-a-reasoningCohere
78.6%
23 Nova Pro 1.0Amazon · nova-pro-v1
78.5%
23 Command ACohere · command-a
78.5%
25 grok-4-1-fast-reasoningxAI
78.5%
26 Qwen3 30B A3B Instruct 2507Alibaba · qwen3-30b-a3b-instruct-2507
78.4%
27 Gemini 2.5 FlashGoogle · gemini-2.5-flash
78.2%
28 Qwen3 30B A3B Instruct 2507Alibaba · qwen3-30b-a3b-instruct-2507
77.9%
28 grok-4-1-fast-non-reasoningxAI
77.9%
30 palmyra-x-004Writer
77.9%
31 toolace-2-8bHuawei Noah And Ustc
77.4%
32 Mistral Small 3.2 24BMistral · mistral-small-3.2-24b-instruct
77.3%
33 Llama 3.3 70B InstructMeta · llama-3.3-70b-instruct
76.6%
34 qwen3-4b-instruct-2507Alibaba
76.4%
35 Claude Opus 4.5Anthropic · claude-opus-4.5
76.0%
35 DeepSeek V3.2 ExpDeepSeek · deepseek-v3.2-exp
76.0%
37 grok-4-0709xAI
75.6%
38 xlam-2-32b-fc-rSalesforce
75.5%
39 falcon3-10b-instructTII
75.4%
40 GPT-4.1 MiniOpenAI · gpt-4.1-mini
74.8%
41 qwen3-4b-instruct-2507Alibaba
74.7%
41 Llama 4 ScoutMeta · llama-4-scout
74.7%
43 qwen3-1.7bAlibaba
74.6%
44 Gemma 3 27BGoogle · gemma-3-27b-it
74.5%
45 Gemini 2.5 FlashGoogle · gemini-2.5-flash
74.4%
46 Gemma 3 12BGoogle · gemma-3-12b-it
74.2%
47 Mistral NemoMistral · mistral-nemo
74.0%
48 Mistral NemoMistral · mistral-nemo
73.8%
49 llama-4-maverick-17b-128e-instruct-fp8Meta
73.7%
50 o3OpenAI
73.2%
51 arch-agent-3bKatanemo
72.9%
52 grok-4-0709xAI
72.5%
53 xlam-2-70b-fc-rSalesforce
72.2%
54 Llama 3.1 8B InstructMeta · llama-3.1-8b-instruct
70.8%
54 o4 MiniOpenAI · o4-mini
70.8%
56 GPT-5 NanoOpenAI · gpt-5-nano
70.7%
57 hammer2.1-3bMadeagents
70.5%
58 GPT-5.2OpenAI · gpt-5.2
70.4%
59 nanbeige3.5-pro-thinkingNanbeige
70.0%
59 GPT-4.1OpenAI · gpt-4.1
70.0%
61 hammer2.1-1.5bMadeagents
69.5%
61 hammer2.1-7bMadeagents
69.5%
63 Command R7B (12-2024)Cohere · command-r7b-12-2024
69.1%
64 Qwen3 235B A22B Instruct 2507Alibaba · qwen3-235b-a22b-2507
68.9%
65 GPT-4.1 MiniOpenAI · gpt-4.1-mini
68.8%
66 falcon3-7b-instructTII
68.3%
67 mistral-large-2411Mistral
68.1%
68 Mistral Medium 3Mistral · mistral-medium-3
68.0%
68 xlam-2-8b-fc-rSalesforce
68.0%
70 bielik-11b-v2.3-instructSpeakleash And Ack Cyfronet Agh
67.8%
71 arch-agent-1.5bKatanemo
67.7%
72 coalm-70bUiuc Oumi
67.3%
73 GPT-5.2OpenAI · gpt-5.2
67.1%
74 coalm-8bUiuc Oumi
66.8%
75 Nova Micro 1.0Amazon · nova-micro-v1
66.3%
76 o3OpenAI
66.2%
77 o4 MiniOpenAI · o4-mini
66.1%
78 Mistral Medium 3Mistral · mistral-medium-3
66.0%
79 Gemini 2.5 Flash LiteGoogle · gemini-2.5-flash-lite
65.8%
80 minicpm3-4bOpenbmb
65.2%
81 xlam-2-3b-fc-rSalesforce
62.9%
82 GPT-5 MiniOpenAI · gpt-5-mini
62.5%
83 Gemma 3 4BGoogle · gemma-3-4b-it
60.8%
84 GPT-4.1 NanoOpenAI · gpt-4.1-nano
60.8%
85 Phi 4Microsoft · phi-4
60.7%
86 granite-3.1-8b-instructIBM
60.3%
86 granite-3.2-8b-instructIBM
60.3%
88 GPT-5 NanoOpenAI · gpt-5-nano
59.4%
89 granite-20b-functioncallingIBM
58.7%
90 GPT-5 MiniOpenAI · gpt-5-mini
58.6%
91 Llama 3.2 3B InstructMeta · llama-3.2-3b-instruct
58.3%
92 qwen3-0.6bAlibaba
56.6%
93 xlam-2-1b-fc-rSalesforce
55.1%
94 Gemini 2.5 Flash LiteGoogle · gemini-2.5-flash-lite
54.9%
95 hammer2.1-0.5bMadeagents
54.6%
96 falcon3-3b-instructTII
54.5%
97 DeepSeek V3.2 ExpDeepSeek · deepseek-v3.2-exp
53.7%
98 Claude Haiku 4.5Anthropic · claude-haiku-4.5
52.5%
99 GPT-4.1 NanoOpenAI · gpt-4.1-nano
50.3%
100 rzn-tPhronetic Ai
49.7%
101 qwen3-0.6bAlibaba
49.4%
102 Claude Sonnet 4.5Anthropic · claude-sonnet-4.5
46.6%
103 granite-4.0-350mIBM
46.1%
104 minicpm3-4bOpenbmb
43.1%
105 gemma-3-1b-itGoogle
11.8%
106 Llama 3.2 1B InstructMeta · llama-3.2-1b-instruct
11.8%
107 falcon3-1b-instructTII
2.9%
108 ministral-8b-2410Mistral
0.0%
108 llama-3.1-nemotron-ultra-253b-v1NVIDIA
0.0%

Results as published by Berkeley Function Calling Leaderboard (BFCL) V4; we do not re-run them.

What it measures

The same task on functions and questions contributed by real users.

What it does not measure

Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.

Berkeley Function Calling Leaderboard (BFCL) V4 by the UC Berkeley Gorilla team, Apache License 2.0.