Models / Qwen3 235B A22B Instruct 2507

Alibaba

Qwen3 235B A22B Instruct 2507

26 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Alibaba
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

1,431
Overall · rank 92 of 177
Unit
Arena rating, higher is better
Range
1,429 to 1,434
Sample
66540 votes
Configuration
Qwen3 235B A22B Instruct 2507
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,432
Business, management and finance · rank 81 of 402
Unit
Arena rating, higher is better
Range
1,427 to 1,437
Sample
18230 votes
Configuration
Qwen3 235B A22B Instruct 2507
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,378
Creative writing · rank 114 of 407
Unit
Arena rating, higher is better
Range
1,372 to 1,383
Sample
13819 votes
Configuration
Qwen3 235B A22B Instruct 2507
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,448
Expert prompts · rank 90 of 359
Unit
Arena rating, higher is better
Range
1,440 to 1,455
Sample
6405 votes
Configuration
Qwen3 235B A22B Instruct 2507
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,414
Instruction following · rank 100 of 409
Unit
Arena rating, higher is better
Range
1,410 to 1,419
Sample
27749 votes
Configuration
Qwen3 235B A22B Instruct 2507
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,422
Overall · rank 110 of 409
Unit
Arena rating, higher is better
Range
1,420 to 1,425
Sample
97815 votes
Configuration
Qwen3 235B A22B Instruct 2507
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,396
Writing, literature and language · rank 101 of 408
Unit
Arena rating, higher is better
Range
1,392 to 1,401
Sample
22059 votes
Configuration
Qwen3 235B A22B Instruct 2507
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
81.7%
Irrelevance detection · rank 50 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl fc)
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
78.9%
Irrelevance detection · rank 65 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl prompt)
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
23.9%
Memory · rank 35 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl fc)
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
19.4%
Memory · rank 46 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl prompt)
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
45.4%
Multi-turn tasks · rank 21 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl fc)
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
44.6%
Multi-turn tasks · rank 23 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl prompt)
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
48.0%
Overall accuracy · rank 31 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl fc)
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
52.1%
Overall accuracy · rank 23 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl prompt)
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
87.5%
Relevance detection · rank 27 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl fc)
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
93.8%
Relevance detection · rank 8 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl prompt)
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
37.4%
Single-turn calls (curated) · rank 104 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl fc)
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
90.3%
Single-turn calls (curated) · rank 2 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl prompt)
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
68.9%
Single-turn calls (user-contributed) · rank 64 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl fc)
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
78.7%
Single-turn calls (user-contributed) · rank 19 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl prompt)
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
54.0%
Web search · rank 25 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl fc)
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
50.5%
Web search · rank 26 of 109
Unit
% correct, higher is better
Configuration
Qwen3 235B A22B Instruct 2507 (bfcl prompt)
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
Reported by UGI Leaderboard
15.0%
Requested-length error · rank 119 of 370
Unit
% off the requested word count, lower is better
Configuration
Qwen3 235B A22B Instruct 2507
Measured
21 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.36
Style adherence · rank 124 of 370
Unit
score from 0 to 1, higher is better
Configuration
Qwen3 235B A22B Instruct 2507
Measured
21 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
49.0
Writing score · rank 164 of 370
Unit
score out of 100, higher is better
Configuration
Qwen3 235B A22B Instruct 2507
Measured
21 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Qwen3 235B A22B Instruct 2507 with