Models / Qwen3.5 397B A17B

Alibaba

Qwen3.5 397B A17B

18 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Alibaba
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

1,445
Overall · rank 61 of 177
Unit
Arena rating, higher is better
Range
1,442 to 1,447
Sample
85678 votes
Configuration
Qwen3.5 397B A17B
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,445
Business, management and finance · rank 60 of 402
Unit
Arena rating, higher is better
Range
1,440 to 1,451
Sample
17374 votes
Configuration
Qwen3.5 397B A17B
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,405
Creative writing · rank 79 of 407
Unit
Arena rating, higher is better
Range
1,400 to 1,411
Sample
14663 votes
Configuration
Qwen3.5 397B A17B
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,476
Expert prompts · rank 45 of 359
Unit
Arena rating, higher is better
Range
1,469 to 1,483
Sample
9308 votes
Configuration
Qwen3.5 397B A17B
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,434
Instruction following · rank 66 of 409
Unit
Arena rating, higher is better
Range
1,429 to 1,439
Sample
29573 votes
Configuration
Qwen3.5 397B A17B
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,442
Overall · rank 77 of 409
Unit
Arena rating, higher is better
Range
1,439 to 1,445
Sample
86319 votes
Configuration
Qwen3.5 397B A17B
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,421
Writing, literature and language · rank 72 of 408
Unit
Arena rating, higher is better
Range
1,416 to 1,426
Sample
20894 votes
Configuration
Qwen3.5 397B A17B
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by tau2-bench
68.0%
Consistency (pass^4) · rank 4 of 8
Unit
% of tasks, higher is better
Configuration
Qwen3.5 397B A17B (thinking reasoning, tau2)
Measured
27 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
81.5%
Task success (pass^1) · rank 5 of 8
Unit
% of tasks, higher is better
Configuration
Qwen3.5 397B A17B (thinking reasoning, tau2)
Measured
27 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
5.2%
Consistency (pass^4) · rank 1 of 2
Unit
% of tasks, higher is better
Configuration
Qwen3.5 397B A17B (thinking reasoning, tau2)
Measured
27 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
9.8%
Task success (pass^1) · rank 1 of 2
Unit
% of tasks, higher is better
Configuration
Qwen3.5 397B A17B (thinking reasoning, tau2)
Measured
27 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
59.6%
Consistency (pass^4) · rank 1 of 8
Unit
% of tasks, higher is better
Configuration
Qwen3.5 397B A17B (thinking reasoning, tau2)
Measured
27 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
84.4%
Task success (pass^1) · rank 1 of 8
Unit
% of tasks, higher is better
Configuration
Qwen3.5 397B A17B (thinking reasoning, tau2)
Measured
27 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
92.1%
Consistency (pass^4) · rank 1 of 8
Unit
% of tasks, higher is better
Configuration
Qwen3.5 397B A17B (thinking reasoning, tau2)
Measured
27 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
97.8%
Task success (pass^1) · rank 1 of 8
Unit
% of tasks, higher is better
Configuration
Qwen3.5 397B A17B (thinking reasoning, tau2)
Measured
27 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by UGI Leaderboard
23.0%
Requested-length error · rank 208 of 370
Unit
% off the requested word count, lower is better
Configuration
Qwen3.5 397B A17B (no reasoning)
Measured
21 Mar 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.33
Style adherence · rank 272 of 370
Unit
score from 0 to 1, higher is better
Configuration
Qwen3.5 397B A17B (no reasoning)
Measured
21 Mar 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
50.1
Writing score · rank 156 of 370
Unit
score out of 100, higher is better
Configuration
Qwen3.5 397B A17B (no reasoning)
Measured
21 Mar 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Qwen3.5 397B A17B with