Models / GLM 5.2

Z.ai

GLM 5.2

19 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Z.ai
Sources
3
Our benchmarks
0
Price
Not yet published

Reported by others

0.04
Confirmed task success · rank 13 of 46
Unit
IPS effect estimate, higher is better
Range
0.02 to 0.05
Sample
53688 observations
Configuration
GLM 5.2
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.10
Praise over complaint · rank 6 of 46
Unit
IPS effect estimate, higher is better
Range
0.07 to 0.12
Sample
20071 observations
Configuration
GLM 5.2
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.04
Steerability · rank 6 of 46
Unit
IPS effect estimate, higher is better
Range
0.03 to 0.06
Sample
67812 observations
Configuration
GLM 5.2
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.00
Sample
5594512 observations
Configuration
GLM 5.2
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,463
Overall · rank 26 of 177
Unit
Arena rating, higher is better
Range
1,459 to 1,466
Sample
43341 votes
Configuration
GLM 5.2 (max)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,463
Business, management and finance · rank 26 of 402
Unit
Arena rating, higher is better
Range
1,456 to 1,471
Sample
8496 votes
Configuration
GLM 5.2 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,455
Creative writing · rank 13 of 407
Unit
Arena rating, higher is better
Range
1,447 to 1,462
Sample
8448 votes
Configuration
GLM 5.2 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,500
Expert prompts · rank 14 of 359
Unit
Arena rating, higher is better
Range
1,491 to 1,509
Sample
4697 votes
Configuration
GLM 5.2 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,470
Instruction following · rank 16 of 409
Unit
Arena rating, higher is better
Range
1,464 to 1,476
Sample
15327 votes
Configuration
GLM 5.2 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,476
Overall · rank 19 of 409
Unit
Arena rating, higher is better
Range
1,471 to 1,480
Sample
43570 votes
Configuration
GLM 5.2 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,459
Writing, literature and language · rank 18 of 408
Unit
Arena rating, higher is better
Range
1,452 to 1,466
Sample
11517 votes
Configuration
GLM 5.2 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by tau2-bench
13.4%
Consistency (pass^4) · rank 13 of 21
Unit
% of tasks, higher is better
Configuration
glm-5.2
Measured
24 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
37.1%
Task success (pass^1) · rank 11 of 21
Unit
% of tasks, higher is better
Configuration
glm-5.2
Measured
24 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by UGI Leaderboard
12.0%
Requested-length error · rank 98 of 370
Unit
% off the requested word count, lower is better
Configuration
glm-5.2
Measured
21 Jun 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
18.0%
Requested-length error · rank 150 of 370
Unit
% off the requested word count, lower is better
Configuration
glm-5.2
Measured
21 Jun 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.37
Style adherence · rank 81 of 370
Unit
score from 0 to 1, higher is better
Configuration
glm-5.2
Measured
21 Jun 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.35
Style adherence · rank 190 of 370
Unit
score from 0 to 1, higher is better
Configuration
glm-5.2
Measured
21 Jun 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
61.2
Writing score · rank 92 of 370
Unit
score out of 100, higher is better
Configuration
glm-5.2
Measured
21 Jun 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
67.0
Writing score · rank 47 of 370
Unit
score out of 100, higher is better
Configuration
glm-5.2
Measured
21 Jun 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare GLM 5.2 with