Models / GLM 5.3

Z.ai

GLM 5.3

15 published results from 2 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Z.ai
Sources
2
Our benchmarks
1
Price
Not yet published

Measured by Spring Prompt

Measured by Spring Prompt · ROASBench
$0.13
Cost of a run · rank 4 of 17
Unit
US dollars, lower is better
Configuration
GLM 5.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
0
Months over budget · rank 1 of 17
Unit
months of 12, lower is better
Configuration
GLM 5.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
24.8
Overall score · rank 15 of 17
Unit
score out of 100, higher is better
Configuration
GLM 5.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
£0.01
Return on ad spend · rank 15 of 17
Unit
profit per £1 spent, higher is better
Configuration
GLM 5.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

Reported by others

0.06
Confirmed task success · rank 7 of 46
Unit
IPS effect estimate, higher is better
Range
0.05 to 0.07
Sample
58590 observations
Configuration
GLM 5.3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.06
Praise over complaint · rank 9 of 46
Unit
IPS effect estimate, higher is better
Range
0.03 to 0.08
Sample
24698 observations
Configuration
GLM 5.3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.03
Steerability · rank 11 of 46
Unit
IPS effect estimate, higher is better
Range
0.02 to 0.04
Sample
81666 observations
Configuration
GLM 5.3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.00
Sample
7552831 observations
Configuration
GLM 5.3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,468
Overall · rank 15 of 177
Unit
Arena rating, higher is better
Range
1,463 to 1,473
Sample
15656 votes
Configuration
GLM 5.3 (max)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,474
Business, management and finance · rank 8 of 402
Unit
Arena rating, higher is better
Range
1,463 to 1,486
Sample
3027 votes
Configuration
GLM 5.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,456
Creative writing · rank 12 of 407
Unit
Arena rating, higher is better
Range
1,445 to 1,467
Sample
3482 votes
Configuration
GLM 5.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,516
Expert prompts · rank 4 of 359
Unit
Arena rating, higher is better
Range
1,501 to 1,531
Sample
1641 votes
Configuration
GLM 5.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,480
Instruction following · rank 7 of 409
Unit
Arena rating, higher is better
Range
1,471 to 1,488
Sample
5485 votes
Configuration
GLM 5.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,480
Overall · rank 13 of 409
Unit
Arena rating, higher is better
Range
1,474 to 1,485
Sample
15904 votes
Configuration
GLM 5.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,470
Writing, literature and language · rank 8 of 408
Unit
Arena rating, higher is better
Range
1,460 to 1,480
Sample
4482 votes
Configuration
GLM 5.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.

Compare GLM 5.3 with