Models / GLM 5.3 Flash

Z.ai

GLM 5.3 Flash

22 published results from 4 sources. Each card shows where the number comes from and what it does not measure. The overall leaderboard combines them; here each stands alone.

Provider
Z.ai
Sources
4
Our benchmarks
0
Price per million tokens
$0.15 in · $0.50 out
OpenRouter list price, 1 Oct 2026 · 1,048,576-token context

Reported by others

Reported by APEX-Agents
52.8%
Tasks passed · rank 19 of 39
Unit
% of tasks, higher is better
Configuration
GLM 5.3 Flash
Measured
1 Oct 2026
Not shown
Not your firm's documents or tools; graded by rubric, not by a client.
0.07
Confirmed task success · rank 13 of 46
Unit
IPS effect estimate, higher is better
Range
0.06 to 0.08
Sample
55182 observations
Configuration
GLM 5.3 Flash
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.00
Praise over complaint · rank 25 of 46
Unit
IPS effect estimate, higher is better
Range
-0.02 to 0.01
Sample
22838 observations
Configuration
GLM 5.3 Flash
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.01
Steerability · rank 29 of 46
Unit
IPS effect estimate, higher is better
Range
-0.02 to -0.00
Sample
74818 observations
Configuration
GLM 5.3 Flash
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 13 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.00
Sample
7664170 observations
Configuration
GLM 5.3 Flash
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,467
Overall · rank 33 of 177
Unit
Arena rating, higher is better
Range
1,462 to 1,471
Sample
18587 votes
Configuration
GLM 5.3 Flash
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,463
Business, management and finance · rank 51 of 402
Unit
Arena rating, higher is better
Range
1,453 to 1,474
Sample
3682 votes
Configuration
GLM 5.3 Flash
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,433
Creative writing · rank 65 of 407
Unit
Arena rating, higher is better
Range
1,423 to 1,444
Sample
3992 votes
Configuration
GLM 5.3 Flash
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,510
Expert prompts · rank 25 of 359
Unit
Arena rating, higher is better
Range
1,497 to 1,524
Sample
2083 votes
Configuration
GLM 5.3 Flash
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,474
Instruction following · rank 26 of 409
Unit
Arena rating, higher is better
Range
1,466 to 1,482
Sample
6780 votes
Configuration
GLM 5.3 Flash
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,474
Overall · rank 35 of 409
Unit
Arena rating, higher is better
Range
1,469 to 1,480
Sample
19103 votes
Configuration
GLM 5.3 Flash
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,450
Writing, literature and language · rank 52 of 408
Unit
Arena rating, higher is better
Range
1,441 to 1,459
Sample
5215 votes
Configuration
GLM 5.3 Flash
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
71.5
Artificial Analysis Coding Index · rank 43 of 151
Unit
index score, higher is better
Configuration
GLM 5.3 Flash
Measured
1 Oct 2026
Not shown
Not your codebase or tools.
41.8
Artificial Analysis Intelligence Index · rank 47 of 314
Unit
index score, higher is better
Configuration
GLM 5.3 Flash
Measured
1 Oct 2026
Not shown
Not business work, and a blend: read the parts for any one task.
91.2%
GPQA Diamond · rank 49 of 276
Unit
% of questions, higher is better
Configuration
GLM 5.3 Flash
Measured
1 Oct 2026
Not shown
Not applied work; multiple-choice science questions.
39.9%
Humanity's Last Exam · rank 75 of 314
Unit
% of questions, higher is better
Configuration
GLM 5.3 Flash
Measured
1 Oct 2026
Not shown
Not everyday work; academic questions at the edge of expertise.
80.0%
Long-context reasoning (AA-LCR) · rank 65 of 309
Unit
% of questions, higher is better
Configuration
GLM 5.3 Flash
Measured
1 Oct 2026
Not shown
Not retrieval over your own document store.
51.6%
SciCode · rank 80 of 153
Unit
% of problems, higher is better
Configuration
GLM 5.3 Flash
Measured
1 Oct 2026
Not shown
Not general software engineering.
84.3%
Terminal-Bench 2.1 · rank 30 of 151
Unit
% of tasks, higher is better
Configuration
GLM 5.3 Flash
Measured
1 Oct 2026
Not shown
Not other harnesses or tools; superseded by 4.0 for newer models.
32.8%
Terminal-Bench 4.0 · rank 35 of 150
Unit
% of tasks, higher is better
Configuration
GLM 5.3 Flash
Measured
1 Oct 2026
Not shown
Not other harnesses or tools; one attempt per task.
47.2%
Τ-bench banking · rank 9 of 142
Unit
% of tasks, higher is better
Configuration
GLM 5.3 Flash
Measured
1 Oct 2026
Not shown
Not your policies or systems; a simulated customer.
Reported by GDP.pdf
14.0%
Professional document tasks · rank 35 of 49
Unit
% of rubric, higher is better
Configuration
GLM 5.3 Flash (max)
Measured
1 Oct 2026
Not shown
Not your documents; 100 tasks, so small differences are noise.

Compare GLM 5.3 Flash with