52.8%
Tasks passed · rank 19 of 39
- Unit
- % of tasks, higher is better
- Configuration
- GLM 5.3 Flash
- Measured
- 1 Oct 2026
- Not shown
- Not your firm's documents or tools; graded by rubric, not by a client.
0.07
Confirmed task success · rank 13 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- 0.06 to 0.08
- Sample
- 55182 observations
- Configuration
- GLM 5.3 Flash
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.00
Praise over complaint · rank 25 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.02 to 0.01
- Sample
- 22838 observations
- Configuration
- GLM 5.3 Flash
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.01
Steerability · rank 29 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- -0.02 to -0.00
- Sample
- 74818 observations
- Configuration
- GLM 5.3 Flash
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 13 of 46
- Unit
- IPS effect estimate, higher is better
- Range
- 0.00 to 0.00
- Sample
- 7664170 observations
- Configuration
- GLM 5.3 Flash
- Measured
- 28 Sep 2026
- Not shown
- Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,467
Overall · rank 33 of 177
- Unit
- Arena rating, higher is better
- Range
- 1,462 to 1,471
- Sample
- 18587 votes
- Configuration
- GLM 5.3 Flash
- Measured
- 25 Sep 2026
- Not shown
- Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,463
Business, management and finance · rank 51 of 402
- Unit
- Arena rating, higher is better
- Range
- 1,453 to 1,474
- Sample
- 3682 votes
- Configuration
- GLM 5.3 Flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,433
Creative writing · rank 65 of 407
- Unit
- Arena rating, higher is better
- Range
- 1,423 to 1,444
- Sample
- 3992 votes
- Configuration
- GLM 5.3 Flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,510
Expert prompts · rank 25 of 359
- Unit
- Arena rating, higher is better
- Range
- 1,497 to 1,524
- Sample
- 2083 votes
- Configuration
- GLM 5.3 Flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,474
Instruction following · rank 26 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,466 to 1,482
- Sample
- 6780 votes
- Configuration
- GLM 5.3 Flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,474
Overall · rank 35 of 409
- Unit
- Arena rating, higher is better
- Range
- 1,469 to 1,480
- Sample
- 19103 votes
- Configuration
- GLM 5.3 Flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,450
Writing, literature and language · rank 52 of 408
- Unit
- Arena rating, higher is better
- Range
- 1,441 to 1,459
- Sample
- 5215 votes
- Configuration
- GLM 5.3 Flash
- Measured
- 25 Sep 2026
- Not shown
- Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
71.5
Artificial Analysis Coding Index · rank 43 of 151
- Unit
- index score, higher is better
- Configuration
- GLM 5.3 Flash
- Measured
- 1 Oct 2026
- Not shown
- Not your codebase or tools.
41.8
Artificial Analysis Intelligence Index · rank 47 of 314
- Unit
- index score, higher is better
- Configuration
- GLM 5.3 Flash
- Measured
- 1 Oct 2026
- Not shown
- Not business work, and a blend: read the parts for any one task.
91.2%
GPQA Diamond · rank 49 of 276
- Unit
- % of questions, higher is better
- Configuration
- GLM 5.3 Flash
- Measured
- 1 Oct 2026
- Not shown
- Not applied work; multiple-choice science questions.
39.9%
Humanity's Last Exam · rank 75 of 314
- Unit
- % of questions, higher is better
- Configuration
- GLM 5.3 Flash
- Measured
- 1 Oct 2026
- Not shown
- Not everyday work; academic questions at the edge of expertise.
80.0%
Long-context reasoning (AA-LCR) · rank 65 of 309
- Unit
- % of questions, higher is better
- Configuration
- GLM 5.3 Flash
- Measured
- 1 Oct 2026
- Not shown
- Not retrieval over your own document store.
51.6%
SciCode · rank 80 of 153
- Unit
- % of problems, higher is better
- Configuration
- GLM 5.3 Flash
- Measured
- 1 Oct 2026
- Not shown
- Not general software engineering.
84.3%
Terminal-Bench 2.1 · rank 30 of 151
- Unit
- % of tasks, higher is better
- Configuration
- GLM 5.3 Flash
- Measured
- 1 Oct 2026
- Not shown
- Not other harnesses or tools; superseded by 4.0 for newer models.
32.8%
Terminal-Bench 4.0 · rank 35 of 150
- Unit
- % of tasks, higher is better
- Configuration
- GLM 5.3 Flash
- Measured
- 1 Oct 2026
- Not shown
- Not other harnesses or tools; one attempt per task.
47.2%
Τ-bench banking · rank 9 of 142
- Unit
- % of tasks, higher is better
- Configuration
- GLM 5.3 Flash
- Measured
- 1 Oct 2026
- Not shown
- Not your policies or systems; a simulated customer.
14.0%
Professional document tasks · rank 35 of 49
- Unit
- % of rubric, higher is better
- Configuration
- GLM 5.3 Flash (max)
- Measured
- 1 Oct 2026
- Not shown
- Not your documents; 100 tasks, so small differences are noise.