Z.ai dropped GLM-5.3 on August 14, and the interesting part is not the benchmarks, it's that they didn't retrain anything. Same base model as GLM-5.2, every single gain came from scaled post-training.
Here's what stood out to me:
Same 743B base model as GLM-5.2, zero new pretraining
Terminal-Bench 3.0 jumped from 4.6 to 28.3, that's long-horizon agent work
DeepSWE v1.1 went from 46.2 to 66.9
CyberGym hit 84.5%, edging past Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%
ExploitBench more than doubled, 24.4% to 54.4%
Z.ai says the cyber jump was unplanned, capability kept compounding as they scaled training
Weights are held back roughly two weeks for safety evaluation and hardening
Live now through the Z.ai API, GLM Coding Plan and ZCode