GLM-5.3
OPENZ.ai's August 14, 2026 release, and an unusually clean natural experiment: the same base model as GLM-5.2, with every gain from post-training alone. Terminal-Bench 3.0 goes 4.6 → 28.3, DeepSWE v1.1 46.2 → 66.9, AutomationBench 26.2 → 48.2, and Z.ai's in-house Code Bench improves ~50% — while consuming FEWER output tokens at every effort level (34.5% at ~75K tokens/task at max, against GLM-5.2's 23.4% at 96K). API-only at launch at $1.40/$4.40 per Mtok with off-peak billing at half rate; Z.ai said weights would follow roughly two weeks later once safety hardening completed, without naming a licence in advance — so `openness` reflects the commitment, and the licence field records that it was still unnamed. The headline finding is emergent cyber capability that Z.ai says developed faster than expected: SOTA on CyberGym at 84.5 (ahead of Mythos 5's 83.8 and GPT-5.6 Sol's 83.6) and more than double GLM-5.2 on exploitation, but the gap to the closed frontier WIDENS further up the exploitation chain — ExploitBench 54.4 against Mythos 5's 78.0 and Sol's 76.5. Its Agents' Last Exam figure (28.5) is omitted: Z.ai reports a pass-rate-style metric incompatible with the Score convention used by the other rows in that file.