GLM-5.3-Flash
OPENZ.ai's August 26, 2026 Flash-tier model and the first natively multimodal GLM-5. A new base rather than a post-training refresh: 320B total / 18B active with hybrid sparse+linear attention, halving both active parameters and layer count against GLM-4.5's 355B/32B while beating GLM-5.2 across the board at about a tenth of the price. Z.ai places it on the Pareto frontier of Artificial Analysis Intelligence Index v4.1.1 at 57 for $0.045 per task (discounted) — a level previously available only at roughly 10x the cost. Notable for being served entirely on domestic Chinese AI chips, and for having been A/B-tested anonymously as `ox-alpha` on OpenCode and OpenRouter before launch. Pricing here is the published per-Mtok rate; peak/off-peak billing applies on the coding plan. Its HLE row (55.3) is the WITH-tools configuration and its CharXiv row is with tools, both noted per row. Agents' Last Exam is omitted for the same metric-incompatibility reason as GLM-5.3.