Skip to main content
2026/9/8 22:17 (UTC+8) · New Model · Zhipu 🚀 glm-5.3 and glm-5.3-flash are live in both the default and svip groups Both are Zhipu’s August releases: glm-5.3 keeps 5.2’s 753B MoE base and scales post-training only, gaining 50% over 5.2 on Z.ai Code Bench and doubling its ExploitBench score; glm-5.3-flash is the first natively multimodal GLM-5, a 320B-A18B model that scores 84.3 on Terminal-Bench 2.1, just 0.7 behind Claude Opus 4.8. Thinking cannot be disabled on either — reasoning_effort offers only low / high / max. Pricing matches Zhipu’s official rates item for item: glm-5.3 at $1.40 in / $4.396 out / $0.259 cached reads, glm-5.3-flash at $0.15 in / $0.50 out / $0.03 cached reads per million tokens. Those are list rates; stacking a recharge bonus brings the effective cost to roughly 83%–91% of list. Before launch we ran streaming tests on glm-5.3-flash with prompts of about 160K tokens: time to first byte was 3–5 seconds and each call cost about $0.005, confirming the long-context path works. 📖 Full benchmarks, specs and code samples: GLM-5.3 and GLM-5.3-Flash Launch
Back to Live Updates · 📚 Monthly archive