> ## Documentation Index
> Fetch the complete documentation index at: https://docs.apiyi.com/llms.txt
> Use this file to discover all available pages before exploring further.

# glm-5.3 and glm-5.3-flash Are Live, Priced at Zhipu's Official Rates

> Zhipu's August releases GLM-5.3 (coding flagship) and GLM-5.3-Flash (multimodal lite) are open for calls on APIYI in the default / svip groups. Priced item for item with the official rates: the flagship at $1.40 in / $4.396 out, Flash at $0.15 in / $0.50 out per million tokens, with recharge bonuses bringing the effective cost to roughly 83%–91% of list. Pre-launch streaming tests with 160K-token prompts confirmed Flash's long-context path.

**2026/9/8 22:17 (UTC+8)** · New Model · Zhipu

🚀 **`glm-5.3` and `glm-5.3-flash` are live in both the `default` and `svip` groups**

Both are Zhipu's August releases: `glm-5.3` keeps 5.2's 753B MoE base and scales post-training only, gaining 50% over 5.2 on Z.ai Code Bench and doubling its ExploitBench score; `glm-5.3-flash` is the first natively multimodal GLM-5, a 320B-A18B model that scores 84.3 on Terminal-Bench 2.1, just 0.7 behind Claude Opus 4.8. Thinking cannot be disabled on either — `reasoning_effort` offers only `low` / `high` / `max`.

Pricing matches Zhipu's official rates item for item: `glm-5.3` at \$1.40 in / \$4.396 out / \$0.259 cached reads, `glm-5.3-flash` at \$0.15 in / \$0.50 out / \$0.03 cached reads per million tokens. Those are list rates; stacking a recharge bonus brings the effective cost to roughly 83%–91% of list.

Before launch we ran streaming tests on `glm-5.3-flash` with prompts of about 160K tokens: time to first byte was 3–5 seconds and each call cost about \$0.005, confirming the long-context path works.

📖 Full benchmarks, specs and code samples: [GLM-5.3 and GLM-5.3-Flash Launch](/en/news/glm-5-3-launch)

***

← [Back to Live Updates](/en/live) · 📚 [Monthly archive](/en/live/archive)
