Limited-time 50% discount via ZAI through September 9, 2026 at 16:00 UTC.
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
Modalities
Price
Free
Context
1M
Released
Aug 26, 2026
Token volume and request traffic to this model over time.
Yes, the stealth model Ox Alpha was revealed to be ZAI's new model, GLM-5.3 Flash.
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
Yes. The pricing shown on this page for GLM 5.3 Flash is zero, so you are not charged for prompt or completion tokens. Free endpoints are rate limited — see the rate limit docs.
GLM 5.3 Flash has a 1,048,576 token context window.
GLM 5.3 Flash accepts text, images and video as input and returns text.
GLM 5.3 Flash was released on August 26, 2026.