GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration. It supports text input and output with a 1M-token context window and up to 128K output tokens, and targets coding and agentic workloads, including long-horizon multi-turn agent orchestration, real-time conversation, and streaming code generation.
Reasoning is always on and cannot be disabled. Reasoning efforts low, high, and max are supported; max is the default.
Modalities
In / Out Price
$2.80 / $8.80per 1M
Context
1.0M
Released
Sep 23, 2026