As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaboration, and has achieved leading performance among open-source models of the same size on several current public benchmark leaderboards.
Modalities
In / Out Price
$0.06 / $0.40per 1M
Context
203K
Released
Jan 19, 2026
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaboration, and has achieved leading performance among open-source models of the same size on several current public benchmark leaderboards.
GLM 4.7 Flash costs $0.06/M input tokens and $0.40/M output tokens, with separate rates for Cache Read at $0.01/M tokens.
GLM 4.7 Flash has a 202,752 token context window. It supports up to 16,384 completion tokens.
Yes. GLM 4.7 Flash accepts tools and tool_choice for function calling. It also supports structured outputs via a JSON schema in response_format.
GLM 4.7 Flash is served by 4 providers on OpenRouter: DeepInfra, Venice, Cloudflare and NovitaAI. Requests are routed to the best available provider, with automatic failover to the others, and you can pin or exclude providers with provider routing.
GLM 4.7 Flash was released on January 19, 2026.