Skip to content
Not available in this workspace
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for z-ai

Z.ai: GLM 4.7 Flash

z-ai/glm-4.7-flash

Model weights
Compare

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaboration, and has achieved leading performance among open-source models of the same size on several current public benchmark leaderboards.

Modalities

In / Out Price

$0.06 / $0.40per 1M

Context

203K

Released

Jan 19, 2026

Compare

About Z.ai: GLM 4.7 Flash

OpenRouter makes Z.ai: GLM 4.7 Flash available through a unified, OpenAI-compatible API using the model ID z-ai/glm-4.7-flash. Requests can be routed across 4 providers, including DeepInfra, Venice, Cloudflare and NovitaAI, with automatic failover when an endpoint is unavailable.

Z.ai: GLM 4.7 Flash accepts text and returns text. It has a 202,752-token context window and a maximum output of 16,384 tokens.

On OpenRouter, Z.ai: GLM 4.7 Flash costs $0.06/M input tokens and $0.40/M output tokens, with separate rates for Cache Read at $0.01/M tokens. Effective pricing can be lower when prompt caching applies. It was released on January 19, 2026.

More models from Z.ai

  • GLM 5.3
  • GLM 5.2
  • GLM 5.1

Frequently asked questions

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaboration, and has achieved leading performance among open-source models of the same size on several current public benchmark leaderboards.

GLM 4.7 Flash costs $0.06/M input tokens and $0.40/M output tokens, with separate rates for Cache Read at $0.01/M tokens.

GLM 4.7 Flash has a 202,752 token context window. It supports up to 16,384 completion tokens.

Yes. GLM 4.7 Flash accepts tools and tool_choice for function calling. It also supports structured outputs via a JSON schema in response_format.

GLM 4.7 Flash is served by 4 providers on OpenRouter: DeepInfra, Venice, Cloudflare and NovitaAI. Requests are routed to the best available provider, with automatic failover to the others, and you can pin or exclude providers with provider routing.

GLM 4.7 Flash was released on January 19, 2026.