Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for z-ai

Z.ai: GLM 5.3 Flash

z-ai/glm-5.3-flash

Model weights
Compare

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Modalities

In / Out Price

50% off

$0.075 / $0.25per 1M

Context

1.3M

Released

Aug 26, 2026

Compare
ProvidersPricingPerformanceUptimeBenchmarksAppsActivityFAQExplore

Providers

Different companies host the same model. OpenRouter routes your request to one of them based on the routing mode you pick — Balanced (price + speed), Nitro (fastest), or Exacto (highest tool-calling accuracy).

Pricing

The average price customers actually pay for this model, next to the prices providers post. Caching and discounts mean the price actually paid is often well below the listed one.

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better).

Uptime

Uptime is the percentage of the past 3 days that at least one provider was responding to requests. Availability is the percentage of time that inference was successfully served. OpenRouter continuously monitors and uses the next-best provider when one returns an error.

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows where this model lands among all models on OpenRouter.

Benchmark score summary for Z.ai: GLM 5.3 Flash (Artificial Analysis and Design Arena)
SourceBenchmarkScore
Artificial AnalysisGLM-5.3-Flash Intelligence Index41.9
Artificial AnalysisGLM-5.3-Flash Coding Index71.5
Artificial AnalysisGLM-5.3-Flash Agentic Index51.2
Artificial AnalysisGLM-5.3-Flash GPQA Diamond91.2%
Artificial AnalysisGLM-5.3-Flash HLE39.9%
Artificial AnalysisGLM-5.3-Flash AA-LCR80.0%
Artificial AnalysisGLM-5.3-Flash GDPval-AA57.7%
Artificial AnalysisGLM-5.3-Flash CritPt15.4%
Artificial AnalysisGLM-5.3-Flash SciCode51.6%
Artificial AnalysisGLM-5.3-Flash AA-Omniscience Accuracy27.5%
Artificial AnalysisGLM-5.3-Flash AA-Omniscience Non-Hallucination Rate72.4%
Design ArenaGLM-5.3-Flash Models Arena 3D Elo1355
Design ArenaGLM-5.3-Flash Models Arena Asciiart Elo1288
Design ArenaGLM-5.3-Flash Models Arena Code Categories Elo1298
Design ArenaGLM-5.3-Flash Models Arena Data Visualization Elo1275
Design ArenaGLM-5.3-Flash Models Arena Game Development Elo1309
Design ArenaGLM-5.3-Flash Models Arena SVG Elo1313
Design ArenaGLM-5.3-Flash Models Arena UI Component Elo1340
Design ArenaGLM-5.3-Flash Models Arena Website Elo1285

Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

Activity

Token volume and request traffic to this model over time.

Quick Start

Drop-in code to call this model. OpenRouter's API is OpenAI-compatible — most SDKs work by just swapping the base URL. The only thing that changes between models is the model slug below.

Explore more models

AI Models with Vision: Multimodal LLMs for Image UnderstandingCollectionAI Model RankingsRanking
50% off
$0.15$0.075$0.50$0.25$0.03$0.0152.44s21 tps
97.61%
$0.10$0.3333$0.021.18s44 tps
97.29%
$0.10$0.35$0.021.16s15 tps
96.82%
25% off
$0.15$0.1124$0.50$0.3745$0.03$0.022471.66s54 tps
97.61%
25% off
$0.15$0.1125$0.50$0.375$0.03$0.02253.34s20 tps
98.65%
12% off
$0.15$0.132$0.50$0.44$0.03$0.02643.33s31 tps
91.48%
$0.14$0.47$0.0240.68s37 tps
34.74%
$0.15$0.50$0.031.73s75 tps
90.38%
$0.15$0.50$0.051.17s39 tps
97.95%
$0.15$0.50$0.031.45s18 tps
98.92%
$0.15$0.50$0.033.16s31 tps
95.11%
$0.15$0.50$0.031.44s64 tps
89.79%
$0.15$0.50$0.033.20s38 tps
97.82%
$0.15$0.50$0.032.04s25 tps
80.18%
$0.15$0.50$0.033.08s16 tps
98.26%
$0.15$0.50$0.030.63s77 tps
86.92%
$0.15$0.50$0.031.22s66 tps
97.96%
$0.15$0.50$0.031.00s65 tps
99.16%
$0.15$0.50$0.037.66s14 tps
98.15%
$0.15$0.50$0.031.19s37 tps
86.05%
$0.15$0.50$0.031.35s52 tps
99.88%
$0.15$0.50$0.033.12s42 tps
99.12%
$0.225$0.75$0.0451.51s63 tps
90.74%
$0.25$0.90$0.054.31s26 tps
99.80%
$0.45$1.50$0.090.69s82 tps
96.55%
$0.10$0.35$0.026.75s23 tps
26.97%
$0.15$0.50$0.031.44s79 tps
77.76%

Throughput

82tok/s

P50, best across providers

Latency

0.63s

P50, best provider

AutoExacto Benchmarks
GPQA DiamondTAU-Benchvgi_benchCoreWeave85.1%----Crusoe82.5%----DeepInfra87.3%76.6%--DigitalOcean89.6%73.3%--Z.ai86.9%74.5%--
+24 more providers
Uptime (3d)The model was reachable. Request routed to a provider.

100.00%

Availability (3d)The model returned inference from any provider. Errors and empty responses count against it.

96.98%

Availability over the last 3 days

Last 72 hours
Availability 96.98%
3 Days Ago2 Days AgoYesterdayNow

Availability over the last 24 hours

OpenRouter Availability
98.24%
Without Routing
76.87%

When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.

1.
Favicon for https://claude.ai/apple-touch-icon.png
Claude Code
Claude Code is Anthropic's agentic coding tool that reads your entire codebase, plans and executes changes across files, runs tests, and iterates on failures, all from natural language prompts.
4.74Ttokens
2.
Favicon for https://nousresearch.com
Hermes Agent
Hermes Agent is an open-source, self-improving AI agent by Nous Research that runs persistently with memory across sessions, and builds reusable skills from experience. It comes with 40+ built-in tools, including web search, browser automation, and vision, plus scheduled automations and subagents.
4.41Ttokens
3.
Favicon for https://cline.bot/
Cline
Cline is an open-source AI coding agent that lives inside your IDE, autonomously exploring your codebase, editing files, running terminal commands, and using browser automation.
2.76Ttokens
4.
Favicon for https://omp.sh/
omp
new
1.15Ttokens
5.
Favicon for https://pi.dev/
pi
There are many coding agents, but this one is yours.
904Btokens

Frequently asked questions

Yes, the stealth model Ox Alpha was revealed to be ZAI's new model, GLM-5.3 Flash.

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

GLM 5.3 Flash costs $0.075/M input tokens and $0.25/M output tokens, with separate rates for Cache Read at $0.015/M tokens.

GLM 5.3 Flash has a 1,310,720 token context window. It supports up to 131,072 completion tokens.

Yes. GLM 5.3 Flash accepts tools and tool_choice for function calling. It also supports structured outputs via a JSON schema in response_format.

GLM 5.3 Flash accepts text, images and video as input and returns text.

GLM 5.3 Flash is served by 27 providers on OpenRouter: DeepInfra, Relace, Morph, Wafer, StreamLake, GMICloud, NovitaAI, Makora and 19 more. Requests are routed to the best available provider, with automatic failover to the others, and you can pin or exclude providers with provider routing.

GLM 5.3 Flash was released on August 26, 2026.

More models from Z.ai

GLM Flash Latest

This model always redirects to the latest model in the GLM Flash family.

Text1.0M context
GLM 5.3 Flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Text1.0M context$0.075 / $0.25
GLM Latest

This model always redirects to the latest GLM model from Z.ai.

Text1.0M context
GLM 5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.

Reasoning is always on and cannot be disabled. Reasoning efforts low, high, and max are supported; max is the default.

Text1.3M context$0.8775 / $2.97
GLM 5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.

Reasoning is always on and cannot be disabled. Reasoning efforts low, high, and max are supported; max is the default.

Text1.0M context$0.70 / $2.20
GLM 5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

Text1.0M context$0.4872 / $1.531
GLM 5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

Text33K contextFree
GLM 5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

Text1.0M context$0.70 / $2.20
GLM 5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on a single task for more than 8 hours, autonomously planning, executing, and improving itself throughout the process, ultimately delivering complete, engineering-grade results.

Text205K context$0.9646 / $3.032
GLM 5V Turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding, and task execution, and works seamlessly with agents to complete the full loop of “perceive → plan → execute“.

Text203K context$1.20 / $4
GLM 5 Turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows involving long execution chains, with improved complex instruction decomposition, tool use, scheduled and persistent execution, and overall stability across extended tasks.

Text203K context$1.20 / $4
GLM 5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading closed-source models. With advanced agentic planning, deep backend reasoning, and iterative self-correction, GLM-5 moves beyond code generation to full-system construction and autonomous execution.

Text205K context$0.60 / $1.92
GLM 4.7 Flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaboration, and has achieved leading performance among open-source models of the same size on several current public benchmark leaderboards.

Text200K context$0.06 / $0.40
GLM 4.7

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while delivering more natural conversational experiences and superior front-end aesthetics.

Text205K context$0.40 / $1.75
GLM 4.6V

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts and charts directly as visual inputs, and integrates native multimodal function calling to connect perception with downstream tool execution. The model also enables interleaved image-text generation and UI reconstruction workflows, including screenshot-to-HTML synthesis and iterative visual editing.

Text131K context$0.30 / $0.90
GLM 4.6

Compared with GLM-4.5, this generation brings several key improvements:

Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks. Superior coding performance: The model achieves higher scores on code benchmarks and demonstrates better real-world performance in applications such as Claude Code、Cline、Roo Code and Kilo Code, including improvements in generating visually polished front-end pages. Advanced reasoning: GLM-4.6 shows a clear improvement in reasoning performance and supports tool use during inference, leading to stronger overall capability. More capable agents: GLM-4.6 exhibits stronger performance in tool using and search-based agents, and integrates more effectively within agent frameworks. Refined writing: Better aligns with human preferences in style and readability, and performs more naturally in role-playing scenarios.

Text205K context$0.43 / $1.75
GLM 4.5V

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding, image Q&A, OCR, and document parsing, with strong gains in front-end web coding, grounding, and spatial reasoning. It offers a hybrid inference mode: a "thinking mode" for deep reasoning and a "non-thinking mode" for fast responses. Reasoning behavior can be toggled via the reasoning enabled boolean. Learn more in our docs

Text66K context$0.60 / $1.80
GLM 4.5

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly enhanced capabilities in reasoning, code generation, and agent alignment. It supports a hybrid inference mode with two options, a "thinking mode" designed for complex reasoning and tool use, and a "non-thinking mode" optimized for instant responses. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs

Text131K context$0.60 / $2.20
GLM 4.5 Air

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter size. GLM-4.5-Air also supports hybrid inference modes, offering a "thinking mode" for advanced reasoning and tool use, and a "non-thinking mode" for real-time interaction. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs

Text131K context$0.13 / $0.85
GLM 4 32B

GLM 4 32B is a cost-effective foundation language model.

It can efficiently perform complex tasks and has significantly enhanced capabilities in tool use, online search, and code-related intelligent tasks.

It is made by the same lab behind the thudm models.

Text128K context