Compare DeepSeek V4 Flash 0731 from DeepSeek and GLM 5.3 Flash from Z.ai on key metrics including benchmarks, price, context length, and other model features. Access both models and hundreds of others through the OpenRouter API.


DeepSeek V4 Flash 0731 and GLM 5.3 Flash are available through the OpenRouter API, so switching between them takes a model slug change rather than a new integration.
DeepSeek V4 Flash 0731, from DeepSeek, has a 1,048,576-token context window and is priced at $0.0047/M tokens input, $1.28/M tokens output on OpenRouter.
GLM 5.3 Flash, from Z.ai, has a 1,048,576-token context window and is priced at $0.04/M tokens input, $0.50/M tokens output on OpenRouter.
| Attribute | DeepSeek V4 Flash 0731 | GLM 5.3 Flash |
|---|---|---|
| Input price | $0.0047/M tokens | $0.04/M tokens |
| Output price | $1.28/M tokens | $0.50/M tokens |
| Context window | 1,048,576 tokens | 1,048,576 tokens |
| Intelligence Index | 34.3 | 41.8 |
| Coding Index | 69.1 | 71.5 |
| Agentic Index | 41.0 | 50.9 |
| Latency (p50) | 0.48 s | 1.50 s |
| Throughput (p50) | 66.0 tok/s | 43.0 tok/s |
Benchmark data updated
DeepSeek V4 Flash 0731 and GLM 5.3 Flash trade off on price: DeepSeek V4 Flash 0731 costs $0.0047/M tokens input and $1.28/M tokens output, while GLM 5.3 Flash costs $0.04/M tokens input and $0.50/M tokens output, so DeepSeek V4 Flash 0731 is cheaper for input tokens and GLM 5.3 Flash is cheaper for output tokens.
DeepSeek V4 Flash 0731 and GLM 5.3 Flash have the same context window of 1,048,576 tokens.
DeepSeek V4 Flash 0731 is faster on OpenRouter, generating a median 66.0 tokens per second compared with 43.0 tokens per second for GLM 5.3 Flash.
GLM 5.3 Flash scores higher on coding, with an Artificial Analysis Coding Index of 71.5 compared with 69.1 for DeepSeek V4 Flash 0731.