
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for tasks like math, coding, and logical inference, and a "non-thinking" mode for faster, general-purpose conversation. The model demonstrates strong performance in instruction-following, agent tool use, creative writing, and multilingual tasks across 100+ languages and dialects. It natively handles 32K token contexts and can extend to 131K tokens using YaRN-based scaling.
Modalities
Price
Free
Context
131K
Released
Apr 28, 2025
Knowledge Cutoff
Mar 2025
Token volume and request traffic to this model over time.
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for tasks like math, coding, and logical inference, and a "non-thinking" mode for faster, general-purpose conversation.
Yes. The pricing shown on this page for Qwen3 32B is zero, so you are not charged for prompt or completion tokens. Free endpoints are rate limited — see the rate limit docs.
Qwen3 32B has a 131,072 token context window.
Qwen3.8 Max Prime, Qwen3.8 Omni Flash, Qwen3.8 Max (0902) and 49 more are other text models from Qwen.
Qwen3 32B was released on April 28, 2025. Its knowledge cutoff is March 31, 2025.