
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math, coding, and logical inference, and "non-thinking" mode for general conversation. The model is fine-tuned for instruction-following, agent integration, creative writing, and multilingual use across 100+ languages and dialects. It natively supports a 32K token context window and can extend to 131K tokens with YaRN scaling.
Modalities
Price
Free
Context
131K
Released
Apr 28, 2025
Knowledge Cutoff
Mar 2025
Token volume and request traffic to this model over time.
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math, coding, and logical inference, and "non-thinking" mode for general conversation.
Yes. The pricing shown on this page for Qwen3 8B is zero, so you are not charged for prompt or completion tokens. Free endpoints are rate limited — see the rate limit docs.
Qwen3 8B has a 131,072 token context window.
Qwen3.8 Max (0902), Qwen3.8 Flash, Qwen3.8 27B and 47 more are other text models from Qwen.
Qwen3 8B was released on April 28, 2025. Its knowledge cutoff is March 31, 2025.