Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token.
The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.
Modalities
Price
Free
Context
262K
Released
Jul 23, 2026
Token volume and request traffic to this model over time.
Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.
Yes. The pricing shown on this page for Ling 3.0 Flash is zero, so you are not charged for prompt or completion tokens. Free endpoints are rate limited — see the rate limit docs.
Ling 3.0 Flash has a 262,144 token context window.
Ling 3.0 Flash VL, Ling 3.0 Flash Sante (free) and Ling 3.0 Flash Fin are other text models from inclusionAI.
Ling 3.0 Flash was released on July 23, 2026.