Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token.
The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.
Modalities
Price
Free
Context
262K
Released
Jul 23, 2026
Token volume and request traffic to this model over time.
Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.
Yes. The pricing shown on this page for Ling 3.0 Flash is zero, so you are not charged for prompt or completion tokens. Free endpoints are rate limited — see the rate limit docs.
Ling 3.0 Flash has a 262,144 token context window.
Ling 3.0 Flash Sante (free) and Ling 3.0 Flash Fin are other text models from inclusionAI.
Ling 3.0 Flash was released on July 23, 2026.