Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at general chat, coding tasks, instruction-following, and function calling. Featuring a 32K context length and optimized through reinforcement learning (RLOO), it provides competitive performance comparable to proprietary models within a smaller parameter footprint. Ideal for low-latency, local, or on-device deployments, Reka Flash 3 is compact, supports efficient quantization (down to 11GB at 4-bit precision), and employs explicit reasoning tags ("<reasoning>") to indicate its internal thought process.
Reka Flash 3 is primarily an English model with limited multilingual understanding capabilities. The model weights are released under the Apache 2.0 license.
Modalities
Price
Free
Context
32K
Released
Mar 12, 2025
Knowledge Cutoff
Jan 2025
Token volume and request traffic to this model over time.
Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at general chat, coding tasks, instruction-following, and function calling. Featuring a 32K context length and optimized through reinforcement learning (RLOO), it provides competitive performance comparable to proprietary models within a smaller parameter footprint.
Yes. The pricing shown on this page for Reka Flash 3 is zero, so you are not charged for prompt or completion tokens. Free endpoints are rate limited — see the rate limit docs.
Reka Flash 3 has a 32,000 token context window.
Reka Edge is another text model from the same author.
Reka Flash 3 was released on March 12, 2025. Its knowledge cutoff is January 31, 2025.