gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deployability on consumer or single-GPU hardware. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.
Modalities
In / Out Price
$0.03 / $0.13per 1M
Context
131K
Released
Aug 5, 2025
Knowledge Cutoff
Jun 2024
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deployability on consumer or single-GPU hardware.
gpt-oss-20b costs $0.03/M input tokens and $0.13/M output tokens, with separate rates for Cache Read at $0.03/M tokens.
gpt-oss-20b has a 131,072 token context window. It supports up to 131,072 completion tokens.
Yes. gpt-oss-20b accepts tools and tool_choice for function calling. It also supports structured outputs via a JSON schema in response_format.
gpt-oss-20b is served by 11 providers on OpenRouter: CoreWeave, DeepInfra, Parasail, Phala, NovitaAI, SiliconFlow, Together, Amazon Bedrock (EU) and 3 more. Requests are routed to the best available provider, with automatic failover to the others, and you can pin or exclude providers with provider routing.
gpt-oss-20b was released on August 5, 2025. Its knowledge cutoff is June 30, 2024.