gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized to run on a single H100 GPU with native MXFP4 quantization. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.
Modalities
In / Out Price
$0.03 / $0.17per 1M
Context
131K
Released
Aug 5, 2025
Knowledge Cutoff
Jun 2024
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized to run on a single H100 GPU with native MXFP4 quantization.
gpt-oss-120b costs $0.03/M input tokens and $0.17/M output tokens, with separate rates for Cache Read at $0.03/M tokens.
gpt-oss-120b has a 131,072 token context window. It supports up to 131,072 completion tokens.
Yes. gpt-oss-120b accepts tools and tool_choice for function calling. It also supports structured outputs via a JSON schema in response_format.
gpt-oss-120b is served by 18 providers on OpenRouter: CoreWeave, DeepInfra (Turbo), AkashML, NovitaAI, SiliconFlow, DigitalOcean, Mancer, Google Vertex and 10 more. Requests are routed to the best available provider, with automatic failover to the others, and you can pin or exclude providers with provider routing.
gpt-oss-120b was released on August 5, 2025. Its knowledge cutoff is June 30, 2024.