Skip to content
Not available in this workspace
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for openai

OpenAI: gpt-oss-120b

openai/gpt-oss-120b

Model weights
Compare

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized to run on a single H100 GPU with native MXFP4 quantization. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.

Modalities

In / Out Price

$0.03 / $0.17per 1M

Context

131K

Released

Aug 5, 2025

Knowledge Cutoff

Jun 2024

Compare

About OpenAI: gpt-oss-120b

OpenRouter makes OpenAI: gpt-oss-120b available through a unified, OpenAI-compatible API using the model ID openai/gpt-oss-120b. Requests can be routed across 18 providers, including CoreWeave, DeepInfra (Turbo), AkashML, NovitaAI, SiliconFlow, DigitalOcean, Mancer, Google Vertex and 10 more, with automatic failover when an endpoint is unavailable.

OpenAI: gpt-oss-120b accepts text and returns text. It has a 131,072-token context window and a maximum output of 131,072 tokens.

On OpenRouter, OpenAI: gpt-oss-120b costs $0.03/M input tokens and $0.17/M output tokens, with separate rates for Cache Read at $0.03/M tokens. Effective pricing can be lower when prompt caching applies. It was released on August 5, 2025; its knowledge cutoff is June 30, 2024.

More models from OpenAI

  • GPT Transcribe
  • GPT-5.6 Luna Pro
  • GPT-5.6 Luna

Frequently asked questions

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized to run on a single H100 GPU with native MXFP4 quantization.

gpt-oss-120b costs $0.03/M input tokens and $0.17/M output tokens, with separate rates for Cache Read at $0.03/M tokens.

gpt-oss-120b has a 131,072 token context window. It supports up to 131,072 completion tokens.

Yes. gpt-oss-120b accepts tools and tool_choice for function calling. It also supports structured outputs via a JSON schema in response_format.

gpt-oss-120b is served by 18 providers on OpenRouter: CoreWeave, DeepInfra (Turbo), AkashML, NovitaAI, SiliconFlow, DigitalOcean, Mancer, Google Vertex and 10 more. Requests are routed to the best available provider, with automatic failover to the others, and you can pin or exclude providers with provider routing.

gpt-oss-120b was released on August 5, 2025. Its knowledge cutoff is June 30, 2024.