Skip to content
Not available in this workspace
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

For the free endpoint, please do not upload any confidential information or personal data (such as voices or faces of people). Your use is logged for security purposes and to improve NVIDIA products and services. The logged session data for improvement purposes is not linked to your identity or any persistent identifier. For more information about NVIDIA's data processing practices, see Privacy Policy(opens in new tab). By using this free endpoint, you consent to NVIDIA's collection, recording, and use of such information and the NVIDIA API Trial Terms of Service(opens in new tab)

Favicon for nvidia

NVIDIA: Nemotron 3 Super

nvidia/nemotron-3-super-120b-a12b

Model weights
Compare

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer Mixture-of-Experts architecture with multi-token prediction (MTP), it delivers over 50% higher token generation compared to leading open models.

The model features a 1M token context window for long-term agent coherence, cross-document reasoning, and multi-step task planning. Latent MoE enables calling 4 experts for the inference cost of only one, improving intelligence and generalization. Multi-environment RL training across 10+ environments delivers leading accuracy on benchmarks including AIME 2025, TerminalBench, and SWE-Bench Verified.

Fully open with weights, datasets, and recipes under the NVIDIA Open License, Nemotron 3 Super allows easy customization and secure deployment anywhere — from workstation to cloud.

Modalities

In / Out Price

$0.085 / $0.40per 1M

Context

1M

Released

Mar 11, 2026

Compare

About NVIDIA: Nemotron 3 Super

OpenRouter makes NVIDIA: Nemotron 3 Super available through a unified, OpenAI-compatible API using the model ID nvidia/nemotron-3-super-120b-a12b. Requests can be routed across 3 providers, including DeepInfra, DigitalOcean and Nebius Token Factory, with automatic failover when an endpoint is unavailable.

NVIDIA: Nemotron 3 Super accepts text and returns text. It has a 1,000,000-token context window and a maximum output of 16,384 tokens.

On OpenRouter, NVIDIA: Nemotron 3 Super costs $0.085/M input tokens and $0.40/M output tokens. It was released on March 11, 2026.

More models from Nvidia

  • Nemotron 3.5 ASR Streaming Multilingual 0.6B
  • Nemotron 3.5 Lightning
  • Nemotron 3 Embed 1B (free)

Frequently asked questions

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer Mixture-of-Experts architecture with multi-token prediction (MTP), it delivers over 50% higher token generation compared to leading open models.

Nemotron 3 Super costs $0.085/M input tokens and $0.40/M output tokens.

Nemotron 3 Super has a 1,000,000 token context window. It supports up to 16,384 completion tokens.

Yes. Nemotron 3 Super accepts tools and tool_choice for function calling. It supports response_format for JSON output, without JSON-schema enforcement.

Nemotron 3 Super is served by 3 providers on OpenRouter: DeepInfra, DigitalOcean and Nebius Token Factory. Requests are routed to the best available provider, with automatic failover to the others, and you can pin or exclude providers with provider routing.

Nemotron 3 Super was released on March 11, 2026.