Skip to content
Not available in this workspace
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Collections/Discounted Models

Discounted AI Models on OpenRouter

Model rankings updated August 2026 based on real usage data.

This collection highlights AI models whose cheapest available provider currently offers a promotional discount. The displayed prices reflect the discounted provider pricing, making it easier to find models available for less through OpenRouter.

Provider discounts and promotional pricing can change, so this collection may update as offers change. OpenAI models are featured first, followed by other models ordered from largest discount to smallest. Every model is accessible through the OpenRouter API, giving you one integration for comparing models and providers.

Browse All ModelsCompare Models

AI Models with Provider Discounts

Favicon for openai

OpenAI: GPT-5.6 Luna

5.79T tokens

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.

by openai1.05M context$0.10/M input tokens$0.60/M output tokens50% off
Favicon for openai

OpenAI: GPT-5.6 Luna Pro

668B tokens

GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks.

Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

by openai1.05M context$0.10/M input tokens$0.60/M output tokens50% off
Favicon for openai

OpenAI: GPT-5.6 Terra

1T tokens

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic tasks where capability and cost need to be balanced, offering strong performance at roughly half the cost of Sol.

by openai1.05M context$1/M input tokens$6/M output tokens50% off
Favicon for openai

OpenAI: GPT-5.6 Terra Pro

50.8B tokens

GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on complex tasks.

Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

by openai1.05M context$1/M input tokens$6/M output tokens50% off
Favicon for inclusionai

inclusionAI: Ling-2.6-flash

213B tokens

Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency. It delivers performance comparable to state-of-the-art models at a similar scale while significantly reducing token usage across coding, document processing, and lightweight agent workflows.

by inclusionai262K context$0.01/M input tokens$0.03/M output tokens90% off
Favicon for upstage

Upstage: Solar Pro 4

292B tokens

Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive work, and coding.

by upstage524K context$0.03/M input tokens$0.12/M output tokens90% off
Favicon for z-ai

Z.ai: GLM 5.2

4.71T tokens

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

by z-ai1.05M context$0.308/M input tokens$0.968/M output tokens78% off
Favicon for deepseek

DeepSeek: DeepSeek V4 Pro 0423

3.03T tokens

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks.

Built on the same architecture as DeepSeek V4 Flash, it introduces a hybrid attention system for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for complex workloads such as full-codebase analysis, multi-step automation, and large-scale information synthesis, where both capability and efficiency are critical

by deepseek1.05M context$0.3969/M input tokens$0.7938/M output tokens77% off
Favicon for inclusionai

inclusionAI: Ling-2.6-1T

3.1B tokens

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast thinking” approach to reduce costs to roughly a quarter of comparable models while maintaining top-tier performance.

The model achieves state-of-the-art results on benchmarks such as AIME26 and SWE-bench Verified, and is well suited for advanced coding, complex reasoning, and large-scale agent workflows where both capability and efficiency are critical.

by inclusionai262K context$0.075/M input tokens$0.625/M output tokens75% off
Favicon for inclusionai

inclusionAI: Ring-2.6-1T

4.31B tokens

Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool use, and long-horizon task execution, delivering leading results on benchmarks including PinchBench, ClawEval, TAU2-Bench, and GAIA2-search.

With adaptive reasoning effort across high and xhigh modes, Ring-2.6-1T dynamically allocates reasoning budget based on task complexity. This enables stronger performance with lower token overhead, especially in tool-heavy and multi-turn agent workflows.

Ring-2.6-1T is designed for advanced coding agents, complex reasoning pipelines, and large-scale autonomous systems where execution quality, latency, and cost efficiency all matter.

by inclusionai262K context$0.075/M input tokens$0.625/M output tokens75% off
Favicon for inclusionai

Ling-3.0-flash

105B tokens

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token.

The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.

by inclusionai262K context$0.021/M input tokens$0.063/M output tokens65% off
Favicon for bytedance

ByteDance: Seedance 2.0 Mini

Seedance 2.0 Mini is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video with image, video, and audio inputs. It supports 480p and 720p output for 4-15 second videos. The number of tokens is given by (height of output video * width of output video * duration * 24) / 1024

by bytedancefrom $0.01345/second60% off
Favicon for meituan

Meituan: LongCat 2.0

8.6B tokens

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic workflows.

by meituan1.05M context$0.30/M input tokens$1.20/M output tokens60% off
Favicon for minimax

MiniMax: MiniMax M2.7

217B tokens

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments.

Trained for production-grade performance, M2.7 handles workflows such as live debugging, root cause analysis, financial modeling, and full document generation across Word, Excel, and PowerPoint. It delivers strong results on benchmarks including 56.2% on SWE-Pro and 57.0% on Terminal Bench 2, while achieving a 1495 ELO on GDPval-AA, setting a new standard for multi-agent systems operating in real-world digital workflows.

by minimax205K context$0.24/M input tokens$0.96/M output tokens60% off
Favicon for deepseek

DeepSeek: DeepSeek V4 Flash 0423

5.38T tokens

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.

The model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.

by deepseek1.05M context$0.06146/M input tokens$0.1229/M output tokens56% off
Favicon for qwen

Qwen: Qwen3 30B A3B Instruct 2507

31.2B tokens

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and agentic tool use. Post-trained on instruction data, it demonstrates competitive performance across reasoning (AIME, ZebraLogic), coding (MultiPL-E, LiveCodeBench), and alignment (IFEval, WritingBench) benchmarks. It outperforms its non-instruct variant on subjective and open-ended tasks while retaining strong factual and coding performance.

by qwen262K context$0.04815/M input tokens$0.1931/M output tokens55% off
Favicon for google

Google Gemini Flash Latest

This model always redirects to the latest model in the Google Gemini Flash family.

by google1.05M context$0.375/M input tokens$1.875/M output tokens50% off
Favicon for google

Google: Gemini 3.7 Flash

387B tokens

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step problem solving.

by google1.05M context$0.375/M input tokens$1.875/M output tokens50% off
Favicon for google

Google: Gemini 3.7 Flash (batch)

1.51B tokens

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step problem solving.

by google1.05M context$0.1875/M input tokens$0.9375/M output tokens50% off
Favicon for moonshotai

MoonshotAI: Kimi K2.6

141B tokens

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and can convert prompts and visual inputs into production-ready interfaces. Its agent swarm architecture scales to hundreds of parallel sub-agents for autonomous task decomposition - delivering documents, websites, and spreadsheets in a single run without human oversight.

by moonshotai262K context$0.5415/M input tokens$2.28/M output tokens43% off
Favicon for poolside

Poolside: Laguna XS 2.1

12.6B tokens

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step forward from their Laguna XS.2 model (released in April 2026). It combines tool calling and reasoning capabilities with a compact footprint, offering a 256K context window and up to 32K output tokens. Quantized to FP8 for fast, cost-efficient agentic coding workflows.

Laguna XS 2.1 is designed for software engineering and agentic coding use cases, and you are responsible for confirming that it is appropriate for your intended application. Laguna XS 2.1 is subject to the OpenMDW-1.1 License, and should be used consistently with Poolside's Acceptable Use Policy. We advise against circumventing Laguna XS 2.1 safety guardrails without implementing substantially equivalent mitigations appropriate for your use case.

Please report security vulnerabilities or safety concerns to [email protected].

If you are using Laguna XS 2.1 for free, we may use your inputs and outputs to train and improve our models.

by poolside262K context$0.06/M input tokens$0.12/M output tokens40% off
Favicon for z-ai

Z.ai: GLM 5

57.2B tokens

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading closed-source models. With advanced agentic planning, deep backend reasoning, and iterative self-correction, GLM-5 moves beyond code generation to full-system construction and autonomous execution.

by z-ai205K context$0.60/M input tokens$1.92/M output tokens40% off
Favicon for z-ai

Z.ai: GLM 5.1

81.9B tokens

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on a single task for more than 8 hours, autonomously planning, executing, and improving itself throughout the process, ultimately delivering complete, engineering-grade results.

by z-ai205K context$0.896/M input tokens$2.816/M output tokens36% off
Favicon for deepseek

DeepSeek V4 Flash Latest

This model always redirects to the latest model in the DeepSeek V4 Flash family.

by deepseek1.05M context$0.0603/M input tokens$0.1206/M output tokens33% off
Favicon for deepseek

DeepSeek: DeepSeek V4 Flash 0731

12.3T tokens

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.

by deepseek1.05M context$0.0603/M input tokens$0.1206/M output tokens33% off
Favicon for xiaomi

Xiaomi: MiMo-V2.5-Pro

528B tokens

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro. It can independently and autonomously complete professional tasks that would take human experts days or weeks, involving more than a thousand tool calls. Its context length of up to 1M makes it well suited for integration with a wide range of agent frameworks.

by xiaomi1.05M context$0.3045/M input tokens$0.609/M output tokens30% off
Favicon for deepseek

DeepSeek: DeepSeek V3.2

459B tokens

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.

Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs

by deepseek164K context$0.2088/M input tokens$0.3096/M output tokens28% off
Favicon for minimax

MiniMax: MiniMax M2

957M tokens

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning, tool use, and multi-step task execution while maintaining low latency and deployment efficiency.

The model excels in code generation, multi-file editing, compile-run-fix loops, and test-validated repair, showing strong results on SWE-Bench Verified, Multi-SWE-Bench, and Terminal-Bench. It also performs competitively in agentic evaluations such as BrowseComp and GAIA, effectively handling long-horizon planning, retrieval, and recovery from execution errors.

Benchmarked by Artificial Analysis, MiniMax-M2 ranks among the top open-source models for composite intelligence, spanning mathematics, science, and instruction-following. Its small activation footprint enables fast inference, high concurrency, and improved unit economics, making it well-suited for large-scale agents, developer assistants, and reasoning-driven applications that require responsiveness and cost efficiency.

To avoid degrading this model's performance, MiniMax highly recommends preserving reasoning between turns. Learn more about using reasoning_details to pass back reasoning in our docs.

by minimax205K context$0.255/M input tokens$1.02/M output tokens15% off
Favicon for xiaomi

Xiaomi: MiMo-V2.5

4.32T tokens

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass, making it ideal for integration with agent frameworks where strong reasoning, rich perception, and cost efficiency all matter.

by xiaomi1.05M context$0.119/M input tokens$0.238/M output tokens15% off
Favicon for deepseek-ai

DeepSeek: DeepSeek V3

29B tokens

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations reveal that the model outperforms other open-source models and rivals leading closed-source models.

For model details, please visit the DeepSeek-V3 repo for more information, or see the launch announcement.

by deepseek-ai164K context$0.2574/M input tokens$1.029/M output tokens10% off
Favicon for poolside

Poolside: Laguna S 2.1

54.8B tokens

Laguna S 2.1 is the latest coding agent model from Poolside. Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and 40.4% on DeepSWE, making it one of the strongest coding models in its category. Open-weight under the OpenMDW-1.1 license.

Laguna S 2.1 is designed for software engineering and agentic coding use cases, and you are responsible for confirming that it is appropriate for your intended application. Laguna S 2.1 is subject to the OpenMDW-1.1 License, and should be used consistently with Poolside's Acceptable Use Policy. We advise against circumventing Laguna S 2.1 safety guardrails without implementing substantially equivalent mitigations appropriate for your use case.

Please report security vulnerabilities or safety concerns to [email protected].

If you are using Laguna S 2.1 for free, we may use your inputs and outputs to train and improve our models.

by poolside1.05M context$0.09/M input tokens$0.18/M output tokens10% off
Favicon for tencent

Tencent: Hy3

11.4T tokens

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort: a direct no-think mode by default, plus low and high chain-of-thought modes for complex math, coding, and multi-step problems. With a 256K context window, Hy3 targets long-horizon tasks, including improved coreference resolution, multi-turn constraint tracking, and stable tool-calling that generalizes across agent scaffoldings.

Tencent positions it as a reliable, cost-effective option across coding, document processing, financial analysis, game development, and frontend design, with a strong emphasis on grounded, anti-hallucination behavior that answers when grounded and flags when evidence is missing rather than fabricating.

by tencent262K context$0.1254/M input tokens$0.5016/M output tokens5% off