Skip to content
  • Models
  • Rankings
  • Ori
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Tools
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status
  • AI Site Map

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

Models

CompareDiscover Models
Favicon for anthropic
Favicon for openai

Models

CompareDiscover Models
Favicon for anthropic
Favicon for openai
  • Favicon for openai
    OpenAI: GPT-6 Luna DecisionsGPT-6 Luna Decisions
    477M tokens

    GPT-6 Luna Decisions is GPT-6 Luna served through OpenAI's Decisions API. Instead of generating text, it reads the content passed as state (text, JSON, or images) and returns typed, probabilistic answers to named questions in a single request: the probability of yes for a yes/no question (noul), a probability for every option plus the most likely one (choice), or a probability for every level of an ordered rubric plus the expected score (score). It is built for classification, routing, moderation, and rubric grading where application code thresholds the returned numbers rather than parsing a chat reply. A request can carry up to 200 questions about the same content.

    by openaiOct 6, 20261.05M context$0.10/M input tokens$0/M output tokens
  • Favicon for x-ai
    SpaceXAI: Grok Imagine Video 1.5 LiteGrok Imagine Video 1.5 Lite
    2 hours

    Grok Imagine Video 1.5 Lite is a faster, lower-cost video generation model from SpaceXAI, distilled from Grok Imagine Video 1.5. It supports text-to-video and image-to-video, trading some quality for speed and price. 1080p output is rendered at 720p and upscaled.

    by x-aiOct 6, 2026from $0.02/second
  • Favicon for google
    Google: Nano Banana 2.1Nano Banana 2.1
    61.8M tokens

    Nano Banana 2.1 (Gemini Nano Banana 2.1) is Google's image generation and editing model on the Flash tier, succeeding Nano Banana 2 and Nano Banana Pro. It improves product recontextualization, mask- and ink-based editing, and factual accuracy, and renders photorealistic skin tones, detailed materials, lighting, and coherent backgrounds. It accepts text and image inputs, returns images with optional text, and supports 1K, 2K, and 4K output plus extended aspect ratios via the image_config API Parameter.

    by googleOct 6, 202666K context$1.50/M input tokens$30/M output tokens
  • Favicon for mistralai
    Mistral: Mistral Large 4Mistral Large 4
    50% off
    11.2B tokens

    Mistral Large 4 is a frontier multimodal (text and image input) model from Mistral AI built for reasoning, coding, and agentic workloads. It offers a 512K-token context window with up to 256K output tokens, and supports tool calling and structured outputs.

    by mistralaiOct 6, 2026524K context$0.68/M input tokens$2.09/M output tokens
  • Favicon for tencent
    Tencent: Hy Image 3.5 PreviewHy Image 3.5 Preview
    84M tokens

    Hy Image 3.5 Preview is a unified image generation and editing model from Tencent. It handles text-to-image, image-to-image, and multi-turn editing through one endpoint, taking up to 20 reference images per request and producing output up to 4K. Built on the 80B mixture-of-experts Hy Image 3.0 base, it is particularly strong at subject consistency across edits, prompt-faithful composition, and rendering Chinese and English text inside images.

    by tencentOct 5, 2026100K context$1.60/M tokens
  • Favicon for inclusionai
    inclusionAI: Ling 3.1 FlashLing 3.1 Flash
    309B tokens

    Ling 3.1 Flash is a hybrid reasoning mixture-of-experts model from inclusionAI, with 25B active parameters out of 560B total.

    by inclusionaiOct 2, 2026262K context$0/M input tokens$0/M output tokens
  • Favicon for bytedance-seed
    ByteDance Seed: Seedream 5.0 FlashSeedream 5.0 Flash
    869M tokens

    Seedream 5.0 Flash is an image generation and editing model from ByteDance Seed. It is the fast, cost-efficient tier of the Seedream 5.0 family, suited for high-volume production and interactive editing workflows that need precise edits at low latency.

    by bytedance-seedOct 1, 2026from $0.018/image
  • Favicon for perplexity
    Perplexity: Decider V1 27BDecider V1 27B
    5.35B tokens
    Legal (#38)
    Trivia (#32)

    Decider V1 27B is a decision model from Perplexity. Instead of generating text, it reads content passed as state and returns typed, probabilistic answers to one or more named questions in a single request: the probability of yes for a yes/no question (noul), a probability for every option plus the most likely one (choice), or a probability for every level of an ordered rubric plus the expected score (score). It is built for classification, routing, moderation, and rubric grading where application code thresholds the returned numbers rather than parsing a chat reply. A request can carry up to 128 questions about the same content. On OpenRouter it currently accepts text and JSON state; image inputs are not yet supported.

    by perplexityOct 1, 2026262K context$0.04/M input tokens$0/M output tokens
  • Favicon for black-forest-labs
    Black Forest Labs: FLUX.3 ImageFLUX.3 Image
    50% off
    210M tokens

    FLUX.3 Image is Black Forest Labs' flagship image generation and editing model. It handles text-to-image and multi-reference editing with up to 10 input images, and renders at fixed resolution tiers from 768 up to 4K with a selectable aspect ratio. Pricing is a flat per-image rate that scales with the chosen resolution tier.

    by black-forest-labsOct 1, 202647K contextfrom $0.0205/image
  • Favicon for liquid
    LiquidAI: d1d1
    5.38B tokens

    d1 is Liquid AI's structured decision model, served as a System One endpoint. Send a state along with typed questions, and it returns a choice, a score, or a yes/no answer, each with a probability taken directly from the model rather than written out as text. It uses the same /v1/systemone schema as other OpenRouter Decisions models, so it suits routing, classification, and policy checks that need a fast, scored answer instead of prose.

    by liquidOct 1, 202666K context$0.04/M input tokens$0/M output tokens
  • Favicon for apodexFavicon for apodex
    Apodex: Apodex 1.1 Mini (free)Apodex 1.1 Mini (free)Free variant
    165B tokens

    Apodex 1.1 Mini is a reasoning-first model from Apodex, built for complex, long-horizon research and forecasting tasks. It works directly with files, data, code, and tools to produce verifiable results, and is designed for agentic research workflows where answers need to be grounded in evidence.

    by apodexOct 1, 2026262K context$0/M input tokens$0/M output tokens
  • Favicon for cloudflare
    Cloudflare: Clef FlashClef Flash
    3.69B tokens
    Trivia (#6)

    Clef-flash is the fast 9B member of Cloudflare's open-source Clef decision model family, a fine-tune of Qwen3.5-9B served on Workers AI. It turns a state (text or structured JSON) plus a schema of typed questions into decisions, returning a calibrated probability for every allowed option of every question in a single forward pass instead of generating tokens. Use it for low-latency classification, routing, scoring, and guardrails through the Decisions API. Note: Workers AI currently truncates long text state to roughly the first 2K tokens, so content beyond that is not read; images are counted separately.

    by cloudflareOct 1, 202666K context$0.09/M input tokens$0/M output tokens
  • Favicon for cloudflare
    Cloudflare: ClefClef
    2.76B tokens

    Clef is Cloudflare's open-source 27B multimodal decision model, a fine-tune of Qwen3.8-27B served on Workers AI. It turns a state (text or structured JSON) plus a schema of typed questions into decisions, returning a calibrated probability for every allowed option of every question in a single forward pass instead of generating tokens. Use it for classification, routing, scoring, guardrails, and agentic control flow through the Decisions API. Note: Workers AI currently truncates long text state to roughly the first 2K tokens, so content beyond that is not read; images are counted separately.

    by cloudflareOct 1, 202666K context$0.24/M input tokens$0/M output tokens
  • Favicon for microsoft
    Microsoft AI: MAI-Voice-2.1-FlashMAI-Voice-2.1-Flash
    2M tokens

    MAI-Voice-2.1-Flash is a low-latency text-to-speech model from Microsoft AI, optimized for real-time responsiveness. It produces natural, expressive speech across 23 languages, with human-like intonation, rhythm, and emotional nuance. It is suited for voice agents, assistants, call centers, and other interactive applications where latency and cost matter most. On OpenRouter, set voice to a full voice ID with the model suffix, such as "en-US-Harper:MAI-Voice-2.1-Flash". A voice's locale sets the synthesis language. Set response_format to "mp3" or "pcm" (24 kHz mono). Harper and Grant support the agent, customer-call-center, educational, and narrator speaking styles, and many locale voices add emotion styles such as excited, happy, sad, and whispering. The full list of voices is in the supported_voices field of the models API. See the text-to-speech guide.

    by microsoftOct 1, 2026$15/M characters
  • Favicon for microsoft
    Microsoft AI: MAI-Voice-2.1MAI-Voice-2.1
    1.33M tokens

    MAI-Voice-2.1 is Microsoft AI's highest-fidelity, most expressive text-to-speech model. It produces natural, studio-grade speech across 23 languages, with detailed prosody, nuanced expressiveness, and speaker consistency over long-form content. It is suited for audiobooks, podcasts, lectures, narration, and brand audio where maximum voice quality matters. The model prioritizes naturalness and expressivity over latency-critical generation. On OpenRouter, set voice to a full voice ID with the model suffix, such as "en-US-Harper:MAI-Voice-2.1". A voice's locale sets the synthesis language. Set response_format to "mp3" or "pcm" (24 kHz mono). Harper and Grant support the agent, customer-call-center, educational, and narrator speaking styles, and many locale voices add emotion styles such as excited, happy, sad, and whispering. The full list of voices is in the supported_voices field of the models API. See the text-to-speech guide.

    by microsoftOct 1, 2026$22/M characters
  • Favicon for unbiased
    Pareto 26.10 PreviewPareto 26.10 Preview
    5.86B tokens

    Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks. This is a preview of the next Pareto version and may change without notice; use pareto-26.9 for stable behaviour.

    by unbiasedOct 1, 20261.05M context$0.80/M input tokens$3.20/M output tokens
  • Favicon for heygen
    HeyGen: HeyGen VideoHeyGen Video
    50% off
    124 hours

    HeyGen Video is a general-purpose video generation model from HeyGen. It renders short clips with synthesized audio (dialogue, ambience, and sound effects) in a single call, working from a text prompt, from a first-frame image, or from a set of image, video, and audio references that the prompt can address directly. It is aimed at short-form business video with a lot of motion in frame: brand moments, product and equipment demos, and training or onboarding content. It is not an avatar product. Unlike HeyGen's Avatar IV and Video Agent, there is no talking-head, script, or lip-sync step; the prompt and references drive the whole scene.

    by heygenSep 30, 2026from $0.01/second
  • Favicon for togethercomputer
    Together: Tev1 4B ExperimentalTev1 4B Experimental
    2.18B tokens

    Tev1 4B Experimental is an experimental decision model from Together AI, a supervised fine-tune of Qwen3.5-4B trained to choose one option from a structured state, question, and list of 2-24 labeled options. It keeps Qwen's standard next-token head, so it is served through the regular chat completions API rather than a dedicated decisions runtime. Send a system instruction followed by a JSON decision containing state, question, and options. The model returns a single option letter, which application code maps back to the option key. Recommended settings are temperature: 0, max_tokens: 8, and thinking disabled. It is intended for routing, classification, and policy checks, not generic chat. Together publishes the full data recipe and training code so teams can fine-tune their own variant.

    by togethercomputerSep 30, 202633K context$0.042/M input tokens$0/M output tokens
  • Favicon for inception
    Inception: Mercury Decide (free)Mercury Decide (free)Free variant
    6.21B tokens

    Mercury Decide is Inception's structured decision model, served as a System One endpoint. Send a state along with typed questions, and it returns a choice, a score, or a yes/no answer, each with a calibrated probability taken directly from the model rather than written out as text, so output tokens are free. It makes up to 14 decisions per second and reports how certain it is, so a decision system can run it on every case and escalate the unsure ones to a human. Mercury Decide uses the same /v1/systemone schema as Jev.

    by inceptionSep 30, 202633K context$0/M input tokens$0/M output tokens
  • Favicon for voyageai
    VoyageAI by MongoDB: rerank-3-litererank-3-lite
    7.07B tokens

    rerank-3-lite is a reranker optimized for both latency and quality and a drop-in upgrade to rerank-2.5-lite, improving on it by 0.94% NDCG@10 on average across domain evaluations and by 1.86% on long-document evaluations, with code retrieval gains of 2.77% atop voyage-3-large and 2.59% atop voyage-4-large. It matches the retrieval quality of rerank-2.5, and across 93 retrieval datasets it outperforms Cohere Rerank v4.0 Pro by 1.44% and Qwen3-Reranker-8B by 2.61%. The model supports a combined context length of 32K tokens per query-document pair, including up to 8K tokens for the query, enabling more accurate retrieval over longer documents. Additionally, rerank-3-lite supports instruction following, allowing users to guide relevance scoring through natural language prompts. Learn more about rerank-3-lite here: blog.voyageai.com/2026/09/30/rerank-3

    by voyageaiSep 30, 202632K context$0.02/M tokens