Skip to content
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

Models

Models

CompareDiscover Models
Favicon for anthropic
Favicon for openai
  • Favicon for meta
    Meta: Muse Spark 1.3 ContributorMuse Spark 1.3 Contributor
    18+
    1.41B tokens

    Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information across extended tasks, work through conflicting inputs, and request clarification or confirmation when needed. Prompts and outputs may be used to improve Meta’s products.

    by metaSep 2, 20261.05M context$0.10/M input tokens$0.20/M output tokens
  • Favicon for meta
    Meta: Muse Spark 1.3Muse Spark 1.3
    18+
    1.68B tokens

    Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through conflicting inputs, and request clarification or confirmation when needed, with an emphasis on concise execution.

    by metaSep 2, 20261.05M context$1.25/M input tokens$4.25/M output tokens
  • Favicon for google
    Google: Gemini 3.8 FlashGemini 3.8 Flash
    50% off
    37.4B tokens

    Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

    by googleSep 2, 20261.05M context$0.75/M input tokens$3.75/M output tokens
  • Favicon for google
    Google: Gemini 3.8 Flash (batch)Gemini 3.8 Flash (batch)
    50% off
    Batch variant
    42.6M tokens

    Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

    by googleSep 2, 20261.05M context$0.375/M input tokens$1.875/M output tokens
  • Favicon for minimax
    MiniMax: H3 MaxH3 Max
    3 hours

    MiniMax H3 Max is a video-generation model from MiniMax, jointly released with fal.ai. Derived through additional training from MiniMax H3, it is designed for faster text-to-video and image-to-video generation with controlled first-frame or last-frame keyframes.

    by minimaxSep 2, 2026from $0.05/second
  • Favicon for anthropic
    Anthropic: Claude Fable 5.1 (batch)Claude Fable 5.1 (batch)Batch variant
    1.24M tokens

    Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual code generation, and finance and analysis tasks in particular. It also tends to be more concise than Fable 5 in its plans and summaries. We recommend testing it as a direct upgrade wherever you use Fable 5 today, and alongside Opus 5 on reasoning-heavy tasks.

    by anthropicSep 1, 20261M context$5/M input tokens$25/M output tokens
  • Favicon for anthropic
    Anthropic: Claude Fable 5.1Claude Fable 5.1
    65.5B tokens

    Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual code generation, and finance and analysis tasks in particular. It also tends to be more concise than Fable 5 in its plans and summaries. We recommend testing it as a direct upgrade wherever you use Fable 5 today, and alongside Opus 5 on reasoning-heavy tasks.

    by anthropicSep 1, 20261M context$10/M input tokens$50/M output tokens
  • Favicon for inception
    Inception: Mercury 2.5 PreviewMercury 2.5 Preview
    80% off
    17.2B tokens

    Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving 1,107 tokens/sec on standard GPUs. It delivers a 10+ point jump in intelligence over Mercury 2, comparable quality to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Mercury 2.5 supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output. It's built for production workloads where latency compounds: search agents, voice pipelines, and coding subagents.

    by inceptionAug 31, 2026260K context$0.04/M input tokens$0.15/M output tokens
  • Favicon for ibm-granite
    IBM: Granite 4.2 8BGranite 4.2 8B
    1.26B tokens

    Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort, and non-thinking modes. The model supports 12 languages: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese.

    by ibm-graniteAug 31, 2026131K context$0.10/M input tokens$0.15/M output tokens
  • Favicon for tencent
    Tencent: Hy4 previewHy4 preview
    7.93T tokens

    Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that require planning, context continuity, and sustained multi-step execution.

    by tencentAug 28, 20261.05M context$0.834/M input tokens$2.501/M output tokens
  • Favicon for alibaba
    Alibaba: Wan 3.0 PrimeWan 3.0 Prime
    3 hours

    Wan 3.0 Prime is a fast-mode variant of Wan 3.0 from Alibaba. It supports text-to-video and first-frame image-to-video generation.

    by alibabaAug 27, 2026from $0.068/second
  • Favicon for inclusionai
    Ling 3.0 Flash Fin (free)Ling 3.0 Flash Fin (free)Free variant
    557B tokens

    Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment workflows that require complex multi-step tasks and long-horizon planning and execution, while retaining general capabilities in reasoning, coding, and mathematics.

    by inclusionaiAug 27, 2026262K context$0/M input tokens$0/M output tokens
  • Favicon for z-ai
    Z.ai: GLM Flash LatestGLM Flash Latest
    50% off

    This model always redirects to the latest model in the GLM Flash family.

    by z-aiAug 27, 20261.31M context$0.075/M input tokens$0.25/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.8 FlashQwen3.8 Flash
    134B tokens

    Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

    by qwenAug 26, 20261M context$0.15/M input tokens$0.47/M output tokens
  • Favicon for meta
    Meta: Muse ImageMuse Image
    18+
    282M tokens

    Muse Image is an agentic image generation model from Meta that generates and edits images from text and reference images. Unlike single-pass image models, it reasons before it renders, breaking down multi-part prompts and refining its output within the chain of thought, and invokes web search for factual accuracy on knowledge-intensive prompts. The model supports text-to-image generation, targeted image editing, multi-image composition, reference-image conditioning for style and subject consistency across a series, and precise text rendering within generated images. Iterative editing works by passing the previous output image back with a new instruction.

    by metaAug 26, 202666K context$0.01/image
  • Favicon for z-ai
    Z.ai: GLM 5.3 FlashGLM 5.3 Flash
    50% off
    11.6T tokens

    GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

    by z-aiAug 26, 20261.31M context$0.075/M input tokens$0.25/M output tokens
  • Favicon for z-ai
    Z.ai: GLM 5.3 Flash (batch)GLM 5.3 Flash (batch)Batch variant
    264M tokens

    GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

    by z-aiAug 26, 20261.05M context$0.15/M input tokens$0.50/M output tokens
  • Favicon for recraft
    Recraft: Recraft V4 Styles ProRecraft V4 Styles Pro
    11.3M tokens

    Recraft V4 Styles Pro is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's rendering technique, colour, texture, and composition rather than editing it. It produces raster images at approximately 2K resolution. Note: 1 to 10 reference images in PNG, JPG or WEBP, each at least 256 px on its shortest edge. Pricing has two parts: $0.10 per image plus a one-time $0.005 style-creation charge per request, regardless of how many images that request returns. A single-image request costs $0.105, and a six-image request costs $0.605.

    by recraftAug 26, 202666K contextfrom $0.10/image
  • Favicon for recraft
    Recraft: Recraft V4 Styles VectorRecraft V4 Styles Vector
    3.14M tokens

    Recraft V4 Styles Vector is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's rendering technique, colour, texture, and composition rather than editing it. It produces SVG output for graphics that need to scale cleanly. Note: 1 to 10 reference images in PNG, JPG or WEBP, each at least 256 px on its shortest edge. Pricing has two parts: $0.05 per image plus a one-time $0.005 style-creation charge per request, regardless of how many images that request returns. A single-image request costs $0.055, and a six-image request costs $0.305.

    by recraftAug 26, 202666K contextfrom $0.05/image
  • Favicon for recraft
    Recraft: Recraft V4 Styles Pro VectorRecraft V4 Styles Pro Vector
    1.27M tokens

    Recraft V4 Styles Pro Vector is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's rendering technique, colour, texture, and composition rather than editing it. It produces SVG output for graphics that need to scale cleanly. Note: 1 to 10 reference images in PNG, JPG or WEBP, each at least 256 px on its shortest edge. Pricing has two parts: $0.12 per image plus a one-time $0.005 style-creation charge per request, regardless of how many images that request returns. A single-image request costs $0.125, and a six-image request costs $0.725.

    by recraftAug 26, 202666K contextfrom $0.12/image