Skip to content
Not available in this workspace
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Collections/Embedding Models

Text Embedding Models

Model rankings updated August 2026 based on real usage data.

Embedding models convert text into dense vectors that power semantic search, retrieval-augmented generation (RAG), clustering, and similarity matching. This collection ranks embedding models by their usage on OpenRouter over the past week. The current top models are Text Embedding 3 Small, Qwen3 Embedding 8B, and Embed V1 0.6B. Access them through one API and compare dimensions, pricing, and performance without managing multiple provider integrations.

Browse All ModelsCompare Models

Embedding Models on OpenRouter

Favicon for openai

OpenAI: Text Embedding 3 Small

128B tokens

text-embedding-3-small is OpenAI's improved, more performant version of the ada embedding model. Embeddings are a numerical representation of text that can be used to measure the relatedness between two pieces of text. Embeddings are useful for search, clustering, recommendations, anomaly detection, and classification tasks.

by openai8K context$0.02/M tokens
Favicon for qwen

Qwen: Qwen3 Embedding 8B

108B tokens

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining.

by qwen33K context$0.01/M tokens
Favicon for perplexity

Perplexity: Embed V1 0.6B

39.6B tokens

pplx-embed-v1-0.6B is one of Perplexity's state-of-the-art text embedding models built for real-world, web-scale retrieval. pplx-embed-v1 is optimized for standard dense text retrieval with the 0.6B parameter model targeting lightweight, low-latency embedding generation.

by perplexity32K context$0.004/M tokens
Favicon for baai

BAAI: bge-m3

30.9B tokens

The bge-m3 embedding model encodes sentences, paragraphs, and long documents into a 1024-dimensional dense vector space, delivering high-quality semantic embeddings optimized for multilingual retrieval, semantic search, and large-context applications.

by baai8K context$0.01/M tokens
Favicon for google

Google: Gemini Embedding 001

22.9B tokens

gemini-embedding-001 provides a unified cutting edge experience across domains, including science, legal, finance, and coding. This embedding model has consistently held a top spot on the Massive Text Embedding Benchmark (MTEB) Multilingual leaderboard since the experimental launch in March.

by google20K context$0.15/M tokens
Favicon for qwen

Qwen: Qwen3 Embedding 4B

22.8B tokens

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining.

by qwen33K context$0.02/M tokens
Favicon for google

Google: Gemini Embedding 2

16.6B tokens

Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports input context up to 8,192 tokens and flexible output dimensions from 128 to 3,072 (recommended: 768, 1536, or 3,072). Designed for cross-modal similarity — you can embed a text query and retrieve the most relevant images, or vice versa — making it well-suited for multimodal search, recommendation, and document understanding pipelines.

by google8K context$0.20/M tokens
Favicon for openai

OpenAI: Text Embedding 3 Large

15.8B tokens

text-embedding-3-large is OpenAI's most capable embedding model for both english and non-english tasks. Embeddings are a numerical representation of text that can be used to measure the relatedness between two pieces of text. Embeddings are useful for search, clustering, recommendations, anomaly detection, and classification tasks.

by openai8K context$0.13/M tokens
Favicon for nvidia

NVIDIA: Nemotron 3 Embed 1B (free)

6.99B tokens

NVIDIA Nemotron 3 Embed 1B is an open text embedding model from NVIDIA, optimized for high-throughput, low-latency retrieval. It is suited for enterprise search, RAG, code retrieval, and agentic retrieval workflows, retaining more than 95% of the 8B model’s accuracy in a smaller deployment footprint.

by nvidia33K context$0/M input tokens$0/M output tokens
Favicon for mistralai

Mistral: Mistral Embed 2312

5.55B tokens

Mistral Embed is a specialized embedding model for text data, optimized for semantic search and RAG applications. Developed by Mistral AI in late 2023, it produces 1024-dimensional vectors that effectively capture semantic relationships in text.

by mistralai8K context$0.10/M tokens