Skip to content
  • Models
  • Rankings
  • Ori
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Tools
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status
  • AI Site Map

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for perplexity

Perplexity: Llama 3.1 Sonar 70B Online

perplexity/llama-3.1-sonar-large-128k-online

Llama 3.1 Sonar is Perplexity's latest model family. It surpasses their earlier Sonar models in cost-efficiency, speed, and performance.

This is the online version of the offline chat model. It is focused on delivering helpful, up-to-date, and factual responses. #online

Modalities
Context
127K
Released
Aug 1, 2024
ActivityFAQExplore

Activity

Token volume and request traffic to this model over time.

Explore more models

AI Model RankingsRanking

Frequently asked questions

Llama 3.1 Sonar is Perplexity's latest model family. It surpasses their earlier Sonar models in cost-efficiency, speed, and performance. This is the online version of the offline chat model. It is focused on delivering helpful, up-to-date, and factual responses. #online

Llama 3.1 Sonar 70B Online has a 127,072 token context window.

Decider V1 27B, Sonar Pro Search, Sonar Reasoning Pro and 3 more are other text models from Perplexity.

Llama 3.1 Sonar 70B Online was released on August 1, 2024.

More models from Perplexity

Decider V1 27B

Decider V1 27B is a decision model from Perplexity. Instead of generating text, it reads content passed as state and returns typed, probabilistic answers to one or more named questions in a single request: the probability of yes for a yes/no question (noul), a probability for every option plus the most likely one (choice), or a probability for every level of an ordered rubric plus the expected score (score).

It is built for classification, routing, moderation, and rubric grading where application code thresholds the returned numbers rather than parsing a chat reply. A request can carry up to 128 questions about the same content. On OpenRouter it currently accepts text and JSON state; image inputs are not yet supported.

Decisions$0.04 / $0
Embed V1 4B

pplx-embed-v1 -4B is one of Perplexity's state-of-the-art text embedding models built for real-world, web-scale retrieval. pplx-embed-v1 is optimized for standard dense text retrieval with the 4B parameter model maximizing retrieval quality.

Embeddings$0.03/M tokens
Embed V1 0.6B

pplx-embed-v1-0.6B is one of Perplexity's state-of-the-art text embedding models built for real-world, web-scale retrieval. pplx-embed-v1 is optimized for standard dense text retrieval with the 0.6B parameter model targeting lightweight, low-latency embedding generation.

Embeddings$0.004/M tokens
Sonar Pro Search

Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based on tokens plus $18 per thousand requests. This model powers the Pro Search mode on the Perplexity platform.

Sonar Pro Search adds autonomous, multi-step reasoning to Sonar Pro. So, instead of just one query + synthesis, it plans and executes entire research workflows using tools.

Text200K context$3 / $15
Sonar Reasoning Pro

Note: Sonar Pro pricing includes Perplexity search pricing. See details here

Sonar Reasoning Pro is a premier reasoning model powered by DeepSeek R1 with Chain of Thought (CoT). Designed for advanced use cases, it supports in-depth, multi-step queries with a larger context window and can surface more citations per search, enabling more comprehensive and extensible responses.

Text128K context$2 / $8
Sonar Pro

Note: Sonar Pro pricing includes Perplexity search pricing. See details here

For enterprises seeking more advanced capabilities, the Sonar Pro API can handle in-depth, multi-step queries with added extensibility, like double the number of citations per search as Sonar on average. Plus, with a larger context window, it can handle longer and more nuanced searches and follow-up questions.

Text200K context$3 / $15
Sonar Deep Research

Sonar Deep Research is a research-focused model designed for multi-step retrieval, synthesis, and reasoning across complex topics. It autonomously searches, reads, and evaluates sources, refining its approach as it gathers information. This enables comprehensive report generation across domains like finance, technology, health, and current events.

Notes on Pricing (Source)

Text128K context$2 / $8
R1 1776

R1 1776 is a version of DeepSeek-R1 that has been post-trained to remove censorship constraints related to topics restricted by the Chinese government. The model retains its original reasoning capabilities while providing direct responses to a wider range of queries. R1 1776 is an offline chat model that does not use the perplexity search subsystem.

The model was tested on a multilingual dataset of over 1,000 examples covering sensitive topics to measure its likelihood of refusal or overly filtered responses. Evaluation Results Its performance on math and reasoning benchmarks remains similar to the base R1 model. Reasoning Performance

Read more on the Blog Post

Text128K context
Sonar Reasoning

Sonar Reasoning is a reasoning model provided by Perplexity based on DeepSeek R1.

It allows developers to utilize long chain of thought with built-in web search. Sonar Reasoning is uncensored and hosted in US datacenters.

Text127K context
Sonar

Sonar is lightweight, affordable, fast, and simple to use — now featuring citations and the ability to customize sources. It is designed for companies seeking to integrate lightweight question-and-answer features optimized for speed.

Text127K context$1 / $1
Llama 3.1 Sonar 8B Online

Llama 3.1 Sonar is Perplexity's latest model family. It surpasses their earlier Sonar models in cost-efficiency, speed, and performance.

This is the online version of the offline chat model. It is focused on delivering helpful, up-to-date, and factual responses. #online

Text127K context
Llama3 Sonar 8B

Llama3 Sonar is Perplexity's latest model family. It surpasses their earlier Sonar models in cost-efficiency, speed, and performance.

This is a normal offline LLM, but the online version of this model has Internet access.

Text33K context
Llama3 Sonar 70B Online

Llama3 Sonar is Perplexity's latest model family. It surpasses their earlier Sonar models in cost-efficiency, speed, and performance.

This is the online version of the offline chat model. It is focused on delivering helpful, up-to-date, and factual responses. #online

Text28K context
Llama3 Sonar 70B

Llama3 Sonar is Perplexity's latest model family. It surpasses their earlier Sonar models in cost-efficiency, speed, and performance.

This is a normal offline LLM, but the online version of this model has Internet access.

Text33K context
Llama3 Sonar 8B Online

Llama3 Sonar is Perplexity's latest model family. It surpasses their earlier Sonar models in cost-efficiency, speed, and performance.

This is the online version of the offline chat model. It is focused on delivering helpful, up-to-date, and factual responses. #online

Text28K context