Skip to content
  • Models
  • Rankings
  • Ori
ElevenLabs launch offer: every ElevenLabs model is 50% off through October 19, 2026. See ElevenLabs models

Footer

OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Ori
  • Collections
  • Providers
  • Tools
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status
  • AI Site Map

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for featherless

Qrwkv 72B

featherless/qwerky-72b:free

Model weights

Qrwkv-72B is a linear-attention RWKV variant of the Qwen 2.5 72B model, optimized to significantly reduce computational cost at scale. Leveraging linear attention, it achieves substantial inference speedups (>1000x) while retaining competitive accuracy on common benchmarks like ARC, HellaSwag, Lambada, and MMLU. It inherits knowledge and language support from Qwen 2.5, supporting approximately 30 languages, making it suitable for efficient inference in large-context applications.

Modalities
Price
Free
Context
33K
Released
Mar 20, 2025
Knowledge Cutoff
Jun 2024
ActivityFAQExplore

Activity

Token volume and request traffic to this model over time.

Explore more models

Free AI Models on OpenRouterCollectionAI Model RankingsRanking

Frequently asked questions

Qrwkv-72B is a linear-attention RWKV variant of the Qwen 2.5 72B model, optimized to significantly reduce computational cost at scale. Leveraging linear attention, it achieves substantial inference speedups (>1000x) while retaining competitive accuracy on common benchmarks like ARC, HellaSwag, Lambada, and MMLU.

Yes. The pricing shown on this page for Qrwkv 72B is zero, so you are not charged for prompt or completion tokens. Free endpoints are rate limited — see the rate limit docs.

Qrwkv 72B has a 32,768 token context window.

Qrwkv 72B was released on March 20, 2025. Its knowledge cutoff is June 30, 2024.

More models from featherless

Qrwkv 72B

Qrwkv-72B is a linear-attention RWKV variant of the Qwen 2.5 72B model, optimized to significantly reduce computational cost at scale. Leveraging linear attention, it achieves substantial inference speedups (>1000x) while retaining competitive accuracy on common benchmarks like ARC, HellaSwag, Lambada, and MMLU. It inherits knowledge and language support from Qwen 2.5, supporting approximately 30 languages, making it suitable for efficient inference in large-context applications.

Text33K context