Skip to content
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Apps
  • Discover
  • Models
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • SDK
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

Benchmarks

Independent, reproducible measurements of the knobs you can actually set on an OpenRouter request: models, providers, search engines, and tool budgets. Every score links to the configuration, costs, and telemetry behind it.

2 benchmarks·2,437,546 task evaluations·last run Aug 5, 2026

Agents & toolsReasoning

Agents & tools

  • τ²-Bench Airline

    Multi-turn service agents making tool calls under strict policy constraints.

    109 models·9,936 runs·last run Aug 5, 2026

Reasoning

  • GPQA Diamond

    Graduate-level science questions that resist retrieval and reward careful reasoning.

    106 models·9,926 runs·last run Aug 5, 2026

For usage-based views of the same models, see the model rankings and the full model list.