Skip to content
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Apps
  • Models
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • SDK
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

OpenRouter Benchmarks

GPQA Diamond

GPQA Diamond is a graduate-level multiple-choice benchmark in biology, physics, and chemistry. Each question is written by a subject-matter expert and designed so that even domain specialists need careful reasoning to identify the correct answer. We run the same fixed question set across provider endpoints to compare model capability, routing, and the practical cost of solving difficult scientific problems.

Last benchmark run Jul 20, 2026, 11:10 PM UTC

PaperGitHub
Model comparisonCost efficiencyLeaderboardExample problemsWhy we run itWhat scores tell youMethodology

Model comparison

Most Accurate

Favicon for minimax
MiniMax: MiniMax M3

91.1%

Best Value

Favicon for deepseek
DeepSeek: DeepSeek V4 Flash

$0.003/question

Fastest

Favicon for openrouterFavicon for openrouter
OpenRouter: Phaser

40s

Accuracy
Representative-run accuracy, best first.
Cost per question
Average cost per question, cheapest first.
Time per question
Average wall-clock time per question, fastest first.

Cost efficiency

Accuracy vs. cost (Pareto frontier)
One point per model, using default routing (not pinned to a provider) when available. The line is the Pareto frontier: no model beats these on both accuracy and cost.

Leaderboard

Top-level rows use default routing where available; click a row to expand provider-pinned results.

#ModelStd dev
1
MiniMax: MiniMax M3
Pareto
91.1%±1.5pp$0.0263.6m18.9k
2
OpenRouter: Phaser
90.0%±3.9pp$0.04640s3.18k
3
Z.ai: GLM 5.2
87.6%±2.4pp$0.0518.4m21.9k
4
DeepSeek: DeepSeek V4 Flash
Pareto
87.0%±1.3pp$0.0033.1m13.5k
5
Qwen: Qwen3.5 397B A17B
86.8%±1.5pp$0.0262.9m8.73k
6
DeepSeek: DeepSeek V4 Pro
86.4%±1.6pp$0.0395.3m16.5k
7
Qwen: Qwen3.5-122B-A10B
85.7%±1.3pp$0.0473.9m21.2k
8
MiniMax: MiniMax M2.5
84.6%±1.5pp$0.0104.4m9.34k
9
MoonshotAI: Kimi K2.5
84.4%±2.9pp$0.0508.9m19.2k
10
MiniMax: MiniMax M2.7
84.4%±1.7pp$0.0165.1m13.7k
11
Xiaomi: MiMo-V2.5-Pro
84.4%±0.5pp$0.0196.9m20.7k
12
Google: Gemma 4 31B
Pareto
84.3%±1.3pp$0.0034.3m7.39k
13
MiniMax: MiniMax M2.1
83.9%±2.2pp$0.0152.7m10.1k
14
Z.ai: GLM 4.7
83.6%±2.1pp$0.0385.7m17.5k
15
Qwen: Qwen3.6 35B A3B
83.5%±1.4pp$0.0314.9m30.9k
16
MoonshotAI: Kimi K2.6
83.4%±3.2pp$0.0969.9m27.7k
17
Z.ai: GLM 5.1
83.3%±2.2pp$0.0917.1m24k
18
MoonshotAI: Kimi K2 Thinking
83.1%±2.6pp$0.0385.5m14.6k
19
DeepSeek: DeepSeek V3.2
82.9%±2.0pp$0.0034.0m7.85k
20
Qwen: Qwen3.5-35B-A3B
82.7%±2.1pp$0.0152.1m11.3k
21
NVIDIA: Nemotron 3 Ultra
82.2%±2.3pp$0.0834.0m23.8k
22
Z.ai: GLM 4.6
81.3%±2.9pp$0.0336.3m15.2k
23
Xiaomi: MiMo-V2-Flash
81.0%±4.1pp$0.00387s10.7k
24
Qwen: Qwen3 235B A22B Thinking 2507
80.7%±2.7pp$0.0204.4m11.3k
25
DeepSeek: DeepSeek V3.1
79.9%±2.3pp$0.0084.1m6.49k
26
Qwen: Qwen3.6 27B
79.8%±3.7pp$0.0607.9m24.5k
27
DeepSeek: DeepSeek V3.2 Exp
78.8%±2.7pp$0.0047.1m8.56k
28
Z.ai: GLM 5
78.7%±6.2pp$0.0506.1m20.7k
29
DeepSeek: R1 0528
78.4%±2.4pp$0.0287.8m11.4k
30
DeepSeek: DeepSeek V3.1 Terminus
78.3%±8.9pp$0.0073.3m6.57k
31
MoonshotAI: Kimi K2.7 Code
77.8%±22.9pp$0.0666.4m18.5k
32
Xiaomi: MiMo-V2.5
76.9%±6.5pp$0.00612.6m22.6k
33
StepFun: Step 3.7 Flash
76.6%--$0.0807.1m69.8k
34
Qwen: Qwen3.5-9B
76.4%±2.9pp$0.00511.5m32.8k
35
OpenAI: gpt-oss-120b
76.2%±2.5pp$0.0072.2m12.1k
36
MoonshotAI: Kimi K2 0905
75.8%±1.9pp$0.0081.8m3k
37
Qwen: Qwen3 235B A22B Instruct 2507
Pareto
75.1%±2.2pp$0.0021.6m3.71k
38
Auto Router (Beta)
74.7%±0.1pp$0.114.8m33.6k
39
Google: Gemma 4 26B A4B
74.5%±2.9pp$0.0055.1m15.2k
40
Qwen: Qwen3 Coder Next
72.9%±2.8pp$0.00653s5.49k
41
Qwen: Qwen3 VL 235B A22B Instruct
71.0%±1.8pp$0.00487s2.68k
42
Qwen: Qwen3 Next 80B A3B Instruct
70.7%±2.2pp$0.00549s4.52k
43
Auto Router
70.3%±8.3pp$0.0321.6m5.79k
44
OpenAI: gpt-oss-20b
67.2%±1.9pp$0.0055.1m26.2k
45
Z.ai: GLM 4.5 Air
66.4%±5.5pp$0.0114.0m11.4k
46
Qwen: Qwen3 VL 30B A3B Instruct
65.9%±2.0pp$0.0052.1m8.07k
47
Qwen: Qwen3 30B A3B Instruct 2507
Pareto
63.7%±2.0pp$0.00153s3.26k
48
DeepSeek: DeepSeek V3 0324
Pareto
63.0%±6.4pp$0.00132s870
49
Qwen: Qwen3 Coder 480B A35B
62.7%±2.8pp$0.00219s843
50
Qwen: Qwen3 32B
62.3%±2.5pp$0.0032.6m7.1k
51
Meta: Llama 4 Maverick
62.1%±3.6pp$0.00128s1.26k
52
Qwen: Qwen3 30B A3B
61.3%±3.0pp$0.0032.1m6.39k
53
NVIDIA: Nemotron 3 Nano 30B A3B
58.4%±2.5pp$0.0055.6m23.5k
54
Qwen: Qwen3 VL 8B Instruct
52.6%±2.1pp$0.0041.6m6.72k
55
Z.ai: GLM 4.7 Flash
51.8%±3.9pp$0.0063.9m14k
56
Meta: Llama 3.3 70B Instruct
Pareto
49.9%±2.2pp$0.00130s971
57
Mistral: Mistral Nemo
Pareto
31.6%±2.1pp$0.0008s363
58
Meta: Llama 3.1 8B Instruct
29.6%±3.6pp$0.00033s1.49k

Example problems

GPQA uses four-choice questions that require more than recalling a definition. These representative examples show the format and the range of scientific domains without reproducing items from the benchmark's protected question pool.

Biology

A researcher observes that a membrane protein is synthesized on ribosomes attached to the rough endoplasmic reticulum. Which destination is most consistent with this protein entering the secretory pathway?

  1. A.The cytosol, where it remains soluble
  2. B.The nucleus, after import through a nuclear pore
  3. C.A membrane of the endomembrane system or the cell surface
  4. D.The mitochondrial matrix through a TOM/TIM complex

Answer: C

Ribosomes on the rough ER synthesize proteins destined for secretion or insertion into the endomembrane system, including the plasma membrane.

Physics

A spacecraft is far from other bodies and fires its engine in the direction opposite to its velocity. Ignoring mass loss during the brief burn, what happens immediately to its speed?

  1. A.It increases because the exhaust carries away backward momentum
  2. B.It decreases because the thrust points opposite to its velocity
  3. C.It remains unchanged because thrust only changes direction
  4. D.It becomes zero because the spacecraft is in free space

Why we run this benchmark

GPQA is a broad graduate-level reasoning test across biology, physics, and chemistry, so it gives us a cheap, high-floor signal that a deployment is healthy. A model that normally clears these questions but suddenly drops usually points to something broken in the endpoint or routing rather than the questions themselves.

Because we run the same fixed question set across provider endpoints, a large accuracy gap between providers serving the same model is a quick way to catch a misconfigured or degraded endpoint. The cost and latency columns show what that reasoning quality costs to serve.

What the scores can and can't tell you

GPQA is a narrow, high-difficulty evaluation, not a complete measure of general intelligence or usefulness. A score reflects performance on expert-written multiple-choice science questions and should be considered alongside coding, instruction-following, factuality, and other evaluations.

Scores can be sensitive to sampling settings, answer-position handling, and the number of repeated runs. Small differences may not be meaningful when models have similar sample counts, so the leaderboard includes run variability and cost context rather than presenting accuracy alone.

The benchmark is publicly described, and some questions may eventually appear in training data. We avoid reproducing the private question pool here, but no public benchmark can guarantee that every future evaluation item is uncontaminated.

Methodology

Scores aggregate successful runs, weighted by question count, with a minimum sample threshold per model-provider pair. A model's headline score uses default routing when available; otherwise it falls back to the median provider. Cost, time, and output-token figures are per-question averages from the same runs. Best value is the cheapest Pareto-optimal model within five percentage points of the top score.

GPQA Diamond is described in the original paper. See the docs for routing details, or browse all models to try one.

Answer: B

An impulse opposite the velocity vector reduces the spacecraft’s momentum and therefore its speed during the burn.

Chemistry

Why does adding a small amount of a common ion generally reduce the solubility of a sparingly soluble ionic solid in water?

  1. A.The common ion increases the solid’s lattice energy
  2. B.The common ion shifts the dissolution equilibrium toward the solid
  3. C.The common ion converts every dissolved ion into a neutral molecule
  4. D.The common ion removes solvent molecules from the solution

Answer: B

The added ion raises the concentration of a dissolution product, so Le Chatelier’s principle shifts the equilibrium toward the undissolved solid.