OpenRouter Benchmarks
GPQA Diamond is a graduate-level multiple-choice benchmark in biology, physics, and chemistry. Each question is written by a subject-matter expert and designed so that even domain specialists need careful reasoning to identify the correct answer. We run the same fixed question set across provider endpoints to compare model capability, routing, and the practical cost of solving difficult scientific problems.
Last benchmark run Jul 20, 2026, 11:10 PM UTC
Top-level rows use default routing where available; click a row to expand provider-pinned results.
| # | Model | Std dev | ||||
|---|---|---|---|---|---|---|
| 1 | MiniMax: MiniMax M3 Pareto | 91.1% | ±1.5pp | $0.026 | 3.6m | 18.9k |
| 2 | 90.0% | ±3.9pp | $0.046 | 40s | 3.18k | |
| 3 | 87.6% | ±2.4pp | $0.051 | 8.4m | 21.9k | |
| 4 | 87.0% | ±1.3pp | $0.003 | 3.1m | 13.5k | |
| 5 | 86.8% | ±1.5pp | $0.026 | 2.9m | 8.73k | |
| 6 | 86.4% | ±1.6pp | $0.039 | 5.3m | 16.5k | |
| 7 | 85.7% | ±1.3pp | $0.047 | 3.9m | 21.2k | |
| 8 | 84.6% | ±1.5pp | $0.010 | 4.4m | 9.34k | |
| 9 | 84.4% | ±2.9pp | $0.050 | 8.9m | 19.2k | |
| 10 | 84.4% | ±1.7pp | $0.016 | 5.1m | 13.7k | |
| 11 | 84.4% | ±0.5pp | $0.019 | 6.9m | 20.7k | |
| 12 | Google: Gemma 4 31B Pareto | 84.3% | ±1.3pp | $0.003 | 4.3m | 7.39k |
| 13 | 83.9% | ±2.2pp | $0.015 | 2.7m | 10.1k | |
| 14 | 83.6% | ±2.1pp | $0.038 | 5.7m | 17.5k | |
| 15 | 83.5% | ±1.4pp | $0.031 | 4.9m | 30.9k | |
| 16 | 83.4% | ±3.2pp | $0.096 | 9.9m | 27.7k | |
| 17 | 83.3% | ±2.2pp | $0.091 | 7.1m | 24k | |
| 18 | 83.1% | ±2.6pp | $0.038 | 5.5m | 14.6k | |
| 19 | 82.9% | ±2.0pp | $0.003 | 4.0m | 7.85k | |
| 20 | 82.7% | ±2.1pp | $0.015 | 2.1m | 11.3k | |
| 21 | 82.2% | ±2.3pp | $0.083 | 4.0m | 23.8k | |
| 22 | 81.3% | ±2.9pp | $0.033 | 6.3m | 15.2k | |
| 23 | 81.0% | ±4.1pp | $0.003 | 87s | 10.7k | |
| 24 | 80.7% | ±2.7pp | $0.020 | 4.4m | 11.3k | |
| 25 | 79.9% | ±2.3pp | $0.008 | 4.1m | 6.49k | |
| 26 | 79.8% | ±3.7pp | $0.060 | 7.9m | 24.5k | |
| 27 | 78.8% | ±2.7pp | $0.004 | 7.1m | 8.56k | |
| 28 | 78.7% | ±6.2pp | $0.050 | 6.1m | 20.7k | |
| 29 | 78.4% | ±2.4pp | $0.028 | 7.8m | 11.4k | |
| 30 | 78.3% | ±8.9pp | $0.007 | 3.3m | 6.57k | |
| 31 | 77.8% | ±22.9pp | $0.066 | 6.4m | 18.5k | |
| 32 | 76.9% | ±6.5pp | $0.006 | 12.6m | 22.6k | |
| 33 | 76.6% | -- | $0.080 | 7.1m | 69.8k | |
| 34 | 76.4% | ±2.9pp | $0.005 | 11.5m | 32.8k | |
| 35 | 76.2% | ±2.5pp | $0.007 | 2.2m | 12.1k | |
| 36 | 75.8% | ±1.9pp | $0.008 | 1.8m | 3k | |
| 37 | 75.1% | ±2.2pp | $0.002 | 1.6m | 3.71k | |
| 38 | 74.7% | ±0.1pp | $0.11 | 4.8m | 33.6k | |
| 39 | 74.5% | ±2.9pp | $0.005 | 5.1m | 15.2k | |
| 40 | 72.9% | ±2.8pp | $0.006 | 53s | 5.49k | |
| 41 | 71.0% | ±1.8pp | $0.004 | 87s | 2.68k | |
| 42 | 70.7% | ±2.2pp | $0.005 | 49s | 4.52k | |
| 43 | 70.3% | ±8.3pp | $0.032 | 1.6m | 5.79k | |
| 44 | 67.2% | ±1.9pp | $0.005 | 5.1m | 26.2k | |
| 45 | 66.4% | ±5.5pp | $0.011 | 4.0m | 11.4k | |
| 46 | 65.9% | ±2.0pp | $0.005 | 2.1m | 8.07k | |
| 47 | 63.7% | ±2.0pp | $0.001 | 53s | 3.26k | |
| 48 | 63.0% | ±6.4pp | $0.001 | 32s | 870 | |
| 49 | 62.7% | ±2.8pp | $0.002 | 19s | 843 | |
| 50 | 62.3% | ±2.5pp | $0.003 | 2.6m | 7.1k | |
| 51 | 62.1% | ±3.6pp | $0.001 | 28s | 1.26k | |
| 52 | 61.3% | ±3.0pp | $0.003 | 2.1m | 6.39k | |
| 53 | 58.4% | ±2.5pp | $0.005 | 5.6m | 23.5k | |
| 54 | 52.6% | ±2.1pp | $0.004 | 1.6m | 6.72k | |
| 55 | 51.8% | ±3.9pp | $0.006 | 3.9m | 14k | |
| 56 | 49.9% | ±2.2pp | $0.001 | 30s | 971 | |
| 57 | Mistral: Mistral Nemo Pareto | 31.6% | ±2.1pp | $0.000 | 8s | 363 |
| 58 | 29.6% | ±3.6pp | $0.000 | 33s | 1.49k |
GPQA uses four-choice questions that require more than recalling a definition. These representative examples show the format and the range of scientific domains without reproducing items from the benchmark's protected question pool.
A researcher observes that a membrane protein is synthesized on ribosomes attached to the rough endoplasmic reticulum. Which destination is most consistent with this protein entering the secretory pathway?
Answer: C
Ribosomes on the rough ER synthesize proteins destined for secretion or insertion into the endomembrane system, including the plasma membrane.
A spacecraft is far from other bodies and fires its engine in the direction opposite to its velocity. Ignoring mass loss during the brief burn, what happens immediately to its speed?
GPQA is a broad graduate-level reasoning test across biology, physics, and chemistry, so it gives us a cheap, high-floor signal that a deployment is healthy. A model that normally clears these questions but suddenly drops usually points to something broken in the endpoint or routing rather than the questions themselves.
Because we run the same fixed question set across provider endpoints, a large accuracy gap between providers serving the same model is a quick way to catch a misconfigured or degraded endpoint. The cost and latency columns show what that reasoning quality costs to serve.
GPQA is a narrow, high-difficulty evaluation, not a complete measure of general intelligence or usefulness. A score reflects performance on expert-written multiple-choice science questions and should be considered alongside coding, instruction-following, factuality, and other evaluations.
Scores can be sensitive to sampling settings, answer-position handling, and the number of repeated runs. Small differences may not be meaningful when models have similar sample counts, so the leaderboard includes run variability and cost context rather than presenting accuracy alone.
The benchmark is publicly described, and some questions may eventually appear in training data. We avoid reproducing the private question pool here, but no public benchmark can guarantee that every future evaluation item is uncontaminated.
Scores aggregate successful runs, weighted by question count, with a minimum sample threshold per model-provider pair. A model's headline score uses default routing when available; otherwise it falls back to the median provider. Cost, time, and output-token figures are per-question averages from the same runs. Best value is the cheapest Pareto-optimal model within five percentage points of the top score.
GPQA Diamond is described in the original paper. See the docs for routing details, or browse all models to try one.
Answer: B
An impulse opposite the velocity vector reduces the spacecraft’s momentum and therefore its speed during the burn.
Why does adding a small amount of a common ion generally reduce the solubility of a sparingly soluble ionic solid in water?
Answer: B
The added ion raises the concentration of a dissolution product, so Le Chatelier’s principle shifts the equilibrium toward the undissolved solid.