Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function calling, and multilingual support across 140+ languages. Strong on coding, reasoning, and document understanding tasks. Apache 2.0 license.
| $0.09 | $0.34 | $0.05 | 0.57s | 20 tps | ||
| $0.10 | $0.34 | $0.10 | 0.48s | 25 tps | ||
| $0.12 | $0.36 | $0.09 | 0.63s | 19 tps | ||
| $0.13 | $0.38 | -- | 2.07s | 9 tps | ||
| $0.14 | $0.40 | $0.14 | 0.44s | 30 tps | ||
| $0.14 | $0.40 | -- | 1.86s | 37 tps | ||
| $0.15 | $0.40 | $0.06 | 2.12s | 17 tps | ||
| $0.38 | $1.15 | -- | 2.32s | 66 tps | ||
| $0.75 | $1.00 | $0.75 | 0.14s | 143 tps | ||
| $0.75 | $1.00 | $0.25 | 1.87s | 17 tps | ||
| $0.14 | $0.40 | -- | 2.01s | 9 tps | ||
| $0.38 | $1.15 | $0.19 | 0.72s | 25 tps | ||
| $0.10 | $0.33 | $0.05 | 0.95s | 38 tps | ||
| $0.12 | $0.37 | $0.012 | 4.53s | 7 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.