
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context training process, reaching up to 128K tokens, and uses FP8 microscaling for efficient inference. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docsOpens in new tab
The model improves tool use, code generation, and reasoning efficiency, achieving performance comparable to DeepSeek-R1 on difficult benchmarks while responding more quickly. It supports structured tool calling, code agents, and search agents, making it suitable for research, coding, and agentic workflows.
It succeeds the DeepSeek V3-0324 model and performs well on a variety of tasks.
| $0.25 | $0.95 | $0.13 | 2.50s | 4 tps | ||
| $0.27 | $1.00 | -- | 1.71s | 17 tps | ||
| $0.30 | $1.00 | $0.135 | 1.54s | 37 tps | ||
| $0.55 | $1.65 | $0.55 | 0.42s | 49 tps | ||
| $0.60 | $1.70 | -- | 0.97s | 88 tps | ||
| $0.65 | $1.50 | -- | 5.45s | 25 tps | ||
| $0.60 | $1.70 | -- | -- | -- |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.