
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of overall performance, this model is second only to Qwen3.5-397B-A17B. Its text capabilities significantly outperform those of Qwen3-235B-2507, and its visual capabilities surpass those of Qwen3-VL-235B.
| $0.26 | $2.08 | -- | 1.20s | 44 tps | ||
| $0.26 | $2.08 | -- | 1.06s | 53 tps | ||
| $0.29 | $2.40 | -- | 0.34s | 86 tps | ||
| $0.30 | $2.40 | $0.30 | 0.87s | 49 tps | ||
| $0.40 | $3.20 | -- | 1.17s | 60 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.