
MAI-Voice-2.1 is Microsoft AI's highest-fidelity, most expressive text-to-speech model. It produces natural, studio-grade speech across 23 languages, with detailed prosody, nuanced expressiveness, and speaker consistency over long-form content. It is suited for audiobooks, podcasts, lectures, narration, and brand audio where maximum voice quality matters. The model prioritizes naturalness and expressivity over latency-critical generation.
On OpenRouter, set voice to a full voice ID with the model suffix, such as "en-US-Harper:MAI-Voice-2.1". A voice's locale sets the synthesis language. Set response_format to "mp3" or "pcm" (24 kHz mono). Harper and Grant support the agent, customer-call-center, educational, and narrator speaking styles, and many locale voices add emotion styles such as excited, happy, sad, and whispering. The full list of voices is in the supported_voices field of the models APIOpens in new tab. See the text-to-speech guideOpens in new tab.
| $22.00 | 1.14s |
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.