Explainer
What is MAI-Voice-2.1?
MAI-Voice-2.1 is Microsoft AI’s in-house text-to-speech model, released on October 1, 2026. It turns text into expressive speech in 23 languages and comes in two versions: MAI-Voice-2.1 for quality and MAI-Voice-2.1-Flash for speed.
Key facts
- Released: October 1, 2026, in public preview on Azure Speech.
- Voices: 97 prebuilt voices across 28 locales (full list).
- Emotion control: styles such as joyful, sad, whispering, audiobook and narrator, set with SSML (all styles).
- Two models: 2.1 at ~550 ms for fidelity, Flash at ~45 ms for real time (comparison).
- Price: $22 / $15 per million characters (API guide).
- Voice cloning: instant cloning from a 5–60 s clip, gated behind Microsoft’s limited-access review and consent requirements.
What it is good at
Microsoft positions MAI-Voice-2.1 for audiobooks, content creation and voice-over, emphasizing natural rhythm, intonation and speaker consistency across long passages. Flash targets call-center agents, voice assistants and IVR, where response time matters more than the last bit of fidelity.
Where you can use it
- Microsoft Foundry / Azure Speech — the full API, with SSML styles and regional endpoints.
- MAI Playground and Copilot — Microsoft’s own consumer surfaces.
- OpenRouter and Vercel AI Gateway — unified API gateways.
- Tonecast — this site: a no-sign-up studio with daily free characters and MP3 download.
Limitations to know
- It is a public preview without an SLA, so behavior and voices may change.
- Style support varies by voice; 21 voices are neutral only.
- Japanese, Arabic and several other major languages are not in the catalog yet.
- Cloning is restricted to consented, approved voices.
Sources: Microsoft Learn, microsoft.ai. Checked 5 October 2026.