TonecastOpen studio

MAI-Voice-2.1 · released October 1, 2026

MAI-Voice-2.1 text to speech,
free online.

Type a line, pick one of 97 voices in 23 languages, and hear Microsoft’s newest voice model in seconds. Switch between 2.1 and Flash, download the MP3. Free characters every day, no sign-up.

STUDIO A —
0 / 500
Model
Delivery · SSML preview

View SSML for your own Azure calls
    97prebuilt voices
    23languages · 28 locales
    19emotion styles on 19 voices
    ~45msFlash latency
    Delivery

    Same words. A different scene.

    MAI-Voice-2.1 changes delivery through SSML style tags. Our free studio renders the default delivery for now — tap a mood to load a voice that supports it and preview the exact SSML that performs it on Azure Speech.

    All 33 styles, and which voices support each →

    For creators

    Voice-overs that don’t sound read aloud.

    Cast Harper, Grant or Sage for explainers and narration, and download the MP3 straight into your editor for videos, ads and podcasts. MAI-Voice-2.1 is tuned for long-form speaker consistency, so a voice stays itself from the first paragraph to the last.

    For developers

    Prototype here, ship on the API.

    Every take shows the exact SSML behind it — voice ID, model suffix and mstts:express-as style — so you can paste it into Azure Speech. API guide and pricing →

    How it works

    Three steps, about ten seconds.

    1. Write your script. Up to 500 characters per take on the free tier.
    2. Cast your voice. Choose language, voice and model.
    3. Generate and download. Listen in the browser, then save the MP3.
    FAQ

    Questions people ask

    What is MAI-Voice-2.1?

    MAI-Voice-2.1 is Microsoft AI’s text-to-speech model, released on October 1, 2026 alongside a faster sibling, MAI-Voice-2.1-Flash. It ships 97 prebuilt voices across 23 languages and lets you steer delivery with styles such as joyful, whispering, or audiobook. Read the full overview.

    Is MAI-Voice-2.1 free to use?

    The API is paid: $22 per million characters for MAI-Voice-2.1 and $15 for Flash. Microsoft also offers a playground. Tonecast gives you free characters every day with no account, and you can download every clip as an MP3.

    How do I make the voice sound happy, sad, or whisper?

    MAI-Voice-2.1 sets emotion with SSML style tags on Azure Speech; 19 voices support the full set, including whispering and shouting. Tonecast’s free studio currently renders the default delivery, but you can pick any style and copy the exact SSML. See every style and which voices support it.

    What is the difference between MAI-Voice-2.1 and Flash?

    Both use the same voices and styles. MAI-Voice-2.1 favors fidelity and long-form consistency (about 550 ms latency); Flash is built for real-time apps (about 45 ms) and costs less. Compare them.

    Which languages are supported?

    23 languages in 28 locales: English (US, UK, Australia, India), Spanish (Spain, Mexico), Portuguese (Brazil, Portugal), French, German, Italian, Dutch, Chinese (Mandarin), Hindi, Korean, Thai, Vietnamese, Indonesian, Turkish, Russian, Polish, Romanian, Hungarian, Czech, Danish, Finnish, Swedish and Norwegian. Browse all voices.

    Can I clone my own voice?

    Microsoft supports instant voice cloning from a 5–60 second clip, but access is gated and requires recorded consent. Tonecast does not offer cloning; it uses the licensed prebuilt voices only.

    Can I use the audio commercially?

    Generated audio is subject to Microsoft’s terms for MAI-Voice. Review them before using clips in paid products, ads, or monetized videos. See our terms.

    Is my text stored?

    No. Your text is sent to the speech provider to create the audio and is not saved by Tonecast. Privacy details.