TonecastOpen studio

Developer guide

MAI-Voice-2.1 API: pricing and code examples

MAI-Voice-2.1 is available through Azure Speech in Microsoft Foundry (public preview), and through gateways such as OpenRouter and Vercel AI Gateway. Here is how each one works and what it costs.

Pricing

ModelPrice1,000 characters
MAI-Voice-2.1$22 / 1M characters$0.022
MAI-Voice-2.1-Flash$15 / 1M characters$0.015

Gateway prices can differ from Microsoft’s list price; check the provider before you commit.

Option 1: Azure Speech REST (full emotion control)

Create a Speech resource in Microsoft Foundry, then send SSML to your region’s endpoint. This is the route that supports mstts:express-as styles.

curl -X POST "https://${SPEECH_REGION}.tts.speech.microsoft.com/cognitiveservices/v1" \
  -H "Content-Type: application/ssml+xml" \
  -H "X-Microsoft-OutputFormat: audio-24khz-160kbitrate-mono-mp3" \
  -H "Ocp-Apim-Subscription-Key: ${SPEECH_KEY}" \
  --data '<speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis" xmlns:mstts="http://www.w3.org/2001/mstts" xml:lang="en-US">
  <voice name="en-US-Harper:MAI-Voice-2.1">
    <mstts:express-as style="audiobook">Welcome back. Let’s pick up where we left off.</mstts:express-as>
  </voice>
</speak>' \
  --output output.mp3

Supported regions include East US, East US 2, West US, West US 2, West US 3, Canada Central, North Europe, West Europe, France Central, Sweden Central, Central India, East Asia, Southeast Asia and Japan East.

Option 2: OpenRouter

An OpenAI-style speech endpoint. Simple JSON in, audio bytes out. Its documented parameters are model, input, voice and response format — there is no style field, so expect neutral delivery.

curl https://openrouter.ai/api/v1/audio/speech \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "microsoft/mai-voice-2.1",
    "input": "Welcome back. Let’s pick up where we left off.",
    "voice": "en-US-Harper:MAI-Voice-2.1",
    "response_format": "mp3"
  }' --output output.mp3

Option 3: Python with the Speech SDK

import os
import azure.cognitiveservices.speech as speechsdk

config = speechsdk.SpeechConfig(subscription=os.environ["SPEECH_KEY"], region=os.environ["SPEECH_REGION"])
config.set_speech_synthesis_output_format(speechsdk.SpeechSynthesisOutputFormat.Audio24Khz160KBitRateMonoMp3)
synth = speechsdk.SpeechSynthesizer(speech_config=config, audio_config=speechsdk.audio.AudioOutputConfig(filename="out.mp3"))

ssml = """<speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis" xmlns:mstts="http://www.w3.org/2001/mstts" xml:lang="en-US">
  <voice name="en-US-Ethan:MAI-Voice-2.1-Flash">
    <mstts:express-as style="joyful">We did it!</mstts:express-as>
  </voice>
</speak>"""
result = synth.speak_ssml_async(ssml).get()

Voice IDs

Format: {locale}-{Name}:{model}, for example zh-CN-Mei:MAI-Voice-2.1-Flash. The full list is on the voices page. Not every voice supports every style; an unsupported style is the most common reason a request sounds flat.

Want to hear a voice before writing code? Use the studio — every take shows its SSML.

Examples adapted from Microsoft Learn and OpenRouter, checked 5 October 2026.