AI News · Models ·
Microsoft releases new transcription and voice models for conversational agents

Microsoft AI announced the availability of its latest MAI-Transcribe and MAI-Voice models for building voice agents. The voice releases include MAI-Voice-2.1 and a faster variant, MAI-Voice-2.1-Flash. Microsoft also said its first streaming transcription model debuted at No. 1 on Artificial Analysis.
Key points
- Microsoft released MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and the faster MAI-Voice-2.1-Flash for conversational agents.
- The streaming transcription model supports 60 languages and returns preliminary text in just over 100 milliseconds, Microsoft said.
- Microsoft said its streaming transcription model ranks No. 1 on Artificial Analysis.
- Both voice models support 23 languages; Flash targets high-volume, latency-sensitive workloads.
- A Chatter demo is available on the MAI Playground; specific pricing was not reported.
What happened: Microsoft AI has released new transcription and voice-generation models for teams building conversational agents, adding options for software that listens to speech and responds aloud. The releases include MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and a faster variant, MAI-Voice-2.1-Flash. In its October 1, 2026 announcement, Microsoft said its first streaming transcription model debuted at No. 1 on Artificial Analysis. Thurrott reported that the model is top-rated for both partial and final transcripts, according to Microsoft.
The details: MAI-Transcribe-2-Streaming produces real-time transcripts in 60 languages with continuous language detection, Thurrott reported. Rather than waiting until a person finishes speaking, it begins returning preliminary text, called partials, just over 100 milliseconds after receiving audio, Microsoft said. It revises those early results as more speech arrives, then commits a stable transcript. That lets voice applications begin acting on speech before a speaker finishes. Microsoft also said its internal evaluations showed words appearing in transcripts twice as fast as with its closest competitor. The competitor's identity and the conditions behind that comparison were not reported.
The details: The two MAI-Voice releases handle the other side of a voice conversation: converting text into spoken audio. Microsoft describes MAI-Voice-2.1 as its strongest multilingual text-to-speech model yet, supporting 23 languages and 26 locales. A single voice can work across languages while changing to a native accent as the language changes, according to Thurrott. MAI-Voice-2.1-Flash supports the same languages and cross-language speakers but is optimized for high-volume workloads where response speed matters. Microsoft said Flash can generate 45 seconds of audio with end-to-end latency of 150 milliseconds. Test conditions for that performance claim were not reported.
Who it affects: The releases are aimed at developers and business teams building voice agents, including customer service agents, multilingual assistants, and interactive learning and media experiences. Microsoft said the streaming transcription model and Flash voice model can work together to keep conversations moving naturally. For buyers, the distinction between the models matters: transcription accuracy concerns what an agent hears, while voice generation concerns how it speaks back. The announcement adds choices for evaluating both parts of that interaction, rather than presenting a finished customer service or assistant application.
What to watch: Teams should test transcription accuracy and conversational response times on their own audio rather than rely solely on rankings or Microsoft's performance claims. Early partial transcripts and stable final transcripts are separate outputs worth evaluating, given that the model revises text as context arrives. Microsoft offers a Chatter demo on the MAI Playground to try the models. Although the company described the releases as low cost, specific pricing was not reported, leaving an important question for teams assessing high-volume use.
Our take
Teams building voice agents have new options to evaluate. Test transcription accuracy and conversational latency on your own audio rather than relying only on benchmark rankings.