Microsoft has released MAI-Transcribe-2-Streaming, which it describes as its first streaming transcription model, alongside MAI-Voice-2.1 and a faster variant, MAI-Voice-2.1-Flash.
The transcription model handles 60 languages with automatic, continuous language detection, so a speaker can switch language mid-conversation without anyone selecting one in advance. Microsoft says it returns first hypotheses in just over 100ms and that words land in the transcript twice as fast as with its closest competitor, and it claims first place for accuracy on both final and partial transcripts on Artificial Analysis, a third-party benchmark. Those are Microsoft’s own figures and we have not reproduced them. The model announcement has the detail.
On the speech side, MAI-Voice-2.1 covers 23 languages and 26 locales. The Flash variant targets high-volume, latency-sensitive work, with quoted end-to-end latency of 150ms and 55% faster inference.
Pricing is $0.54 per hour of audio for the transcription model through the end of the year, $22 per million characters for MAI-Voice-2.1 and $15 for Flash. All three are available through Microsoft Foundry, the MAI Playground, Vercel and OpenRouter.
Why it matters: Microsoft continues to build a first-party model line rather than relying solely on OpenAI, and voice is the segment where latency, not reasoning, decides whether a product is usable at all.
