Microsoft AI releases new transcription and text-to-speech models for voice agents

Microsoft AI has released MAI-Transcribe-2-Streaming, a model for real-time transcription, along with two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash.
According to Microsoft, MAI-Transcribe-2-Streaming ranks first for accuracy on Artificial Analysis. The model transcribes 60 languages and returns its first partial results in just over 100 milliseconds. Microsoft says this allows voice agents to respond while a person is still mid-sentence. Through the end of the year, an hour of audio costs $0.54 at the introductory price.
MAI-Voice-2.1 is designed to speak 23 languages in the same voice, with a native accent in each one. Microsoft says the MAI-Voice-2.1-Flash variant reaches a latency of 150 milliseconds and costs $15 per million characters, compared with $22 for the other model.
Both voice models can clone a voice from a few seconds of reference audio, and built-in safeguards are intended to prevent misuse. The models are available through Microsoft Foundry and the MAI Playground, among other platforms, and the two voice models are also listed on OpenRouter. In one test cited by the publisher, about half of 4,000 participants believed the voices belonged to a real person.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Microsoft AI has released MAI-Transcribe-2-Streaming, a new model for real-time transcription. The article Microsoft AI releases new transcription and text-to-speech models for voice agents appeared first on The Decoder .