Speaking of Voxtral
Mistral AI announced Voxtral TTS, its first text-to-speech model, describing it as state-of-the-art in multilingual voice generation. The model is lightweight at 4B parameters.
Mistral said Voxtral TTS produces realistic, emotionally expressive speech in 9 languages — English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic — with support for diverse dialects, and offers very low latency for time-to-first-audio. It is available to test in Mistral Studio and via API at $0.016 per 1k characters. A model with several reference voices is available as open weights on Hugging Face under a CC BY-NC 4.0 license, Mistral said.
According to Mistral, the model can adapt to a custom voice from a reference as short as 3 seconds, capturing accent, inflections, intonations and disfluencies. It also demonstrates zero-shot cross-lingual voice adaptation despite not being explicitly trained for it, Mistral said, giving the example of generating English speech from a French voice prompt that sounds naturally French-accented.
On latency, Mistral reported a model latency of 70ms for a typical input voice sample of 10 seconds and 500 characters, with a real-time factor of approximately 9.7x. The model natively generates up to two minutes of audio, and the API handles arbitrarily long generations with interleaving, Mistral said.
Mistral described the architecture as a transformer-based, autoregressive, flow-matching model built on Ministral 3B, comprising a 3.4B-parameter transformer decoder backbone, a 390M flow-matching acoustic transformer and a 300M neural audio codec. It takes a voice prompt of 5 to 25 seconds and a text prompt in the 9 supported languages.
Mistral said comparative human evaluations by native speakers found Voxtral TTS achieves superior naturalness compared to ElevenLabs Flash v2.5 while maintaining similar Time-to-First-Audio, and performs at parity with the quality of ElevenLabs v3, supporting emotion-steering. In a zero-shot custom voice evaluation against ElevenLabs v2.5 Flash using two recognizable voices per language for each of the 9 languages, with 3 annotators performing side-by-side preference tests on naturalness, accent adherence and acoustic similarity, Mistral said Voxtral TTS widened the quality gap.
Mistral listed intended enterprise uses including customer support, financial services, manufacturing, public services, compliance, supply chain, automotive, sales and marketing, and real-time translation.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Voxtral TTS: A frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents.