AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Voxtral transcribes at the speed of sound.

Collected Oct 1, 2026

Mistral AI announced the release of Voxtral Transcribe 2, a family of two speech-to-text models. Voxtral Mini Transcribe V2 is intended for batch transcription, while Voxtral Realtime targets live applications and is released as open weights under the Apache 2.0 license. Mistral also launched an audio playground in Mistral Studio for testing transcription with diarization and timestamps.

Voxtral Mini Transcribe V2 adds speaker diarization, context biasing, and word-level timestamps across 13 languages. Mistral states the model reaches approximately 4% word error rate on the FLEURS benchmark at $0.003 per minute and processes audio about 3x faster than ElevenLabs' Scribe v2 while matching on quality at one-fifth the cost. It supports recordings up to 3 hours in a single request. Mistral notes that with overlapping speech, the model typically transcribes one speaker, and that context biasing is optimized for English with other languages experimental.

Voxtral Realtime uses a streaming architecture with latency configurable down to sub-200ms. Mistral reports that at 2.4 seconds of delay it matches Voxtral Mini Transcribe V2, and at 480ms delay it stays within 1-2% word error rate. It has a 4B parameter footprint and is natively multilingual across 13 languages: English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian, and Dutch. Weights are published on the Hugging Face Hub.

Voxtral Mini Transcribe V2 is available via API at $0.003 per minute, and Voxtral Realtime at $0.006 per minute. Both models are described as supporting GDPR and HIPAA-compliant deployments through on-premise or private cloud setups.

Read at Mistral AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt