Mistral AI has announced 'Voxtral TTS,' a text-to-speech AI model that allows users to clone their own voice. It supports 9 languages, is incredibly fast, lightweight, and open-source.

French AI company Mistral AI has announced ' Voxtral TTS ,' a text-to-speech model that can generate natural and emotionally expressive voices. It supports nine major languages and features 'zero-shot clone voice playback' that requires no prior training, allowing it to generate voices that understand context and express emotions skillfully at lightning speed.
Speaking of Voxtral | Mistral AI
https://mistral.ai/news/voxtral-tts

🔊Introducing Voxtral TTS: our new frontier open-weight model for natural, expressive, and ultra-fast text-to-speech
— Mistral AI (@MistralAI) March 26, 2026
🎭Realistic, emotionally expressive speech.
🌍Supports 9 languages and accurately captures diverse dialects.
⚡Very low latency for time-to-first-audio.
🔄Easily… pic.twitter.com/Q2mdo8UBVo
Mistral releases a new open source model for speech generation | TechCrunch
https://techcrunch.com/2026/03/26/mistral-releases-a-new-open-source-model-for-speech-generation/
Mistral AI just released a text-to-speech model it says beats ElevenLabs — and it's giving away the weights for free | VentureBeat
https://venturebeat.com/orchestration/mistral-ai-just-released-a-text-to-speech-model-it-says-beats-elevenlabs-and
Voxtral TTS supports nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic, and enables the creation of voice agents that include even minor dialects. According to Mistral AI, it can generate custom voices that capture subtle accents, intonation, and irregularities in speech flow from voice samples less than 5 seconds long.
Voxtral TTS is designed with real-time performance in mind, boasting a fast Time-to-First Audio (TTFA) of 90 milliseconds for a 10-second sample audio of 500 characters, indicating the time from receiving input to the model starting to play the voice. It also has a Real-Time Factor (RTF) of 6x, meaning a 10-second clip can be rendered in approximately 1.6 seconds.
Voxtral TTS also features a 'cascading speech-to-speech translation' function, allowing you to, for example, use a voice sample created from French audio to input English text and generate English speech. On the official website, you can run a demo where you can select speakers of American English, French, and British English, and then select prompts in English, French, Spanish, and German to generate speech.

The following video demonstrates how to record a voice and create a cloned voice.
Voice Customization 101 with Voxtral in Mistral Studio - YouTube
In zero-shot custom voice tests, Voxtral TTS has been reported to outperform ElevenLabs v2.5 Flash, ElevenLabs' high-speed text-to-speech model, and perform comparably to ElevenLabs v3, a more advanced model with higher latency, based on metrics evaluated by native speakers regarding naturalness, accent accuracy, and similarity to the original voice.
State-of-the-art performance.
— Mistral AI (@MistralAI) March 26, 2026
In zero-shot custom voice tests, Voxtral TTS outperformed ElevenLabs v2.5 Flash - judged by native speakers for naturalness, accent accuracy, and similarity to the original voice. pic.twitter.com/ZY7PcRZGY3
Pierre Stock, Vice President of Scientific Operations at Mistral AI, said, 'We received requests from users for a voice model. So we developed a small voice model that can be installed in smartwatches, smartphones, laptops, and other edge devices. We aimed for a human-like voice, not a robotic one. It is priced considerably lower than other products on the market, but it offers cutting-edge performance.'
Another key feature of the model released this time is that it is provided as 'open weights.' This means that developers can freely download the model and run and modify it in their own environment, eliminating privacy concerns such as 'sending their voice samples to an external service.' Voxtral TTS can be downloaded from Hugging Face.
mistralai/Voxtral-4B-TTS-2603 · Hugging Face
https://huggingface.co/mistralai/Voxtral-4B-TTS-2603
Mr. Stock outlined two directions for the next development of Voxtral TTS. The first is to expand language and dialect support, with the goal of functioning while taking all cultural nuances into account. The second direction is the vision for Mistral AI, which not only generates speech from text but also understands the intonation, rhythm, and manner of speaking of spoken language, and responds by reading intentions and nuances.
Related Posts:







