Mistral AI has announced 'Voxtral TTS,' a text-to-speech AI model that allows users to clone their own voice. It supports 9 languages, is incredibly fast, lightweight, and open-source.



French AI company Mistral AI has announced ' Voxtral TTS ,' a text-to-speech model that can generate natural and emotionally expressive voices. It supports nine major languages and features 'zero-shot clone voice playback' that requires no prior training, allowing it to generate voices that understand context and express emotions skillfully at lightning speed.

Speaking of Voxtral | Mistral AI
https://mistral.ai/news/voxtral-tts






Mistral releases a new open source model for speech generation | TechCrunch
https://techcrunch.com/2026/03/26/mistral-releases-a-new-open-source-model-for-speech-generation/

Mistral AI just released a text-to-speech model it says beats ElevenLabs — and it's giving away the weights for free | VentureBeat
https://venturebeat.com/orchestration/mistral-ai-just-released-a-text-to-speech-model-it-says-beats-elevenlabs-and

Voxtral TTS supports nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic, and enables the creation of voice agents that include even minor dialects. According to Mistral AI, it can generate custom voices that capture subtle accents, intonation, and irregularities in speech flow from voice samples less than 5 seconds long.

Voxtral TTS is designed with real-time performance in mind, boasting a fast Time-to-First Audio (TTFA) of 90 milliseconds for a 10-second sample audio of 500 characters, indicating the time from receiving input to the model starting to play the voice. It also has a Real-Time Factor (RTF) of 6x, meaning a 10-second clip can be rendered in approximately 1.6 seconds.

Voxtral TTS also features a 'cascading speech-to-speech translation' function, allowing you to, for example, use a voice sample created from French audio to input English text and generate English speech. On the official website, you can run a demo where you can select speakers of American English, French, and British English, and then select prompts in English, French, Spanish, and German to generate speech.



The following video demonstrates how to record a voice and create a cloned voice.

Voice Customization 101 with Voxtral in Mistral Studio - YouTube


In zero-shot custom voice tests, Voxtral TTS has been reported to outperform ElevenLabs v2.5 Flash, ElevenLabs' high-speed text-to-speech model, and perform comparably to ElevenLabs v3, a more advanced model with higher latency, based on metrics evaluated by native speakers regarding naturalness, accent accuracy, and similarity to the original voice.




Pierre Stock, Vice President of Scientific Operations at Mistral AI, said, 'We received requests from users for a voice model. So we developed a small voice model that can be installed in smartwatches, smartphones, laptops, and other edge devices. We aimed for a human-like voice, not a robotic one. It is priced considerably lower than other products on the market, but it offers cutting-edge performance.'

Another key feature of the model released this time is that it is provided as 'open weights.' This means that developers can freely download the model and run and modify it in their own environment, eliminating privacy concerns such as 'sending their voice samples to an external service.' Voxtral TTS can be downloaded from Hugging Face.

mistralai/Voxtral-4B-TTS-2603 · Hugging Face
https://huggingface.co/mistralai/Voxtral-4B-TTS-2603

Mr. Stock outlined two directions for the next development of Voxtral TTS. The first is to expand language and dialect support, with the goal of functioning while taking all cultural nuances into account. The second direction is the vision for Mistral AI, which not only generates speech from text but also understands the intonation, rhythm, and manner of speaking of spoken language, and responds by reading intentions and nuances.

in AI,   Video, Posted by log1e_dh