Mistral unveils open-source speech generation model
French AI company Mistral unveiled a new open-source text-to-speech model on Thursday, designed for voice AI assistants and enterprise applications like customer support. The model enables businesses to build voice agents for sales and customer engagement, positioning Mistral as a direct competitor to ElevenLabs, Deepgram, and OpenAI.
Called Voxtral TTS, the model supports nine languages, including English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic.
"Our customers have been asking for a speech model. So we built a small-sized speech model that can fit on a smartwatch, a smartphone, a laptop, or other edge devices. The cost is a fraction of anything else on the market, yet it delivers state-of-the-art performance," said Pierre Stock, VP of science operations at Mistral AI, in a phone interview with TechCrunch.

Image credit: Mistral
Mistral states the new model can adapt to a custom voice using a sample shorter than five seconds, capturing subtle accents, inflections, intonations, and irregularities in speech flow. Built on Ministral 3B, it can switch between languages smoothly while preserving voice characteristics, making it ideal for dubbing or real-time translation. Stock noted the company aimed to make the model sound human, not robotic.
According to the company, the model is built for real-time performance. Its time-to-first-audio (TTFA) — the time between receiving input and beginning to 'speak' — is 90ms for a 10-second sample of 500 characters. The model also achieves a real-time factor (RTF) of 6x, meaning it can generate a 10-second clip in roughly 1.6 seconds.

Image credit: Mistral AI
Earlier this year, Mistral launched two transcription models — one for large-scale batch processing, the other for low-latency real-time use cases. With the new speech model, the company appears to be building a comprehensive suite of voice products for enterprises.
Stock added, "We plan to create an end-to-end platform capable of handling multimodal input streams — audio, text, and image — as well as output. The key advantage is that an end-to-end agentic system supporting audio input and output provides much richer information."
Mistral positions its open-source nature and customization capabilities as key differentiators, allowing enterprises to tune the model to their specific needs, thus favoring it over competitor solutions.
Related article
ElevenLabs Unveils Music AI That Switches Genres Mid-Song
ElevenLabs, a voice AI company, has launched Music v2, the latest version of its music-generation model, which can shift genres mid-track. The model is designed to handle complex vocals and composition. This release comes nearly ten months after the
Google rolls out voice-based prompts in Docs and Keep
At the Google I/O developer conference, the company announced it's introducing voice-based prompting to Workspace apps like Docs, Keep, and Gmail. These capabilities let users create drafts, take notes, and search for emails using voice.In Docs, you
Ex-Goldman, Meta founders build voice AI for overlooked markets
Customer support and service are among the hottest areas in voice AI today. But building a product that sounds human and responds with minimal delay is far more challenging in some markets than in others — and most major players weren’t designed with
Related Special Topic Recommendations
Comments (1)
0/500
I just tried out Mistral’s new open-source TTS model yesterday, and its pronunciation is way more natural than most free options I’ve used before. It’s super useful for companies that want to build their own voice assistants without spending tons on licensed services. Do you think this will change the game in the enterprise voice AI market?
French AI company Mistral unveiled a new open-source text-to-speech model on Thursday, designed for voice AI assistants and enterprise applications like customer support. The model enables businesses to build voice agents for sales and customer engagement, positioning Mistral as a direct competitor to ElevenLabs, Deepgram, and OpenAI.
Called Voxtral TTS, the model supports nine languages, including English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic.
"Our customers have been asking for a speech model. So we built a small-sized speech model that can fit on a smartwatch, a smartphone, a laptop, or other edge devices. The cost is a fraction of anything else on the market, yet it delivers state-of-the-art performance," said Pierre Stock, VP of science operations at Mistral AI, in a phone interview with TechCrunch.

Image credit: Mistral
Mistral states the new model can adapt to a custom voice using a sample shorter than five seconds, capturing subtle accents, inflections, intonations, and irregularities in speech flow. Built on Ministral 3B, it can switch between languages smoothly while preserving voice characteristics, making it ideal for dubbing or real-time translation. Stock noted the company aimed to make the model sound human, not robotic.
According to the company, the model is built for real-time performance. Its time-to-first-audio (TTFA) — the time between receiving input and beginning to 'speak' — is 90ms for a 10-second sample of 500 characters. The model also achieves a real-time factor (RTF) of 6x, meaning it can generate a 10-second clip in roughly 1.6 seconds.

Image credit: Mistral AI
Earlier this year, Mistral launched two transcription models — one for large-scale batch processing, the other for low-latency real-time use cases. With the new speech model, the company appears to be building a comprehensive suite of voice products for enterprises.
Stock added, "We plan to create an end-to-end platform capable of handling multimodal input streams — audio, text, and image — as well as output. The key advantage is that an end-to-end agentic system supporting audio input and output provides much richer information."
Mistral positions its open-source nature and customization capabilities as key differentiators, allowing enterprises to tune the model to their specific needs, thus favoring it over competitor solutions.
ElevenLabs Unveils Music AI That Switches Genres Mid-Song
ElevenLabs, a voice AI company, has launched Music v2, the latest version of its music-generation model, which can shift genres mid-track. The model is designed to handle complex vocals and composition. This release comes nearly ten months after the
Google rolls out voice-based prompts in Docs and Keep
At the Google I/O developer conference, the company announced it's introducing voice-based prompting to Workspace apps like Docs, Keep, and Gmail. These capabilities let users create drafts, take notes, and search for emails using voice.In Docs, you
Ex-Goldman, Meta founders build voice AI for overlooked markets
Customer support and service are among the hottest areas in voice AI today. But building a product that sounds human and responds with minimal delay is far more challenging in some markets than in others — and most major players weren’t designed with
I just tried out Mistral’s new open-source TTS model yesterday, and its pronunciation is way more natural than most free options I’ve used before. It’s super useful for companies that want to build their own voice assistants without spending tons on licensed services. Do you think this will change the game in the enterprise voice AI market?





Home






