Alibaba Tongyi unveils voice model with 'FreeStyle' natural language control
Today, Alibaba Tongyi Lab's Speech Team introduced two groundbreaking voice generation models: Fun-CosyVoice3.5 and Fun-AudioGen-VD. The standout feature of these models is their support for "FreeStyle" commands. Instead of complex parameter adjustments, users can precisely control vocal expression styles or build intricate audio scenes from scratch using simple natural language descriptions.

Each model serves distinct purposes:
Fun-CosyVoice3.5: Multilingual Replication and Fine-Grained Control
This enhanced version of CosyVoice achieves core breakthroughs in understanding speech expression nuances.
Command-Driven Generation: Users can input instructions like "speak more confidently" or "slow down with emotional variation" for real-time vocal adjustments.
Language Expansion: Added support for Thai, Indonesian, Portuguese, and Vietnamese maintains industry-leading performance in transcription accuracy (WER) and voice similarity across 13 languages.
Rare Character Optimization: Specialized training reduced error rates for uncommon characters from 15.2% to 5.3%.
Performance Boost: First packet latency decreased by 35%, significantly enhancing real-time interaction fluidity.
Fun-AudioGen-VD: Comprehensive Sound Design
This model acts as an "audio director," generating integrated audio combining "characters + environments."
Voice Customization: Specify gender, age, accent, and detailed characteristics like "hoarse, deep, or low-pitched" voices.
Emotion and Role Play: Simulates roles including customer service agents, broadcasters, and children, even conveying complex states like "outward calm with internal tension."
Immersive Environments: Adds background sounds (battlefield chaos, café murmurs) and spatial effects (cathedral reverb, underwater acoustics) for full spatial simulation.
Tongyi Lab notes these models will democratize high-quality voice creation, offering powerful AI support for podcasting, game development, and film post-production.
Related article
ByteDance’s Seed launches global campus drive, offering virtual shares to win top large model talent
In the competitive landscape of large language models, securing top-tier talent remains the most critical strategic asset.On April 1st, ByteDance announced the launch of its Seed global campus recruitment initiative, part of its large model talent de
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
Related Special Topic Recommendations
Comments (0)
0/500
Today, Alibaba Tongyi Lab's Speech Team introduced two groundbreaking voice generation models: Fun-CosyVoice3.5 and Fun-AudioGen-VD. The standout feature of these models is their support for "FreeStyle" commands. Instead of complex parameter adjustments, users can precisely control vocal expression styles or build intricate audio scenes from scratch using simple natural language descriptions.

Each model serves distinct purposes:
Fun-CosyVoice3.5: Multilingual Replication and Fine-Grained Control
This enhanced version of CosyVoice achieves core breakthroughs in understanding speech expression nuances.
Command-Driven Generation: Users can input instructions like "speak more confidently" or "slow down with emotional variation" for real-time vocal adjustments.
Language Expansion: Added support for Thai, Indonesian, Portuguese, and Vietnamese maintains industry-leading performance in transcription accuracy (WER) and voice similarity across 13 languages.
Rare Character Optimization: Specialized training reduced error rates for uncommon characters from 15.2% to 5.3%.
Performance Boost: First packet latency decreased by 35%, significantly enhancing real-time interaction fluidity.
Fun-AudioGen-VD: Comprehensive Sound Design
This model acts as an "audio director," generating integrated audio combining "characters + environments."
Voice Customization: Specify gender, age, accent, and detailed characteristics like "hoarse, deep, or low-pitched" voices.
Emotion and Role Play: Simulates roles including customer service agents, broadcasters, and children, even conveying complex states like "outward calm with internal tension."
Immersive Environments: Adds background sounds (battlefield chaos, café murmurs) and spatial effects (cathedral reverb, underwater acoustics) for full spatial simulation.
Tongyi Lab notes these models will democratize high-quality voice creation, offering powerful AI support for podcasting, game development, and film post-production.
ByteDance’s Seed launches global campus drive, offering virtual shares to win top large model talent
In the competitive landscape of large language models, securing top-tier talent remains the most critical strategic asset.On April 1st, ByteDance announced the launch of its Seed global campus recruitment initiative, part of its large model talent de
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a





Home






