Tongyi Qianwen Qwen-Audio-3.0-TTS Goes Live With 300ms First Packet Delay, 20 Dialects
The Alibaba Qwen team has launched the new real-time speech synthesis model Qwen-Audio-3.0-TTS, moving speech synthesis from basic generation to true expression. The Plus version now ranks first globally in the Speech Arena, a respected benchmark by Artificial Analysis, outperforming Gemini 3.1 TTS, ElevenLabs v3, and other major models.

Two versions are available: the Flash version targets real-time interaction with an initial latency of around 300ms, ideal for smart assistants and other low-latency use cases; the Plus version prioritizes high-quality generation, delivering improved naturalness and voice similarity.
Four key breakthroughs in core capabilities:
First, multilingual and dialect support has been expanded to 16 languages, with the Plus version achieving an average speaker similarity of 82.75% across all 16 languages — the highest in the industry. It also covers 20 Chinese dialects, maintaining authentic dialect characteristics rather than weakening them, for more native-like expressions.
Second, flexible natural language instructions allow users to generate accurate speech with prompts like "gentle customer service tone" or "live-streaming host style," without needing professional annotations.
Third, fine-grained tag control supports structured tags such as [gasp] and [angry], enabling precise management of non-verbal details like breathing and laughter — suitable for games, audiobooks, and similar scenarios.
Fourth, strong acoustic robustness: even if the reference audio contains high noise or reverberation, the model automatically filters out background noise while preserving voice quality, ensuring stable synthesis output.
The accompanying premium voice library includes various voice types — instructions, dialects, and lesser-used languages — supports 48K high-definition audio output (expected to open on July 24th), and can synthesize up to 3 minutes of long text per session. The model is now fully available on the Alibaba Cloud BaiLian platform, where developers can access and try it, opening up new possibilities in voice interaction.

Related article
ByteDance’s Seed launches global campus drive, offering virtual shares to win top large model talent
In the competitive landscape of large language models, securing top-tier talent remains the most critical strategic asset.On April 1st, ByteDance announced the launch of its Seed global campus recruitment initiative, part of its large model talent de
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
Related Special Topic Recommendations
Comments (0)
0/500
The Alibaba Qwen team has launched the new real-time speech synthesis model Qwen-Audio-3.0-TTS, moving speech synthesis from basic generation to true expression. The Plus version now ranks first globally in the Speech Arena, a respected benchmark by Artificial Analysis, outperforming Gemini 3.1 TTS, ElevenLabs v3, and other major models.

Two versions are available: the Flash version targets real-time interaction with an initial latency of around 300ms, ideal for smart assistants and other low-latency use cases; the Plus version prioritizes high-quality generation, delivering improved naturalness and voice similarity.
Four key breakthroughs in core capabilities:
First, multilingual and dialect support has been expanded to 16 languages, with the Plus version achieving an average speaker similarity of 82.75% across all 16 languages — the highest in the industry. It also covers 20 Chinese dialects, maintaining authentic dialect characteristics rather than weakening them, for more native-like expressions.
Second, flexible natural language instructions allow users to generate accurate speech with prompts like "gentle customer service tone" or "live-streaming host style," without needing professional annotations.
Third, fine-grained tag control supports structured tags such as [gasp] and [angry], enabling precise management of non-verbal details like breathing and laughter — suitable for games, audiobooks, and similar scenarios.
Fourth, strong acoustic robustness: even if the reference audio contains high noise or reverberation, the model automatically filters out background noise while preserving voice quality, ensuring stable synthesis output.
The accompanying premium voice library includes various voice types — instructions, dialects, and lesser-used languages — supports 48K high-definition audio output (expected to open on July 24th), and can synthesize up to 3 minutes of long text per session. The model is now fully available on the Alibaba Cloud BaiLian platform, where developers can access and try it, opening up new possibilities in voice interaction.

ByteDance’s Seed launches global campus drive, offering virtual shares to win top large model talent
In the competitive landscape of large language models, securing top-tier talent remains the most critical strategic asset.On April 1st, ByteDance announced the launch of its Seed global campus recruitment initiative, part of its large model talent de
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a





Home






