Home
Tencent, Tsinghua Launch SongGeneration 2 With 8.55% Error Rate, Intensifying Pressure on Suno
The AI music landscape experienced another significant shift in early 2026. On March 9, the music foundation model SongGeneration2, developed jointly by Tencent and Tsinghua University's Human-Computer Speech Interaction Lab , was officially released. This model not only marked a qualitative leap in technical architecture but also outperformed current mainstream open-source models across multiple core dimensions, even rivaling top commercial models in overall quality.

Three Breakthroughs: No More 'Plastic' AI Music
SongGeneration2 ’s core advantages come from a comprehensive upgrade of its underlying architecture, addressing three key pain points that plagued earlier AI music:
High musicality: Unlike basic melody stacking, this model handles complex multi-track arrangements with impressive spatial depth.
High lyric accuracy: Unclear pronunciation and hallucinated pitch shifts are now history. Its phoneme error rate (PER) is as low as 8.55%, notably outperforming the leading commercial model Suno v5 (12.4%), and trailing only MiniMax2.5 .
Strong controllability: Whether through text descriptions or audio prompts, it precisely follows instructions, enabling deep customization of style and emotion.

'Dual-Core' Drive: A Dream Collaboration Between LLM and Diffusion Models
Architecturally, SongGeneration2 employs an innovative hybrid LLM-diffusion architecture:
Composing Brain (LeLM): Handles global structure planning and vocal details, solving the challenge of "how to sing".
High-Fidelity Renderer (Diffusion): Synthesizes extremely complex acoustic details under the guidance of the language model.
Hierarchical Representation: The first to adopt parallel modeling of mixed and multi-track representations, balancing melodic stability with sound quality refinement.
True Open Source, Low Barrier: Ordinary Computers Can Also 'Write Songs'
What excited developers most was Tencent's remarkable openness. The SongGeneration-v2-large model with 4 billion parameters is now fully open-source, supporting multilingual generation including Chinese and English. Remarkably, it runs smoothly on consumer-grade hardware with just 22GB VRAM, enabling local and private music creation.
To let users try it right away, the team also released the SongGeneration-v2-Fast version on HuggingFace, trading a minimal drop in audio quality for ultra-fast generation—producing a complete song in under a minute.
Based on the performance of SongGeneration2 , AI music has officially evolved from a 'geek toy' into the realm of 'commercial-grade applications.' With the upcoming open-sourcing of a Medium model supporting 12GB VRAM and an automated evaluation framework, the era of everyone becoming a 'composer' may finally be upon us.
Related article
ByteDance’s Seed launches global campus drive, offering virtual shares to win top large model talent
In the competitive landscape of large language models, securing top-tier talent remains the most critical strategic asset.On April 1st, ByteDance announced the launch of its Seed global campus recruitment initiative, part of its large model talent de
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
Related Special Topic Recommendations
Comments (0)
0/500
The AI music landscape experienced another significant shift in early 2026. On March 9, the music foundation model SongGeneration2, developed jointly by

Three Breakthroughs: No More 'Plastic' AI Music
High musicality: Unlike basic melody stacking, this model handles complex multi-track arrangements with impressive spatial depth.
High lyric accuracy: Unclear pronunciation and hallucinated pitch shifts are now history. Its phoneme error rate (PER) is as low as 8.55%, notably outperforming the leading commercial model
Strong controllability: Whether through text descriptions or audio prompts, it precisely follows instructions, enabling deep customization of style and emotion.

'Dual-Core' Drive: A Dream Collaboration Between LLM and Diffusion Models
Architecturally,
Composing Brain (LeLM): Handles global structure planning and vocal details, solving the challenge of "how to sing".
High-Fidelity Renderer (Diffusion): Synthesizes extremely complex acoustic details under the guidance of the language model.
Hierarchical Representation: The first to adopt parallel modeling of mixed and multi-track representations, balancing melodic stability with sound quality refinement.
True Open Source, Low Barrier: Ordinary Computers Can Also 'Write Songs'
What excited developers most was Tencent's remarkable openness. The SongGeneration-v2-large model with 4 billion parameters is now fully open-source, supporting multilingual generation including Chinese and English. Remarkably, it runs smoothly on consumer-grade hardware with just 22GB VRAM, enabling local and private music creation.
To let users try it right away, the team also released the SongGeneration-v2-Fast version on HuggingFace, trading a minimal drop in audio quality for ultra-fast generation—producing a complete song in under a minute.
Based on the performance of
ByteDance’s Seed launches global campus drive, offering virtual shares to win top large model talent
In the competitive landscape of large language models, securing top-tier talent remains the most critical strategic asset.On April 1st, ByteDance announced the launch of its Seed global campus recruitment initiative, part of its large model talent de
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a











