Home
Alibaba's Tongyi Unveils Fun-CineForge: Open-Source AI Model Achieves Film-Grade Voice Synthesis
Alibaba Tongyi Lab officially launched and open-sourced the film-grade, multi-scenario voice synthesis multimodal model Fun-CineForge on March 16. This model addresses core challenges in AI dubbing, including lip-sync mismatch, lack of emotional expression, and inconsistent voice characteristics across multiple characters. It also introduces a high-quality method for dataset construction.

Technically, Fun-CineForge pioneers the concept of "temporal modality." Unlike conventional models that focus solely on text or visuals, it ensures voice synthesis occurs within precise time intervals through accurate timestamp control. Even in complex film scenes with occluded characters, frequent camera cuts, or blurred faces, the model maintains a high degree of audio-visual synchronization and adherence to instructions.
The accompanying open-source CineDub dataset construction pipeline is another key innovation. Tongyi Lab employed large language model chain-of-thought reasoning to automatically transform raw film footage into structured data, significantly reducing the need for manual annotation. This process achieves a word error rate of approximately 1% and a speaker diarization error rate of just 1.20%, providing a highly competitive training foundation for large models.

Fun-CineForge is now available on GitHub, HuggingFace, and the ModelScope community, supporting inference for video clips up to 30 seconds long. It excels not only in single-speaker monologues but also offers professional-grade support for duet and multi-speaker dialogue scenarios. This advancement signals AI voice technology's evolution from basic customer service and assistant roles into high-standard animation and film post-production.
GitHub: https://github.com/FunAudioLLM/FunCineForge
HuggingFace: https://huggingface.co/FunAudioLLM/Fun-CineForge
ModelScope: https://www.modelscope.cn/models/FunAudioLLM/Fun-CineForge/
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (1)
0/500
Just tried the demo and honestly blown away by how natural the lip-sync feels now! 😮 Always thought AI dubbing sounded a bit robotic, but this seems like a huge leap. Wonder if this will start being used in indie films or even gaming soon? The open-source move is pretty bold too—curious to see how other companies respond.
Alibaba Tongyi Lab officially launched and open-sourced the film-grade, multi-scenario voice synthesis multimodal model Fun-CineForge on March 16. This model addresses core challenges in AI dubbing, including lip-sync mismatch, lack of emotional expression, and inconsistent voice characteristics across multiple characters. It also introduces a high-quality method for dataset construction.

Technically, Fun-CineForge pioneers the concept of "temporal modality." Unlike conventional models that focus solely on text or visuals, it ensures voice synthesis occurs within precise time intervals through accurate timestamp control. Even in complex film scenes with occluded characters, frequent camera cuts, or blurred faces, the model maintains a high degree of audio-visual synchronization and adherence to instructions.
The accompanying open-source CineDub dataset construction pipeline is another key innovation. Tongyi Lab employed large language model chain-of-thought reasoning to automatically transform raw film footage into structured data, significantly reducing the need for manual annotation. This process achieves a word error rate of approximately 1% and a speaker diarization error rate of just 1.20%, providing a highly competitive training foundation for large models.

Fun-CineForge is now available on GitHub, HuggingFace, and the ModelScope community, supporting inference for video clips up to 30 seconds long. It excels not only in single-speaker monologues but also offers professional-grade support for duet and multi-speaker dialogue scenarios. This advancement signals AI voice technology's evolution from basic customer service and assistant roles into high-standard animation and film post-production.
GitHub: https://github.com/FunAudioLLM/FunCineForge
HuggingFace: https://huggingface.co/FunAudioLLM/Fun-CineForge
ModelScope: https://www.modelscope.cn/models/FunAudioLLM/Fun-CineForge/
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Just tried the demo and honestly blown away by how natural the lip-sync feels now! 😮 Always thought AI dubbing sounded a bit robotic, but this seems like a huge leap. Wonder if this will start being used in indie films or even gaming soon? The open-source move is pretty bold too—curious to see how other companies respond.











