Home
27B Math SOTA and 3-Second Emotional Cloning: Youdao Open-Sources Zi Yue 4 Multimodal and TTS Engine
NetEase Youdao recently announced a comprehensive upgrade of its "Confucius4" large model to version 4.0. Now fully multimodal, Confucius4 supports integrated text, image, and audio interactions. Youdao also open-sourced its core multimodal and text-to-speech (TTS) models, while the translation model underwent a deep technical overhaul, delivering improvements in both quality and efficiency.
The multimodal model achieves SOTA in vision and mathematics, with superior performance on pure-text math problems
According to the announcement, the open-source Confucius4 multimodal model, with 27 billion parameters, has brought visual-input-based math capabilities to an industry-leading level (SOTA) in educational scenarios. Among models of similar scale, Confucius4 excels at handling advanced visual math and physics problems that incorporate charts. It also shows significant improvement on Chinese pure-text math problems, achieving an accuracy of 81.4%, which is industry-leading.

▲ Confucius4 achieves best-in-class results on multiple visual math benchmarks among models of the same scale
Image source: https://huggingface.co/netease-youdao/Confucius4
A more critical breakthrough lies in practical cost-effectiveness. According to relevant officials, the new model employs a refined reasoning chain reconstruction scheme. By aggregating a large volume of high-quality, concise reasoning samples for deep optimization, it compresses the output length of the reasoning chain by 43.2%.
This means it can deliver answers faster with fewer tokens and shorter reasoning paths, significantly reducing inference costs in real-world business scenarios for enterprises and developers.

▲ Confucius4 significantly reduces output tokens on multiple visual math benchmarks
Image source: https://huggingface.co/netease-youdao/Confucius4
Additionally, the Confucius4 research team deeply optimized the model for real-world homework, exams, and problem-solving scenarios faced by Chinese students. This enables it to address authentic learning challenges, making it a more empathetic digital assistant.
Open-source TTS: supports 14 languages, clones original voice in 3 seconds with no accent issues across languages
Alongside the multimodal model, the speech synthesis (TTS) engine was also open-sourced. Built on a cutting-edge "speech encoder + LLM" architecture, it offers developers and content creators zero-shot, low-barrier voice cloning and emotional synthesis capabilities.
Currently, it fully supports 14 languages: Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, and Vietnamese. The system can naturally transfer a speaker's voice across different languages without additional training, maintaining voice consistency while ensuring the synthesized results sound native and fluent, with no accent leakage during cross-language cloning.
For voice cloning, Confucius4 achieves full "upload and clone" support. Users simply provide any audio material, and the system replicates the original voice within three seconds. According to the announcement, the engine's accuracy on cloning tasks exceeds 97%, and the similarity between the cloned voice and the original voice is over 85%. It preserves the speaker's unique vocal characteristics while accurately reproducing their emotional tone, with comprehensive capabilities that rank among the top in the field.
Furthermore, this open-source model demonstrates strong robustness in real multilingual scenarios. It can handle various synthesis needs, including daily conversations, news broadcasts, corporate promotions, and complex emotional expressions.
Translation model quality upgraded comprehensively, with 80% faster inference speed
As one of Youdao's most deeply rooted technological assets, the translation model also received significant technical upgrades in this update, further enhancing its performance on translation tasks.
Regarding data, the Confucius team collected and cleaned billions of multilingual items and hired professionals with TEM-8 certification for multidimensional manual evaluation, ensuring high-quality corpora from the start.
On the algorithmic side, the model uses an innovative "multi-expert OPD" mode, adopting a smarter "soft approach" to leverage the strengths of various experts. It also introduces format rewards and language detection mechanisms through reinforcement learning, effectively resolving common issues like out-of-context translations and mixed-language outputs in machine translation.
To meet the demands of high-frequency, high-concurrency industrial applications, the updated translation model is equipped with an efficient acceleration mechanism that directly boosts overall inference speed by 80%. Combined with a customized solution of automated large model evaluation and random manual sampling, the new generation translation model demonstrates extremely high standards of speed and quality across multiple scenarios, including text, image, and document translation.
Reflecting on Youdao's journey in AI, from the initial launch of Confucius as the first education-focused large model, introducing the "virtual speaking coach Hi Echo" that overturned traditional oral practice methods, to the comprehensive integration of Confucius 2.0 and 3.0 into software and hardware ecosystems, Youdao has consistently led AI empowerment in real-world scenarios. In 2026, Youdao accelerated AI application with a series of AI Agent products such as LobsterAI, Youdao Treasure, Youdao Conference Agent, and Thinkflow, realizing a forward-looking full-scenario AI Agent matrix.
The upgrade of Confucius4 and the full open-sourcing of core models not only significantly lower barriers for developers in multimodal and speech synthesis fields but also demonstrate an ecological closed loop where underlying core technologies nurture upper-layer Agent matrices. Youdao hopes that, with the joint contributions of global developers and the open-source community, this full-modal large model ecosystem will unleash true productivity transformation across a wide range of industries.
Appendix: Open-source addresses:
"Confucius4" multimodal model: https://huggingface.co/netease-youdao/Confucius4
"Confucius4" TTS model: https://github.com/netease-youdao/Confucius4-TTS
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (1)
0/500
NetEase Youdao recently announced a comprehensive upgrade of its "Confucius4" large model to version 4.0. Now fully multimodal, Confucius4 supports integrated text, image, and audio interactions. Youdao also open-sourced its core multimodal and text-to-speech (TTS) models, while the translation model underwent a deep technical overhaul, delivering improvements in both quality and efficiency.
The multimodal model achieves SOTA in vision and mathematics, with superior performance on pure-text math problems
According to the announcement, the open-source Confucius4 multimodal model, with 27 billion parameters, has brought visual-input-based math capabilities to an industry-leading level (SOTA) in educational scenarios. Among models of similar scale, Confucius4 excels at handling advanced visual math and physics problems that incorporate charts. It also shows significant improvement on Chinese pure-text math problems, achieving an accuracy of 81.4%, which is industry-leading.

▲ Confucius4 achieves best-in-class results on multiple visual math benchmarks among models of the same scale
Image source: https://huggingface.co/netease-youdao/Confucius4
A more critical breakthrough lies in practical cost-effectiveness. According to relevant officials, the new model employs a refined reasoning chain reconstruction scheme. By aggregating a large volume of high-quality, concise reasoning samples for deep optimization, it compresses the output length of the reasoning chain by 43.2%.
This means it can deliver answers faster with fewer tokens and shorter reasoning paths, significantly reducing inference costs in real-world business scenarios for enterprises and developers.

▲ Confucius4 significantly reduces output tokens on multiple visual math benchmarks
Image source: https://huggingface.co/netease-youdao/Confucius4
Additionally, the Confucius4 research team deeply optimized the model for real-world homework, exams, and problem-solving scenarios faced by Chinese students. This enables it to address authentic learning challenges, making it a more empathetic digital assistant.
Open-source TTS: supports 14 languages, clones original voice in 3 seconds with no accent issues across languages
Alongside the multimodal model, the speech synthesis (TTS) engine was also open-sourced. Built on a cutting-edge "speech encoder + LLM" architecture, it offers developers and content creators zero-shot, low-barrier voice cloning and emotional synthesis capabilities.
Currently, it fully supports 14 languages: Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, and Vietnamese. The system can naturally transfer a speaker's voice across different languages without additional training, maintaining voice consistency while ensuring the synthesized results sound native and fluent, with no accent leakage during cross-language cloning.
For voice cloning, Confucius4 achieves full "upload and clone" support. Users simply provide any audio material, and the system replicates the original voice within three seconds. According to the announcement, the engine's accuracy on cloning tasks exceeds 97%, and the similarity between the cloned voice and the original voice is over 85%. It preserves the speaker's unique vocal characteristics while accurately reproducing their emotional tone, with comprehensive capabilities that rank among the top in the field.
Furthermore, this open-source model demonstrates strong robustness in real multilingual scenarios. It can handle various synthesis needs, including daily conversations, news broadcasts, corporate promotions, and complex emotional expressions.
Translation model quality upgraded comprehensively, with 80% faster inference speed
As one of Youdao's most deeply rooted technological assets, the translation model also received significant technical upgrades in this update, further enhancing its performance on translation tasks.
Regarding data, the Confucius team collected and cleaned billions of multilingual items and hired professionals with TEM-8 certification for multidimensional manual evaluation, ensuring high-quality corpora from the start.
On the algorithmic side, the model uses an innovative "multi-expert OPD" mode, adopting a smarter "soft approach" to leverage the strengths of various experts. It also introduces format rewards and language detection mechanisms through reinforcement learning, effectively resolving common issues like out-of-context translations and mixed-language outputs in machine translation.
To meet the demands of high-frequency, high-concurrency industrial applications, the updated translation model is equipped with an efficient acceleration mechanism that directly boosts overall inference speed by 80%. Combined with a customized solution of automated large model evaluation and random manual sampling, the new generation translation model demonstrates extremely high standards of speed and quality across multiple scenarios, including text, image, and document translation.
Reflecting on Youdao's journey in AI, from the initial launch of Confucius as the first education-focused large model, introducing the "virtual speaking coach Hi Echo" that overturned traditional oral practice methods, to the comprehensive integration of Confucius 2.0 and 3.0 into software and hardware ecosystems, Youdao has consistently led AI empowerment in real-world scenarios. In 2026, Youdao accelerated AI application with a series of AI Agent products such as LobsterAI, Youdao Treasure, Youdao Conference Agent, and Thinkflow, realizing a forward-looking full-scenario AI Agent matrix.
The upgrade of Confucius4 and the full open-sourcing of core models not only significantly lower barriers for developers in multimodal and speech synthesis fields but also demonstrate an ecological closed loop where underlying core technologies nurture upper-layer Agent matrices. Youdao hopes that, with the joint contributions of global developers and the open-source community, this full-modal large model ecosystem will unleash true productivity transformation across a wide range of industries.
Appendix: Open-source addresses:
"Confucius4" multimodal model: https://huggingface.co/netease-youdao/Confucius4
"Confucius4" TTS model: https://github.com/netease-youdao/Confucius4-TTS
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur











