Home
Tencent Hunyuan and Partners Launch MMAE Benchmark, Revealing Under 5% Precision in AI Audio Editing

Artificial intelligence has made significant strides in audio generation, but editing existing audio remains a major challenge. Recently, Tencent Hy, in partnership with leading research institutions including Shanghai Jiao Tong University (SJTU), Nanyang Technological University (NTU), Tianjin University (TJU), Peking University (PKU), and Fudan University (FDU), introduced MMAE (Massive Multitask Audio Editing Benchmark)—the first large-scale, multi-task benchmark designed for general instruction-driven audio editing. This release establishes a systematic evaluation standard for the AI audio editing field, revealing clear gaps in the ability of current technology to perform precise modifications.
From "Generation" to "Editing": The Real Challenge for AI Audio
Traditional audio AI focuses on generating new content from text or prompts. The core of the MMAE benchmark, however, requires models to comprehend existing audio clips and make precise adjustments based on natural language instructions: modify only the parts that need changing, while leaving everything else untouched. This "editing, not reconstruction" approach demands higher audio fidelity, better instruction following, and deeper context understanding—making it far more aligned with real-world applications like podcast post-production, music mixing, and voice personalization.
Tests reveal that current mainstream models achieve an overall Exact Match Rate (EMR) below 5%, highlighting significant gaps in reliable audio editing technology. This means AI tends to over-modify, miss instructions, or degrade original audio quality when tackling practical editing tasks.
MMAE Benchmark Highlights: Multi-Dimensional Evaluation for Real-World Scenarios
The MMAE benchmark is designed with thoroughness and rigor, featuring the following key elements:
2,000 high-fidelity samples: All sourced from real-world scenarios, ensuring practical and diverse evaluation.17,741 fine-grained evaluation metrics: Provide a detailed scoring rubric for objective quantification.7 modal settings: Cover sound, music, speech, and their combinations, supporting testing in complex audio environments.6 levels of task complexity: Range from basic modifications to multi-hop reasoning and multi-round editing, enabling comprehensive assessment of model capabilities.8 types of operations: Support editing tasks at different granularities, both local and global, challenging the model's fine control.AIbase Comment: MMAE is not just a technical evaluation tool—it is a milestone in driving audio AI from "generative" to "editable." It provides a unified standard for researchers and developers, and is expected to accelerate the next generation of audio editing models.
Future Outlook: Audio Editing Could Become a Core Strength of AI Multimodal Systems
As multimodal large models evolve rapidly, precise audio editing will play a vital role in content creation, film post-production, and accessibility tools. The recent collaboration between Tencent Hy and other institutions highlights China's leading position in audio AI research. The industry anticipates more open-source resources and follow-up models to collectively fill this technological gap.
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (0)
0/500

Artificial intelligence has made significant strides in audio generation, but editing existing audio remains a major challenge. Recently, Tencent Hy, in partnership with leading research institutions including Shanghai Jiao Tong University (SJTU), Nanyang Technological University (NTU), Tianjin University (TJU), Peking University (PKU), and Fudan University (FDU), introduced MMAE (Massive Multitask Audio Editing Benchmark)—the first large-scale, multi-task benchmark designed for general instruction-driven audio editing. This release establishes a systematic evaluation standard for the AI audio editing field, revealing clear gaps in the ability of current technology to perform precise modifications.
From "Generation" to "Editing": The Real Challenge for AI Audio
Traditional audio AI focuses on generating new content from text or prompts. The core of the MMAE benchmark, however, requires models to comprehend existing audio clips and make precise adjustments based on natural language instructions: modify only the parts that need changing, while leaving everything else untouched. This "editing, not reconstruction" approach demands higher audio fidelity, better instruction following, and deeper context understanding—making it far more aligned with real-world applications like podcast post-production, music mixing, and voice personalization.
Tests reveal that current mainstream models achieve an overall Exact Match Rate (EMR) below 5%, highlighting significant gaps in reliable audio editing technology. This means AI tends to over-modify, miss instructions, or degrade original audio quality when tackling practical editing tasks.
MMAE Benchmark Highlights: Multi-Dimensional Evaluation for Real-World Scenarios
The MMAE benchmark is designed with thoroughness and rigor, featuring the following key elements:
2,000 high-fidelity samples: All sourced from real-world scenarios, ensuring practical and diverse evaluation.17,741 fine-grained evaluation metrics: Provide a detailed scoring rubric for objective quantification.7 modal settings: Cover sound, music, speech, and their combinations, supporting testing in complex audio environments.6 levels of task complexity: Range from basic modifications to multi-hop reasoning and multi-round editing, enabling comprehensive assessment of model capabilities.8 types of operations: Support editing tasks at different granularities, both local and global, challenging the model's fine control.AIbase Comment: MMAE is not just a technical evaluation tool—it is a milestone in driving audio AI from "generative" to "editable." It provides a unified standard for researchers and developers, and is expected to accelerate the next generation of audio editing models.
Future Outlook: Audio Editing Could Become a Core Strength of AI Multimodal Systems
As multimodal large models evolve rapidly, precise audio editing will play a vital role in content creation, film post-production, and accessibility tools. The recent collaboration between Tencent Hy and other institutions highlights China's leading position in audio AI research. The industry anticipates more open-source resources and follow-up models to collectively fill this technological gap.
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur











