Home
Open-Source Visual Large Models Upgrade Multimodal Generation, 2K HD Audio and Video Set to Revolutionize Industry

On August 3, AI company MiniMax announced the open-source release of its next-generation general video model, MiniMax H3 . This full-modal generation system transcends traditional single-task boundaries, deeply integrating and understanding multimodal contexts of text, images, video, and audio, showcasing strong task generalization capabilities.
For generation, H3 supports output durations from 4 to 15 seconds at 24 FPS and 32 kHz stereo audio. Users can choose from common aspect ratios like 21:9, 16:9, 4:3, or 1:1, with a default shorter side of 768 pixels for basic creation. The dedicated H3-Regenerate-2K solution enables stunning images up to 2K resolution. For multilingual interaction, it reliably supports 11 major languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Spanish, Portuguese, and Russian.
To accommodate diverse creative scenarios, H3 offers flexible model versions and rich input options. The H3-Base-FL2VA first-last frame mode accepts 0 to 2 images, handling tasks like text-to-video, first-frame-to-video, last-frame-to-video, and first-last frame collaboration. The H3-Base-Ref2VA full-modal reference mode, meanwhile, supports mixed input of up to 9 images, 3 video clips (total duration under 15 seconds), and 3 audio segments—up to 12 files total—giving professional creators unprecedented fine control.
Under the hood, the MiniMax H3 system comprises three core modules. First, the H3-Context-IR (context intermediate representation) deeply analyzes complex multi-instructions and converts them into model-friendly structures, serving as a key link for high-quality output. Next, the H3-Base module generates base audio and video at 768p resolution. Finally, the H3-Regenerate-2K module achieves 2K super-resolution reconstruction by feeding back initial results and original context. This combination preserves vivid details while greatly enhancing visual fidelity.
The model is now officially released under the MiniMax H3 Community License. Developers and tech enthusiasts can access the full open-source code and resources via Hugging Face . Industry experts believe this open-source release will further lower the barriers to multimodal video generation, propelling film production, visual design, and AI interaction to new heights.
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (0)
0/500

On August 3, AI company MiniMax announced the open-source release of its next-generation general video model,
For generation, H3 supports output durations from 4 to 15 seconds at 24 FPS and 32 kHz stereo audio. Users can choose from common aspect ratios like 21:9, 16:9, 4:3, or 1:1, with a default shorter side of 768 pixels for basic creation. The dedicated H3-Regenerate-2K solution enables stunning images up to 2K resolution. For multilingual interaction, it reliably supports 11 major languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Spanish, Portuguese, and Russian.
To accommodate diverse creative scenarios, H3 offers flexible model versions and rich input options. The H3-Base-FL2VA first-last frame mode accepts 0 to 2 images, handling tasks like text-to-video, first-frame-to-video, last-frame-to-video, and first-last frame collaboration. The H3-Base-Ref2VA full-modal reference mode, meanwhile, supports mixed input of up to 9 images, 3 video clips (total duration under 15 seconds), and 3 audio segments—up to 12 files total—giving professional creators unprecedented fine control.
Under the hood, the MiniMax H3 system comprises three core modules. First, the H3-Context-IR (context intermediate representation) deeply analyzes complex multi-instructions and converts them into model-friendly structures, serving as a key link for high-quality output. Next, the H3-Base module generates base audio and video at 768p resolution. Finally, the H3-Regenerate-2K module achieves 2K super-resolution reconstruction by feeding back initial results and original context. This combination preserves vivid details while greatly enhancing visual fidelity.
The model is now officially released under the MiniMax H3 Community License. Developers and tech enthusiasts can access the full open-source code and resources via
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur











