Home
Cosmos3, First Fully Open-Source Multimodal Physical AI Model, Debuts; NVIDIA Co-Founds Cosmos Alliance
Significant progress has been made in physical artificial intelligence. On June 1, NVIDIA officially launched Cosmos3, an open-world foundational large model for physical AI. As the first fully open-source and multimodal physical AI large model in the world, it uses an innovative hybrid Transformer architecture that combines visual reasoning, world generation, and action prediction into a single system. This could dramatically shorten the physical AI training and evaluation cycle from months to just days.
Cosmos3 offers a fresh solution to the long-standing industry challenge of "difficulty generalizing in real-world scenarios with limited data and fragmented simulation frameworks." The model was trained on a massive physical AI dataset containing billions of text, images, videos, audio, and motion trajectories. It can naturally understand and generate cross-modal content, achieving industry-leading accuracy in physical simulation.

On the technical side, Cosmos3 innovatively combines a reasoning Transformer with a generative Transformer. The model first deeply analyzes object interaction rules, motion states, and spatiotemporal relationships, then accurately generates video and predicts action trajectories. This design gives it strong multimodal understanding of images and text, the ability to simulate and predict physical environments, and action strategy capabilities to help robots complete specific tasks. In several mainstream physical AI benchmarks, including Artificial Analysis, Physics-IQ, and RoboLab, Cosmos3 ranks among the top open-source models.
To accommodate different development stages, NVIDIA has released multiple versions: Cosmos3Super, focused on secondary training of robot and autonomous driving models with extreme precision, and Cosmos3Nano, which can perform high-quality video parsing and action reasoning in seconds. Both are now officially available. The Cosmos3Edge version, designed for real-time inference on edge devices, is also on the release roadmap.
Alongside the large model launch, NVIDIA co-founded the "NVIDIA Cosmos Coalition" with leading world model research teams and AI developers worldwide, including Agile Robots, Black Forest Labs, Generalist, LTX, Runway, and Skild AI. NVIDIA founder and CEO Jensen Huang stated that with continuous breakthroughs in multimodal reasoning and world models, the transformative era of physical AI has arrived. The release of these cutting-edge open-source models will help developers around the world achieve technological leaps and create the next generation of intelligent systems capable of perceiving, reasoning, and acting in the real world.
Related article
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
How to fix Core Web Vitals for better SEO rankings
Streamline Report Card Comments with AI ToolsIntroductionAI Tools for Generating Report Card CommentsMagic SchoolAlmanac AIChat GPTUsing Magic School to Generate Report Card CommentsLogging into Magic SchoolSelecting the Report Card Comments ToolCust
Related Special Topic Recommendations
Comments (0)
0/500
Significant progress has been made in physical artificial intelligence. On June 1, NVIDIA officially launched Cosmos3, an open-world foundational large model for physical AI. As the first fully open-source and multimodal physical AI large model in the world, it uses an innovative hybrid Transformer architecture that combines visual reasoning, world generation, and action prediction into a single system. This could dramatically shorten the physical AI training and evaluation cycle from months to just days.
Cosmos3 offers a fresh solution to the long-standing industry challenge of "difficulty generalizing in real-world scenarios with limited data and fragmented simulation frameworks." The model was trained on a massive physical AI dataset containing billions of text, images, videos, audio, and motion trajectories. It can naturally understand and generate cross-modal content, achieving industry-leading accuracy in physical simulation.

On the technical side, Cosmos3 innovatively combines a reasoning Transformer with a generative Transformer. The model first deeply analyzes object interaction rules, motion states, and spatiotemporal relationships, then accurately generates video and predicts action trajectories. This design gives it strong multimodal understanding of images and text, the ability to simulate and predict physical environments, and action strategy capabilities to help robots complete specific tasks. In several mainstream physical AI benchmarks, including Artificial Analysis, Physics-IQ, and RoboLab, Cosmos3 ranks among the top open-source models.
To accommodate different development stages, NVIDIA has released multiple versions: Cosmos3Super, focused on secondary training of robot and autonomous driving models with extreme precision, and Cosmos3Nano, which can perform high-quality video parsing and action reasoning in seconds. Both are now officially available. The Cosmos3Edge version, designed for real-time inference on edge devices, is also on the release roadmap.
Alongside the large model launch, NVIDIA co-founded the "NVIDIA Cosmos Coalition" with leading world model research teams and AI developers worldwide, including Agile Robots, Black Forest Labs, Generalist, LTX, Runway, and Skild AI. NVIDIA founder and CEO Jensen Huang stated that with continuous breakthroughs in multimodal reasoning and world models, the transformative era of physical AI has arrived. The release of these cutting-edge open-source models will help developers around the world achieve technological leaps and create the next generation of intelligent systems capable of perceiving, reasoning, and acting in the real world.
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
How to fix Core Web Vitals for better SEO rankings
Streamline Report Card Comments with AI ToolsIntroductionAI Tools for Generating Report Card CommentsMagic SchoolAlmanac AIChat GPTUsing Magic School to Generate Report Card CommentsLogging into Magic SchoolSelecting the Report Card Comments ToolCust











