Home
AntGroup Open-Sources LingBot-Video, the First Video Foundation Model for Embodied Intelligence
On July 9, Ant Group open-sourced LingBot-Video, the world's first open-source video generation foundation model built on the Mixture-of-Experts (MoE) architecture and purpose-built for embodied intelligence. The model redefines video pre-training by centering on the core needs of robots and embodied intelligence, delivering systematic gains in reasoning efficiency, physical plausibility, action understanding, and task completion. It offers a new open-source foundation for video foundation models to move beyond digital content creation into embodied intelligence.
On the RBench benchmark, jointly developed by Peking University and ByteDance, LingBot-Video scored 0.620, outperforming Wan2.6 (0.607), Seedance1.5Pro (0.584), and Cosmos3Super (0.581). RBench, a comprehensive evaluation benchmark for robotic manipulation videos, specifically assesses whether a model can generate robot behaviors that follow real-world physics. This result indicates that LingBot-Video is better at preserving action rationality and task execution integrity when generating robot-related videos.

(Figure caption: LingBot-Video achieves the best performance on RBench)
To further verify LingBot-Video's ability to model changes in the physical world, Ant Group evaluated it on an internal benchmark across two dimensions: general quality and embodied domain. The results show that compared to five open-source models—NVIDIA Cosmos3, Wan2.2A14B, LongCat-Video, Hunyuan Video1.5, and LTX-2.3—LingBot-Video outperforms the major baseline models in the embodied domain.

(Figure caption: Comprehensive evaluation shows LingBot-Video exhibits stronger physical understanding and action consistency in embodied scenarios)
In recent years, video generation models have advanced rapidly in visual quality, smoothness, and creative expression. However, for embodied intelligence, a video that looks realistic and moves smoothly may not reflect real physical laws, making it difficult to support continuous prediction, planning, and task execution for robots. At the same time, embodied intelligence also demands higher reasoning efficiency to adapt to real-time interaction and control loops.
This has driven video generation to evolve along two distinct paths: one toward cinemas, serving content creation; the other toward robots, serving the understanding, prediction, and interaction with the physical world. LingBot-Video represents Ant Group's significant exploration of a new path for video generation tailored to embodied intelligence.
LingBot-Video introduces systematic innovations in architecture, data, and training.
Architecturally, LingBot-Video employs a DiT + MoE design, replacing traditional dense architectures with MoE. This allows the model to scale capacity while keeping inference costs manageable. With 30 billion parameters, the model activates only about 3 billion during generation, offering roughly three times the reasoning efficiency of a dense architecture of the same size. This design gives the model visual expressiveness from large-scale parameters while better meeting the efficiency demands of embodied intelligence.
For data, LingBot-Video built a data profiling engine, incorporating robot-related data such as VLA, VLN, and Ego on top of massive internet videos, covering scenarios including dexterous manipulation, robot mobility, and first-person interactions, totaling 70,000 hours of embodied data. This data helps the model learn the relationship between actions and environmental changes, rather than just the surface texture and visual style of videos.

In training, LingBot-Video introduced a multi-dimensional reinforcement learning reward system. Beyond conventional metrics like aesthetics, prompt following, and motion consistency, the model also aligns with physical plausibility and task completion, making the generated results more consistent with real-world physics and closer to the needs of robots performing tasks in the real world.
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (0)
0/500
On July 9, Ant Group open-sourced LingBot-Video, the world's first open-source video generation foundation model built on the Mixture-of-Experts (MoE) architecture and purpose-built for embodied intelligence. The model redefines video pre-training by centering on the core needs of robots and embodied intelligence, delivering systematic gains in reasoning efficiency, physical plausibility, action understanding, and task completion. It offers a new open-source foundation for video foundation models to move beyond digital content creation into embodied intelligence.
On the RBench benchmark, jointly developed by Peking University and ByteDance, LingBot-Video scored 0.620, outperforming Wan2.6 (0.607), Seedance1.5Pro (0.584), and Cosmos3Super (0.581). RBench, a comprehensive evaluation benchmark for robotic manipulation videos, specifically assesses whether a model can generate robot behaviors that follow real-world physics. This result indicates that LingBot-Video is better at preserving action rationality and task execution integrity when generating robot-related videos.

(Figure caption: LingBot-Video achieves the best performance on RBench)
To further verify LingBot-Video's ability to model changes in the physical world, Ant Group evaluated it on an internal benchmark across two dimensions: general quality and embodied domain. The results show that compared to five open-source models—NVIDIA Cosmos3, Wan2.2A14B, LongCat-Video, Hunyuan Video1.5, and LTX-2.3—LingBot-Video outperforms the major baseline models in the embodied domain.

(Figure caption: Comprehensive evaluation shows LingBot-Video exhibits stronger physical understanding and action consistency in embodied scenarios)
In recent years, video generation models have advanced rapidly in visual quality, smoothness, and creative expression. However, for embodied intelligence, a video that looks realistic and moves smoothly may not reflect real physical laws, making it difficult to support continuous prediction, planning, and task execution for robots. At the same time, embodied intelligence also demands higher reasoning efficiency to adapt to real-time interaction and control loops.
This has driven video generation to evolve along two distinct paths: one toward cinemas, serving content creation; the other toward robots, serving the understanding, prediction, and interaction with the physical world. LingBot-Video represents Ant Group's significant exploration of a new path for video generation tailored to embodied intelligence.
LingBot-Video introduces systematic innovations in architecture, data, and training.
Architecturally, LingBot-Video employs a DiT + MoE design, replacing traditional dense architectures with MoE. This allows the model to scale capacity while keeping inference costs manageable. With 30 billion parameters, the model activates only about 3 billion during generation, offering roughly three times the reasoning efficiency of a dense architecture of the same size. This design gives the model visual expressiveness from large-scale parameters while better meeting the efficiency demands of embodied intelligence.
For data, LingBot-Video built a data profiling engine, incorporating robot-related data such as VLA, VLN, and Ego on top of massive internet videos, covering scenarios including dexterous manipulation, robot mobility, and first-person interactions, totaling 70,000 hours of embodied data. This data helps the model learn the relationship between actions and environmental changes, rather than just the surface texture and visual style of videos.

In training, LingBot-Video introduced a multi-dimensional reinforcement learning reward system. Beyond conventional metrics like aesthetics, prompt following, and motion consistency, the model also aligns with physical plausibility and task completion, making the generated results more consistent with real-world physics and closer to the needs of robots performing tasks in the real world.
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur











