LPM1.0 Model Generates Real-Time Interactive Digital Human Video from Single Image

Researchers have officially unveiled the LPM1.0 model, a project designed to generate real-time videos of people speaking, listening, and singing from a single reference image. Its key breakthrough lies in multimodal processing—synchronizing text, audio, and image inputs to produce dynamic scenes with accurate lip sync, nuanced facial expressions, and natural emotional transitions. The model integrates directly with leading voice AI platforms like ChatGPT and Doubao, transforming conventional voice conversations into real-time interactive experiences with visual feedback.
Technically, LPM1.0 employs "multi-granularity identity conditioning" to extract detailed features from reference materials across multiple angles and expressions, eliminating the need for the model to independently generate complex elements like teeth, wrinkles, or side profiles. This greatly boosts cross-style processing—enabling instant drive for photorealistic faces, animations, or 3D game characters without retraining. The model also supports streaming transmission, ensuring system stability even when generating videos up to 45 minutes long.
Regarding interaction logic, LPM1.0 accurately detects three conversational states: while listening, it produces responsive expressions like nodding or shifting gaze; while speaking, it drives body and lip movements based on audio; while idle, it generates natural resting behaviors according to text prompts. Project manager Zeng Ailing noted that LPM1.0 is suitable not only for real-time conversation but also for offline audio-driven video generation, offering technical redundancy for podcasts and film production.
Despite its strong application potential, the development team stresses that LPM1.0 remains a research project with no current plans to release code or weights publicly. Researchers acknowledge a qualitative gap between generated videos and real footage, and the inherent deepfake risks cannot be overlooked. The study's significance lies in charting the future evolution of AI systems—shifting from single logical interaction to multidimensional interaction incorporating emotional responses, eye contact, and visual embodiment.
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (0)
0/500

Researchers have officially unveiled the LPM1.0 model, a project designed to generate real-time videos of people speaking, listening, and singing from a single reference image. Its key breakthrough lies in multimodal processing—synchronizing text, audio, and image inputs to produce dynamic scenes with accurate lip sync, nuanced facial expressions, and natural emotional transitions. The model integrates directly with leading voice AI platforms like ChatGPT and Doubao, transforming conventional voice conversations into real-time interactive experiences with visual feedback.
Technically, LPM1.0 employs "multi-granularity identity conditioning" to extract detailed features from reference materials across multiple angles and expressions, eliminating the need for the model to independently generate complex elements like teeth, wrinkles, or side profiles. This greatly boosts cross-style processing—enabling instant drive for photorealistic faces, animations, or 3D game characters without retraining. The model also supports streaming transmission, ensuring system stability even when generating videos up to 45 minutes long.
Regarding interaction logic, LPM1.0 accurately detects three conversational states: while listening, it produces responsive expressions like nodding or shifting gaze; while speaking, it drives body and lip movements based on audio; while idle, it generates natural resting behaviors according to text prompts. Project manager Zeng Ailing noted that LPM1.0 is suitable not only for real-time conversation but also for offline audio-driven video generation, offering technical redundancy for podcasts and film production.
Despite its strong application potential, the development team stresses that LPM1.0 remains a research project with no current plans to release code or weights publicly. Researchers acknowledge a qualitative gap between generated videos and real footage, and the inherent deepfake risks cannot be overlooked. The study's significance lies in charting the future evolution of AI systems—shifting from single logical interaction to multidimensional interaction incorporating emotional responses, eye contact, and visual embodiment.
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur





Home






