PhiZero pioneers physical language, enabling AI world models to think like humans
In today’s landscape of large models, video generation is increasingly viewed as the foundational simulator for embodied intelligence. Yet, most mainstream approaches rely on direct pixel-space prediction, often resulting in unstable or inconsistent physical laws. To solve this critical industry challenge, our research team has officially introduced PhiZero , a new physical world model. Moving beyond traditional blind prediction, PhiZero innovatively incorporates "physical language," empowering AI to conduct explicit physical reasoning prior to video generation.
At its core, PhiZero introduces a discrete representation known as "physical language." Through self-supervised learning on vast outdoor video datasets, this language extracts state transition patterns, capturing dynamic evolution rules across diverse visual contexts. Leveraging this, the model adopts a "reasoning-first, then-rendering" architecture: it first autonomously infers future world evolution trajectories within the physical language space, then passes these to a specialized diffusion decoder for high-definition video synthesis. This separation of underlying dynamics simulation from pixel-level synthesis significantly enhances physical consistency.

Technically, the system comprises two key modules: a tokenizer and a reasoner. The physical language tokenizer employs a transition-level Q-Former to capture changes between adjacent states, generating compact discrete symbols via limited scalar quantization. The diffusion decoder then reconstructs detailed visuals based on the initial frame and these discrete symbols. Meanwhile, the physical language reasoner, initialized with a pre-trained vision-language model, accurately predicts future physical language sequences given the first frame and text-based action intents.

Benchmark evaluations confirm that PhiZero achieves state-of-the-art performance across multiple physical reasoning and video generation tests, including Physics-IQ Verified, PhyGround, and WorldModelBench. Whether handling object displacement from collisions, gravity-driven deformation, or complex chain reactions, the model exhibits highly realistic physical plausibility.
As embodied intelligence and robotics accelerate, the release of PhiZero paves a new path for controllable, interactive world modeling. It is poised to play a pivotal role in long-term simulation and motion transfer scenarios.
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (0)
0/500
In today’s landscape of large models, video generation is increasingly viewed as the foundational simulator for embodied intelligence. Yet, most mainstream approaches rely on direct pixel-space prediction, often resulting in unstable or inconsistent physical laws. To solve this critical industry challenge, our research team has officially introduced
At its core,

Technically, the system comprises two key modules: a tokenizer and a reasoner. The physical language tokenizer employs a transition-level Q-Former to capture changes between adjacent states, generating compact discrete symbols via limited scalar quantization. The diffusion decoder then reconstructs detailed visuals based on the initial frame and these discrete symbols. Meanwhile, the physical language reasoner, initialized with a pre-trained vision-language model, accurately predicts future physical language sequences given the first frame and text-based action intents.

Benchmark evaluations confirm that
As embodied intelligence and robotics accelerate, the release of
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur





Home






