Home
Google breaks multimodal switching barrier by integrating native computer operations into Gemini 3.5 Flash

The Google DeepMind team announced a major breakthrough, integrating native computer-use capabilities directly into the Gemini 3.5 Flash model. Developers can now build AI agents that autonomously view and perform actions on screens across browsers, mobile devices, and desktops using a single model.
Previously, this capability was only available as a separate model, requiring developers to handle complex switching and context transfers between different models. With native integration, the AI no longer needs to manually pass information when executing long-running cross-platform tasks, greatly simplifying the development process.
Eliminating Context Loss to Improve Agent Reliability
Google's team believes the core bottleneck for AI agents is not the limitation of individual tools, but the loss of context information when switching between multiple tools. By unifying search, maps, and computer operations within a single model architecture, context flows continuously, significantly reducing the likelihood of failure in complex tasks.
This "multi-tool integration" design is akin to constructing a single building with internal connections, eliminating the lengthy and error-prone communication between separate buildings. This architectural shift has the potential to bring substantial improvements in the reliability and response latency of agent-based tasks.
Focusing on Three Core Scenarios with Multi-Layered Security Defenses
This native capability will primarily be applied to three core scenarios: automated tasks that require continuous operation for hours or days, continuous software testing for automatic UI consistency verification, and knowledge-intensive work spanning multiple applications. These scenarios rely heavily on the continuity of context across tasks and can effectively replace humans in performing repetitive, high-energy operations.
In terms of security, Google has adopted a multi-layered defense strategy, including targeted adversarial training, enterprise security safeguards for sensitive operations, and indirect prompt injection detection. These mechanisms collectively help enterprise users establish a relatively complete security boundary in open and uncontrolled computer environments.
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (0)
0/500

The Google DeepMind team announced a major breakthrough, integrating native computer-use capabilities directly into the Gemini 3.5 Flash model. Developers can now build AI agents that autonomously view and perform actions on screens across browsers, mobile devices, and desktops using a single model.
Previously, this capability was only available as a separate model, requiring developers to handle complex switching and context transfers between different models. With native integration, the AI no longer needs to manually pass information when executing long-running cross-platform tasks, greatly simplifying the development process.
Eliminating Context Loss to Improve Agent Reliability
Google's team believes the core bottleneck for AI agents is not the limitation of individual tools, but the loss of context information when switching between multiple tools. By unifying search, maps, and computer operations within a single model architecture, context flows continuously, significantly reducing the likelihood of failure in complex tasks.
This "multi-tool integration" design is akin to constructing a single building with internal connections, eliminating the lengthy and error-prone communication between separate buildings. This architectural shift has the potential to bring substantial improvements in the reliability and response latency of agent-based tasks.
Focusing on Three Core Scenarios with Multi-Layered Security Defenses
This native capability will primarily be applied to three core scenarios: automated tasks that require continuous operation for hours or days, continuous software testing for automatic UI consistency verification, and knowledge-intensive work spanning multiple applications. These scenarios rely heavily on the continuity of context across tasks and can effectively replace humans in performing repetitive, high-energy operations.
In terms of security, Google has adopted a multi-layered defense strategy, including targeted adversarial training, enterprise security safeguards for sensitive operations, and indirect prompt injection detection. These mechanisms collectively help enterprise users establish a relatively complete security boundary in open and uncontrolled computer environments.
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur











