Alibaba Unveils Qwen-UI-Agent, a Real-World Multimodal Foundation Model
Alibaba has officially introduced Qwen-UI-Agent, a foundation model for GUI agents designed to operate in real-world settings. Covering mobile, desktop, web, and deep search environments, this model moves beyond traditional simulation testing to function seamlessly across complex physical devices.

Performance That Surpasses Industry Leaders
Qwen-UI-Agent has demonstrated exceptional capabilities in both authoritative benchmarks and live environments:
Mobile Performance: It achieved an impressive 82.1% score in MobileWorld, outperforming several major international flagship models. In its custom benchmark, MobileWorld-Real, which involves over 100 actual devices, the success rate reached 92.2%, with near-perfect results in Android Daily tasks.Desktop Performance: Scoring 79.5% in OSWorld-Verified, the model significantly reduced execution steps across multiple tasks compared to baseline models.Web and General Capabilities: It ranked first in the WebArena web test among all comparison models. Its general and agentic capabilities exceed those of the base training model, allowing it to handle long-tail requests with ease.Bridging the Gap from Simulation to Reality
To close the divide between simulation and real-world application, Qwen-UI-Agent utilizes a live mobile environment featuring over 100 physical phones and 150+ applications. This setup supports task creation, trajectory collection, and model training. The team also introduced MobileWorld-Real, a real-device benchmark with over 400 tasks, enabling precise model selection based on actual device performance.

Streamlined Instructions and Robust Security
Beyond standard GUI operations, the model can execute command-line commands directly. By outputting batch actions in a single decision, it shortens execution trajectories and boosts efficiency. A comprehensive security mechanism is integrated throughout the process: the model refuses and terminates tasks involving illegal or high-risk requests. For sensitive actions like payments, data deletion, or privacy authorization, it pauses at critical steps to await user confirmation.
Furthermore, the model supports online reinforcement learning on trajectories exceeding 100 steps. Combined with adaptive curriculum learning, it continuously improves its ability to tackle challenging long-term tasks.
Project Homepage: https://tongyi-mai.github.io/Qwen-UI-Agent/
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (0)
0/500
Alibaba has officially introduced Qwen-UI-Agent, a foundation model for GUI agents designed to operate in real-world settings. Covering mobile, desktop, web, and deep search environments, this model moves beyond traditional simulation testing to function seamlessly across complex physical devices.

Performance That Surpasses Industry Leaders
Qwen-UI-Agent has demonstrated exceptional capabilities in both authoritative benchmarks and live environments:
Mobile Performance: It achieved an impressive 82.1% score in MobileWorld, outperforming several major international flagship models. In its custom benchmark, MobileWorld-Real, which involves over 100 actual devices, the success rate reached 92.2%, with near-perfect results in Android Daily tasks.Desktop Performance: Scoring 79.5% in OSWorld-Verified, the model significantly reduced execution steps across multiple tasks compared to baseline models.Web and General Capabilities: It ranked first in the WebArena web test among all comparison models. Its general and agentic capabilities exceed those of the base training model, allowing it to handle long-tail requests with ease.Bridging the Gap from Simulation to Reality
To close the divide between simulation and real-world application, Qwen-UI-Agent utilizes a live mobile environment featuring over 100 physical phones and 150+ applications. This setup supports task creation, trajectory collection, and model training. The team also introduced MobileWorld-Real, a real-device benchmark with over 400 tasks, enabling precise model selection based on actual device performance.

Streamlined Instructions and Robust Security
Beyond standard GUI operations, the model can execute command-line commands directly. By outputting batch actions in a single decision, it shortens execution trajectories and boosts efficiency. A comprehensive security mechanism is integrated throughout the process: the model refuses and terminates tasks involving illegal or high-risk requests. For sensitive actions like payments, data deletion, or privacy authorization, it pauses at critical steps to await user confirmation.
Furthermore, the model supports online reinforcement learning on trajectories exceeding 100 steps. Combined with adaptive curriculum learning, it continuously improves its ability to tackle challenging long-term tasks.
Project Homepage: https://tongyi-mai.github.io/Qwen-UI-Agent/
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur





Home






