Ant Group releases LingBot-Depth 2.0 spatial perception model, advancing robot vision.
On July 7, Lingbo Technology, an embodied intelligence company under Ant Group, launched the LingBot-Depth 2.0 spatial perception model. Trained on a dataset of 150 million samples, it delivers comprehensive improvements in edge clarity, fine object recognition, long-range depth estimation, and robustness in complex scenarios.
LingBot-Depth is a spatial perception model independently developed by Lingbo, serving as the eyes of robots in the physical world. The first version addressed spatial perception challenges in difficult scenarios like transparent and reflective surfaces. Compared to version 1.0, the training data for LingBot-Depth 2.0 expanded from 3 million to 150 million samples, with overall performance gains: it achieved 12 first-place finishes out of 16 evaluation items in depth completion benchmarks. In the most challenging indoor scenario with large-scale depth missing, the depth error was halved compared to the previous generation (RMSE dropped from 0.132 to 0.062). It performs particularly well in situations where traditional depth cameras often fail, such as glass, mirrors, and transparent objects.
At the same time, LingBot-Depth 2.0 introduced its visual foundation model, LingBot-Vision, building a capability chain from "understanding" to "accurate perception." This aims to address core challenges in robot vision, including spatial perception, fine recognition, and adaptation to complex environments.

(Figure 1: LingBot-Depth 2.0 completes a full and flat 3D structure in difficult scenarios like mirrors and glass)
The breakthrough of LingBot-Depth 2.0 relies on the outstanding visual representation capabilities of LingBot-Vision. As a general-purpose visual model, LingBot-Vision is the first visual foundation model in the industry to use "boundary structure" as a pre-training objective, achieving a breakthrough in the spatial perception training paradigm. It offers sub-pixel-level boundary positioning and spatial structure understanding, delivering higher precision and more stable spatial perception.
The pre-training corpus for LingBot-Vision consists of only 160 million images, an order of magnitude smaller than DINOv3, yet its depth estimation accuracy surpasses DINOv3. Moreover, LingBot-Vision's judgment of object boundaries is sufficiently stable to enable continuous boundary tracking in videos. This release includes four versions: ViT-G, ViT-L, ViT-B, and ViT-S.
According to available information, LingBot-Vision not only supports the training of LingBot-Depth 2.0 but also has the ability to be used in multiple other ways.

(Figure 2: LingBot-Depth 2.0 performs well in real sensor depth completion tests)

(Figure 3: Compared to mainstream visual foundation models, LingBot-Vision identifies object boundaries and spatial structures more clearly and stably)
Currently, LingBot-Depth 2.0 has passed the professional certification of Orbbec Depth Vision Lab. Practical scenario tests show that, based on the chip-level 3D raw data provided by Orbbec's Gemini330 series stereo 3D camera, LingBot-Depth 2.0 significantly improves edge clarity, object contour integrity, fine object recognition, long-range depth estimation, and robustness in complex lighting and material scenarios.

(Figure 4: LingBot Depth 2.0 passed the professional evaluation of Orbbec Depth Vision Lab, demonstrating extremely high accuracy and stability in spatial and temporal depth estimation tasks across multiple types of sensors)
In terms of commercialization, Ant Lingbo has engaged in deep cooperation with Orbbec in many areas. According to the information, the latest RGB-D version of the EGO device in Orbbec's new no-body data acquisition product matrix will be adapted to the LingBot-Depth version optimized by Lingbo for data acquisition scenarios. In the future, it will further integrate higher-level commercial versions, continuously fill depth gaps, optimize object edges and spatial structure details, and provide more accurate, stable, and usable real-world data bases for embodied intelligence model training.
Additionally, Orbbec will launch an SDK product integrating the latest model capabilities of LingBot-Depth, available for use by robot customers at the edge, allowing robots using the Gemini330 series camera to achieve better depth effects. They also plan to launch an integrated camera product with the commercial version of LingBot-Depth by the end of the year, realizing the integration of "3D camera + spatial perception capabilities." With the release of the two models, their collaboration is expected to expand into more fields.
Related article
Slackbot Becomes an AI Agent
Slackbot, the automated assistant embedded in Salesforce’s corporate messaging platform Slack, is evolving into an AI agent. Salesforce CTO Parker Harris envisions it achieving viral status comparable to OpenAI’s ChatGPT.The cloud software giant laun
ByteDance Boosts Core AI Incentives as Doubao Surges 14.6%
ByteDance recently convened a DouBao equity briefing to unveil fresh incentive policies for staff involved in the DouBao division. The strike price for DouBao shares has been lifted from $14.85 in June 2026 to $17.02, marking an approximate 14.6% inc
MiniMax Unveils 10x Team Program to Incentivize Global AI Experts
MiniMax (Xiyu Technology), the General Artificial Intelligence Lab, has officially launched "10x Team," a global talent collaboration initiative. This program aims to recruit top experts across industries to explore the deep application of large mode
Related Special Topic Recommendations
Comments (0)
0/500
On July 7, Lingbo Technology, an embodied intelligence company under Ant Group, launched the LingBot-Depth 2.0 spatial perception model. Trained on a dataset of 150 million samples, it delivers comprehensive improvements in edge clarity, fine object recognition, long-range depth estimation, and robustness in complex scenarios.
LingBot-Depth is a spatial perception model independently developed by Lingbo, serving as the eyes of robots in the physical world. The first version addressed spatial perception challenges in difficult scenarios like transparent and reflective surfaces. Compared to version 1.0, the training data for LingBot-Depth 2.0 expanded from 3 million to 150 million samples, with overall performance gains: it achieved 12 first-place finishes out of 16 evaluation items in depth completion benchmarks. In the most challenging indoor scenario with large-scale depth missing, the depth error was halved compared to the previous generation (RMSE dropped from 0.132 to 0.062). It performs particularly well in situations where traditional depth cameras often fail, such as glass, mirrors, and transparent objects.
At the same time, LingBot-Depth 2.0 introduced its visual foundation model, LingBot-Vision, building a capability chain from "understanding" to "accurate perception." This aims to address core challenges in robot vision, including spatial perception, fine recognition, and adaptation to complex environments.

(Figure 1: LingBot-Depth 2.0 completes a full and flat 3D structure in difficult scenarios like mirrors and glass)
The breakthrough of LingBot-Depth 2.0 relies on the outstanding visual representation capabilities of LingBot-Vision. As a general-purpose visual model, LingBot-Vision is the first visual foundation model in the industry to use "boundary structure" as a pre-training objective, achieving a breakthrough in the spatial perception training paradigm. It offers sub-pixel-level boundary positioning and spatial structure understanding, delivering higher precision and more stable spatial perception.
The pre-training corpus for LingBot-Vision consists of only 160 million images, an order of magnitude smaller than DINOv3, yet its depth estimation accuracy surpasses DINOv3. Moreover, LingBot-Vision's judgment of object boundaries is sufficiently stable to enable continuous boundary tracking in videos. This release includes four versions: ViT-G, ViT-L, ViT-B, and ViT-S.
According to available information, LingBot-Vision not only supports the training of LingBot-Depth 2.0 but also has the ability to be used in multiple other ways.

(Figure 2: LingBot-Depth 2.0 performs well in real sensor depth completion tests)

(Figure 3: Compared to mainstream visual foundation models, LingBot-Vision identifies object boundaries and spatial structures more clearly and stably)
Currently, LingBot-Depth 2.0 has passed the professional certification of Orbbec Depth Vision Lab. Practical scenario tests show that, based on the chip-level 3D raw data provided by Orbbec's Gemini330 series stereo 3D camera, LingBot-Depth 2.0 significantly improves edge clarity, object contour integrity, fine object recognition, long-range depth estimation, and robustness in complex lighting and material scenarios.

(Figure 4: LingBot Depth 2.0 passed the professional evaluation of Orbbec Depth Vision Lab, demonstrating extremely high accuracy and stability in spatial and temporal depth estimation tasks across multiple types of sensors)
In terms of commercialization, Ant Lingbo has engaged in deep cooperation with Orbbec in many areas. According to the information, the latest RGB-D version of the EGO device in Orbbec's new no-body data acquisition product matrix will be adapted to the LingBot-Depth version optimized by Lingbo for data acquisition scenarios. In the future, it will further integrate higher-level commercial versions, continuously fill depth gaps, optimize object edges and spatial structure details, and provide more accurate, stable, and usable real-world data bases for embodied intelligence model training.
Additionally, Orbbec will launch an SDK product integrating the latest model capabilities of LingBot-Depth, available for use by robot customers at the edge, allowing robots using the Gemini330 series camera to achieve better depth effects. They also plan to launch an integrated camera product with the commercial version of LingBot-Depth by the end of the year, realizing the integration of "3D camera + spatial perception capabilities." With the release of the two models, their collaboration is expected to expand into more fields.
Slackbot Becomes an AI Agent
Slackbot, the automated assistant embedded in Salesforce’s corporate messaging platform Slack, is evolving into an AI agent. Salesforce CTO Parker Harris envisions it achieving viral status comparable to OpenAI’s ChatGPT.The cloud software giant laun
ByteDance Boosts Core AI Incentives as Doubao Surges 14.6%
ByteDance recently convened a DouBao equity briefing to unveil fresh incentive policies for staff involved in the DouBao division. The strike price for DouBao shares has been lifted from $14.85 in June 2026 to $17.02, marking an approximate 14.6% inc
MiniMax Unveils 10x Team Program to Incentivize Global AI Experts
MiniMax (Xiyu Technology), the General Artificial Intelligence Lab, has officially launched "10x Team," a global talent collaboration initiative. This program aims to recruit top experts across industries to explore the deep application of large mode





Home






