LingBot-Depth 2.0: Sensor Data-Based Depth Estimation
Key point
Research results on LingBot-Depth 2.0, which maximizes depth estimation performance by using the sensor's missing regions as a training signal.
Details
Instead of the conventional random block dropout method, we propose a Sensor-validity masking technique that uses sensor failure regions (Specular highlights, transparent surfaces, textureless regions, etc.) as masking signals. This induces the model to learn the failure distribution it will actually encounter during inference.
According to research on LingBot-Depth 2.0 by Robbyant, an Embodied AI company under Ant Group, when only the Encoder-init was changed within the same pipeline for performance comparison, the LingBot-Vision initialization method showed superior performance on most benchmarks. (However, DINOv2 was superior on Hammer capture data.)
Key results are as follows:
- Achieved best RMSE on 7 out of 8 block mask and sparse depth benchmarks
- Demonstrated superior performance across 6 real camera configurations (Hammer, ClearGrasp, and the team's own dataset)
- Showed particularly strong performance on ClearGrasp captures targeting transparent objects, with DIODE-Indoor RMSE improved to approximately half the level of version 1.0
While the Depth 2.0 weights have not yet been released, the 4 LingBot-Vision backbone models are publicly available under the Apache-2.0 license.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.