A depth-guided lightweight multi-task 3D scene understanding method based on a lightweight U-Net architecture, uses RGB and depth four-channel dual- modal input, and constructs a network structure with a shared encoder and independent decoders, allowing a single inference to simultaneously accomplish the three core tasks of semantic segmentation, depth completion, and obstacle detection.
Abstract
To address the core issues faced by edge-side 3D scene perception of mobile robots, such as the difficulty of balancing accuracy and real-time performance, insufficient multi-task fusion, and deviations in depth geometric consistency, this paper proposes a depth-guided lightweight multi-task 3D scene understanding method. This method is based on a lightweight U-Net architecture, uses RGB and depth four-channel dual- modal input, and constructs a network structure with a shared encoder and independent decoders, allowing a single inference to simultaneously accomplish the three core tasks of semantic segmentation, depth completion, and obstacle detection. It also Introduce depth consistency error (DCE) to construct a weighted joint loss function, strengthening scene geometry structure learning; through a dual lightweight strategy of structured pruning and INT8 quantization, precisely adapt to NVIDIA Jetson Xavier NX edge hardware. Experimental results on the NYU Depth V2 and KITTI datasets show that this method achieves a semantic segmentation mIoU ≥ 60%, depth completion RMSE ≤ 1.0m, obstacle detection F1-Score ≥ 85%, and edge- end inference speed ≥ 15 FPS, effectively balancing perception accuracy and real-time performance, providing a highly practical lightweight solution for autonomous perception in mobile robots.
A 3D Local- Global Linear Attention Mechanism (LG-LAM) is devised that efficiently captures long-range contextual information with linear complexity, enabling a comprehensive understanding of the 3D scene without heavy computational burdens.
Jie Li, Jia-Heng Xu, Laiyan Ding et al.· International Conference on...· 0 citations
This paper proposes π³-LEGS, an efficient geometric inference system designed for scalable long-sequence 3D reconstruction that maintains stable performance on thousand-frame sequences without runtime failures, highlighting its effectiveness in achieving a favorable accuracy efficiency trade-off for large-scale 3D reco...
Faline Fu, Xiaoli Cao, Can Tang et al.· International Conference on...· 0 citations
XiDepth, a lightweight architecture based on the XiNet operator block, designed to enhance feature extraction while maintaining low computational complexity and energy demand, is proposed.
Elena Izzo, Riccardo Toniolo, Lamberto Ballan· 1 citation
Semantically-guided progressive network (SGP-Net) is proposed, a semantically guided progressive refinement framework for MDE based on multi-task learning that improves key relative-error and accuracy metrics over the DCDepth baseline and remains competitive with recent methods.
Henan Hu, Xu Cheng, Rong-Hua Li et al.· Measurement science and tech...· 0 citations
Camera-based 3D semantic scene completion (SSC) provides comprehensive scene understanding for autonomous driving and robotics. However, existing methods often treat stereo depth estimates as deterministic geometric constraints, causing depth uncertainty and local correspondence errors to propagate directly into voxel...
Meng Wang, Hongxia Yu, Wenzhe He et al.· 0 citations
Monocular depth estimation is an important dense prediction task for autonomous driving, robotic perception, and unmanned aerial vehicle navigation. Although recent deep networks have achieved impressive accuracy, many of them depend on large backbones and expensive context modeling modules, making deployment on resour...
Dun-Yu Hu, Han-Xiang Zhang, J. Gu et al.· 2026 IEEE Canadian Atlantic...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.