Skip to content
Open access

LDANet: a lightweight depth-aware framework for RGB-D salient object detection

Sep 2026 · Journal of King Saud University: Computer and Information Sciences · Vol 38 · 0 citations · 35 references

Abstract

As a foundational research in computer vision, salient object detection (SOD) has received widespread attention from researchers. However, existing methods still have notable limitations, mainly reflected in two aspects. (1) Crude multi-modal fusion strategies fail to emphasize consistent multi-modal features and preserve target details, resulting in weak feature representations and redundant computation. (2) The absence of effective edge-direction modeling, together with the significant scale mismatch between high- and low-level features, further contributes to blurred edges, low localization precision, and suboptimal decoding accuracy. To address these challenges, this paper proposes a depth-aware framework (LDANet) for RGB-D SOD. The architecture consists of a dual-branch encoder and four elaborately designed collaborative modules for progressive feature optimization. First, we adopt a dual-branch encoder to extract hierarchical multi-level features, laying a solid foundation for subsequent feature processing and optimization. Second, a modality fusion module is introduced to achieve lightweight and high-integrity multi-modal feature fusion. To strengthen multi-modal dependencies and capture multi-scale contextual information, we design a deep guided attention module that achieves this objective through depth-aware sparse attention. Subsequently, an adaptive edge refinement module is employed to suppress non-edge noise and refine edge contour representations, thereby enhancing edge localization accuracy. Finally, a hierarchical feature fusion decoder is utilized to align multi-level feature scales and generate high-quality saliency maps. Extensive empirical results demonstrate the superiority of LDANet over recent competitive methods in accurately identifying salient objects in complex scenes. It achieves MAEs of 0.025, 0.051, 0.018, 0.042, 0.034, and 0.034 on six RGB-D datasets, and MAEs of 0.031, 0.020, and 0.032 on three RGB-T datasets. The quantitative and qualitative results demonstrate that the proposed method achieves excellent performance while maintaining a lightweight architecture with only 7.4M parameters and 4.5G FLOPs.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.