Underwater salient object detection is a critical task in computer vision, relying heavily on data from underwater sensors, with wide-ranging applications in object tracking, content-aware editing, and object recognition. To tackle the challenges inherent in underwater multimodal information fusion, this paper introduces a novel underwater salient object detection framework based on an information cross-fusion network. The proposed approach integrates a cross-attention feature injection module and an information embedding module to facilitate efficient multimodal feature aggregation and refinement across both channel and spatial dimensions. By modeling the complementarity between RGB and depth data at global and local scales, these modules enhance the representation of salient regions while effectively suppressing background noise. Furthermore, the architecture employs multi-level and multimodal information fusion, which mitigates the effects of depth-related noise and reduces uncertainty in predictions. Extensive experiments conducted on multiple underwater datasets demonstrate that the proposed method achieves superior performance compared to state-of-the-art approaches, highlighting its efficacy in multimodal feature integration and salient object detection.
Yan Mou, Zhao-Long Gao, Jinjiang Li· PLoS ONE· 0 citations
Remote sensing change detection (RSCD) aims to conduct difference analysis on RS images obtained in different time phases of the same area. It plays a critical role in applications, such as disaster monitoring and forest cover analysis, and has evolved rapidly in recent years. However, how to suppress false changes while enhancing the response to minor real changes and maintaining fine boundaries under the interference of complex backgrounds and imaging differences remains a key challenge in high-resolution RSCD. To address this, this article proposes the Discrepancy-Invariant Boundary-Guided Network (DIBNet). Specifically, to address pseudochanges caused by imaging condition differences, this article proposes a discrepancy-invariant recalibration module, which explicitly utilizes invariant features between different time phases to calibrate the difference features and enhance the ability to suppress false changes. Second, a cross-granularity boundary modeling module is proposed, which uses the details in shallow difference features to provide precise boundary positioning, and introduces semantic context in deep difference features to calibrate the boundary response at the regional level. Through the mutual guidance between semantic granularity and detail granularity, clear and semantically consistent boundary features are generated. Subsequently, a boundary-guided image fusion module is constructed, which injects boundary features into the multiscale difference feature fusion process to improve the response completeness and boundary precision of tiny change regions. Experimental results on multiple public datasets show that the performance of DIBNet is superior to existing mainstream methods.
Semantic segmentation of high-resolution remote sensing images remains challenging due to complex spatial structures, multiscale object variations, fine-grained category differences, and high interclass similarities. Conventional segmentation methods usually rely on fixed convolutional heads or single feature representations, which makes it difficult to effectively model both intraclass appearance variations and interclass texture similarities, often leading to category confusion, missed objects, and incomplete segmentation in complex scenes. To address these challenges, we propose a state-aware prototype learning network, termed SAPLNet. Specifically, a cross-stage state refiner is introduced to progressively refine multilevel features by integrating the input features with the outputs of different stages through state-aware gated normalization. Then, a weighted feature pyramid decoder performs top-down fusion of the refined hierarchical features, combining high-level semantic information with low-level spatial details. Furthermore, a state-aware multiprototype classifier is designed to construct multiple semantic prototypes for each class via ground-truth-guided local class-center extraction and momentum-based prototype memory updating. A global state vector derived from the refined cross-stage features is used to adaptively modulate decoder features, improving the matching reliability between pixel features and class prototypes. In addition, prototype compactness loss, prototype diversity loss, and lightweight boundary loss are employed to enhance intraclass consistency, prototype discriminability, and boundary awareness. Experimental results demonstrate the effectiveness and superiority of SAPLNet.
Zeyu Zhao, Zhaolong Gao, Jun Feng· IEEE Journal of Selected Top...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.