A novel underwater target detection framework that integrates feature enhancement with semantic-spatial guided fusion, built upon the RT-DETR architecture, that significantly reduces false positives and missed detections while maintaining real-time performance is proposed.
Abstract
Underwater object detection remains a challenging task due to severe image degradation, scale variation, and frequent occlusions, which often result in high false and missed detection rates. To address these issues, this paper proposes a novel underwater target detection framework that integrates feature enhancement with semantic-spatial guided fusion. Built upon the RT-DETR architecture, the proposed model introduces three key innovations. First, a Multi-Scale Edge Feature Injection (MSI-Edge) module is designed to incorporate fine-grained edge information from shallow layers into deeper representations, significantly improving the detection of small and low-contrast targets. Second, a Global–Local Feature Enhancement (GLF-Enhance) module replaces conventional multi-head self-attention, enabling efficient and balanced learning of both global contextual and local structural information while reducing computational overhead. Third, a Semantic–Location Path Aggregation Network (SL-PAN) is proposed to enable bidirectional interaction between semantic and spatial features, effectively mitigating information degradation during multi-scale feature fusion.Extensive experiments on benchmark underwater datasets demonstrate that the proposed method consistently outperforms the baseline RT-DETR (ResNet50 backbone), achieving improvements of 3.2%, 3.0%, and 2.7% in AP, AP50, and AP75 on the URPC dataset, and 2.9%, 2.7%, and 3.0% on the DUO dataset, respectively. Moreover, the model significantly reduces false positives and missed detections while maintaining real-time performance. Comprehensive ablation studies further validate the effectiveness of each module. Overall, the proposed approach establishes a robust and efficient solution, advancing the state-of-the-art in underwater object detection.
Experimental results demonstrate that the proposed Bidirectional weighted Concat with Efficient multi-scale Attention You Only Look Once (BCEA-YOLO) method outperforms competing methods for underwater object detection, and edge deployment tests on the NVIDIA Jetson platform validate its real-time inference efficiency.
Xing-Yu Wang, Yu-Han Lin, Quan J. Wang et al.· Intelligent Marine Technolog...· 0 citations
Underwater object detection is of significant practical importance for marine resource exploration, underwater robotic navigation, and marine ecological monitoring. However, underwater images are often severely degraded by light attenuation and scattering, suspended particulates, and complex background interference. Th...
Feng Zou, Botong Zhou, Jia-Qi Ma et al.· Journal of Real-Time Image P...· 0 citations
The proposed UW-D-FINE, an enhanced real-time detector addressing underwater object detection challenges through three key innovations, enhances the backbone by integrating parallel multi-scale convolutional branches with omnidirectional depthwise convolutions, enabling more effective extraction of discriminative featu...
Han-Jie Ma, Tingting Wan, Hui-Jun Dong et al.· Journal of Real-Time Image P...· 0 citations
Results support the effectiveness of the proposed framework for mixed-scale underwater object detection, including small-scale objects, and compare with the YOLOv11 baseline.
Underwater object detection plays a crucial role in fisheries resource assessment and ecological environment protection. Current underwater object detection models are characterized by large parameter sizes and high computational costs, which hinder the simultaneous achievement of lightweight deployment and high detect...
Xue-Feng Zhao, Yong-Jie Guo, Zhao-Man Zhong et al.· Measurement science and tech...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.