Skip to content

Enhancing underwater object detection through bidirectional cross-scale gated feature fusion

Jul 2026 · Engineering Research Express · Vol 8, pp. 155235 · 0 citations · 45 references
Physics

TL;DR

Results indicate that RCG-YOLO provides a practical accuracy improvement over YOLOv8n for underwater object detection, and its higher GFLOPs and lower frames per second compared with YOLOv8n indicate that the method prioritizes accuracy over maximum inference efficiency.

Abstract

Underwater target detection is hindered by light scattering, low contrast, blur, occlusion, and background clutter, which particularly affect small objects. To address these limitations, we propose RCG-YOLO, a YOLOv8n-based detector that improves feature preservation, multi-scale feature extraction, and cross-scale feature fusion. Specifically, the proposed model introduces a residual down-sampling (ResDown) module to preserve fine spatial information during downsampling, a gated attention and multi-scale extractor (GAME) module to strengthen multi-scale feature extraction and positional sensitivity, and a cross-scale gated fusion (CGF) module to selectively fuse shallow spatial features and deep semantic features. Experiments on three public underwater datasets shows consistent improvements over YOLOv8n. RCG-YOLO achieves mAP50 scores of 85.2 ± 0.03%, 86.5 ± 0.02%, and 85.0 ± 0.03% on UTDAC2020, DUO, and RUOD datasets, corresponding to absolute mAP50 gains of 3.4%, 3.9%, and 1.7% over the baseline, respectively. Ablation studies confirm that ResDown, GAME, and CGF each contribute to the final detection performance. Although RCG-YOLO maintains real-time inference on an NVIDIA RTX 4090, its higher GFLOPs and lower frames per second compared with YOLOv8n indicate that the method prioritizes accuracy over maximum inference efficiency. These results indicate that RCG-YOLO provides a practical accuracy improvement over YOLOv8n for underwater object detection.

View source

Similar papers

Open access Sep 2026

Underwater tiny object detection network based on multi-scale attention and adaptive feature fusion

Experimental results demonstrate that the proposed Bidirectional weighted Concat with Efficient multi-scale Attention You Only Look Once (BCEA-YOLO) method outperforms competing methods for underwater object detection, and edge deployment tests on the NVIDIA Jetson platform validate its real-time inference efficiency.

Xing-Yu Wang, Yu-Han Lin, Quan J. Wang et al. · 0 citations
Open access Jul 2026

A lightweight underwater biological object detector with enhanced cross-scale feature interaction

Underwater biological object detection is important for intelligent marine monitoring, yet its performance is often limited by severe image degradation, background clutter, and large variations in target scale. These factors can weaken feature representation during multi-scale fusion and reduce localization reliabili...

Xiao-Long Zhu, Jia-Yu Wang, Yukang Wang et al. · 0 citations
Open access Sep 2026

Underwater target detection based on target feature enhancement and semantic-position path aggregation

A novel underwater target detection framework that integrates feature enhancement with semantic-spatial guided fusion, built upon the RT-DETR architecture, that significantly reduces false positives and missed detections while maintaining real-time performance is proposed.

P. Parashar, A. Kushwah · 0 citations
Aug 2026

TRIDEN-YOLO: a reparameterized interactive network for underwater object detection

TRIDEN-YOLO, a lightweight detector built upon YOLOv11n, provides the primary reparameterized contextual representation design through multi-branch training and inference-time fusion, while HFFE and GCD loss are incorporated to enhance hierarchical feature fusion and boundary-aware localization.

Xi Chen, Yuping Sun, Kaibin Zeng · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.