Jul 2026· Engineering Research Express· Vol 8, pp. 155235· 0 citations· 45 references
Physics
TL;DR
Results indicate that RCG-YOLO provides a practical accuracy improvement over YOLOv8n for underwater object detection, and its higher GFLOPs and lower frames per second compared with YOLOv8n indicate that the method prioritizes accuracy over maximum inference efficiency.
Abstract
Underwater target detection is hindered by light scattering, low contrast, blur, occlusion, and background clutter, which particularly affect small objects. To address these limitations, we propose RCG-YOLO, a YOLOv8n-based detector that improves feature preservation, multi-scale feature extraction, and cross-scale feature fusion. Specifically, the proposed model introduces a residual down-sampling (ResDown) module to preserve fine spatial information during downsampling, a gated attention and multi-scale extractor (GAME) module to strengthen multi-scale feature extraction and positional sensitivity, and a cross-scale gated fusion (CGF) module to selectively fuse shallow spatial features and deep semantic features. Experiments on three public underwater datasets shows consistent improvements over YOLOv8n. RCG-YOLO achieves mAP50 scores of 85.2 ± 0.03%, 86.5 ± 0.02%, and 85.0 ± 0.03% on UTDAC2020, DUO, and RUOD datasets, corresponding to absolute mAP50 gains of 3.4%, 3.9%, and 1.7% over the baseline, respectively. Ablation studies confirm that ResDown, GAME, and CGF each contribute to the final detection performance. Although RCG-YOLO maintains real-time inference on an NVIDIA RTX 4090, its higher GFLOPs and lower frames per second compared with YOLOv8n indicate that the method prioritizes accuracy over maximum inference efficiency. These results indicate that RCG-YOLO provides a practical accuracy improvement over YOLOv8n for underwater object detection.
Results support the effectiveness of the proposed framework for mixed-scale underwater object detection, including small-scale objects, and compare with the YOLOv11 baseline.
Experimental results demonstrate that the proposed Bidirectional weighted Concat with Efficient multi-scale Attention You Only Look Once (BCEA-YOLO) method outperforms competing methods for underwater object detection, and edge deployment tests on the NVIDIA Jetson platform validate its real-time inference efficiency.
Xing-Yu Wang, Yu-Han Lin, Quan J. Wang et al.· Intelligent Marine Technolog...· 0 citations
Underwater biological object detection is important for intelligent marine monitoring, yet its performance is often limited by severe image degradation, background clutter, and large variations in target scale. These factors can weaken feature representation during multi-scale fusion and reduce localization reliabili...
Xiao-Long Zhu, Jia-Yu Wang, Yukang Wang et al.· Scientific Reports· 0 citations
A novel underwater target detection framework that integrates feature enhancement with semantic-spatial guided fusion, built upon the RT-DETR architecture, that significantly reduces false positives and missed detections while maintaining real-time performance is proposed.
P. Parashar, A. Kushwah· Discover Computing· 0 citations
TRIDEN-YOLO, a lightweight detector built upon YOLOv11n, provides the primary reparameterized contextual representation design through multi-branch training and inference-time fusion, while HFFE and GCD loss are incorporated to enhance hierarchical feature fusion and boundary-aware localization.
Xi Chen, Yuping Sun, Kaibin Zeng· Signal, Image and Video Proc...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.