Underwater Acoustic–Optical Multimodal Fusion Detection Algorithm for UUVs with Cross-Domain Validation
Abstract
Underwater object detection is a core technology for environmental perception and autonomous operation of unmanned underwater vehicles (UUVs). However, optical and acoustic sensing alone suffer from physical limitations, leading to missed and false detections in turbid, low-light, or long-range conditions. To overcome these limitations, this paper develops an acoustic–optical multimodal fusion detection module (AOMFDM) tailored for UUV deployment. The module employs dual YOLOv5 models for separate processing of sonar and optical images. An interference source quantification estimation network is introduced to extract environmental degradation features, including noise, blur, illumination, contrast, and color cast. A heterogeneous feature map matching network and a deep sparse autoencoder are further designed to achieve cross-modal alignment and fusion of acoustic and optical features. Additionally, attention mechanisms, anchor-based box annotation, and weighted boxes fusion (WBF) are incorporated to enhance detection robustness. For model training and evaluation, we construct the Underwater Sonar Detection (USD) and Underwater Optical Detection (UOD) datasets, covering diverse water qualities, illumination levels, target materials, and interference scenarios. Experimental results demonstrate that, by exploiting the complementarity of acoustic and optical modalities together with adaptive alignment strategies, the proposed module significantly boosts both detection reliability and generalization capability for UUVs in challenging underwater environments.