Gearbox multimodal fault diagnosis based on microphone array and cross-attention
Abstract
In order to solve the problems of limited vibration signal acquisition, incomplete single-modal features, and sound signals being vulnerable to noise interference in gearbox fault diagnosis, a Multi-Channel Convolutional-Transformer Cross- Attention Diagnosis Model is proposed. Its core innovations are as follows: combining prior knowledge of strong noise sources, directional noise reduction is performed on sound signals through a beamforming algorithm to retain key acoustic fault features; a cross-attention fusion module is constructed, which takes time-domain vibration features as queries and frequency-domain sound features as keys and values to realize deep interaction and adaptive weight allocation of acoustic-vibration bimodal features. Experimental results show that the diagnostic accuracy of the model reaches 99.91%, which is significantly better than classical models such as CNN and LSTM. Control and ablation experiments verify the effectiveness of the directional noise reduction and cross-attention mechanism, confirming the feasibility of the "using sound to supplement vibration" strategy in scenarios where vibration acquisition is limited. This provides a new path for gearbox multimodal fusion fault diagnosis.