Performance Optimization of Digital Music Multi-Instrument Recognition Based on an Improved DCNN Model and Bi-LSTM
Abstract
To solve the problems of low accuracy and insufficient robustness in instrument recognition in multi-instrument scenarios of digital music, a parallel structure combining a standard Convolution Neural Network and a multi-scale Dilated Convolution Neural Network (CNN-DCNN) is designed to extract multi-scale spatial acoustic features from Mel spectrograms. The feature sequence extracted by CNN-DCNN is fed into a Bidirectional Long Short-Term Memory (Bi-LSTM) network for bidirectional context modeling, and a self-attention mechanism is introduced to dynamically weight temporal features, capturing the inherent long-range temporal dependencies of music signals. The findings demonstrate that the proposed fusion-improved model achieves optimal performance across multiple key metrics, with mean accuracy, macro-average F1 score, and mean precision reaching 88.92%, 87.11% and 88.34%, respectively, and significantly outperforming other comparative models. Ablation experiments further confirm the effectiveness of the parallel convolutional structure, temporal modeling, and attention mechanism. The proposed model, by deeply fusing multi-scale spatial features with bidirectional temporal contextual information, provides a high-performance solution for multi-instrument recognition in complex acoustic scenarios and offers valuable references for intelligent music analysis and retrieval.