The results demonstrate that the contribution lies in music-specific cross-module coupling and constrained symbolic reconstruction rather than in introducing residual, graph, recurrent, or dilated convolution as isolated operators.
Abstract
Automatic understanding of multi-instrument music requires the joint interpretation of score images, harmonic relations, and audio-side instrument identities, yet these tasks are commonly optimized in isolation. This study proposes a modular, interface-coupled framework for score transcription, harmony-aware symbolic reconstruction, and multi-instrument recognition. The framework consists of a multi-scale residual neural network (MSR-NN), a temporal harmonic graph convolutional network (THGCN), and a dual-branch CNN-DCNN. MSR-NN preserves staff-position and note-contour information through bottom-up residual extraction and top-down shallow–deep feature fusion. THGCN constructs graph edges according to integer harmonic ratios in the log-frequency domain and uses a GRU to model temporal continuity. CNN-DCNN combines standard and dilated convolutions to capture local timbral cues and broader spectro-temporal context without reducing Mel-spectrogram resolution. The three outputs are coupled through a constrained symbolic decoder rather than a fully joint end-to-end network. Under the PrIMuS and GrandStaff protocols, MSR-NN reduces the symbol error rate to 3.82%, representing an improvement of 0.40–4.09 percentage points over the selected baselines. Component ablation increases SER from 3.82% to 4.37% after removing multi-scale fusion and to 4.68% after removing residual learning. Replacing THGCN with temporal-only modeling increases F0-RMSE from 3.52 to 4.07, while replacing CNN-DCNN with a standard CNN reduces macro-F1 from 0.742 to 0.712. Symbolic refinement improves sequence consistency from 88.3% to 94.3%. On IRMAS, CNN-DCNN obtains an accuracy of 0.929, a macro-F1 of 0.742, and a macro-AUC of 0.934. External and perturbation tests obtain an AUC of 0.846 on the common-label OpenMIC-2018 transfer setting and 0.911 under 10-dB additive noise. Five-seed paired tests remain significant after Holm correction (adjusted p < 0.05). These results demonstrate that the contribution lies in music-specific cross-module coupling and constrained symbolic reconstruction rather than in introducing residual, graph, recurrent, or dilated convolution as isolated operators.
To solve the problems of low accuracy and insufficient robustness in instrument recognition in multi-instrument scenarios of digital music, a parallel structure combining a standard Convolution Neural Network and a multi-scale Dilated Convolution Neural Network (CNN-DCNN) is designed to extract multi-scale spatial acou...
This paper introduces Harmonica, a family of instrument-agnostic music transcription models built around multi-depth harmonic convolution. At each model scale, Harmonica achieves the best performance among the evaluated models: the x-large model attains state-of-the-art performance in instrument-agnostic transcription,...
Long-Shen Ou, Héctor Martel, Joe Hennessy-Priest et al.· 0 citations
V2A-AlignNet, a genre-aware cross-modal deep learning framework, provides a computational basis for audio-visual temporal relationship analysis and offers valuable insights for multimodal signal interpretation and intelligent information processing in advanced electromagnetic sensing and communication-related applicati...
Vision-to-music generation transforms visual input into structured musical output, but many recent systems rely on end-to-end neural models whose internal cross-modal decisions are difficult to explain. This paper studies an interpretable alternative based on explicit visual analysis and rule-based symbolic generation....
The issue of music source separation is discussed, focusing on its use in many audio processing tasks like remixing, target sound extraction, and audio quality improvement. This work addresses the problem through a recurrent inference framework that uses a pretrained deep neural network model, specifically within m a s...
M. Monastyrskyi, L. Lyubchyk· Automation, Control, and Inf...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduJul 15, 2026
Assistant Professor Pat Pataranutaporn describes a new interface that lets everyday users glimpse inside an AI's neural network before their chatbot ever says a word.
Microsoft Research Blog· microsoft.comJul 13, 2026
Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduJul 6, 2026
PhD student Rachel Sava, winner of the Envisioning the Future of Computing Prize, explores transformative improvements and dystopian risks of neural technology.