Skip to content

Multi-instrument music score transcription, symbolic generation, and harmony analysis based on multi-scale residual neural networks

Sep 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 36 references

TL;DR

The results demonstrate that the contribution lies in music-specific cross-module coupling and constrained symbolic reconstruction rather than in introducing residual, graph, recurrent, or dilated convolution as isolated operators.

Abstract

Automatic understanding of multi-instrument music requires the joint interpretation of score images, harmonic relations, and audio-side instrument identities, yet these tasks are commonly optimized in isolation. This study proposes a modular, interface-coupled framework for score transcription, harmony-aware symbolic reconstruction, and multi-instrument recognition. The framework consists of a multi-scale residual neural network (MSR-NN), a temporal harmonic graph convolutional network (THGCN), and a dual-branch CNN-DCNN. MSR-NN preserves staff-position and note-contour information through bottom-up residual extraction and top-down shallow–deep feature fusion. THGCN constructs graph edges according to integer harmonic ratios in the log-frequency domain and uses a GRU to model temporal continuity. CNN-DCNN combines standard and dilated convolutions to capture local timbral cues and broader spectro-temporal context without reducing Mel-spectrogram resolution. The three outputs are coupled through a constrained symbolic decoder rather than a fully joint end-to-end network. Under the PrIMuS and GrandStaff protocols, MSR-NN reduces the symbol error rate to 3.82%, representing an improvement of 0.40–4.09 percentage points over the selected baselines. Component ablation increases SER from 3.82% to 4.37% after removing multi-scale fusion and to 4.68% after removing residual learning. Replacing THGCN with temporal-only modeling increases F0-RMSE from 3.52 to 4.07, while replacing CNN-DCNN with a standard CNN reduces macro-F1 from 0.742 to 0.712. Symbolic refinement improves sequence consistency from 88.3% to 94.3%. On IRMAS, CNN-DCNN obtains an accuracy of 0.929, a macro-F1 of 0.742, and a macro-AUC of 0.934. External and perturbation tests obtain an AUC of 0.846 on the common-label OpenMIC-2018 transfer setting and 0.911 under 10-dB additive noise. Five-seed paired tests remain significant after Holm correction (adjusted p < 0.05). These results demonstrate that the contribution lies in music-specific cross-module coupling and constrained symbolic reconstruction rather than in introducing residual, graph, recurrent, or dilated convolution as isolated operators.

Read PDF

Similar papers

Open access Sep 2026

Performance Optimization of Digital Music Multi-Instrument Recognition Based on an Improved DCNN Model and Bi-LSTM

To solve the problems of low accuracy and insufficient robustness in instrument recognition in multi-instrument scenarios of digital music, a parallel structure combining a standard Convolution Neural Network and a multi-scale Dilated Convolution Neural Network (CNN-DCNN) is designed to extract multi-scale spatial acou...

Jun-Hui Zhao, Xiao-Hang Jia, Yuxuan Zhou · 0 citations
Preprint Sep 2026

Harmonica: Accurate and Lightweight Instrument-Agnostic Music Transcription

This paper introduces Harmonica, a family of instrument-agnostic music transcription models built around multi-depth harmonic convolution. At each model scale, Harmonica achieves the best performance among the evaluated models: the x-large model attains state-of-the-art performance in instrument-agnostic transcription,...

Long-Shen Ou, Héctor Martel, Joe Hennessy-Priest et al. · 0 citations
Open access Aug 2026

Combining CRNN Modeling to Dynamically Align the Relationship Between Music Rhythm and Plot Twists in Different Film Genres

V2A-AlignNet, a genre-aware cross-modal deep learning framework, provides a computational basis for audio-visual temporal relationship analysis and offers valuable insights for multimodal signal interpretation and intelligent information processing in advanced electromagnetic sensing and communication-related applicati...

S.-G. Chi, H.-Y. Liang · 0 citations
Conference Open access 2026

Interpretable Vision-to-Music Mapping from Low-Level Visual Statistics to Mode, Rhythm, and Harmonic Structure

Vision-to-music generation transforms visual input into structured musical output, but many recent systems rely on end-to-end neural models whose internal cross-modal decisions are difficult to explain. This paper studies an interpretable alternative based on explicit visual analysis and rule-based symbolic generation....

Shu-Han Yang · 0 citations
Conference Sep 2026

Interference Reduction in Music Source Separation Through Recurrent Inference with Deep Neural Network Models

The issue of music source separation is discussed, focusing on its use in many audio processing tasks like remixing, target sound extraction, and audio quality improvement. This work addresses the problem through a recurrent inference framework that uses a pretrained deep neural network model, specifically within m a s...

M. Monastyrskyi, L. Lyubchyk · 0 citations

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.