Skip to content
Conference

Style-aware data augmentation for deep learning on symbolic music

Jul 2026 · International Conference on Image, Video and Signal Processing · Vol 14268, pp. 1426811 - 1426811-13 · 0 citations · 22 references
Engineering

TL;DR

A style-aware data augmentation framework that combines rule-based design with statistical constraints, applied to small-scale, highly constrained Jiangnan symbolic music generation tasks is proposed, demonstrating the need for carefully designed augmentation strategies in highly constrained symbolic music generation tasks.

Abstract

This study proposes a style-aware data augmentation framework that combines rule-based design with statistical constraints, applied to small-scale, highly constrained Jiangnan symbolic music generation tasks. By leveraging YNote representation, fixed rhythmic frameworks, and Markov-style local transition statistics, we systematically expand the training data while maintaining musical structural plausibility. Using the augmented data, we fine-tune a GPT-2 model to analyze how different training data scales affect generation behavior. Experimental results show that Bilingual Evaluation Understudy (BLEU) -based reference-overlap metrics exhibit only minor fluctuations across different training scales and are insufficient to directly reflect style improvement. In contrast, Kullback-Leibler (KL) divergence and bigram statistics effectively characterize the style consistency of the generated set in terms of overall distribution proximity and local transition plausibility. Further analysis indicates that a medium-scale training set (approximately 3,000–12,000 samples) achieves the best balance between distribution consistency and transition coverage, whereas excessive augmentation may lead to distribution calibration drift. Overall, the study demonstrates that data augmentation has a positive but non-monotonic effect on style consistency, highlighting the need for carefully designed augmentation strategies in highly constrained symbolic music generation tasks.

View source

Similar papers

Preprint Aug 2026

UniVerse: Benchmarking and Enhancing LALMs on Culturally Inclusive Low-Resource Music Understanding

Experimental results indicate that fully automated data curation combined with imbalance-aware training yields non-trivial improvements, but models still struggle to capture fine-grained acoustic features, indicating a gap between surface-level alignment and deep musical comprehension.

Zi-Ya Zhou, Shangda Wu, Shenyang Xu et al. · 0 citations
Preprint Aug 2026

Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping

A cross-modal framework that learns implicit music styles from raw audio and applies them to symbolic music generation and generates piano performances jointly conditioned on a lead sheet and a reference audio example, enabling controllable and stylistically faithful arrangement.

Jing-Wei Zhao, Gus G. Xia, Ziyu Wang et al. · 0 citations
Jul 2026

Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances

Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes. However, these methods often disregard dependencies between notes, which poses a potential limitation in assessing the similarity between two sets of performanc...

D. Gavrilev, Ilya Borovik, Vladimir Viro · 0 citations
Open access Jul 2026

Uncertainty-aware music genre classification using evidential deep learning

Music is known to play a primary role in stress reduction, thereby enhancing our mental well-being. Retrieval of musical category tailored to specific needs of an individual is gaining traction but remains challenging and quite cumbersome. This is because of hazy and ambiguous nature of music classification, due to h...

V. Sidharrth, Jayan Sarada, Bilal Alataş · 0 citations
#natural language process... Preprint Aug 2026

MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques

MuseCritic is introduced, a semi-scalar reward model that generates a natural-language critique covering five aesthetic dimensions and uses it as an intermediate representation to predict continuous reward scores, showing that critique-conditioned reward modeling reduces scoring error and provides an effective optimiza...

Jiabao Zhuang, Changhao Jiang, Hanchen Wang et al. · 0 citations
Preprint Aug 2026

SPHERE: Automatic Music Upmixing via Audio Language Model Post-Training with Spatial Heuristic Rewards

In this paper, we study the task of automatic music upmixing, wherein a system predicts spatial mixing parameters from a multi-stem recording. Different from existing methods that rely on task-specific music encoders, we approach this task via audio language model (ALM) post-training, leveraging rich representations fr...

Zi-Xun Guo, Calvin Murdock, Sanjeel Parekh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.