Skip to content

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling

Jul 2026 · arXiv.org · Vol abs/2607.19776 · 0 citations · 51 references
Computer Science

TL;DR

Experiments show that melodies generated by RPPNet are superior in both long-term structure and musicality, with significant improvements across all subjective evaluation dimensions, and an interdisciplinary perspective for music generation, integrating music theory, computational modeling, and music psychology.

Abstract

Existing symbolic music generation models typically use bars as the basic structural unit. However, human perception of musical phrases often does not align with notated bar lines, leading to long-term structural fragmentation. This paper proposes RPPNet-a two-stage deep learning architecture with variable structural boundaries. It first generates variable-length Rhythm-Pitch Primitive (RPP) sequences, where each RPP encodes note count, rhythm, and contour; then decodes the RPP sequences into concrete notes. The grouping of RPPs is automatically derived from acoustic cues, auditory inertia, and similarity perception based on music psychology. Experiments show that melodies generated by RPPNet are superior in both long-term structure and musicality, with significant improvements across all subjective evaluation dimensions. Ablation studies confirm that the performance gain stems from the structural correctness of the psychological representation, rather than from model capacity. This work offers an interdisciplinary perspective for music generation, integrating music theory, computational modeling, and music psychology.

View source

Similar papers

Open access Aug 2026

Combining CRNN Modeling to Dynamically Align the Relationship Between Music Rhythm and Plot Twists in Different Film Genres

V2A-AlignNet, a genre-aware cross-modal deep learning framework, provides a computational basis for audio-visual temporal relationship analysis and offers valuable insights for multimodal signal interpretation and intelligent information processing in advanced electromagnetic sensing and communication-related applicati...

S.-G. Chi, H.-Y. Liang · 0 citations
Preprint Aug 2026

Pitch Contour Tokenization using VQ-VAE and Its Application on Korean Traditional Music Analysis

Computational analysis of music often relies on discrete representations, yet many musical traditions are organized around continuous pitch movement that resists segmentation into note-like units. For such traditions, the discrete units that analysis would build on are not given in advance. We address this gap by learn...

Seonguk Ju, Seola Cho, Sooin Chung et al. · 0 citations
Preprint Sep 2026

StepAudio 3 Music Technical Report

This work introduces StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation, and achieves the highest AudioBox Content Enjoyment, Content Usefulness, and Production Quality scores and the highest MuQ-MuLan similarity among th...

Chengli Feng, Zhi-Yue Wu, Jia-Hao Song et al. · 0 citations
Conference Open access 2026

Interpretable Vision-to-Music Mapping from Low-Level Visual Statistics to Mode, Rhythm, and Harmonic Structure

Vision-to-music generation transforms visual input into structured musical output, but many recent systems rely on end-to-end neural models whose internal cross-modal decisions are difficult to explain. This paper studies an interpretable alternative based on explicit visual analysis and rule-based symbolic generation....

Shu-Han Yang · 0 citations
Preprint Aug 2026

Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding

Conventional music representations describe acoustic energy over time and frequency but do not explicitly expose relations among simultaneous frequency components. We introduce the \emph{Dissonance Spectrum} (DS), a nonnegative time--frequency representation that applies a tolerance-based rational pitch-relation kernel...

Tian-Le Wang, Xin-Yi Tong, Liang Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.