This paper proposes the Equivariant Music Transformer (EMT), which enforces equivariance through self-distillation by jointly optimizing a next-token-prediction and an auxiliary equivariance regularization loss, and finds that the additional equivariance loss acts as a beneficial regularizer, simultaneously improving next-token prediction and producing equivariant latent representations.
Abstract
Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation space. Our analysis, however, shows that standard music transformers map such time-shifted or pitch-transposed inputs onto uncorrelated representations: these models become progressively less equivariant as they scale in size or train longer. This suggests that in standard music transformers, additional model capacity is allocated to memorizing absolute patterns rather than capturing shared musical structures. In this paper, we propose the Equivariant Music Transformer (EMT), which enforces equivariance through self-distillation by jointly optimizing a next-token-prediction and an auxiliary equivariance regularization loss. We find that the additional equivariance loss acts as a beneficial regularizer, simultaneously improving next-token prediction and producing equivariant latent representations. Through both objective and subjective evaluations, EMT demonstrates superior equivariance and generative capability compared to data augmentation, feature engineering, and state-of-the-art (SOTA) baselines. More broadly, our findings reveal that standard language modeling methods alone do not capture music's translational symmetries, and dedicated inductive biases are required to produce better music representations. The code, weights and demos are available online.
We propose a framework for the alignment and interpolation of pitch-aligned time-frequency representations. Building on the tonal interval vector, we introduce a series of extensions that reformulate it as an invertible operator, culminating in a new feature extractor that embeds perceptual consonance priors within a p...
Shahan C. Nercessian, Jeff Sontag, Alejandro Koretzky· 0 citations
Conventional music representations describe acoustic energy over time and frequency but do not explicitly expose relations among simultaneous frequency components. We introduce the \emph{Dissonance Spectrum} (DS), a nonnegative time--frequency representation that applies a tolerance-based rational pitch-relation kernel...
Tian-Le Wang, Xin-Yi Tong, Liang Zhao et al.· 0 citations
This work introduces StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation, and achieves the highest AudioBox Content Enjoyment, Content Usefulness, and Production Quality scores and the highest MuQ-MuLan similarity among th...
Chengli Feng, Zhi-Yue Wu, Jia-Hao Song et al.· 0 citations
Compositionality, the ability to represent complex acoustic scenes as combinations of simpler sound sources, is central to auditory perception and classical additive signal models. Still, it remains unclear whether modern pre-trained audio representations internalize additive structure without compositional supervision...
Chen-Hao Xue, Zhi-Jin Guo, Joyraj Chakraborty et al.· 0 citations
RelFx is proposed, a contrastive learning framework that learns relative effect transformations from general audio collections without requiring dry references during representation training, and demonstrates state-of-the-art performance under the standard Fx-Encoder++ MUSDB18 evaluation protocol.
Xinlu Liu, Huibin Lin, Weixing Wei et al.· 0 citations
Discrete audio tokenization has become the critical interface between raw waveforms and autoregressive modeling in recent music generation. As a result, music tokenizers must simultaneously support high-fidelity reconstruction and produce discrete sequences that remain amenable to language modeling. Existing reconstruc...
Hua-Kang Chen, Guo-Bin Ma, Yue-Peng Jiang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.