Skip to content
Preprint

Beyond Dry References: Learning Relative Audio Effects Representations via Contrastive Distance Learning

Aug 2026 · 0 citations · 31 references
Computer Science

TL;DR

RelFx is proposed, a contrastive learning framework that learns relative effect transformations from general audio collections without requiring dry references during representation training, and demonstrates state-of-the-art performance under the standard Fx-Encoder++ MUSDB18 evaluation protocol.

Abstract

Audio effects (Fx) representation learning plays a key role in intelligent music production, including automatic mixing and Fx style transfer. Existing methods typically rely on dry or nearly dry references for effect modeling, yet truly unprocessed audio is rarely available in practice, as real recordings inevitably reflect the microphone, room acoustics, and preceding signal processing. Instead of pursuing absolute effect encodings, we argue that the relative effect distance between audio signals is more meaningful for real-world music production. Motivated by this, we propose RelFx, a contrastive learning framework that learns relative effect transformations from general audio collections without requiring dry references during representation training. Our approach uses a dual-branch Siamese encoder equipped with cross-attention and differential gating fusion to infer the shared effect transformation from a reference clip and an effect-processed, content-related clip. We further propose an antisymmetric fusion variant for bidirectional effect encoding, such that swapping the input order directly produces a nearly sign-reversed embedding, a property not explored in earlier work. Moreover, our dry-reference-free formulation eliminates the reliance on dry multitrack datasets and enables training on effect-bearing audio. Experiments on Fx style transfer demonstrate state-of-the-art performance under the standard Fx-Encoder++ MUSDB18 evaluation protocol, consistently outperforming existing approaches across all four instrument categories.

View source

Similar papers

Preprint Aug 2026

Exploring the Design Space of Representation Learning for Audio Transformations

This framework produces both a transformation embedding and a processed-audio embedding, and it finds that the two play complementary roles: distance-based tasks favor the former, while probe-based tasks favor the latter.

Sungho Lee, Marco A. Martínez-Ramírez, Junghyun Koo et al. · 0 citations
Sep 2026

Self-supervised audio representation learning model based on time-frequency decoupling and masked reconstruction

A self-supervised model that integrates a time-frequency decoupled (TF-D) stem with masked latent reconstruction with the strongest overall transfer results among the internal variants and a competitive balance among representation quality, encoder scale, and inference efficiency is proposed.

Jie Xu, Yu-Hao Dai, Zhi-Feng Wang · 0 citations
Preprint Aug 2026

Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners

NAPE (Next-Audio-Patch-Embedding prediction), a self-supervised framework in which a causal Transformer predicts each next patch embedding of a log-mel spectrogram from the previous ones, using causal masking and stop-gradient as its sole training signal is introduced.

Umberto Cappellazzo, Xu-Bo Liu, Stavros Petridis et al. · 0 citations
Preprint Aug 2026

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching

MiDashengLM-Gen is an end-to-end framework that couples a pre-trained Large Language Model (LLM) with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation and drastically improves speech intelligibility over existing unified models.

Xingwei Sun, Heinrich Dinkel, Gang Li et al. · 0 citations
Preprint Aug 2026

DINO-A: Adapting Self-Distillation Vision Transformers to General Audio Representation Learning

DINO-A is presented, an adaptation of self-distillation from vision to general audio representation learning, and it is traced to two mechanisms: the interaction between DINO's high-dimensional projection space and FSD50K's limited scale, and the additional cost of multi-crop augmentation, which DINO uses but BYOL-A v2...

Tomasz Radzikowski, M. Modrzejewski, Przemysław Rokita · 0 citations
Preprint Aug 2026

Beyond Reconstruction: Full-Context Generative DiT for Music Generation

FullDiT is introduced, a conditional DiT that fuses eight frame-aligned RVQ streams with independently encoded captions and lyrics and uses non-causal self-attention over the complete acoustic latent sequence and outperforms five commercial systems on 15 of 18 automatic metrics.

Yun-Jia Li, Meng-Li Wu, Jun-Yu Dai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.