Skip to content
Preprint

Decoupled Latent Flow Matching for Few-Step Joint Vocal-Accompaniment Separation

Aug 2026 · 0 citations · 25 references
Computer Science

Abstract

Generative modeling provides a flexible way to model mixture-conditioned source distributions, but iterative diffusion and flow matching models are costly for long music signals. This paper studies joint vocal-accompaniment separation through latent flow matching, where a pretrained variational autoencoder (VAE) maps mixtures and sources into a compact latent space and a flow matching model generates vocal and accompaniment latents jointly. The proposed framework decouples semantic separation from acoustic velocity prediction through a Separation Encoder and a Velocity Decoder. To reduce sampling cost, we further apply latent adversarial post-training inspired by Flow2GAN for few-step generation. Experiments show that latent adversarial refinement can improve perceptual and separation metrics under a reduced sampling budget.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.