Skip to content

Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

This work proposes masked autoregressive SE (MARSE), a method for SE based on iterative decoding of masked clean speech frames using continuous NAC representations of speech, which enables a flexible trade-off between SE performance and computational cost.

Abstract

Most previous work on speech enhancement (SE) based on masked generative modeling relied on discrete token representations of audio signals, obtained using neural audio codecs (NACs). However, a recent study has shown that continuous latent representations of NACs can be advantageous for SE in terms of speech quality and intelligibility. In this work, we propose masked autoregressive SE (MARSE), a method for SE based on iterative decoding of masked clean speech frames using continuous NAC representations of speech. In particular, we investigate a set of different decoding policies, ceteris paribus, that is, using the same DNN (a Conformer model), the same NAC (the DAC codec) and the same training setup. The results show that MARSE enables a flexible trade-off between SE performance and computational cost. Audio examples and code are available online.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Test-time adaptation for speech enhancement with an autoregressive speech prior

Test-time adaptation (TTA) offers a promising direction for improving speech enhancement models under mismatched acoustic conditions, without requiring access to labeled target data. In this work, we propose a single-utterance TTA method that regularizes a pretrained speech enhancement model using an autoregressive pri...

S. Kammoun, Simon Leglaive, Xavier Alameda-Pineda et al. · 1 citation
Preprint Aug 2026

Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec

Language model (LM)-based speech enhancement (SE) has recently emerged rapidly using latent space features of neural audio codecs (NACs). In this paper, first, we present a unified framework covering six popular LM-based generative SE modeling paradigms based on discrete/continuous latent NAC features: discrete or cont...

Yihui Fu, Zhengyang Li, Tim Fingscheidt · 0 citations
Preprint Aug 2026

Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition

Large language model (LLM)-based audio-visual speech recognition (AVSR) systems are robust under noise. Contrastive decoding (CD), originally introduced to stabilize LLM generation by contrasting a weaker model against a stronger one at inference time, adjusts predictions without additional training. In this work, we a...

Youngchae Kim, Dali Yang, Joon-Hyuk Chang · 0 citations
Preprint Aug 2026

A Hybrid Classical-Learning Framework for Adaptive Decision Directed Speech Enhancement

An Adaptive Beta-Constrained Decision-Directed (ABCDD) speech enhancement framework that extends the conventional DD method through a frame-dependent lower gain bound and combines interpretable classical enhancement structure with lightweight machine-learning-based parameter adaptation provides an effective and practic...

Ali Rajabi, Xiang-Wei Zhou · 0 citations
#artificial intelligence Preprint Sep 2026

What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection

The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes....

Jia-Jun Xu, Meng-Lu Li, Xiao-Ping Zhang · 0 citations
Preprint Sep 2026

Mask-Based Speech Enhancement for Spatial Audio: A Comparison of Ambisonics, Beamforming, and Microphone Channels

Mask-based speech enhancement is widely used for suppressing noise and interference, but its performance in spatial audio algorithms with multichannel output has not been studied extensively. In such settings, speech enhancement must improve speech quality while preserving spatial cues that are essential for localizati...

Sheli Hendel, B. Rafaely, Dorothea Kolossa · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.