We propose a self-supervised approach for learning room impulse response (RIR) representations from single-channel noisy-reverberant speech. It consists of first training on reverberant data, then on noisy-reverberant data, and finally with a teacher-student approach, where the student learns to replicate the teacher's embeddings when given a noisy version of the reverberant input. We assess their representational capabilities by estimating acoustic room parameters from them. Conditioning a discriminative speech enhancement model on the derived embeddings yields consistent gains across all evaluated metrics, including downstream word error rate, for both reverberant and noisy-reverberant speech.
Increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments and shows that increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments.
Alessia Milo, G. Götz, S. Guðjónsson et al.· 1 citation
Test-time adaptation (TTA) offers a promising direction for improving speech enhancement models under mismatched acoustic conditions, without requiring access to labeled target data. In this work, we propose a single-utterance TTA method that regularizes a pretrained speech enhancement model using an autoregressive pri...
S. Kammoun, Simon Leglaive, Xavier Alameda-Pineda et al.· 1 citation
This paper addresses unsupervised speech enhancement in the unpaired setting using drifting methods, where training relies on separate collections of degraded and clean audio without corresponding pairs. While recent drifting approaches enable unpaired training, they do so at a heavy cost: because the objective optimiz...
Diego Caviedes-Nozal, Liang Xu, R. Olsson et al.· 0 citations
An EnCodec-based neural RIR compression method, which incorporates RIR structure-aware constraints at two levels, which achieves lower RIR reconstruction error and better reverberant-speech perceptual consistency than audio-oriented codecs.
Chen-Yuan Ning, Yang Ai, Hui-Peng Du et al.· 0 citations
Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder's continuous representations a...
G. Botté, Séverin Baroudi, Samir Sadok et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.