We adapt Reinforce Adjoint Matching (RAM), a reward-based post-training method, to generative speech enhancement (SE). Starting from a pretrained SE model, RAM tilts the model's conditional distribution toward outputs with higher reward. During training, the current model generates enhanced speech on-policy, evaluates...
Julius Richter, Christoph Boeddeker, Yoshiki Masuyama et al.· 0 citations
In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio recordings, where one microphone is located close to the target machine and the other is located farther away to capture noise. For this task,...
Takuya Fujimura, G. Wichern, Yoshiki Masuyama et al.· 0 citations
MERL's submission to the Real-TSE Challenge achieved first place in the second track, demonstrating the critical importance of high-quality data preparation and observing that DNSMOS and speaker similarity are susceptible to over-optimization.
Dominik Klement, Yoshiki Masuyama, Christoph Boeddeker et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.