Increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments and shows that increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments.
Abstract
We investigate how the realism of synthetic room impulse response (RIR) datasets affects the training of DeepFilterNet3 for single-channel speech enhancement. We compare a DNS4 image-source-method (ISM) RIR dataset with a higher-acoustic-fidelity dataset generated using hybrid wave-based and geometrical acoustics simulation. Rather than isolating individual simulation factors, we compare complete RIR generation pipelines while keeping the enhancement model unchanged. Models are evaluated on unseen measured RIRs using objective speech enhancement metrics and downstream automatic speech recognition (ASR). Training with the higher-fidelity dataset consistently yields modest improvements in objective metrics and substantially lower ASR word error rates than the ISM dataset. Although the experiments do not attribute these gains to individual modelling components, they show that increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments.
In this paper, we propose an advanced speech enhancement model capable of effectively separating clean speech from noisy audio signals. The primary objective here is to improve speech intelligibility and quality in noisy environments while preserving critical speech components. We propose a GAN based novel residual lea...
Debabrata Gogoi, Sushanta Kabir Dutta· Engineering Research Express· 0 citations
An EnCodec-based neural RIR compression method, which incorporates RIR structure-aware constraints at two levels, which achieves lower RIR reconstruction error and better reverberant-speech perceptual consistency than audio-oriented codecs.
Chen-Yuan Ning, Yang Ai, Hui-Peng Du et al.· 0 citations
We propose a self-supervised approach for learning room impulse response (RIR) representations from single-channel noisy-reverberant speech. It consists of first training on reverberant data, then on noisy-reverberant data, and finally with a teacher-student approach, where the student learns to replicate the teacher's...
Compact acoustic models face a challenging quality-capacity trade-off. We investigate two factors in this regime: encoder context and Mel-spectrogram supervision. A receptive-field-scaling study shows that expanding self-attention beyond 15 phonemes provides no consistent gains in pitch, energy, or duration prediction....
Over the past several decades, numerous methods have been developed to improve the signal-to-noise ratio, perceptual quality, and intelligibility of speech. In practice, no single method is universally optimal, as each category exhibits distinct strengths and limitations under specific acoustic conditions. The proposed...
Pushpraj Tanwar, A. Somkuwar, Rakesh Kumar Gumasta· Engineer· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.