Noise and sampling schedules in score-based diffusion models for speech enhancement
Abstract
As diffusion models continue to demonstrate significant potential in the field of speech enhancement, it has become increasingly important to investigate the factors that influence reconstruction quality and to evaluate their robustness and generalization capabilities across various real-world environments. This paper explores various noise scheduling and sampling scheduling schemes based on the score-based diffusion model, and analyzes their differing impacts on the process of recovering clean speech. The results show that different noise scheduling schemes affect the perceptual speech quality during the forward process. Specifically, with cosine scheduling, our method improved the PESQ score to 2.96, achieving a 10% improvement over the baseline. Meanwhile, it was demonstrated that sampling scheduling has a negligible impact on model performance when the number of reverse steps is sufficient. Additionally, we constructed a dataset of noisy speech samples featuring various noise types and intensities to evaluate the robustness and generalization capabilities of the best-performing cosine noise scheduling model, and the results confirmed the ability to generate intelligible speech even at low SNR levels.