EMD-Based Speech Denoising with Instantaneous Noise Modeling
Abstract
This paper introduces a fully data-driven speech denoising framework based on Empirical Mode Decomposition (EMD) combined with instantaneous, frame-dependent noise modeling. The proposed method departs from classical EMD-shrinkage approaches by (i) estimating noise directly from the noisy signal without requiring voice activity detection, (ii) applying an SNR-driven, IMF-specific adaptive threshold, and (iii) using a smooth sigmoid-based adaptation function to ensure gradual transitions between speech-dominant and noise-dominant frames. This mechanism enables more accurate discrimination between mixed speech-noise structures while preserving essential temporal details. The method is evaluated on speech signals corrupted by additive white Gaussian noise (AWGN) over input SNR levels ranging from -10 to 10 dB. Experimental results demonstrate consistent improvements over conventional EMD-shrinkage and adaptive wavelet thresholding, with average gains of up to 13.2 dB in output SNR and noticeable perceptual improvements measured by PESQ. The proposed framework remains training-free and computationally efficient, making it suitable for real-time or resource-constrained applications. Although the present study focuses on AWGN, the proposed instantaneous noise modeling strategy provides a solid foundation for future extensions to colored and nonstationary noise environments.