Skip to content
Preprint

A Hybrid Classical-Learning Framework for Adaptive Decision Directed Speech Enhancement

Aug 2026 · 0 citations · 15 references
Engineering

TL;DR

An Adaptive Beta-Constrained Decision-Directed (ABCDD) speech enhancement framework that extends the conventional DD method through a frame-dependent lower gain bound and combines interpretable classical enhancement structure with lightweight machine-learning-based parameter adaptation provides an effective and practical direction for robust speech enhancement.

Abstract

Speech enhancement aims to recover clean speech signals from noisy observations while preserving speech quality and intelligibility. Classical methods such as Spectral Subtraction and Decision-Directed (DD) enhancement remain widely used because of their interpretability and low computational complexity, but they may suffer from musical-noise artifacts or excessive attenuation of weak speech components under low signal-to-noise ratio (SNR) conditions. This paper proposes an Adaptive Beta-Constrained Decision-Directed (ABCDD) speech enhancement framework that extends the conventional DD method through a frame-dependent lower gain bound. The introduced beta parameter controls the tradeoff between noise suppression and speech preservation. To automate parameter selection for large and diverse datasets, a lightweight multilayer perceptron (MLP) model is further developed to predict frame-level beta values directly from noisy-speech features. The proposed framework is evaluated using both a representative speech example and large-scale testing on the VoiceBank-DEMAND dataset. In the representative example, ABCDD outperformed conventional Spectral Subtraction and classical DD across multiple objective metrics, including SNR, Log-Spectral Distance (LSD), Root-Mean-Square Error (RMSE), correlation, and Scale-Invariant Signal-to-Distortion Ratio (SI-SDR). On 100 unseen VoiceBank-DEMAND test files, the proposed MLP-beta ABCDD method improved average scale-aligned SNR from 9.41 dB to 13.82 dB, corresponding to an average gain of 4.41 dB. The results indicate that combining interpretable classical enhancement structure with lightweight machine-learning-based parameter adaptation provides an effective and practical direction for robust speech enhancement.

View source

Similar papers

Open access Aug 2026

A deep residual complex learning framework with long-range temporal context for phase-aware speech enhancement

In this paper, we propose an advanced speech enhancement model capable of effectively separating clean speech from noisy audio signals. The primary objective here is to improve speech intelligibility and quality in noisy environments while preserving critical speech components. We propose a GAN based novel residual lea...

Debabrata Gogoi, Sushanta Kabir Dutta · 0 citations
Open access Aug 2026

EMD-Based Speech Denoising with Instantaneous Noise Modeling

This paper introduces a fully data-driven speech denoising framework based on Empirical Mode Decomposition (EMD) combined with instantaneous, frame-dependent noise modeling. The proposed method departs from classical EMD-shrinkage approaches by (i) estimating noise directly from the noisy signal without requiring voice...

Kais Khaldi, Anis Mohamed · 0 citations
Review Open access Sep 2026

Advances in Speech Enhancement: A Comprehensive Review of Noise Suppression Techniques

Over the past several decades, numerous methods have been developed to improve the signal-to-noise ratio, perceptual quality, and intelligibility of speech. In practice, no single method is universally optimal, as each category exhibits distinct strengths and limitations under specific acoustic conditions. The proposed...

Pushpraj Tanwar, A. Somkuwar, Rakesh Kumar Gumasta · 0 citations
Open access Aug 2026

Efficient speech denoising using CleanUNet optimized with Mamba and hybrid spectral loss

Speech enhancement aims to recover clean speech from signals degraded by noise and adverse acoustic conditions, such as background interference and echo. CleanUNet has emerged as an effective solution for causal speech denoising, but its high computational demands limit deployment on resource-constrained devices. In th...

Matheus Vieira da Silva, João Fernando Mari, A. Backes · 0 citations
Sep 2026

A dual-stream spectral network for versatile full-band speech enhancement.

Full-band speech restoration aims to recover high-fidelity speech from signals degraded by noise, reverberation, bandwidth limitation, or their combinations. Existing methods are typically designed for a specific type of degradation, which limits their flexibility when multiple distortions coexist. To address this prob...

Jun-Kang Yang, Hiromitsu Nishizaki, C. Leow et al. · 0 citations
Open access Aug 2026

Single-Channel Speech Enhancement Method Based on Deep Neural Networks

In practical voice communication and processing, speech signals are extremely sensitive to background noise, equipment noise, and interference from complex environments, leading to a decline in speech quality and clarity. This, in turn, affects the overall performance of downstream systems such as speech recognition, s...

Wen Fan, Wei-Yu Liang, Duo-Duo Han et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.