Skip to content
Review Open access

Deep learning for speech enhancement: architectures, paradigms, and emerging trends

Aug 2026 · Scientific and Technical Journal of Information Technologies, Mechanics and Optics · 0 citations · 11 references

TL;DR

The need for further research in the development of lightweight models for mobile devices, multi-distortion suppression methods, and the integration of neural network noise suppression with generative models to achieve a new level of speech signal restoration quality is demonstrated.

Abstract

A systematic review of modern neural network methods for Speech Enhancement is presented, aimed at improving speech intelligibility and quality under acoustic distortions. The reviewed methods can be applied in voice control systems, telecommunications, hearing aids, and human–machine interaction interfaces. Key architectural approaches are considered, including classical recurrent and convolutional networks as well as modern hybrid architectures with attention mechanisms (Transformer, Conformer), state-space models (Mamba), and advanced recurrent blocks (xLSTM). The advantages and disadvantages of different architectures are shown in terms of speech restoration quality and computational efficiency. Specific problems of existing methods are highlighted, including high computational cost and insufficient generalization capability under non-stationary noise conditions. The need for further research in the development of lightweight models for mobile devices, multi-distortion suppression methods, and the integration of neural network noise suppression with generative models to achieve a new level of speech signal restoration quality is demonstrated.

Read PDF

Similar papers

Review Open access Sep 2026

Advances in Speech Enhancement: A Comprehensive Review of Noise Suppression Techniques

Overall, this review provides a unified perspective on speech enhancement techniques, identifies their strengths, and highlights emerging research opportunities for the development of robust, efficient, and intelligent speech enhancement systems.

Pushpraj Tanwar, A. Somkuwar, Rakesh Kumar Gumasta · 0 citations
Open access Aug 2026

A deep residual complex learning framework with long-range temporal context for phase-aware speech enhancement

In this paper, we propose an advanced speech enhancement model capable of effectively separating clean speech from noisy audio signals. The primary objective here is to improve speech intelligibility and quality in noisy environments while preserving critical speech components. We propose a GAN based novel residual lea...

Debabrata Gogoi, Sushanta Kabir Dutta · 0 citations
Open access Aug 2026

Single-Channel Speech Enhancement Method Based on Deep Neural Networks

In practical voice communication and processing, speech signals are extremely sensitive to background noise, equipment noise, and interference from complex environments, leading to a decline in speech quality and clarity. This, in turn, affects the overall performance of downstream systems such as speech recognition, s...

Wen Fan, Wei-Yu Liang, Duo-Duo Han et al. · 0 citations
Preprint Aug 2026

A Hybrid Classical-Learning Framework for Adaptive Decision Directed Speech Enhancement

An Adaptive Beta-Constrained Decision-Directed (ABCDD) speech enhancement framework that extends the conventional DD method through a frame-dependent lower gain bound and combines interpretable classical enhancement structure with lightweight machine-learning-based parameter adaptation provides an effective and practic...

Ali Rajabi, Xiang-Wei Zhou · 0 citations
Open access Sep 2026

Speech intelligibility comparison of standalone and two-stage deep learning architectures for behind-the-ear-to-binaural enhancement

The relative perceptual performance of alternative deep learning based speech processing architectures remains uncertain under complex acoustic conditions. This study compared speech intelligibility, in listeners with normal hearing and with mild-to-moderate hearing loss, obtained with two processing models: a speech...

R. Viveros-Muñoz, Carla E. Contreras-Saavedra, Sebastián Guajardo-Herrera et al. · 0 citations
Aug 2026

DMAWNet: Multi-Objective Based Distributed Attention Enabled WNet Framework for Speech Enhancement

Speech enhancement is significant with the advancement of communication systems, and in real-world circumstances, ambient noise is often present in speech data, emphasizing the vital role of enhancement techniques. The existing methods employed for enhancement are susceptible to various drawbacks, including the lack of...

Anil Garg, A. Singhal, Deepti Chaudhary et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.