Skip to content

A dual-stream spectral network for versatile full-band speech enhancement.

Sep 2026 · Journal of the Acoustical Society of America · Vol 160 3, pp. 2272-2287 · 0 citations · 51 references
Medicine

Abstract

Full-band speech restoration aims to recover high-fidelity speech from signals degraded by noise, reverberation, bandwidth limitation, or their combinations. Existing methods are typically designed for a specific type of degradation, which limits their flexibility when multiple distortions coexist. To address this problem, this paper proposes a dual-stream versatile speech enhancement network (DSVN)-a unified task-conditioned system for full-band speech restoration. The model introduces a dual-stream spectral block to jointly model frequency-wise spectral structures and temporal dynamics. A band-refill decoder is further designed to recover missing or degraded high-frequency components. Task-conditioned modulation is introduced to enable the same network to adapt to different objectives. Evaluation results demonstrate that DSVN improves restoration quality under both single- and mixed-degradation scenarios. The model shows clear advantages in non-intrusive quality assessment, spectral distance, and automatic speech recognition evaluation, indicating that the enhanced speech is not only cleaner but also more useful for downstream systems. Additional analysis of internal responses and ablation studies provides evidence that the proposed modules contribute to different aspects of restoration.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.