Skip to content
Preprint

U-PAST: A Phase-Aware Audio Spectrogram Transformer-U-Net for Single-Channel Speech Enhancement

Aug 2026 · 0 citations · 38 references
Engineering Computer Science

TL;DR

U-PAST is a hybrid transformer-U-Net architecture that addresses self-attention dependency-modeling in the complex spectrogram domain through self-attention dependency-modeling in the complex spectrogram domain, offering an attractive performance-to-cost trade-off at a small parameter footprint.

Abstract

Convolutional neural networks (CNNs), used widely and successfully in audio enhancement, capture long-range time-frequency dependencies only indirectly, through successive convolution and pooling. Here, we present U-PAST, a hybrid transformer-U-Net architecture that addresses this limitation through self-attention dependency-modeling in the complex spectrogram domain. U-PAST tokenizes a complex STFT representation, similarly to the magnitude spectrogram tokenization of the Audio Spectrogram Transformer (AST), applies a multi-layer transformer encoder, and reconstructs the enhanced complex spectrogram with a U-Net-style decoder. We evaluate four architectural variants with between 1.17M and 2.40M parameters on the DNS Challenge, VoiceBank-DEMAND, and LibriMix corpora under matched, acoustic mismatch, and two-dataset mismatch conditions. U-PAST attains the best SI-SDR of any evaluated model under acoustic mismatch and closely trails substantially larger convolutional and time-domain baselines by 0.26 dB to 0.63 dB SI-SDR under the remaining three conditions while achieving the strongest perceptual (DNSMOS) quality under dataset mismatch. The largest evaluated configuration, U-PAST-H (2.40M parameters), is consistently the strongest variant of the family, offering an attractive performance-to-cost trade-off at a small parameter footprint.

View source

Similar papers

Preprint Aug 2026

Beyond Speech: Dual-Domain SSL Fusion for Unified All-Type Audio Deepfake Detection

Unified all-type audio deepfake detection aims to determine whether an input clip is real or fake when its audio type may be speech, environmental sound, singing voice, or music. Existing speech-centric or type-dependent solutions are insufficient for this setting because the test-time audio type is unknown, while the...

Cunhang Fan, Jun-Qin Cao, Tian Gao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

A receptive-field-scaling study shows that expanding self-attention beyond 15 phonemes provides no consistent gains in pitch, energy, or duration prediction, and introduces a fixed-receptive-field convolutional encoder that reduces the respective prediction errors.

Zi-Tao Liang, Chang Gao · 0 citations
Preprint Sep 2026

HAMMER: Harmonic-Aware Parallel Context Modeling and Discriminator-Free Perceptual Optimization for Speech Enhancement

Recent speech enhancement systems combine self-attention and Mamba to capture global interactions and long-range dependencies. Yet these hybrids usually operate as sequence mixers and do not explicitly exploit harmonic periodicity, a strong cue for preserving voiced speech under noise. Perceptual optimization poses ano...

Shang-Fu Chen, Szu-Wei Fu, Sung-Feng Huang et al. · 0 citations
Open access Aug 2026

A deep residual complex learning framework with long-range temporal context for phase-aware speech enhancement

In this paper, we propose an advanced speech enhancement model capable of effectively separating clean speech from noisy audio signals. The primary objective here is to improve speech intelligibility and quality in noisy environments while preserving critical speech components. We propose a GAN based novel residual lea...

Debabrata Gogoi, Sushanta Kabir Dutta · 0 citations
Preprint Sep 2026

Disentangled Global-Local Feature Learning with E-Branchformer for Audio Deepfake Detection

The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervis...

Phuong Dat, Học Thủ, T. Nguyễn et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.