Jul 2026· International Conference on Signal Processing and Communications· pp. 1-5· 0 citations· 17 references
Abstract
Deep learning speech enhancement models are trained without grounding in acoustic physics, and evaluations remain confined almost exclusively to English. We address both gaps with MRAN-UNet, which embeds Harmonic Frequency Attention (HFA) - a parameter-free module derived from the source-filter model that aggregates spectral features at candidate $F_{0}$ positions and their harmonic overtones. On VoiceBank-DEMAND, MRAN-UNet achieves CSIG 4.77 (the highest among compared CNN/UNet/RNN baselines), STOI 0.927, and RTF 0.24 with only 3.1 M parameters. PESQ (2.42) trails the strongest convolutional baseline due to decoder spectral coloration, not the HFA mechanism - an effect confirmed by ablation. Complementing the architecture, we release Vaakdhara-DLSE-TE, the first paired enhancement corpus for Telugu (32,000 utterances). Zero-shot transfer improves Telugu STOI from 0.65 to 0.74; 20epoch fine-tuning reaches PESQ 1.91 and STOI 0.92 at 20 dB SNR, outperforming zero-shot DCCRN by 0.56 PESQ at 20 dB SNR.
Text-to-speech systems are improving fast, but measuring how natural they sound still requires expensive human listening tests. Existing automatic methods struggle to generalise well across different datasets. We present QMOS, a MOS prediction framework that extracts hierarchical speech quality features from a frozen Q...
Ravi Sastry Kolluru, S. Devarakonda, S. Radhe Shyam Salopanthula et al.· International Conference on...· 0 citations
CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.
Nhan Phan, Ilona Lähteenmäki, Anna von Zansen et al.· 0 citations
Deepfake voice detection suffers from poor generalization across unseen domains. While Audio Large Language Models (ALLMs) show promise, the modality gap between continuous audio embeddings which capture the subtle acoustic details necessary for deepfake detection and the semantic space of LLMs remains a critical, unde...
Y. El Kheir, Xin Wang, Wan-Ying Ge et al.· 1 citation
U-PAST is a hybrid transformer-U-Net architecture that addresses self-attention dependency-modeling in the complex spectrogram domain through self-attention dependency-modeling in the complex spectrogram domain, offering an attractive performance-to-cost trade-off at a small parameter footprint.
The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervis...
Phuong Dat, Học Thủ, T. Nguyễn et al.· 0 citations
This work proposes Phoenix TTS, a unified framework that tightly couples representation learning with generative acoustic modeling and is optimized to reconstruct self-supervised features to maintain semantic richness, while simultaneously receiving direct supervision from a Flow Matching training loss.
Pei-Jie Chen, Zhuanling Zha, Zhipeng Nie et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.