Skip to content

CoFi-Lite: Pushing the Limits of Ultra-Lightweight Speech Enhancement

Jul 2026 · IEEE Signal Processing Letters · Vol 33, pp. 2954-2958 · 0 citations · 35 references
Computer Science Engineering

TL;DR

CoFi-Lite is proposed, a highly efficient model that decouples spectral modeling into coarse- and fine-grained streams that outperforms the ultra-lightweight baseline GTCRN while requiring only 40.26% of its computational complexity.

Abstract

Ultra-lightweight models are essential for the deployment of deep learning-based speech enhancement algorithms on edge devices. Although recent approaches have achieved a certain balance between computational complexity and performance, pushing the complexity limits further demands more sophisticated designs. In this letter, we propose CoFi-Lite, a highly efficient model that decouples spectral modeling into coarse- and fine-grained streams. By leveraging two parallel and symmetric encoder-decoder paths, it simultaneously extracts full-band envelopes and low-frequency details for complementary enhancement. In addition, a novel Cross-Path Fusion (CPF) module is introduced to bridge the distinct paths, facilitating efficient feature interaction. Remarkably, CoFi-Lite requires extremely low computational resources, featuring only 12.87 M MACs/s and 83.12 k parameters. Experimental results demonstrate that our proposed model outperforms the ultra-lightweight baseline GTCRN while requiring only 40.26% of its computational complexity. Its scaled-up variant also delivers performance on par with that of the SOTA ultra-lightweight model AdaptCRN alongside a 19.34% reduction in computational cost.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

Compact acoustic models face a challenging quality-capacity trade-off. We investigate two factors in this regime: encoder context and Mel-spectrogram supervision. A receptive-field-scaling study shows that expanding self-attention beyond 15 phonemes provides no consistent gains in pitch, energy, or duration prediction....

Zi-Tao Liang, Chang Gao · 0 citations
Open access Aug 2026

EMBNet: Multi-scale feature learning with efficient channel attention for deepfake speech detection.

With the rapid advancement of generative artificial intelligence, deepfake speech has emerged as a significant threat to digital audio authenticity, posing challenges for forensic analysis and legal applications. In this study, we propose EMBNet, a task-oriented deepfake speech detection framework that integrates effic...

Haitao Yang, Fen Li, Xin Cai et al. · 0 citations
Preprint Sep 2026

Disentangled Global-Local Feature Learning with E-Branchformer for Audio Deepfake Detection

The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervis...

Phuong Dat, Học Thủ, T. Nguyễn et al. · 0 citations
Preprint Sep 2026

PhaseGAN: High-Fidelity Vocoder via Decoupled Amplitude and GAN-Driven Phase Reconstruction

A vocoder is a pivotal component of modern text-to-speech (TTS) systems. Despite the significant progress of neural network-based vocoders, accurate phase reconstruction remains the main challenge limiting both audio quality and modeling efficiency. We introduce PhaseGAN, a lightweight vocoder that addresses this limit...

Wenzheng Zhang, Xue-Liang Zhang, Shu-Lin He et al. · 0 citations
Preprint Sep 2026

EConv-TasNet: Efficient Conv-TasNet for Effective Speech Separation

Conv-TasNet has served as a strong baseline for time-domain speech separation, and many studies have extended it with advanced architectures such as dual-path networks, U-Nets, and attention mechanisms. However, these methods often introduce high computational cost and complexity, limiting their deployment in resource-...

Pei-Chun Chang, Chuan-Yi Liu · 0 citations
#artificial intelligence Preprint Sep 2026

CRAF: Cross-View Residual-Aware Fusion for Deepfake Speech Detection

Recent advances in speech synthesis and voice conversion have made deepfake speech increasingly realistic, making generalization to unseen spoofing attacks a critical challenge. Pretrained speech and audio models offer a promising direction for improving robustness to such unseen attacks. Self-supervised learning (SSL)...

M. Phan, Khalid Zaman, C. Mawalim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.