Diffusion- and flow-based generative models have achieved strong performance in speech enhancement, but they typically rely on an explicit time variable to specify the generation stage. Since the noisy speech condition provides a strong reference to the underlying clean speech, enhancement progress could be reflected b...
Hao Liang, Wei Liu, Gong-Ping Huang et al.· IEEE Signal Processing Lette...· 0 citations
Audio anti-spoofing systems increasingly combine self-supervised learning, parameter-efficient fine-tuning, and graph-attention-based backends. However, performance gains in such systems are often entangled with concurrent changes in the backbone, fine-tuning strategy, and training protocol, making the independent cont...
Hao-Yu Wang, Jing Yang, Chen-Yu Liu et al.· 0 citations
Experiments show that ER-EDF consistently improves empathetic response quality in both automatic and human evaluations, highlighting the importance of jointly modeling emotion perception and regulation in spoken empathetic dialogue systems, paving a new direction for psychologically grounded empathetic AI.
Hong-Yu Jin, Wen-Da Zhang, Run-Qiu Fei et al.· 0 citations
An evidence-grounded generative SE framework that uses a deterministic estimate as imperfect evidence as imperfect evidence is proposed and SNR-Conditioned CSG (SNR-CSG), which maps a calibrated residual-SNR estimate to an utterance-level strength and constructs an adaptive grounded anchor is introduced.
Hao Shi, Yuan Gao, Zhao-Heng Ni et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.