Modern text-to-speech (TTS) systems are rarely deployed as unchanged pre-trained models. They are often adapted through supervised fine-tuning (SFT) or preference optimization such as DPO and GRPO. This raises a practical question for audio deepfake forensics: do fingerprints learned from a foundation generator remain...
Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic...
Yu-Xiang Wang, Kun-Yu Feng, Yuan-Cheng Wang et al.· 2 citations
Tffic-Audio is presented, a general speech deepfake detection system designed for comprehensive evaluation environment that achieves a pooled EER of 1.454% on the 14 test sets of Speech-DF-Arena, outperforming all currently public systems on the leaderboard.
Wan Lin, Li Wang, Jin-Dong Wang et al.· arXiv.org· 1 citation
RecurTrace introduces Loop Memory Attention, which lets each looped layer attend to its own states from previous iterations along the loop-time axis, so the model can revisit earlier computations instead of relying on the latest state alone.
Yu-Xiang Wang, Kun-Yu Feng, Ying-Da Shen et al.· 6 citations
MSEditor is proposed, the first framework designed specifically for consistent multi-shot video editing, which significantly outperforms existing methods on the authors' curated multi-shot video editing benchmark in terms of identity preservation, temporal stability, and overall visual quality.
Kun-Yu Feng, Yue Ma, Bing-Yuan Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.