Dynamic-frame-rate neural speech codecs replace a uniform frame grid with variable-duration tokens, making boundary placement part of the representation itself. Yet it is unclear what these boundaries encode and whether interpretable boundaries are also useful for neural speech reconstruction. This work combines bounda...
Han Wang, Jia-Qi Li, Ying Shen et al.· 0 citations
While Chain-of-Thought (CoT) reasoning has improved the capability of language models, directly applying it to Spoken Language Models (SLMs) may introduce long silent intervals under the serial"think-then-speak"paradigm, disrupting real-time spoken interaction. To address this issue, we propose an asynchronous think-al...
Junyi Ao, Kainan Peng, Mingbo Ma et al.· 0 citations
Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoofing artifacts can be further shaped by variability in source speech, recording environments, and tra...
Wan Lin, Li Wang, Jindong Wang et al.· arXiv.org· 1 citation
ParaBridge is proposed, an on-policy self-distillation method that turns a brittle inference-time scaffold into stable model behavior and generalizes to unseen paralinguistic cues, transfers from safety-oriented training to empathy-oriented dialogue, and works on a different SLM backbone.
Yuxiang Wang, Qin-Ke Ni, Sheng-Bo Cai et al.· arXiv.org· 2 citations
RecurTrace introduces Loop Memory Attention, which lets each looped layer attend to its own states from previous iterations along the loop-time axis, so the model can revisit earlier computations instead of relying on the latest state alone.
Yu-Xiang Wang, Kunyu Feng, Ying-Da Shen et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.