As human--AI interactions become more conversational, full-duplex speech language models capable of natural real-time dialogue are growing in importance. Beyond generating appropriate responses, these models must coordinate turn-taking, backchanneling, and floor management in real time. Reinforcement learning (RL) prov...
Kyudan Jung, Hyunsin Park, Yoonhyung Lee et al.· 0 citations
Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-BigVGAN pipeline. Since vocoder reconstruction of a real mel ca...
Eunji Shin, Kyudan Jung, Jihwan Kim et al.· 0 citations
SNAP, a speaker-nulling framework, is introduced that reduces speaker entanglement and encourages detectors to focus on artifact-related patterns, leading to state-of-the-art performance.
Kyudan Jung, Ji-Hoon Kim, Minwoo Lee et al.· arXiv.org· 0 citations
OmniACBench, a benchmark for evaluating context-grounded acoustic control in omni-modal models, is introduced and three common failure modes are identified-weak direct control, failed implicit inference, and failed multimodal grounding-providing insights for developing models that can verbalize responses effectively.
Seunghee Kim, B. Park, Kyudan Jung et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.