Mar 2026· arXiv.org· Vol abs/2603.16783· 0 citations· 58 references
Computer Science
TL;DR
SpokenUS is presented, a spoken user simulator grounded in TOD that decides when to speak through a dedicated turn-taking head that achieves comparable goal coverage to much larger models while substantially outperforming all baselines in human MOS.
Abstract
Robust voice agents require exposure to the full diversity of how people interact through speech. However, obtaining enough spoken interactions is prohibitively expensive. Building spoken user simulators that address this requires large-scale spoken task-oriented dialogue (TOD) data encompassing spoken user behaviors, yet existing datasets are limited in scale and domain coverage, with no systematic pipeline for augmenting them. To address this, we introduce SpokenTOD, a spoken TOD dataset of 52,390 dialogues and 1,034 hours of speech augmented with four spoken user behaviors---cross-turn slots, barge-in, disfluency, and emotional prosody---across diverse speakers and domains. Building on SpokenTOD, we present SpokenUS, a spoken user simulator grounded in TOD that decides when to speak through a dedicated turn-taking head. SpokenUS achieves comparable goal coverage to much larger models while substantially outperforming all baselines in human MOS, disclosing slot values gradually across the dialogue as humans do rather than front-loading them. Further analysis confirms that SpokenUS's spoken behaviors pose meaningful challenges to voice agents, making it a practical tool for evaluating more robust spoken dialogue systems. Our code is available at https://github.com/holi-lab/SpokenUS.
This work presents AVERT, which scores each candidate value by combining cross-turn agreement with a trained audio-conditioned verifier and resolves the three error types with three operators, vote, add, and swap, each restricted to the slots where its error is common.
Full-duplex spoken dialogue models support low-latency turn taking, interruption handling, and backchanneling, yet a key capability remains underexplored: steerability, the ability to reliably shift conversational behavior along attributes such as tone, persona, speaking rate, and voice style in response to user instru...
Utkarsh Tyagi, Ramaneswaran Selvakumar, Advait Gosai et al.· 2 citations
Full-duplex spoken dialogue models aim to make voice agents more natural by allowing them to listen, speak, pause, and respond during ongoing conversation. However, it is not clear whether full-duplex benchmarks behave the same way when models are evaluated in a different language. To investigate this, we introduce a F...
RoleBreak, an open benchmark for long-horizon role-playing robustness in spoken dialogue, is introduced and four key patterns are found, highlighting persistent gaps in both long-horizon robustness and vocal expressiveness in spoken role-playing systems.
Yu-Qi Wang, Feng-Yuan Liu, Hao-Chen Luo et al.· 0 citations
Many voice interaction applications require exact control over both the content and delivery of responses, typically using pre-recorded lines. Recent full-duplex models respond with low latency but cannot guarantee exact content or reproduce a specific recorded performance, while cascaded systems can be constrained to...
Sathvik Udupa, Naveen Kumar, Ryan Folmsbee· 0 citations
TurnBench, a multi-domain benchmark that pairs a 30-hour, hand-labeled corpus of dyadic human conversation with a standardized evaluation protocol for end-of-turn and interruption detection, finds end-of-turn recall stable across types, while interruption false positives are strongly type-dependent and concentrated in...
Freeman Jiang, Ramon Sanabria, Soham Deshmukh et al.· 5 citations· ⚡1
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.