Skip to content

SpokenUS: A Spoken User Simulator for Task-Oriented Dialogue

Mar 2026 · arXiv.org · Vol abs/2603.16783 · 0 citations · 58 references
Computer Science

TL;DR

SpokenUS is presented, a spoken user simulator grounded in TOD that decides when to speak through a dedicated turn-taking head that achieves comparable goal coverage to much larger models while substantially outperforming all baselines in human MOS.

Abstract

Robust voice agents require exposure to the full diversity of how people interact through speech. However, obtaining enough spoken interactions is prohibitively expensive. Building spoken user simulators that address this requires large-scale spoken task-oriented dialogue (TOD) data encompassing spoken user behaviors, yet existing datasets are limited in scale and domain coverage, with no systematic pipeline for augmenting them. To address this, we introduce SpokenTOD, a spoken TOD dataset of 52,390 dialogues and 1,034 hours of speech augmented with four spoken user behaviors---cross-turn slots, barge-in, disfluency, and emotional prosody---across diverse speakers and domains. Building on SpokenTOD, we present SpokenUS, a spoken user simulator grounded in TOD that decides when to speak through a dedicated turn-taking head. SpokenUS achieves comparable goal coverage to much larger models while substantially outperforming all baselines in human MOS, disclosing slot values gradually across the dialogue as humans do rather than front-loading them. Further analysis confirms that SpokenUS's spoken behaviors pose meaningful challenges to voice agents, making it a practical tool for evaluating more robust spoken dialogue systems. Our code is available at https://github.com/holi-lab/SpokenUS.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

SteerDuplex: Steerable Duplex Speech Dialogue Models

Full-duplex spoken dialogue models support low-latency turn taking, interruption handling, and backchanneling, yet a key capability remains underexplored: steerability, the ability to reliably shift conversational behavior along attributes such as tone, persona, speaking rate, and voice style in response to user instru...

Utkarsh Tyagi, Ramaneswaran Selvakumar, Advait Gosai et al. · 2 citations
Preprint Sep 2026

From Metrics to Natural Dialogue: French Full-Duplex Benchmark for Spoken Dialogue Models

Full-duplex spoken dialogue models aim to make voice agents more natural by allowing them to listen, speak, pause, and respond during ongoing conversation. However, it is not clear whether full-duplex benchmarks behave the same way when models are evaluated in a different language. To investigate this, we introduce a F...

Hamid Soltani, Gilles Boulianne · 0 citations
#artificial intelligence Preprint Sep 2026

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

RoleBreak, an open benchmark for long-horizon role-playing robustness in spoken dialogue, is introduced and four key patterns are found, highlighting persistent gaps in both long-horizon robustness and vocal expressiveness in spoken role-playing systems.

Yu-Qi Wang, Feng-Yuan Liu, Hao-Chen Luo et al. · 0 citations
Preprint Sep 2026

RePlay: Retrieval-Based Voice Playback for Multi-Turn spoken dialogue

Many voice interaction applications require exact control over both the content and delivery of responses, typically using pre-recorded lines. Recent full-duplex models respond with low latency but cannot guarantee exact content or reproduce a specific recorded performance, while cascaded systems can be constrained to...

Sathvik Udupa, Naveen Kumar, Ryan Folmsbee · 0 citations
#natural language process... Preprint Aug 2026

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

TurnBench, a multi-domain benchmark that pairs a 30-hour, hand-labeled corpus of dyadic human conversation with a standardized evaluation protocol for end-of-turn and interruption detection, finds end-of-turn recall stable across types, while interruption false positives are strongly type-dependent and concentrated in...

Freeman Jiang, Ramon Sanabria, Soham Deshmukh et al. · 5 citations · ⚡1

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.