Skip to content

Human-robot conversation with multiple participants in noisy public spaces

Sep 2026 · 0 citations · 34 references
Computer Science

TL;DR

This work proposes an audio system that can be used for both an attentive listening system with the android ERICA and a conversation support system with mobile Teleco robots, with one of them acting as an avatar for a remote operator.

Abstract

For noisy real-world environments such as those in open public spaces, spoken dialogue systems for both autonomous robots and avatars should be carefully designed to provide enhanced speech signals. These signals can be used either for speech recognition or, in the case of an avatar system, transmitted as clean speech to a remote operator. This work proposes an audio system that can be used for both these scenarios and was demonstrated as a proof-of-concept at the 2025 World Expo in Osaka. The first scenario is an attentive listening system with the android ERICA, and the second is a conversation support system with mobile Teleco robots, with one of them acting as an avatar for a remote operator. Both systems feature multi-party conversation and use a single multi-channel microphone array. We describe how our audio system not only enhances the speech of multiple speakers in a noisy environment, but provides a form of spatial audio which allows for more immersiveness in avatar-based conversational interactions.

View source

Similar papers

Preprint Aug 2026

From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios

Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby conversations. This"cocktail party"scenario still presents severe challenges to speech recognition systems. The CHiME-9 MCoRec task provides a...

Thai-Binh Nguyen, Zhaolin Li, Jan Niehues et al. · 0 citations
#natural language process... Preprint Sep 2026

KABURI-TTS: Phoneme-Keyed Activity-conditioned Bi-channel Utterance Rendering for Interaction

Realizing full-duplex spoken dialogue requires large amounts of two-channel, one-speaker-per-channel conversational speech data. Although conversational text-to-speech (TTS) engines have been developed, they are not necessarily robust to two-party simultaneous phenomena such as backchannels, interruptions, and overlaps...

Ryuichiro Higashinaka, Shinnosuke Takamichi, Tetsuji Ogawa · 0 citations
Preprint Aug 2026

SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning

Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR around a proprietary voice API, where gradients cannot flow and per-call cost makes on-policy reinforcement learning prohibitive, or stay in te...

Jia-Jun Fan, Jing-Yuan Li, Prashanth Gurunath Shivakumar et al. · 1 citation · ⚡1
Conference Aug 2026

Real-Time Multilingual Speech-to-Text AR Captioning Glasses for the Deaf and Hard-of-Hearing

For people who are deaf or hard-of-hearing (HoH), everyday conversations can be difficult to navigate, even with modern hearing aids. Augmented reality (AR) smart glasses offer a practical way to provide live captions directly in the user’s field of view. However, current models are often too expensive, rely exclusivel...

Mahesh Paul J, Soundharesh M, Vaissnave V et al. · 0 citations
Preprint Aug 2026

VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-speech and speech-to-text models reduce latency by replacing multi-stage pipelines, but often compromis...

Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor et al. · 3 citations
Preprint Aug 2026

Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio handed to them. The standard remedy is a cascade that first gen...

Chengqian Ma, Wei Tao, Hao-Yu Zhang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.