This work proposes an audio system that can be used for both an attentive listening system with the android ERICA and a conversation support system with mobile Teleco robots, with one of them acting as an avatar for a remote operator.
Abstract
For noisy real-world environments such as those in open public spaces, spoken dialogue systems for both autonomous robots and avatars should be carefully designed to provide enhanced speech signals. These signals can be used either for speech recognition or, in the case of an avatar system, transmitted as clean speech to a remote operator. This work proposes an audio system that can be used for both these scenarios and was demonstrated as a proof-of-concept at the 2025 World Expo in Osaka. The first scenario is an attentive listening system with the android ERICA, and the second is a conversation support system with mobile Teleco robots, with one of them acting as an avatar for a remote operator. Both systems feature multi-party conversation and use a single multi-channel microphone array. We describe how our audio system not only enhances the speech of multiple speakers in a noisy environment, but provides a form of spatial audio which allows for more immersiveness in avatar-based conversational interactions.
Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby conversations. This"cocktail party"scenario still presents severe challenges to speech recognition systems. The CHiME-9 MCoRec task provides a...
Thai-Binh Nguyen, Zhaolin Li, Jan Niehues et al.· 0 citations
Realizing full-duplex spoken dialogue requires large amounts of two-channel, one-speaker-per-channel conversational speech data. Although conversational text-to-speech (TTS) engines have been developed, they are not necessarily robust to two-party simultaneous phenomena such as backchannels, interruptions, and overlaps...
Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR around a proprietary voice API, where gradients cannot flow and per-call cost makes on-policy reinforcement learning prohibitive, or stay in te...
For people who are deaf or hard-of-hearing (HoH), everyday conversations can be difficult to navigate, even with modern hearing aids. Augmented reality (AR) smart glasses offer a practical way to provide live captions directly in the user’s field of view. However, current models are often too expensive, rely exclusivel...
Mahesh Paul J, Soundharesh M, Vaissnave V et al.· 2026 International Conferenc...· 0 citations
Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-speech and speech-to-text models reduce latency by replacing multi-stage pipelines, but often compromis...
Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor et al.· 3 citations
An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio handed to them. The standard remedy is a cascade that first gen...
Chengqian Ma, Wei Tao, Hao-Yu Zhang et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 30, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.