Jul 2026· International Journal of Precision Engineering and Manufacturing· 2 citations· 18 references
TL;DR
An LLM-based Voice-to-Action (VTA) system that converts spoken user commands into grounded robot behaviors for an indoor mobile manipulator, LeeAhn 2, and suggests that LLM-grounded spoken interfaces can reduce operator burden and improve accessibility for indoor service robots.
Abstract
Voice interaction has been actively studied in human–robot interaction (HRI) for decades, yet deploying spoken interfaces on physical mobile manipulators remains challenging because language is ambiguous, tasks are long-horizon, and robot actions must be grounded to perception and motion under real-world uncertainties. Recent large language models (LLMs) offer a practical way to interpret open-ended spoken requests, but their non-deterministic outputs and limited transparency can hinder safe and reproducible robot execution. This paper presents an LLM-based Voice-to-Action (VTA) system that converts spoken user commands into grounded robot behaviors for an indoor mobile manipulator, LeeAhn 2. The system combines speech transcription with an LLM that produces structured, skill-level action plans aligned with a predefined library of robot capabilities, including vision-based seeking, wheeled navigation, and manipulation. To improve reliability, we incorporate interface constraints that restrict generated actions to executable skills and enable recovery from common failures during execution. We evaluate the proposed system through simulation experiments and real-world trials on a representative search task, reporting component-level performance across seeking, navigation, and manipulation. The simulation experiments provide repeatable analysis under controlled conditions, while the real-world trials demonstrate practical applicability on the LeeAhn 2 mobile manipulator and reveal limitations such as latency and occasional plan/execution failures. Our results suggest that LLM-grounded spoken interfaces can reduce operator burden and improve accessibility for indoor service robots.
Initial experiences with the application of a Mistral AI language model with a Pepper robot for Human-Robot Interaction (HRI) in dialogue and an investigation of the effects of additional visual information on response time in different models show that incorporating visual information adds context to the dialogue with...
The World-Cognition Model is presented, a human-centered embodied agent built on the SLAK architecture (Sensing, Logic, Action, and Knowledge) and an asynchronous runtime and introduces a human-in-the-loop teaching mode that enables users to interactively teach the robot difficult or long-horizon tasks.
Natural language interfaces can lower the expertise barrier for operating robotic manipulators by allowing users to express goals in everyday speech. This paper presents RANLP-Arm, a modular bilingual voice-to-manipulation system that converts spoken Thai or English instructions into real-time robotic arm actions for t...
Supakorn Thavornvong, A. Kitsommart, Mahannop Thabua et al.· 2026 23rd International Conf...· 0 citations
This paper proposes an HRI agentic artificial intelligence system—integrating a construction domain-specific, noise-robust automatic speech recognition (ASR) agent and a vision-language model (VLM)-based robotic control agent—to reliably transcribe speech in noisy environments and parse instructions for robot navigatio...
Oscar Poudel, Rayan H. Assaad, Mohamad Awada· Journal of computing in civi...· 0 citations
AnthroDial is presented, a closed-loop framework that formulates anthropomorphic dialogue as a joint problem of system architecture, executable evaluation, and diagnostic alignment and shows that anthropomorphic dialogue benefits when generation, evaluation, and reward shaping share the same behavioral dimensions.
Wentao Liu, Si-Yu Song, Xi Chen et al.· arXiv.org· 0 citations
A ready-to-deploy intent-aware system in which a social robot conveys active listening through non-verbal backchannels grounded in interactional intents to enable active listening for robots.
Yang Sun, Jan Leusmann, Michael A. Hedderich· Message Understanding Confer...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.