Skip to content

The World According to a Social Robot - Augmenting Human-Robot Dialogue With Vision Language Models

Jul 2026 · arXiv.org · Vol abs/2607.16318 · 0 citations · 16 references
Computer Science

TL;DR

Initial experiences with the application of a Mistral AI language model with a Pepper robot for Human-Robot Interaction (HRI) in dialogue and an investigation of the effects of additional visual information on response time in different models show that incorporating visual information adds context to the dialogue with only a moderate increase in response time.

Abstract

Vision Language Models (VLMs) enable robots to visually perceive their environment as well as the actions and characteristics of their conversation partner or humans in collaboration. Especially for social robots deployed in everyday settings and for uncomplicated, natural use, it is essential that the robot has an understanding of situations that is appropriate to human customs. This paper presents initial experiences with the application of a Mistral AI language model with a Pepper robot for Human-Robot Interaction (HRI) in dialogue, as well as an investigation of the effects of additional visual information on response time in different models. The results show that incorporating visual information adds context to the dialogue with only a moderate increase in response time, enabling both the robot and the human to take into account unspoken elements of the situation. Furthermore, using an LLM hosted in Europe offers a solution that complies with European data protection regulations and can therefore facilitate real-life applications more easily.

View source

Similar papers

Book Open access Oct 2026

Exploring Strategies for Conceptual Alignment in LLM-Based Human-Robot Dialogue

Successful conversations require speakers to align on conceptual understanding, a challenging but crucial task in human-robot interaction. With the increasing use of large language models for dialogue, robots move from passively acquiring human conceptualizations to actively shaping alignment. However, the design space...

Sheng-Chen Zhang, Mei-Ying Li, Zi-Xuan Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LLM-based Conversational AI Knowledge Assistant for MyBuddy Humanoid Robot

Humanoid robots are increasingly being popular and developed for human-centered applications, yet their ability to provide intelligent conversations and natural interactive knowledge assistance remains constrained by traditional rule-based dialogue systems, pre-defined responses and limited knowledge repositories. Larg...

Han-Xiao Chen · 0 citations
Preprint Sep 2026

Learning to Plan in Human-Robot Collaboration: Multimodal Reinforcement Learning for Adaptive Interaction

Robot assistants for older adults and people with disabilities need to perform collaborative tasks with users effectively. The core component of these systems is an interaction manager whose job is to observe and assess the task and infer the state of the human and their intent for the robot to choose the best course o...

Afagh Mehri Shervedani, S. Li, Natawut Monaikul et al. · 0 citations
#machine learning Review Sep 2026

STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nu...

Nathan Tsoi, M. Munje, Tejas Oberoi et al. · 0 citations
Preprint Sep 2026

From Wizard-of-Oz Human-Robot Dialogue Collection to a Taxonomy of Robot Response Decisions: A Retrospective Analysis of Assistive Pilot Interactions

Robots that follow natural-language instructions in everyday indoor environments must act on incomplete human utterances. Instructions often omit essential information, such as the identity of an out-of-view object, an intended destination, or the user's goal. Existing datasets contain little real-world situated dialog...

Guang-Ping Liu, Nicholas Hawkins, Tipu Sultan et al. · 0 citations
Review Sep 2026

EmoPose: Vision-Language Model Guided Emotion-Aware Gesture Generation for Humanoid Robots

Socially competent humanoid robots must communicate affect and intent through gesture as well as speech, yet open-ended interaction must become motion that is both expressive and executable on a specific body. This demands semantic flexibility for contextual social intent while preserving deterministic, embodiment-awar...

Dao-Jie Peng, Bing-Tao Wang, Fu-Long Ma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.