Jul 2026· International Conference on Control, Decision and Information Technologies· pp. 1853-1858· 0 citations· 19 references
Abstract
Over the last decade, significant advances have been made in the field of human robot interaction (HRI), such as collaborative workcells in smart factories and autonomous mobile robots. However, for these systems to achieve widespread societal acceptance, more natural and efficient communication interfaces must be developed. This work proposes a computer vision-based human–robot cooperative system, composed of a set of gestures that, when combined into composite sequences, allow intuitive interaction with the environment. Gesture identification relies on a human pose analysis model based on deep learning, which earned the 1st place in the Flying Robots Demo at RoboCup 2025. The semantic and contextual interpretation of these sequences is conducted by Large Language Models (LLMs). Integration between visual perception and linguistic understanding enables a more expressive and adaptable form of interaction. The developed system achieved a high-fidelity gesture detection model, with average precision and recall of 0.94.
This study presents a real-time vision-based gesture control framework for mobile robots using a monocular RGB camera. The goal is to enable intuitive and low-cost human– robot interaction without requiring specialized sensing hardware. The proposed system combines geometric rule-based reasoning with a Support Vector M...
Chuyu Guo· 2026 IEEE International Conf...· 0 citations
Collaborative robot or COBOT has emerged as a key advancement in robotics, significantly transforming industrial automation. Consequently, real-time gesture recognition for human-COBOT interaction has recently garnered significant attention. However, both datasets designed for practical robot control scenarios and end-...
Ha-Anh Nguyen, A. Nguyen, Duy-Anh Doan et al.· IEEE International Conferenc...· 0 citations
The feasibility of combining refined gesture detection with multimodal agents for resource-constrained robotic interaction is demonstrated, and an enhanced YOLO-based gesture detector with coordinated large language model (LLM) and vision-language model (VLM) agents is proposed.
Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot interaction. However, this task faces three critical barriers: the scarcity of semantically rich datasets, the"modality eclipse"where models ignore audio cues in favor of kine...
Zi-Fan Wang, Ziang Ren, Pengteng Shi et al.· 0 citations
Initial experiences with the application of a Mistral AI language model with a Pepper robot for Human-Robot Interaction (HRI) in dialogue and an investigation of the effects of additional visual information on response time in different models show that incorporating visual information adds context to the dialogue with...
Preliminary experiments on several representative arm gestures indicate that the proposed method can produce meaningful imitative motions from monocular RGB input only, while also highlighting limitations in more complex poses and wrist-related movements.
Anastasiya Ihnatovich, Igor Farkas· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.