Skip to content
Conference

Gesture-aware adaptive interface design using vision-language models for enhanced human–computer interaction

Aug 2026 · International Conference on Machine Vision and Deep Learning · Vol 14326, pp. 143262F - 143262F-7 · 0 citations · 19 references
Engineering

Abstract

Gesture-based interaction represents a natural and intuitive paradigm for human-computer interaction (HCI), yet current systems struggle to bridge the semantic gap between low-level gesture recognition and high-level interface adaptation. This paper proposes GAUID (Gesture-Aware UI Design), a novel framework that leverages vision-language models to jointly understand hand gestures, voice commands, and interaction context for adaptive interface generation. The system integrates MediaPipe hand tracking with a Vision Transformer (ViT) for gesture encoding, a Whisper-BERT pipeline for voice command understanding, and a CLIP-based vision-language fusion module that aligns gesture-text representations in a shared embedding space. A cross-attention mechanism combines multi-modal features to generate context-aware UI layout recommendations. We evaluate GAUID on three benchmark datasets: NVGesture, HaGRID, and a custom GestureHCI corpus comprising 8,680 gesture-command-UI triplets. Experimental results demonstrate that GAUID achieves gesture recognition accuracy of 92.4% on GestureHCI and 89.7% on NVGesture, outperforming six baseline methods. A user study with 40 participants confirms significant improvements in task completion rate (91.2%), user satisfaction (4.8/5), and learnability (4.7/5) compared to static and rule-based interfaces.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.