Gesture-aware adaptive interface design using vision-language models for enhanced human–computer interaction
Abstract
Gesture-based interaction represents a natural and intuitive paradigm for human-computer interaction (HCI), yet current systems struggle to bridge the semantic gap between low-level gesture recognition and high-level interface adaptation. This paper proposes GAUID (Gesture-Aware UI Design), a novel framework that leverages vision-language models to jointly understand hand gestures, voice commands, and interaction context for adaptive interface generation. The system integrates MediaPipe hand tracking with a Vision Transformer (ViT) for gesture encoding, a Whisper-BERT pipeline for voice command understanding, and a CLIP-based vision-language fusion module that aligns gesture-text representations in a shared embedding space. A cross-attention mechanism combines multi-modal features to generate context-aware UI layout recommendations. We evaluate GAUID on three benchmark datasets: NVGesture, HaGRID, and a custom GestureHCI corpus comprising 8,680 gesture-command-UI triplets. Experimental results demonstrate that GAUID achieves gesture recognition accuracy of 92.4% on GestureHCI and 89.7% on NVGesture, outperforming six baseline methods. A user study with 40 participants confirms significant improvements in task completion rate (91.2%), user satisfaction (4.8/5), and learnability (4.7/5) compared to static and rule-based interfaces.