Skip to content
Conference

Multiview Consistency Learning with Synthetic Camera Views for Hand-in-the-Wild Gesture Recognition

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 706-711 · 0 citations · 15 references

Abstract

Hand gesture recognition is an important problem in human–computer interaction, virtual reality, augmented reality, robotics, and multimedia systems. Despite strong performance in previous studies, recognition in uncontrolled environments remains challenging due to viewpoint variation, self-occlusion, articulation ambiguity, and cross-dataset domain shift. To address these issues, we propose a 3D skeleton-based synthetic multi-view framework. A Blender-based animation pipeline reconstructs hand poses from 3D skeleton annotations, while multiple virtual cameras generate camera-space skeleton views from different viewpoints. These synthetic views are used only during training, preserving single-view inference. The original skeleton and its generated views are processed by a shared GRU encoder, and an original-anchored consistency loss aligns their representations. Experiments are conducted under a challenging 7class cross-dataset setting, using CanonicalSet for training and HandinWildSet for external evaluation. The proposed 8-camera original-anchored consistency framework achieves the highest accuracy of 75.85% on HandinWildSet, outperforming the GRU baseline (62.96%) and rotation-based augmentation (70.22%). The corresponding 8-camera model without consistency achieves 68.44%, confirming the additional benefit of original-anchored consistency. Experimental results demonstrate that combining synthetic camera-space views with consistency learning improves viewpoint robustness and cross-dataset generalization.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.