Jul 2026· International Conference on Computer Vision, Al and Intelligent Automation· Vol 14260, pp. 142600O - 142600O-6· 0 citations· 10 references
Engineering
TL;DR
This paper systematically reviews the evolution of computer vision-based human motion recognition technology, from traditional handcrafted feature methods to convolutional neural networks and recurrent neural networks in the deep learning era, and then to the recently emerging Transformer architectures and vision-language models.
Abstract
Human motion recognition is an important research direction in computer vision, aiming to automatically analyze human pose changes and action semantics from images or videos captured by visual sensors. This paper systematically reviews the evolution of computer vision-based human motion recognition technology, from traditional handcrafted feature methods to convolutional neural networks and recurrent neural networks in the deep learning era, and then to the recently emerging Transformer architectures and vision-language models. The core ideas and technical characteristics of various approaches are comprehensively analyzed. On this basis, key challenges in current research are deeply explored, including the complexity of spatiotemporal feature extraction, robustness issues with occlusion and viewpoint changes, difficulties in fine-grained action and few-shot learning, and the trade-off between multimodal fusion and computational efficiency. Finally, future development trends are discussed, pointing out that cutting-edge directions such as 3D human pose estimation, fisheye lens adaptation, micro-action detection, and attribute-aware generation will drive the field toward higher precision, stronger generalization capabilities, and broader application scenarios. This paper aims to provide systematic technical reference and forward-looking insights for researchers in human motion analysis and intelligent human-computer interaction.
A compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling is proposed, indicating a robust, real-time-capable solution for video understanding in both offline a...
H. Khan, Altaf Hussain· ICCK Transactions on Advance...· 0 citations
A critical review of computer vision, illustrating how architectural design, learning paradigms, and evaluation practices have co-evolved over time to facilitate more flexible and scalable systems, and outlining new research directions.
This study provides a replicable technical path for the validation of rehabilitation evaluation algorithms without clinical data collection through the adaptive fusion mechanism to dynamically integrate the confidence of the deep network and the matching score of dynamic time warping template.
Mingxiang Yang· Journal of Discovery Core· 0 citations
The article compares the performance of traditional machine learning techniques with recent deep learning architectures such as CNNs, RNNs, TCNs, and Transformers, based on accuracy, computational cost, and suitability for real-world disorderly plotting.
Disha Deotale, Madhushi Verma, P. Suresh et al.· Discover Artificial Intellig...· 0 citations
A comparative analysis of existing studies is presented to highlight the evolution of deep learning techniques and their effectiveness in improving recognition accuracy and computational efficiency and emerging research directions are outlined to provide insights for future research.
Patel Bhautika Ronak· International journal of res...· 0 citations