Skip to content
Open access

Deep Features Evaluation Method of Human Action Recognition Based on Convolutional Neural Network

Jul 2026 · ICCK Transactions on Advanced Computing and Systems · 0 citations · 89 references

TL;DR

A compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling is proposed, indicating a robust, real-time-capable solution for video understanding in both offline analytics and online deployment.

Abstract

Human Action Recognition (HAR) in unconstrained video remains difficult due to cluttered backgrounds, camera motion, and long-range temporal dependencies. The recognition of human action is the most complex study in the area of Artificial Intelligence (AI) and Computer Vision (CV). Machine vision for online and offline video processing is typically employed in the development of human behavior recognition systems. In video broadcasting and analysis, identifying the type and content of human actions present in the footage is a fundamental requirement. In this article, we propose a compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling. The approach is evaluated on three representative benchmarks KTH (6 classes), Hollywood-2 (12 actions), and UCF-50 (50 actions). The system attains strong aggregate performance, reaching 98.5% accuracy on KTH, 93.0% accuracy on UCF-50, and 96.0% mAP on Hollywood-2, with improvements distributed broadly across classes. Error–length analyses show that recognition quality rises steadily with longer clips and that the recurrent module extracts the largest gains, while precision–recall and ROC curves with superior Area Under the Curve (AUC) and reliability demonstrate improved ranking and probability calibration across operating points. Despite a larger parameter count than several baselines, the design achieves the most favorable efficiency profile measured at 18ms per frame latency, 135 mJ per frame energy, and the highest throughput placing it on the accuracy–latency–energy Pareto frontier. An ablation study confirms the centrality of temporal modeling (6-12 percentage-point drops without the LSTM), the benefit of longer temporal windows, and the usefulness of stepped learning-rate schedules for stable optimization. The results indicate a robust, real-time-capable solution for video understanding in both offline analytics and online deployment.

Read PDF

Similar papers

Conference Jul 2026

Computer vision-based human motion recognition: technological evolution, key challenges, and cutting-edge trends

This paper systematically reviews the evolution of computer vision-based human motion recognition technology, from traditional handcrafted feature methods to convolutional neural networks and recurrent neural networks in the deep learning era, and then to the recently emerging Transformer architectures and vision-langu...

Xin Sun · 0 citations
Review Open access Aug 2026

AI-based vision techniques for human activity recognition in surveillance videos

The article compares the performance of traditional machine learning techniques with recent deep learning architectures such as CNNs, RNNs, TCNs, and Transformers, based on accuracy, computational cost, and suitability for real-world disorderly plotting.

Disha Deotale, Madhushi Verma, P. Suresh et al. · 0 citations
Open access Aug 2026

Research on Rehabilitation Training Movement Recognition and Real-time Feedback Model Based on Computer Vision

This study provides a replicable technical path for the validation of rehabilitation evaluation algorithms without clinical data collection through the adaptive fusion mechanism to dynamically integrate the confidence of the deep network and the matching score of dynamic time warping template.

Mingxiang Yang · 0 citations
2026

Explainable 3D Convolutional Neural Networks Spatiotemporal Learning for Human Handshake Interaction Recognition

Human Activity Recognition (HAR) has gained significant attention in computer vision due to its wide range of applications in surveillance, social behaviour analysis, and human–computer interaction. Among various human-to-human interactions, handshake recognition is particularly important as it represents social intent...

S. Kumaravel, S. Veni · 0 citations
Conference Jul 2026

AVT-PAC: A Pipeline for Multimodal Action Prediction and Captioning

Action prediction from frames and videos is a well-studied problem. Models trained with a single modality, mostly vision, will fail in low-light conditions. Recent works have attempted to predict action categories using vision-language and audio-visual models. A challenge, however, is that some dataset annotations lack...

A. R, Ambarish Parthasarathy, Sucharitha Devarakonda et al. · 0 citations
Open access Jul 2026

LIP Reading Using Neural Network and Deep Learning

An automated lip reading system built using Convolutional Neural Networks to classify spoken words from video sequences of a speaker's mouth region and is able to reliably identify isolated words such as “Hello”, “Start”, “Stop”, and “Previous” in real time.

B. Arjun, M. Saad, Aaron Biju et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.