Skip to content
Open access

ActNet: focus-aware multi-scale CNN for human activity recognition from images

Sep 2026 · PeerJ Computer Science · 0 citations · 51 references

Abstract

Recognizing human actions from still images is a challenging task due to the absence of temporal information and the need to infer actions from subtle pose and contextual cues. In this article, we propose ActNet, a novel deep convolutional neural network (CNN) architecture that combines multi-scale feature learning with a focus-aware attention mechanism to address this problem. ActNet integrates a Multi-Feature Network (MFNet) backbone for extracting rich features from multiple receptive fields, an Activity Multi-scale Block (AMB) for learning spatially diverse action patterns, and a Focus-Aware Recognition Module (FARM) that adaptively highlights the most informative regions of the image. We evaluate ActNet on the Stanford 40 Actions and PASCAL Visual Object Classes (VOC) 2012 datasets and show that it outperforms several state-of-the-art CNN and transformer-based models, achieving superior accuracy, precision, recall, and F1-score. Extensive ablation studies confirm the effectiveness of both AMB and FARM components. ActNet demonstrates robust generalization to a wide range of human actions, making it a strong candidate for still-image-based action recognition tasks in practical applications.

Read PDF

Similar papers

Sep 2026

MICA-Net: A Multimodal Cross-Attention Network for Human Action Recognition

A novel action recognition method, named MICA-Net, which combines data from multiple sensors to improve the efficiency of the HAR model, and a new compact version of a wrist-worn sensor device with Wi-Fi connectivity to an edge device, enhancing usability in human-machine interaction applications.

Trung-Hieu Le, Thai-Khanh Nguyen, T. Tran et al. · 0 citations
Conference Aug 2026

MSCALNet: a multiscale convolutional attention LSTM network for IMU-based human activity recognition

Wearable devices play an increasingly pivotal role in human activity recognition (HAR), particularly driven by the urgent demand in medical applications ranging from rehabilitation monitoring to fine-grained gait analysis. However, existing methods still struggle with insufficient exploration of cross-modal information...

Zi-Bo Wang, Runyang Lyu, Bin Xiao · 0 citations
Open access Aug 2026

Hand Gesture Recognition Based on Multi-Scale Attention Graph Convolutional Network

Advances in artificial intelligence have made hand gesture recognition an important human–computer interaction modality. Graph convolutional networks (GCNs) are widely used for skeleton-based hand gesture recognition, yet their performance can be limited by weak semantic topology modeling, underused feature channels, a...

Xiaowei Han, Ting-Shan Yan, Yunjing Lu et al. · 0 citations
Open access Sep 2026

Learn the interactions: Weakly supervised video anomaly detection with human-object interactions

Weakly supervised video anomaly detection remains a challenging problem, primarily due to the scarcity of abnormal training samples and the lack of diverse feature representations, which hamper the learning of discriminative models. To address these issues, we introduce a novel weakly supervised cross-domain framework...

Mao-Wen Zhou, Erma Rahayu Mohd Faizal Abdullah, Aznul Qalid Md Sabri et al. · 0 citations
Open access Sep 2026

Action Recognition in Sports Videos Using Multiscale Convolutional Networks with Long Short-Term Memory (LSTM)-Based Temporal Modeling.

Action recognition in sports videos remains challenging because of complex motion dynamics, occlusion, and high intra-class variability. Although existing deep learning approaches, including CNN-BiLSTM and transfer learning-based models, have demonstrated effectiveness in human activity recognition, their performance m...

Hai-Ming Yang, Hafiz Mohd Sarim, Xiao-Juan Ma et al. · 0 citations
Open access Aug 2026

Action Recognition Method Based on Multi-Scale Dilated Feature Fusion and Decoupled Spatiotemporal Attention Pooling

Video action recognition requires the joint modeling of spatial appearance information and temporal dynamics. However, existing efficient action recognition methods based on two-dimensional convolution still have limitations in representing multi-scale spatial cues and aggregating key spatiotemporal information. To add...

Han-Bo Zhang, Jing Huang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.