Occlusion-Aware GCN-Transformer Framework for Robust Skeleton-Based Human Activity Recognition
Abstract
There has been a lot of development in “skeleton-based human activity recognition” due to its emphasis on human joint actions, which eliminates the effect of background noises such as lighting, body appearance changes, or other environmental aspects. However, there are limitations in the skeleton data that can be caused by various factors, including occlusions, missing joints, appearance loss, skeleton overlap, and inaccurate estimates of human actions. In general, existing Graph Convolutional Network (GCN) and Transformer approaches perform poorly with incomplete skeletons. To solve the problem, this paper introduces “Occlusion-Aware GCN-Transformer Framework for robust skeleton-based human activity recognition”. The framework uses an adaptive spatial GCN to detect local joint relationships and a spatiotemporal transformer to identify motion dependencies between frames. To augment model robustness, the training and evaluation processes encompass various occlusion scenarios, including random joint masking, upper-body occlusion, lower-body occlusion, and corrupted temporal frames.