Deep pose estimation-based action recognition and performance analysis for intelligent motion understanding
Abstract
Human motion understanding requires not only accurate action recognition but also interpretable performance evaluation capable of reflecting motion quality. This paper proposes Pose-ARPA, a unified framework that combines deep pose estimation with spatiotemporal representation learning for comprehensive action recognition and performance analysis. The proposed model first extracts 2D joint keypoints from video frames and lifts them to 3D skeletal sequences using geometry-constrained reconstruction. A hybrid GCN–Transformer encoder is then designed to capture spatial coordination and temporal dynamics across body joints, while a dual-head output performs both action classification and performance scoring. The performance analysis head introduces four quantitative metrics—Fluency, Stability, Symmetry, and Range of Motion—which align closely with human evaluation and offer interpretable insight into movement proficiency. Experimental results demonstrate that Pose-ARPA achieves superior recognition accuracy, robust generalization under noisy or occluded conditions, and real-time inference capability suitable for deployment in sports analytics and rehabilitation applications. The proposed framework provides a pathway toward explainable and human-centered motion understanding in intelligent systems.