Joint Smoothed Cross-Entropy and Soft-DTW Learning for TimeSformer-Based Video Action Recognition
Abstract
: Video-based action recognition remains challenging because of variations in action execution speed and overfitting to static spatial cues. Although the TimeSformer architecture effectively captures long-range spatiotemporal dependencies, the standard cross-entropy loss lacks mechanisms to align dynamic temporal sequences. In this study, we propose a novel joint optimization framework that combines label-smoothed cross-entropy with Soft Dynamic Time Warping (Soft-DTW). Each input frame is divided into fixed-size patches that are embedded as tokens. Spatial attention captures structures and object appearance, while temporal attention learns how these patches evolve over time, enabling robust modeling of gesture dynamics. By factorizing the attention into separate space and time operations, TimeSformer efficiently captures subtle motion cues, and label smoothing reduces prediction overconfidence and reliance on static visual cues. Our primary contribution is the use of Soft-DTW as a temporal regularizer to reduce temporal misalignment between predicted temporal features and class-specific temporal prototypes. To evaluate the framework in a dynamic, domain-specific setting, we constructed a virtual reality (VR) gesture dataset comprising tapping and directional swiping actions. The evaluation results obtained using five-fold cross-validation demonstrate improved classification accuracy. We also conducted an extensive evaluation across standard action-recognition datasets such as Joint-annotated Human Motion Data Base (JHMDB), Human Motion Database 51 (HMDB51), University of Central Florida action recognition dataset (UCF101), and a 5% subset of the Kinetics human action video datasets (Kinetics-400/600). Experimental results show significant improvements in top-1 recognition accuracy on these public datasets. Furthermore, our proposed method demonstrates robust generalization in low-data regimes where TimeSformer typically underperforms.