Joint Smoothed Cross-Entropy and Soft-DTW Learning for TimeSformer-Based Video Action Recognition
: Video-based action recognition remains challenging because of variations in action execution speed and overfitting to static spatial cues. Although the TimeSformer architecture effectively captures long-range spatiotemporal dependencies, the standard cross-entropy loss lacks mechanisms to align dynamic temporal seque...