Efficient compressed video action recognition via frame selection and token merging in vision transformers
This work proposes a compressed-video-oriented framework, the Frame Selection and Token Merging for Efficient Compressed Video Transformer (FSTM-ECVT), which follows a dual-stream Transformer architecture equipped with a Global Multi-Modal Fusion module to effectively leverage the distinctive and complementary characte...