Exploring Multimodal Adapters with Multi-Task Decoders for Video Action Recognition.
Large-scale vision-language pretrained models like CLIP, paired with parameter-efficient fine-tuning techniques, have emerged as promising solutions for image-to-video transfer in video action recognition. However, existing methods often prioritize strong supervised performance at the expense of transferability and gen...