Spatial Knowledge Distillation in Video Models via Vision-Language Guided Zero-Shot Pretraining
Video action recognition tasks require large-scale labeled data to achieve high performance when pretraining is not used. However, annotating video data is costly and time-consuming. To reduce this dependency on labeled data, transfer learning approaches are commonly employed. Vision–language models, which learn genera...