Advancing Video-Text Pretraining with Multi-View Captions
This work proposes a large-scale multimodal large language model-based supervision generation framework that improves supervision diversity, fidelity, and semantic coverage, and introduces a granularity-aware text representation with separate CLS tokens for summary and detailed views.
F. M. Thoker, Renaud Vandeghen, Karen Sanchez et al.
· 0 citations