SynMulti: Synthetic-to-Real Learning for Multimodal Video Understanding
This work proposes a VQA-based fine-tuning strategy that trains models to answer structured questions about visual content rather than relying solely on captions or simple instructions, which encourages deeper visual grounding and reasoning.
Tanzila Rahman, Renjie Liao, Leonid Sigal
· 0 citations