DeCAF: Decoupling Core-Frame Semantic Abstraction from Chain-of-Thought Generation for Video Reasoning
The rapid progress of multimodal large language models (MLLMs) has significantly advanced multimodal understanding. However, video reasoning remains challenging due to the lack of scalable data construction methods with explicit reasoning supervision. Existing VideoQA datasets either rely on costly manual annotations,...