DeCAF: Decoupling Core-Frame Semantic Abstraction from Chain-of-Thought Generation for Video Reasoning
Abstract
The rapid progress of multimodal large language models (MLLMs) has significantly advanced multimodal understanding. However, video reasoning remains challenging due to the lack of scalable data construction methods with explicit reasoning supervision. Existing VideoQA datasets either rely on costly manual annotations, model videos at the frame level with substantial redundancy, or generate reasoning traces in an unconstrained manner, limiting their effectiveness for complex temporal and causal reasoning.In this work, we propose Decoupling Core-Frame Semantic Abstraction from chain-of-thought generation for video reasoning (DeCAF), a scalable framework that decouples video perception from reasoning trace construction. DeCAF first reduces semantic redundancy and abstracts videos into temporally ordered segment-level semantic abstractions (SSA) that capture key events. These event-level observations then serve as intermediate evidence for generating structured reasoning traces with relative event indices.We instantiate DeCAF on NExT-QA and conduct extensive evaluations across multiple MLLMs. Further analyses suggest that DeCAF improves reasoning quality by separating answer-agnostic observations from supervised trace construction, leading to more faithful and temporally coherent rationales. Experimental results demonstrate consistent improvements under different DeCAF configurations, validating the effectiveness of core-frame semantic abstraction and SSA-guided CoT supervision for video reasoning.