Skip to content
Conference

DeCAF: Decoupling Core-Frame Semantic Abstraction from Chain-of-Thought Generation for Video Reasoning

Aug 2026 · 2026 12th International Conference on Big Data and Information Analytics (BigDIA) · pp. 466-475 · 0 citations · 20 references

Abstract

The rapid progress of multimodal large language models (MLLMs) has significantly advanced multimodal understanding. However, video reasoning remains challenging due to the lack of scalable data construction methods with explicit reasoning supervision. Existing VideoQA datasets either rely on costly manual annotations, model videos at the frame level with substantial redundancy, or generate reasoning traces in an unconstrained manner, limiting their effectiveness for complex temporal and causal reasoning.In this work, we propose Decoupling Core-Frame Semantic Abstraction from chain-of-thought generation for video reasoning (DeCAF), a scalable framework that decouples video perception from reasoning trace construction. DeCAF first reduces semantic redundancy and abstracts videos into temporally ordered segment-level semantic abstractions (SSA) that capture key events. These event-level observations then serve as intermediate evidence for generating structured reasoning traces with relative event indices.We instantiate DeCAF on NExT-QA and conduct extensive evaluations across multiple MLLMs. Further analyses suggest that DeCAF improves reasoning quality by separating answer-agnostic observations from supervised trace construction, leading to more faithful and temporally coherent rationales. Experimental results demonstrate consistent improvements under different DeCAF configurations, validating the effectiveness of core-frame semantic abstraction and SSA-guided CoT supervision for video reasoning.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.