Skip to content
Conference

Component-Aware Spatio-Temporal Adaptation of Frozen Foundation Models for Video Deepfake Detection

2026 · Poster Volume 0007 The 2026 Twenty-Second International Conference on Intelligent Computing July 23-26, 2026 Toronto, Canada · 0 citations

Abstract

The advancement of deep generative models facilitates realistic synthetic facial videos, threatening social trust and digital security. Existing detection methods achieve strong in-domain performance but suffer from cross-dataset degradation, primarily due to overfitting to dataset-specific spatial artifacts. To address these challenges, we propose a parameter-efficient, video-based deepfake detection framework that leverages a frozen foundation model encoder coupled with a lightweight spatio-temporal decoder. First, we introduce a Component-Aware Spatial Enhancement (CASE) module that selectively accentuates manipulation-prone facial regions, such as the eyes, mouth and nose, while capturing global-local structural inconsistencies, thereby enabling the detection of subtle artifacts. Second, it is complemented by a Bidirectional Spatio-Temporal (Bi-ST) decoder that models local temporal transitions and bidirectional temporal dependencies across sampled frames, enabling robust temporal reasoning within each video clip. Without fine-tuning the backbone network, our framework achieves robust cross-dataset generalization by jointly reasoning about spatial and temporal anomalies. Finally, extensive experiments demonstrate that the proposed method performs competitively against strong baselines, particularly under cross-dataset evaluation.

View source