A Comprehensive Investigation on Image-Level Face Forgery Detection in the Spatial Domain
Abstract
Deepfake technology, powered by deep learning models, enables the synthesis of highly realistic facial images and videos. However, in recent years, the misuse of deepfakes has posed severe challenges to both individual privacy and social trust. Consequently, this paper systematically reviews research pertaining to deepfake detection in the facial spatial domain. First, based on differences in neural network feature extraction paradigms, this paper categorizes existing spatial-domain detection methods into three groups: physical-statistical methods that focus on low-level pixel-wise noise and lighting distributions; boundary-fusion methods that utilize self-supervised learning to capture splicing artifacts; and high-order semantic consistency discrimination methods based on ViT and feature decoupling. Second, this paper focuses on examining the robustness of these models when confronted with degraded data—specifically, data subjected to social media transcoding and high-intensity compression. Furthermore, this paper provides an in-depth analysis of the advantages offered by cutting-edge techniques—such as self-supervised feature decoupling and Diffusion Model Reconstruction (DIRE)—in identifying the intrinsic logic of the generative manifold. Finally, this paper summarizes the current limitations of spatial-domain detection techniques regarding cross-dataset generalization and environmental robustness, and this paper offers a forward-looking perspective on the future development of multi-dimensional consistency verification and anti-interference defense systems.