This work introduces the first large-scale benchmark for this task by collecting real portraits and fake portraits from multiple sources and evaluates representative existing detectors on this benchmark, revealing their lack of explicit mechanisms for retaining fine-grained information and maintaining feature consistency across rendered views.
Abstract
Recent advances in single-image 3D Gaussian head reconstruction have enabled highly realistic and freely renderable digital heads from a single portrait. However, reconstruction and rendering can weaken the forgery traces in the source portrait, making the resulting 3D face difficult to classify whether its underlying face is real or fake, and thereby posing risks to identity authentication and face privacy. To study this problem, we introduce the first large-scale benchmark for this task by collecting real portraits and fake portraits from multiple sources and evaluate representative existing detectors on this benchmark, revealing their lack of explicit mechanisms for retaining fine-grained information and maintaining feature consistency across rendered views. To directly address these two limitations, we propose a detector trained with a two-stage strategy. In Stage I, masked autoencoding encourages the visual backbone to retain the fine-grained appearance information required for local reconstruction, while multi-view contrastive learning enforces feature consistency across rendered views of the same head. Since CLS tokens at different depths exhibit complementary spatial attention patterns, Stage II freezes the adapted backbone and concatenates low-, middle-, and high-level CLS tokens for classification. Experiments show that our method achieves the highest accuracy and ranks first across all reported metrics among the evaluated detectors.
Deepfake technology, powered by deep learning models, enables the synthesis of highly realistic facial images and videos. However, in recent years, the misuse of deepfakes has posed severe challenges to both individual privacy and social trust. Consequently, this paper systematically reviews research pertaining to deep...
The proposed sketch-guided face generation pipeline based on a Conditional Variational Autoencoder designed to generate facial reconstructions from sketches and conditional attributes uses a stochastic preprocessing pipeline to extract edge maps from facial photographs, reducing dependence on manually paired sketch-pho...
Edson M. Odake, E. P. Ribeiro· Multimedia tools and applica...· 0 citations
This paper proposes a novel transfer framework that addresses the stochastic nature of generative priors, and results are consistently preferred over diffusion and GAN-based baselines for ocular consistency, temporal stability, and overall restoration quality.
Radim Spetlík, David Futschik, Radek Danecek et al.· 0 citations
The role of GANs in overcoming occlusion by synthesizing realistic facial textures in the masked regions, thereby restoring the identity cues is focused on.
Payal Parekh, Hina Choksi, Mahesh Goyani et al.· ITEGAM- Journal of Engineeri...· 0 citations
Generalizable face forgery detection has become a critical problem in multimedia forensics as modern face manipulation techniques can generate increasingly realistic facial content. Vision foundation models provide a promising basis for this problem, but existing detectors usually rely on the last-layer visual feature,...
Yi-Meng Zhao, Shuo Zhu, Jia-Lang Liu et al.· 2026 12th International Conf...· 0 citations
The rapid development of generative portrait models has raised growing concerns about privacy leakage and identity misuse. In particular, audio-driven 3D talking face generation can reconstruct a reusable 3D portrait of a target person from a monocular video and animate it with arbitrary speech, making realistic identi...
Rui-qing Sun, Chenxing Cui, Hui Zhao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.