FIT: Fine-Grained Identity-Aware Transformer for Generalizable Diffusion Face Forgery Detection
Abstract
Face forgery detection (FFD) is essential for the security and authenticity verification of digital media. Current FFD methods suffer from poor generalization to unseen face images created by diffusion models. Besides, they tend to rely on coarse-grained prior information interaction paradigms. In this paper, we propose a novel fine-grained identity-aware transformer (FIT) method for generalizable diffusion face forgery detection (DFFD). Specifically, we are motivated by the novel observation that the preservation of target identity in facial images generated by GAN and diffusion models varies significantly. We employ the inherent identity preservation differences between GAN and diffusion face images to capture identity-aware forgery representations in a fine-grained learning manner. We employ the learned identity forgery embeddings as prior information to facilitate DFFD. We propose a fine-grained identity-aware transformer block (FITB) to mine fine-grained global identity-appearance forgery features based on intra-patch identity-aware relations as well as inter-patch global identity-perceptual relationships in diffusion face images. An identity contrastive center loss is devised to achieve intra-class identity forgery embedding compaction and inter-class identity forgery representation separation, to study discriminative and general diffusion face forgery patterns. Extensive experimental results demonstrate that FIT outperforms the state-of-the-art via cross-generator, cross-dataset, and robustness evaluation.