Semantic-Conditioned Forensic Consistency Model for Open-World Image Authenticity Verification
Abstract
The recent upsurge in the development of sophisticated generative models has significantly improved the visual realism and semantic coherence of synthetic images, thus presenting a major challenge to the field of multimedia forensics. The conventional approaches often rely on either artifacts or high level semantic cues, limiting their robustness when handling images/videos generated by models that were previously unseen. The proposed work addresses this problem by developing a novel forensic consistency learning framework that is conditioned on semantic content. Specifically, the model leverages pretrained DINO (self-DIstillation with NO labels) encoder for visual content, forensic features capturing latent acquisition characteristics and image-level CLIP (Contrastive Language-Image Pretraining) features for global semantic content. A forensic predictor module estimates the expected forensic features conditioned on semantic information, enabling capture of inconsistencies between visual content and underlying artifacts. Additionally, a patch-level anomaly score module enables robust image-level prediction. The method was evaluated under cross-generator setting, by training on GenImage dataset augmented with ProGAN, and evaluated on UniversalFakeDetect (UFD) benchmark. The model achieves 91.5% Area Under ROC Curve, 92.2% Average Precision and 84.1% classification accuracy, significantly outperforming existing baselines like UFD and ResNet (upto 6-10% improvement). Extensive ablations further validate the effectiveness of proposed semantic-conditioned forensic modeling for open-world image authentication.