The recent upsurge in the development of sophisticated generative models has significantly improved the visual realism and semantic coherence of synthetic images, thus presenting a major challenge to the field of multimedia forensics. The conventional approaches often rely on either artifacts or high level semantic cues, limiting their robustness when handling images/videos generated by models that were previously unseen. The proposed work addresses this problem by developing a novel forensic consistency learning framework that is conditioned on semantic content. Specifically, the model leverages pretrained DINO (self-DIstillation with NO labels) encoder for visual content, forensic features capturing latent acquisition characteristics and image-level CLIP (Contrastive Language-Image Pretraining) features for global semantic content. A forensic predictor module estimates the expected forensic features conditioned on semantic information, enabling capture of inconsistencies between visual content and underlying artifacts. Additionally, a patch-level anomaly score module enables robust image-level prediction. The method was evaluated under cross-generator setting, by training on GenImage dataset augmented with ProGAN, and evaluated on UniversalFakeDetect (UFD) benchmark. The model achieves 91.5% Area Under ROC Curve, 92.2% Average Precision and 84.1% classification accuracy, significantly outperforming existing baselines like UFD and ResNet (upto 6-10% improvement). Extensive ablations further validate the effectiveness of proposed semantic-conditioned forensic modeling for open-world image authentication.
The rapid development in deep learning-based generative softwares and image rendering tools has led togeneration of massive photorealistic digital content – Fake images, fake videos, fake speech, etc. Such fake digitalmedia may result in communication of misinformation, forgery of digital data, and losing trustworthiness in theinformation source. This poses a significant challenge to the field of digital forensics’ techniques, to which our presentwork attempts to make a contribution, by addressing the problem of differentiating AI-generated images from realphotographs, using transfer learning and multi-branch fusion model. We propose a multi-branch model that integratestwo pre-trained Vision Transformer models (DINO (self-distillation with no labels) and Contrastive Language – ImagePretraining (CLIP)) to extract complementary global features, along with a forensic and a hand-crafted feature branch,which extract low-level discriminating cues. These features are complimentary to each other and hence contribute inimproving the robustness and performance of the model. These features from the four branches are adaptively weightedand combined by a cross-attention module, to give a fused and rich embedding. The model is further optimized by usingaugmentation-invariant loss, center loss and supervised contrastive loss in addition to the cross-entropy loss function.This framework achieves improved accuracy of 95.83% on PRCG dataset, 96.79% on CIFAKE dataset and 99.15% onGenImage dataset as compared to baselines. It also achieved stable cross-generator performance and enhancedrobustness against real world corruptions like Blur, Noise, Compression, and others. The experimental results show agood separability between the classes, and enhanced performance on publicly available datasets.
Venkata Satya Renuka Devi Bhamidipati, Srinivasa Rao Chanamallu, Sudheer Gopinathan· International Journal of Ima...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.