The proliferation of algorithmically synthesized visual con-tent broadly labelled as deepfake poses a mounting threat to theintegrityofdigitalinformationecosystems. Despiterapid advancesingenerativemodelling, robustautomateddetection remains challenging, as synthesis quality now routinely ex-ceedsthethresholdofreliablehumaninspection.Thispaper fine-tunesaVisionTransformer(ViT)classifierfromthepub-licly released dima806/deepfake_vs_real_image_detection checkpoint via the Hugging Face Trainer API on a balanced corpus of 190,081 images drawn from Kaggle, comprising equalproportionsofauthenticphotographsandAI-generated samples spanning GAN-based and latent diffusion architec-tures.The fine-tuned model achieves an overall classifi-cation accuracy of 99.20% and a macro-averaged F1-score of 0.9920 on a held-out evaluation set of 38,081 images, withsymmetricper-classerrorrates(128falsepositives;175 false negatives).These results demonstrate that the Vision Transformer, whose globally unconstrained multi-head self-attentionmechanismenablesdetectionofthelong-rangespa-tial incoherence characteristic of synthetic imagery, consti-tutesacomputationallytractableandhigh-performingarchi-tectureforsynthetic-imageforensics,surpassingallsurveyed CNN and frequency-domain baselines on the same bench-mark.
Kanwerjit, Deep Gaurav, Kaur Sumandeep· Zenodo (CERN European Organi...· 0 citations
The proliferation of algorithmically synthesized visual con-tent broadly labelled as deepfake poses a mounting threat to theintegrityofdigitalinformationecosystems. Despiterapid advancesingenerativemodelling, robustautomateddetection remains challenging, as synthesis quality now routinely ex-ceedsthethresholdofreliablehumaninspection.Thispaper fine-tunesaVisionTransformer(ViT)classifierfromthepub-licly released dima806/deepfake_vs_real_image_detection checkpoint via the Hugging Face Trainer API on a balanced corpus of 190,081 images drawn from Kaggle, comprising equalproportionsofauthenticphotographsandAI-generated samples spanning GAN-based and latent diffusion architec-tures.The fine-tuned model achieves an overall classifi-cation accuracy of 99.20% and a macro-averaged F1-score of 0.9920 on a held-out evaluation set of 38,081 images, withsymmetricper-classerrorrates(128falsepositives;175 false negatives).These results demonstrate that the Vision Transformer, whose globally unconstrained multi-head self-attentionmechanismenablesdetectionofthelong-rangespa-tial incoherence characteristic of synthetic imagery, consti-tutesacomputationallytractableandhigh-performingarchi-tectureforsynthetic-imageforensics,surpassingallsurveyed CNN and frequency-domain baselines on the same bench-mark.
Kanwerjit, Deep Gaurav, Kaur Sumandeep· Zenodo (CERN European Organi...· 0 citations