Skip to content
Open access

DeepFake Image Detection using Swin Transformer with Attention-guided Feature Fusion and Explainable AI Techniques

Sep 2026 · ELCVIA Electronic Letters on Computer Vision and Image Analysis · 0 citations · 35 references

Abstract

Deepfakes, synthetic media created using advanced machine learning techniques, pose significant societal challenges by spreading misinformation and undermining trust in media. With the increasing sophistication of deepfake technologies, distinguishing between genuine and synthetic media has become increasingly difficult. This paper presents a robust deepfake image detection framework using the Swin-B Transformer, a pre-trained model fine-tuned for our application. By integrating a hybrid dataset that combines real images from the FFHQ dataset and synthetically generated fake images from a publicly available Kaggle dataset, we simulate real-world media scenarios. Our model achieves an impressive accuracy of 97.47\% on the test set, demonstrating superior generalization to both real and synthetic visual data. Using Grad-CAM, we visualize the spatial segments of the image that the model focuses on during classification, providing insight into the decision-making process. This work contributes to enhancing content authenticity, controlling fake news, and ensuring digital trust and safety.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.