Skip to content
Conference

Vision Transformer with Attention Rollout for Deepfake Face Image Detection and Localization

K. V. Sai Phani Shaik Mahaboob Jailan F. Mahammad M. Subramanyam
Jul 2026 · 2026 International Conference on Intelligent and Sustainable AI Systems (ICOSAAS) · pp. 1571-1580 · 0 citations · 17 references

Abstract

Generative AI and synthetic media generation tools have enabled widespread media manipulation tools and raised important privacy concerns with misinformation, identity fraud and the verification of authenticity of media. Most of the current convolution-based deepfake detection methods are hard to be deployed in real scenarios and hard to be interpretable, especially because they have limited ability to capture long-range spatial dependency. It introduces an explainable deepfake face image detection framework based on a vision transformer network for performing powerful binary classification of manipulated and real facial images and an explainable face image localization framework for localizing deepfake image faces. The proposed system involves a transformer-encoder backbone for extracting features through a patches-wise process, which proves suitable for modeling the subtle changes of features when the processes of manipulating the image are designed. To make the network more interpretable, and aid the understanding of the transformer attention distribution as well as localization of manipulated facial regions, a dedicated attention rollout mechanism is embedded. A dedicated rollout mechanism for attention distribution of the transformer and heatmap generating and attention spatial localization are incorporated to improve the interpretability of the network. The framework comprises an end-to-end inference pipeline, such as image preprocessing, estimation of confidence scores, fake-real classification, generation of explainable visualization and storage of prediction history using an integrated database system. An experimental evaluation shows the system can effectively detect deepfakes while also providing accurate local information as justification for classification decisions, contributing to transparency, reliability and trust towards automated synthetic media detection systems.

View source

Similar papers

Aug 2026

Hybrid Deep Learning Model for Fake Image Detection with Advanced Face Object Segmentation Method

The fast development of deepfake technologies for image generation produces more and more realistic manipulated facial images that are harder to distinguish from the real content. In this paper, we propose a novel Hybrid CNN–Vision Transformer (HybCNNViT) framework for robust deepfake image detection by combining discr...

Saurabh Kumar Jain, Mohd. Akbar · 0 citations
Aug 2026

Image Forgery Detection Based on Fusion of Lightweight Deep Learning Models

A fusion-based lightweight deep learning framework for copy-move image forgery detection and localization that offers an efficient and practical solution for digital image authentication and is applicable to digital forensics, journalism, law enforcement, cyber security, and multimedia content verification.

K. Sumalini, K. B. Maruthiram · 0 citations
#artificial intelligence Preprint Sep 2026

Swin Meets EfficientNet: Lightweight Architectures for GAN-Based Face Forensics

Modern generative models, such as GANs, diffusion architectures, and autoregressive systems, now produce facial images that are nearly indistinguishable from authentic photographs. This capability makes detecting forged images increasingly difficult, raising serious concerns about identity theft, fraud, and misinformat...

S. Basu, Ashima Sood, Vijay Kumar et al. · 0 citations
Conference Aug 2026

Explainable Deep Learning Framework for Accurate Detection and Interpretation of Copy-Move Forgery in Digital Images

In the era of advanced digital image generation techniques such as copy-move forgery, the NDV is having increasing concerns regarding image authenticity particularly with the spread of digital images through social media, media and courts. This type of forgery, where a part of an image is replicated and pasted on the s...

Shaheena K. V., D. S · 0 citations
Open access Aug 2026

Attention-Guided Cross-Connected Filters Convolutional Neural Network with Surrogate-Based Interpretability for Image Splicing Forgery Detection

Background: Image splicing forgery detection is one of the most challenging problems in the field of image forensics as it involves identifying and localizing suspicious regions that are created by integrating contents from one or more different sources. The accurate detection and classification of splicing forgery sti...

Aruna Srinivasan, Surabhi Narayan, Aarnav Sandeep Deshmukh · 0 citations
Open access Aug 2026

Vision Transformer Based Digital Image Forgery Detection and Localization Using Global Contextual Feature Learning

The proposed Vision Transformer (ViT)-based framework provides a robust and scalable solution for modern digital image forensics and can be extended to hybrid transformer architectures and video forgery detection in future work.

G. Mary Pushpa, Dr. K. Sravan Adbhilash · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.