Jul 2026· International Journal for Research in Applied Science and Engineering Technology· Vol 14, pp. 1641-1647· 0 citations
TL;DR
BlinkNet, an explainable deepfake-detection framework that examines the spatial appearance and temporal kinematics of eye blinks, indicates that ocular dynamics can complement spatial evidence while improving efficiency and interpretability.
Abstract
The increasing realism of synthetic facial videos has reduced the reliability of detectors that depend only on visible
artifacts in isolated frames. This paper presents BlinkNet, an explainable deepfake-detection framework that examines the
spatial appearance and temporal kinematics of eye blinks. The system detects a face, localizes 68 facial landmarks, extracts
normalized ocular crops, and computes the Eye Aspect Ratio (EAR) for each frame. Overlapping sequences of 20 frames are
processed by a dual-stream Temporal-Spatial Physiological Blink Anomaly Network (TPBAN). A lightweight MobileNetV2
encoder models local visual inconsistencies, while a bidirectional gated recurrent unit models the forward and backward
dynamics of eyelid motion. Temporal attention assigns a relevance weight to every frame and supports frame-level anomaly
visualization. Training uses a multi-task objective for authenticity classification and blink-phase recognition, together with a
class-weighted binary cross-entropy term to address the imbalance between genuine and manipulated sequences. On the
FaceForensics++ c23 test partition, BlinkNet obtained 83.16% accuracy, 87.31% ROC-AUC, 96.02% average precision, and a
21.08% equal error rate. The manipulated class achieved 0.90 precision and 0.88 recall. The implementation processed video at
approximately 52 frames per second on a consumer laptop GPU and was integrated into a Flask-based forensic dashboard.
These results indicate that ocular dynamics can complement spatial evidence while improving efficiency and interpretability.
This work proposes a spatial-temporal model with two key components: one targeting artifacts within individual frames and the other analyzing inconsistencies across consecutive frames, both leveraging a bidirectional long short-term memory (Bi-LSTM) mechanism.
Ahmed Tammam, H. Abdelkader, Amira Abdelatey et al.· Signal, Image and Video Proc...· 0 citations
A detection framework is proposed that extracts per-video rPPG wave- forms via RhythmFormer and trains a suite of lightweight classifiers to distinguish real from synthesized physiologi- cal signals and shows that detec- tion difficulty is strongly method-dependent.
DeepFakes pose significant risks to digital security by enabling realistic facial manipulations that can evade conventional visual inspection. This study presents an attention-enhanced EfficientNet-B7 framework with a Custom Soft Spatial Attention (CSSA) module designed to localize manipulation-sensitive facial regions...
Kislay Raj, Raja Vavekanand, Aditya Singh· Journal of Computer Virology...· 0 citations
The rapid development of artificial intelligence has led to the emergence of deepfakes, which pose serious threats to information security and public trust in digital media. This study develops a facial deepfake detection system that integrates YOLOv11 for face detection and the Xception architecture for classifying re...
Fachril Akbar, N. Nurdin, Kurniawati Kurniawati· JOURNAL OF APPLIED INFORMATI...· 0 citations
A novel method for video deepfake detection that assimilates the Pelican Optimization algorithm with a DL model jointly named as Pelican Attention Stacked Bidirectional Long-Short Term Memory (PAttSBiL), aimed at improving recognition accuracy and efficacy is presented.
D. R. Agrawal, Farha Haneef· Multimedia tools and applica...· 0 citations
The existing emotion detection systems are either audio or video centric. While such approaches have proven effective in controlled environments like a lab or a studio, these systems fail to perform under real-world conditions such as low light, audio interference, mispronunciation, and other environmental factors. The...