Skip to content
Open access

Multimodal Deepfake Detection for Digital Forensics: A Robust Audio-Visual Inconsistency Approach for Evidence Integrity

Aug 2026 · Journal of Science and Technology on Information security · 0 citations · 35 references

TL;DR

This research provides a resilient forensic layer for digital identity verification, ensuring evidence integrity in the GenAI era by proposing a novel Multimodal Deep Learning framework designed to detect high-fidelity Deepfakes by exploiting audio visual temporal inconsistencies.

Abstract

The rapid proliferation of Generative AI (GenAI) has democratized the creation of hyper-realistic multimedia forgeries, posing severe threats to electronic Know Your Customer (eKYC) systems and digital forensic investigations. While visual synthesis has reached near-perfection, maintaining precise synchronization between lip movements (visemes) and speech signals (phonemes) remains a formidable challenge. To address this, we propose a novel Multimodal Deep Learning framework designed to detect high-fidelity Deepfakes by exploiting audio visual temporal inconsistencies. Beyond traditional feature fusion, our architecture integrates a Contrastive Synchronization Loss with a Transformer based Cross-Modal Attention mechanism. This hybrid objective explicitly enforces intra-class compactness for authentic pairs while amplifying the distance for asynchronous forgeries. Extensive experiments on FaceForensics++, DFDC, and a custom Vietnamese dataset (Vn-eKYC-Aug) demonstrate that our model achieves state-of-the-art performance, maintaining high robustness against video compression and environmental noise, though operational efficacy remains sensitive to extreme low-light conditions and diverse regional dialects. This research provides a resilient forensic layer for digital identity verification, ensuring evidence integrity in the GenAI era.

Read PDF

Similar papers

Preprint Aug 2026

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet existing research either focuses on visual ma...

Yan-Qiu Li, Yang Xiao, Jisheng Bai et al. · 0 citations
Open access Aug 2026

Deepguardnet: A Resnet-Based Hybrid Framework for Intelligent Deepfake Image and Video Authentication

The evolution of sophisticated generative artificial intelligence has led to the rapid development of very realistic manipulated images and videos, posing substantial risks for digital trust, cyber security, and multimedia authenticity. Advanced Deepfake generation technologies result in the creation of believable forg...

Bella Inba Suganthi V, S. Jose · 0 citations
Open access 2026

Multimodal Fusion and Explainable Deep Learning for Synthetic Voice and Vishing Detection

A multimodal fusion framework for synthetic voice detection that combines handcrafted acoustic features with deep spectrogram representations to improve detection robustness and generalization and highlights the effectiveness of combining deep learning, multimodal feature fusion, and explainable artificial intelligence...

Mahima Bg, Pallavi Gb · 0 citations
Jul 2026

Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection

A detection framework is proposed that extracts per-video rPPG wave- forms via RhythmFormer and trains a suite of lightweight classifiers to distinguish real from synthesized physiologi- cal signals and shows that detec- tion difficulty is strongly method-dependent.

Othmane Harraq, Tamer Aldwairi · 0 citations
Review Open access Aug 2026

Advancements of Audio Unimodal Deep Faking Detection Technology

This paper systematically reviews the types of audio forgery, detection principles, authoritative datasets and evaluation indicators, compares and analyzes traditional detection methods with deep learning detection techniques, and points out the core challenges in generalization, robustness, etc. of current methods.

Ming-Lei Zhu · 0 citations
Preprint Aug 2026

Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow

A generative detector that localizes tampering by estimating the local restoration cost required to align a query image with authentic visual-text statistics, rather than by learning forgery-specific decision boundaries is proposed, and Sparse-Constraint Rectified Flow is introduced, a detector-oriented adaptation of F...

Jiangling Zhang, Shuxuan Gao, Zeyu Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.