Skip to content
Open access

Multimodal Fusion and Explainable Deep Learning for Synthetic Voice and Vishing Detection

2026 · International journal of research and scientific innovation · Vol 13, pp. 218-229 · 0 citations

TL;DR

A multimodal fusion framework for synthetic voice detection that combines handcrafted acoustic features with deep spectrogram representations to improve detection robustness and generalization and highlights the effectiveness of combining deep learning, multimodal feature fusion, and explainable artificial intelligence.

Abstract

Synthetic speech generation and voice-cloning technologies have achieved unprecedented levels of realism, enabling numerous applications in accessibility, virtual assistants, and media production. However, these advancements also introduce significant risks, including identity fraud, impersonation attacks, misinformation, and security breaches. This paper proposes a multimodal fusion framework for synthetic voice detection that combines handcrafted acoustic features with deep spectrogram representations to improve detection robustness and generalization. The proposed architecture employs a Convolutional Neural Network–Bidirectional Long ShortTerm Memory (CNN-BiLSTM) network to capture both spectral artifacts and temporal inconsistencies characteristic of AI-generated speech. To enhance transparency and interpretability, an explainability module incorporating attention visualization and feature attribution techniques is integrated into the detection pipeline. Furthermore, the framework is deployed through a real-time inference interface, demonstrating its practical applicability in cybersecurity, digital forensics, and media authentication scenarios. The findings highlight the effectiveness of combining deep learning, multimodal feature fusion, and explainable artificial intelligence to address the growing challenge of synthetic speech detection.

Read PDF

Similar papers

Preprint Sep 2026

Disentangled Global-Local Feature Learning with E-Branchformer for Audio Deepfake Detection

The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervis...

Phuong Dat, Học Thủ, T. Nguyễn et al. · 0 citations
Open access Aug 2026

Multimodal Deepfake Detection for Digital Forensics: A Robust Audio-Visual Inconsistency Approach for Evidence Integrity

This research provides a resilient forensic layer for digital identity verification, ensuring evidence integrity in the GenAI era by proposing a novel Multimodal Deep Learning framework designed to detect high-fidelity Deepfakes by exploiting audio visual temporal inconsistencies.

H. Truong, D. T. Luong, Tuan Tran · 0 citations
Open access Sep 2026

Audio Deepfake Detection Using Dual-Branch CNN with Shared Weights

The detection of audio deepfakes has emerged as a significant problem in the field of voice biometrics systems, aiming to distinguish real human voices from those generated by Artificial Intelligence (AI). With synthetic voice becoming increasingly high-quality, it is more likely that such a voice will be abused for il...

Zainab A. Jawad, Ahmed J. Obaid · 0 citations
Open access Aug 2026

EMBNet: Multi-scale feature learning with efficient channel attention for deepfake speech detection.

With the rapid advancement of generative artificial intelligence, deepfake speech has emerged as a significant threat to digital audio authenticity, posing challenges for forensic analysis and legal applications. In this study, we propose EMBNet, a task-oriented deepfake speech detection framework that integrates effic...

Haitao Yang, Fen Li, Xin Cai et al. · 0 citations
Conference Aug 2026

Multilingual Deepfake Audio Detection Using a Hybrid CNN-Bi LSTM with Multi-Head Attention Architecture

Deepfake audio is a critical problem affecting information integrity, offering synthesized speech capable of fooling both humans and machines. This research proposes an approach towards deepfake audio detection across multiple languages, namely English, Hindi, and Marathi. It focuses on a novel approach, Hybrid CNN-BiL...

Priya Yadav, S. B. Patil · 0 citations
Open access Aug 2026

A Hybrid MFCC–WavLM Feature Fusion for Audio Deepfake Detection

Recent advances in generative artificial intelligence have enabled highly realistic speech synthesis using text-to-speech (TTS), voice conversion (VC), and neural voice cloning techniques, posing significant security threats to Automatic Speaker Verification (ASV) systems. Conventional handcrafted features such as Mel-...

K. S. Kumar, Madduluri Suneetha, K. R. Anudeep Laxmi Kanth et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.