Multilingual Deepfake Audio Detection Using a Hybrid CNN-Bi LSTM with Multi-Head Attention Architecture
Abstract
Deepfake audio is a critical problem affecting information integrity, offering synthesized speech capable of fooling both humans and machines. This research proposes an approach towards deepfake audio detection across multiple languages, namely English, Hindi, and Marathi. It focuses on a novel approach, Hybrid CNN-BiLSTM with Multi-Head Attention Model, addressing the current shortcomings in the literature. Its model architecture consists of CNN layers to extract spatial features, BiLSTM layers to capture time dependency, and multi-head attention layers for context modeling, trained on Mel Frequency Cepstral Coefficients combined with delta and delta-delta coefficients. The dataset includes LibriSpeech for English language and OpenSLR for Indic languages with synthetic fakes created by augmentations like adding noise and changing the pitch. Robust preprocessing allows handling different lengths up to 120 seconds along with z-score normalization. The system is designed as a Flask web application, allowing users to record their voice and upload it, detect fake content, create a mimicry, and visualize spectrogram. The results show high accuracy (95.2%, 96.1% for English, 93.8% for Hindi, 92.5% for Marathi) with 4.8% equal error rate and high robustness against noise (85%+). Ablation study shows importance of each layer, filling research gaps in multilingual capabilities and hybrid model