Skip to content
Conference

Multilingual Deepfake Audio Detection Using a Hybrid CNN-Bi LSTM with Multi-Head Attention Architecture

Aug 2026 · International Conference on Computing Communication Control and automation · pp. 1-11 · 0 citations · 16 references

Abstract

Deepfake audio is a critical problem affecting information integrity, offering synthesized speech capable of fooling both humans and machines. This research proposes an approach towards deepfake audio detection across multiple languages, namely English, Hindi, and Marathi. It focuses on a novel approach, Hybrid CNN-BiLSTM with Multi-Head Attention Model, addressing the current shortcomings in the literature. Its model architecture consists of CNN layers to extract spatial features, BiLSTM layers to capture time dependency, and multi-head attention layers for context modeling, trained on Mel Frequency Cepstral Coefficients combined with delta and delta-delta coefficients. The dataset includes LibriSpeech for English language and OpenSLR for Indic languages with synthetic fakes created by augmentations like adding noise and changing the pitch. Robust preprocessing allows handling different lengths up to 120 seconds along with z-score normalization. The system is designed as a Flask web application, allowing users to record their voice and upload it, detect fake content, create a mimicry, and visualize spectrogram. The results show high accuracy (95.2%, 96.1% for English, 93.8% for Hindi, 92.5% for Marathi) with 4.8% equal error rate and high robustness against noise (85%+). Ablation study shows importance of each layer, filling research gaps in multilingual capabilities and hybrid model

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.