Skip to content
Open access

Web-Integrated Deepfake Detection of Manipulated Political Audio Using CNN and Wav2Vec2 Transformer Models

Sep 2026 · Applied Sciences · 0 citations · 19 references

Abstract

Deepfake audio poses increasing risks to political forensics by enabling scalable misinformation, speaker impersonation, and identity spoofing. In response to this challenge, this paper presents a web-integrated deepfake detection framework for manipulated political speech together with a politically grounded multilingual dataset designed to support reproducible evaluation. The STS Political Deepfake Dataset contains authentic and manipulated political speech samples generated through retrieval-based voice conversion (RVC), covering multiple languages, speakers, and recording contexts. Convolutional neural network baselines based on MFCC and Mel-spectrogram representations were evaluated against an enhanced transformer-based architecture built on Wav2Vec 2.0. The proposed transformer pipeline incorporates weighted layer fusion, partial backbone freezing, and waveform-level augmentation to improve representation quality and robustness. Experimental results show that the transformer-based approach consistently outperforms the CNN baselines under the main held-out evaluation setting. Beyond the internal benchmark, the study also includes cross-dataset evaluation on public deepfake audio corpora and comparison against external publicly available models, providing additional evidence on generalization and revealing important differences between modern and legacy generative speech artifacts. Finally, the selected Wav2Vec2-based detector was deployed through Hugging Face inference and a Django-based web application, and its practical use in political audio verification scenarios was demonstrated.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.