Audio Deepfake Detection Using Dual-Branch CNN with Shared Weights
Abstract
The detection of audio deepfakes has emerged as a significant problem in the field of voice biometrics systems, aiming to distinguish real human voices from those generated by Artificial Intelligence (AI). With synthetic voice becoming increasingly high-quality, it is more likely that such a voice will be abused for illicit purposes like identity theft and impersonation. The dual-branch CNN with shared weights architecture presented here is augmented with self-attention modules to detect audio deepfakes with greater efficiency. Convolutional operations and dual branches are used to extract complex characteristics from raw audio signals in our module in order to directly compare the unprocessed original audio with the modified audio. Afterward, residual connections improve network performance. Designed alongside these fundamental layers, self-attention modules are trained in a layered manner to detect multi-headed attention within audio frames. This feature helps the network distinguish between original and modified audio and improves feature extraction compared with the standard method. A range of audio modifications have been analyzed to assess the effectiveness of the method, and comprehensive testing across all possible audio manipulation situations has been conducted on the Controlled Singing Voice Deepfake Detection Challenge (CtrSVDD) dataset to assess its resilience. Both deep learning (DL) and machine learning (ML) models were outperformed by the proposed dual-branch CNN with shared weights. With an accuracy of 97.26%, precision of 99%, recall of 99.27%, and an F1 score of 98.88%, this model has achieved a remarkable performance.