Advancements of Audio Unimodal Deep Faking Detection Technology
Abstract
In recent years, the generative artificial intelligence technology has undergone rapid iterations, leading to the widespread abuse of audio deep forgery methods such as speech synthesis and speech conversion, which seriously threaten information security and social credibility. Among various countermeasures, audio single-modal deep forgery detection has advantages such as lightweight and strong real-time performance, and is a key technology for security protection in pure audio scenarios, with significant research and application value. This paper systematically reviews the types of audio forgery, detection principles, authoritative datasets and evaluation indicators, compares and analyzes traditional detection methods with deep learning detection techniques, focuses on elaborating the characteristics and applicable scenarios of four mainstream detection models, and points out the core challenges in generalization, robustness, etc. of current methods. Finally, it looks forward to the development trends of this field towards generalization, high robustness, lightweight and forgery traceability, which can provide references for related research.