Defending against deepfake speech attacks via feature dropout and dropin regularization in large audio language models
Abstract
Deepfakes have raised widespread concern owing to their threats to privacy, security, and societal trust, driving growing research interest in effective detection methods. Audio deepfake detection has moved from handcrafted-feature classifiers to end-to-end deep learning architectures and, more recently, to self-supervised speech foundation models and general-purpose speech recognizers such as Whisper. The use of instruction-following speech large language models (LLMs), of which Speech Audio Language Music Open Neural Network (SALMONN) is a representative example, remains comparatively less explored, and forms the setting of this study. We argue that large language models (LLMs), with their powerful representational and information retrieval capabilities, hold significant potential for this task—provided that relevant audio features are carefully selected and prioritized. To this end, we employ a speech LLM, specifically SALMONN, for deepfake audio detection, incorporating a two-stage training strategy in which Stage 1 injects acoustic-feature awareness into the Low-Rank Adaptation (LoRA) adapters via an auxiliary feature-prediction pretext task (referred to as feature dropin), and Stage 2 performs windowed random retention on the encoder token sequence (referred to as feature dropout) before the final real/fake decision. Our approach achieves an absolute accuracy improvement of 0.188 on the FakeOrReal benchmark and attains state-of-the-art performance on the Automatic Speaker Verification Spoofing and Countermeasures (ASVspoof) 2019 LA subset with an Equal Error Rate (EER) of 0.00496, representing a 33.0% relative improvement over prior methods. Furthermore, the model transfers non-trivially to an out-of-domain artificial intelligence (AI)-generated music dataset (M6), reaching accuracy comparable to a ResNet18 baseline reported in prior study, which we interpret as evidence of transferable audio representations rather than broad cross-domain robustness.