Skip to content

Large Audio Language Models for Spoofing-Aware Speaker Verification

Jul 2026 · arXiv.org · Vol abs/2607.14753 · 0 citations · 52 references
Computer Science

TL;DR

This work systematically evaluates LALMs for SASV against conventional pipelines under zero-shot prompting, supervised adaptation, reasoning-oriented training, and reinforcement-learning-based optimization, and finds that competitive SASV performance can be achieved through several distinct routes.

Abstract

Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV). Existing defenses mainly address this threat through binary countermeasures (CMs) for deepfake detection or spoofing-aware speaker verification (SASV), where current systems are dominated by modular ASV-CM fusion and cascaded pipelines. Although large audio language models (LALMs) have shown promise on related audio tasks, including CM and ASV, their use for SASV remains unexplored, despite their capacity to produce natural-language rationales for auditing and robustness beyond discriminative predictions. This work systematically evaluates LALMs for SASV against conventional pipelines under zero-shot prompting, supervised adaptation, reasoning-oriented training, and reinforcement-learning-based optimization. Our results show that pretrained LALMs are near chance in the zero-shot setting, confirming that they are not natively suited to SASV, but that task-specific adaptation closes this gap. We further find that competitive SASV performance can be achieved through several distinct routes. These findings position LALMs as a promising and auditable foundation for unified SASV, while clarifying where conventional cascade systems still lead.

View source

Similar papers

Open access 2026

Multi-Prototype Variational Information Bottleneck for SSL-Based Speech Spoofing Detection

Rapid advances in text-to-speech (TTS) and voice conversion (VC) have substantially lowered the barrier to generating zero-shot spoofed speech, posing critical challenges for identity authentication, media forensics, and speech security. Current audio deepfake detection methods typically employ large-scale self-supervised learning (SSL) models as front-end feature extractors to improve detection performance and generalizability. However, SSL representations often encode information irrelevant to spoofing discrimination, such as speaker identity, semantic content, emotional prosody, and channel conditions. These redundant factors interfere with downstream detectors’ ability to model discriminative features of spoofed speech and impair generalization to unknown attacks and cross-domain scenarios. To address this limitation, we propose the Multi-Prototype Variational Information Bottleneck (MP-VIB), a method that compresses representations to retain spoofing-relevant information while discarding spoofing-irrelevant redundancy. Critically, MP-VIB assigns a single-prototype prior for real speech to capture a compact genuine distribution. For spoofed speech, a multi-prototype mixture prior explicitly models the latent structure arising from diverse spoofing types. This asymmetric design prevents the over-compression of spoofing-related information that results from standard single-prior formulations, thereby yielding more robust latent representations for deepfake detection. Experiments with a WavLM front-end demonstrate that MP-VIB achieves a pooled equal error rate (Pooled EER) of 16.07% across 14 evaluation sets. This represents a 14.4% relative improvement in Pooled EER over the standard variational information bottleneck. On ASVspoof 2019 LA, ASVspoof 2021 LA and ASVspoof 2021 DF, the method attains EERs of 0.33%, 2.47% and 4.13%, respectively. The code is available at https://github.com/Hench-Ho/MP-VIB

Heng-Chang Hou, Ruo-Hua Zhou, Qing-Sheng Yuan · 0 citations
Jul 2026

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition

ParaASR is introduced, an ASR system that leverages Multi-Token Prediction (MTP) to let a 4B LLM decoder emit multiple tokens per forward step and shows that decoder scaling, low-latency inference, and long-context transcription need not be competing goals when future-token proposals are anchored by the acoustic signal and guarded by autoregressive verification.

Qing-Jian Lin, Yuxin Li, Haoyang Zhang et al. · 2 citations
#machine learning Preprint Aug 2026

Machine Unlearning for Speech Question Answering in Large Audio-Language Models

Large Audio-Language Models (LALMs) have recently shown strong capabilities in speech understanding and question answering (QA), but they also inherit privacy risks from large-scale training data, including the unintended memorization of sensitive information. In this work, we study machine unlearning for speech QA in LALMs, a setting that is more challenging than prior work on text-based Large Language Models (LLMs) or Automatic Speech Recognition (ASR) due to the tight coupling between acoustic perception and factual knowledge. We present and evaluate multiple unlearning strategies, including gradient ascent, task arithmetic, and alignment-based fine-tuning methods that enforce safe refusal responses, to remove private knowledge while still preserving performance on core capabilities. Through extensive experiments on speech QA datasets, we show that these unlearning methods can reduce the privacy leakage rate by up to 80% while maintaining near-neutral performance on non-private speech QA and general speech understanding benchmarks.

Zhe Liu · 0 citations
Preprint Aug 2026

Evaluating Pre-trained Speech Encoders for Spontaneous Speech Detection and Out of Domain Synthetic Speech Generalisation in Indic Languages

Centroid analysis shows that out-of-domain generalisation is predicted by a training system's proximity to unseen TTS embeddings, not its distance from natural speech, a finding with direct implications for training data selection in real-world deepfake detectors.

Varun Rai, J. PavanKumar, Sujith Pulikodan et al. · 0 citations
Jul 2026

Teffic-Audio: Tell Fact from Fiction

Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoofing artifacts can be further shaped by variability in source speech, recording environments, and transmission channels. This variability makes robust generalization across heterogeneous conditions a central requirement for practical detection systems. This report presents Teffic-Audio, a general speech deepfake detection system designed for comprehensive evaluation environment. Teffic-Audio adopts a straightforward detector architecture consisting of a Conformer-based speech encoder, multi-head attentive statistics pooling, and a binary classifier. Rather than relying on additional architectural complexity, the system improves generalization through its training recipe, which integrates multi-source data, attack- and source-balanced sampling, and diverse audio augmentation. Trained only with open-source data, Teffic-Audio achieves a pooled EER of 1.454% on the 14 test sets of Speech-DF-Arena, outperforming all currently public systems on the leaderboard. It also obtains the lowest EER on five individual test sets and shows a favorable performance-complexity trade-off compared with larger leading systems. Overall, Teffic-Audio provides a strong and practical reference system for general speech deepfake detection.

Wan Lin, Li Wang, Jindong Wang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.