Jun 2026· arXiv.org· Vol abs/2606.06037· 1 citation· 42 references
Computer ScienceEngineering
TL;DR
SpeechJBB is introduced, an audio jailbreak dataset for benchmarking state-of-the-art LALMs across five European languages: English, French, German, Italian, and Spanish, as well as code-switched variants combining pairs of these languages, and the extent of safety weaknesses is probed.
Abstract
Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evaluated on monolingual, text-based harmful prompts. This leaves their generalizability under multilingual and spoken settings, particularly code-switched speech, largely underexplored. To address this gap, we introduce SpeechJBB, an audio jailbreak dataset for benchmarking state-of-the-art LALMs across five European languages: English, French, German, Italian, and Spanish, as well as code-switched variants combining pairs of these languages. The extent of safety weaknesses is further probed by introducing an augmented setting where phonologically plausible pseudo-words are inserted around safety-critical terms to simulate localized obfuscation. Across models, code-switched harmful audio yields substantially high jailbreak success rates (JSR), with non-English monolingual and non-English code-switched pairs exhibiting the highest attack success. Pseudo-word insertion monotonically reduces refusal as insertion density increases, even though models rarely attribute harmful meaning to the inserted tokens. Comprehension benchmarks show that these failures are not reducible to multilingual misunderstanding, as several models with strong ASR, spoken language understanding, and spoken reasoning performance are among the most vulnerable.
AfriSwitch is presented, a 61.36-hour human-transcribed benchmark of in-the-wild code-switched speech spanning 16 African languages and language varieties, released with switch-level English span tags, perutterance Code-Mixing Index (CMI), and switch-point counts.
MuLA-Bench exposes conditional failure patterns that a single long-context score does not capture, and evaluates ten audio-language models and conducts pooled diagnostics on a fixed eight-model cohort.
Ze-Yu Yang, Xin-Yu Zhang, Zi-Bo Bi et al.· 1 citation
This survey provides the first comprehensive review of research on code-switched LLM safety and robustness evaluation, and aims to guide researchers and practitioners in developing more robust evaluation frameworks and alignment techniques for real-world multilingual deployments.
Pulagam Naveen Kumar, Sowjanya Bojja, K. Anoosha et al.· International Journal for Re...· 0 citations
SEA-SpeechBench is introduced, the first large-scale multitask benchmark that evaluates speech understanding in 11 SEA languages through 97,194 samples across 99 evaluation sets and 597 hours of curated audio data, exposing critical model limitations and underscore the need for inclusive model development.
Jingyi Liao, Wenyu Zhang, Zhuo-Han Liu et al.· 1 citation
This work presents and evaluates multiple unlearning strategies, including gradient ascent, task arithmetic, and alignment-based fine-tuning methods that enforce safe refusal responses, to remove private knowledge while still preserving performance on core capabilities.
The central finding is that aggregate WER hides code switching behavior, and the best system by WER (an ASR model) is statistically indistinguishable from a leading audio LM on WER, yet the audio LM is significantly better on every switch localized metric.
C. Okocha, Christan Grant· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.