Skip to content

Mechanistic Insights into Context Failures of Audio-Language Models for Impaired Speech

· 0 citations · 18 references

TL;DR

This work introduces Self-Contrastive Residual Alignment (SCRA), a one-pass inference-time edit that learns a low-rank linear corrector for the prompt-induced residual shift and slightly but significantly improves over the audio-only baseline on difficult samples.

View source

Similar papers

Jul 2026

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

Overall, matched-text speech delivery should be treated as a first-class factor in Audio LLM safety evaluation by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary.

Jiachen Qian, Junyu Li · 0 citations
#artificial intelligence Preprint Sep 2026

Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

A mechanistic analysis of paralinguistic information in four open source models using the Expresso dataset with controlled speaking styles identifies a gap between what models encode and what they use, highlighting a key limitation in current audio language models.

Bhuvan Koduru, Dareen Alharthi, Rita Singh et al. · 2 citations · ⚡1
Preprint Aug 2026

Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec

A unified framework covering six popular LM-based generative SE modeling paradigms based on discrete/continuous latent NAC features, and a fine-tuning strategy with auxiliary losses on reconstructed speech to improve both intrusive and non-intrusive metrics is presented.

Yihui Fu, Zhengyang Li, Tim Fingscheidt · 0 citations
Jul 2026

ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions

As large audio language models (LALMs) advance, robust evaluation frameworks have become essential. In this context, Spanish speech understanding under realistic acoustic conditions has received particularly little attention. We introduce ESCUCHA, the first Spanish speech understanding benchmark designed to evaluate LA...

Fernando López, A. Ayala, Guillermo Segovia et al. · 0 citations
Preprint Sep 2026

AudioICL-Bench: A Benchmark for Large Audio Language Model In-Context Learning

AudioICL-Bench is introduced, a diagnostic benchmark whose per-episode rules are resampled so that no correct answer is recoverable from prior knowledge, and its nine tasks are organized along two axes that separate what must be learned from demonstrations from what must be perceived in the signal, enabling failures to...

Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.