Listen-to-Reason (L2R), an interpretable-by-design pipeline that passes audio to the LLM through an explicit, human-readable tree that outperforms all LALMs the authors compare against on SAKURA and trails them by 6-12 points on MMAU and MMAR, is proposed.
Pooneh Mousavi, M. Ravanelli, Cem Subakan· 0 citations
Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and q...
Pooneh Mousavi, Amir Ivry, M. Ravanelli et al.· 0 citations
Large Audio Language Models (LALMs) perform well on complex question answering but often show a reasoning gap, where explicit Chain-of-Thought (CoT) reduces accuracy compared to direct answers. We hypothesize that long reasoning sequences shift attention away from the audio input. To address this, we propose SPARE (Sem...
Francesco Bonzi, Pooneh Mousavi, Cem Subakan et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.