This work develops ParA-LLM, a framework of 22 paralinguistic characteristics that surpasses state-of-the-art Audio LLMs like GPT-4o-Audio by 7.5% on ParA-Bench and releases ParA-Bench, a benchmark of 6,000 multiple-choice questions across speaker-speech, acoustic, and mixed categories.
Abstract
Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralinguistic characteristics and create a dataset of over 1.2M Audio-QA pairs. We develop ParA-LLM, trained with a two-stage curriculum: first on single-attribute questions to build foundational knowledge, then on multi-attribute questions for joint reasoning over speaker and acoustic characteristics. We also release ParA-Bench, a benchmark of 6,000 multiple-choice questions across speaker-speech, acoustic, and mixed categories, where frontier models like GPT-4o-Audio achieve only 36% accuracy. ParA-LLM surpasses state-of-the-art Audio LLMs like GPT-4o-Audio by 7.5% on ParA-Bench, with additional gains of 1.13% on MMAU-Pro Speech and 7.49% on MMAR Speech.
CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.
Nhan Phan, Ilona Lähteenmäki, Anna von Zansen et al.· 0 citations
Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the complexity of natural conversation, which involves overlapping speech, spontaneous turn-taking, unscripted vocabulary and variable acoustic conditions. To shift the field...
Rishabh Jain, A. Papadopoulos, Zhao-Feng Lin et al.· 0 citations
Mizar, a 159.3M-parameter ALM, is introduced, a recipe that brings together architecture, data, and three-stage training to build Mizar, a 159.3M-parameter ALM that surpasses the previous best-performing ALM below 200M parameters on all three benchmarks.
Kai-Yang Li, Shaobo Han, Yue Tian et al.· 0 citations
FireRedAudio is introduced, a general-purpose audio language model with a shared 9B-parameter LLM that achieves competitive or leading performance in audio understanding and multilingual ASR, strong content accuracy and speaker preservation in zero-shot TTS, leading instruction following in Instruct TTS, and substantia...
Fei-Yu Shen, Fenglong Xie, Junjie Li et al.· 3 citations· ⚡1
Automatic dysarthric speech detection approaches can support traditional clinical diagnosis, which relies on costly and time-consuming evaluation by a speech and language pathologist. Existing automatic approaches predominantly rely on deep learning (DL). More recently, Large Audio Language Models (LALMs) have emerged...
Mahdi Amiri, H. O. Shahreza, Pascal Frossard et al.· 0 citations
In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility r...
Chuan-Meng Bian, Da-Ren Chen, Pei-Xin Chen et al.· 2 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 29, 2026
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.