Skip to content

ParA-LLM: A Unified Approach to Paralinguistic and Acoustic Speech Understanding

Sep 2026 · 1 citation · 33 references
Computer Science Engineering

TL;DR

This work develops ParA-LLM, a framework of 22 paralinguistic characteristics that surpasses state-of-the-art Audio LLMs like GPT-4o-Audio by 7.5% on ParA-Bench and releases ParA-Bench, a benchmark of 6,000 multiple-choice questions across speaker-speech, acoustic, and mixed categories.

Abstract

Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralinguistic characteristics and create a dataset of over 1.2M Audio-QA pairs. We develop ParA-LLM, trained with a two-stage curriculum: first on single-attribute questions to build foundational knowledge, then on multi-attribute questions for joint reasoning over speaker and acoustic characteristics. We also release ParA-Bench, a benchmark of 6,000 multiple-choice questions across speaker-speech, acoustic, and mixed categories, where frontier models like GPT-4o-Audio achieve only 36% accuracy. ParA-LLM surpasses state-of-the-art Audio LLMs like GPT-4o-Audio by 7.5% on ParA-Bench, with additional gains of 1.13% on MMAU-Pro Speech and 7.49% on MMAR Speech.

View source

Similar papers

Preprint Aug 2026

CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.

Nhan Phan, Ilona Lähteenmäki, Anna von Zansen et al. · 0 citations
Preprint Sep 2026

Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition

Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the complexity of natural conversation, which involves overlapping speech, spontaneous turn-taking, unscripted vocabulary and variable acoustic conditions. To shift the field...

Rishabh Jain, A. Papadopoulos, Zhao-Feng Lin et al. · 0 citations
#natural language process... Preprint Sep 2026

Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding

Mizar, a 159.3M-parameter ALM, is introduced, a recipe that brings together architecture, data, and three-stage training to build Mizar, a 159.3M-parameter ALM that surpasses the previous best-performing ALM below 200M parameters on all three benchmarks.

Kai-Yang Li, Shaobo Han, Yue Tian et al. · 0 citations
Preprint Aug 2026

FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation

FireRedAudio is introduced, a general-purpose audio language model with a shared 9B-parameter LLM that achieves competitive or leading performance in audio understanding and multilingual ASR, strong content accuracy and speaker preservation in zero-shot TTS, leading instruction following in Instruct TTS, and substantia...

Fei-Yu Shen, Fenglong Xie, Junjie Li et al. · 3 citations · ⚡1
Preprint Oct 2026

Augmenting Large Audio Language Models with Low-Level Acoustic Features for Dysarthric Speech Detection

Automatic dysarthric speech detection approaches can support traditional clinical diagnosis, which relies on costly and time-consuming evaluation by a speech and language pathologist. Existing automatic approaches predominantly rely on deep learning (DL). More recently, Large Audio Language Models (LALMs) have emerged...

Mahdi Amiri, H. O. Shahreza, Pascal Frossard et al. · 0 citations
#natural language process... Preprint Sep 2026

Qwen-Audio-3.0-ASR Technical Report

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility r...

Chuan-Meng Bian, Da-Ren Chen, Pei-Xin Chen et al. · 2 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.