Skip to content

ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models

Jun 2026 · arXiv.org · Vol abs/2606.10581 · 2 citations · 45 references
Computer Science Engineering

TL;DR

ParaBridge is proposed, an on-policy self-distillation method that turns a brittle inference-time scaffold into stable model behavior and generalizes to unseen paralinguistic cues, transfers from safety-oriented training to empathy-oriented dialogue, and works on a different SLM backbone.

Abstract

Speech carries more information than just words: a child's voice, a fearful tone, or a noisy background should all lead a sufficiently competent spoken-dialogue assistant to different replies. Current Speech Language Models (SLMs) can recognize such paralinguistic cues but often ignore them in open-ended dialogue. We observe that a simple paralinguistic instruction scaffold at the inference stage narrows this perception-behavior gap, suggesting that the relevant cues are already latent in the model. Such scaffolds, however, remain brittle under multi-turn context and competing instructions. Therefore, we propose \textbf{ParaBridge}, an on-policy self-distillation method that turns a brittle inference-time scaffold into stable model behavior. During training, the scaffold serves only as a temporary privileged view; the scaffold-free model rolls out its own response, while the scaffolded view supplies dense, full-vocabulary next-token targets along its trajectory. This supervision teaches when non-lexical cues should affect the reply without the need for curated dialogues, human labels, or external reward models. On Qwen3-Omni-thinking, ParaBridge raises scaffold-free VoxSafeBench SAR from $14.6\%$ to $40.3\%$ and improves EchoMind average rating from $3.27$ to $3.92$. It also preserves general ability, with MMAU-Pro, VoiceBench, and GPQA all within $0.4$ points of the original model. Beyond the training distribution, ParaBridge generalizes to unseen paralinguistic cues, transfers from safety-oriented training to empathy-oriented dialogue, and works on a different SLM backbone.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Thinking in Depth, Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue

Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguistic perception and make acoustic cues more explicit in replies, yet does not ensure their effective use in response planning. We call this mismatch the perception-reasonin...

Sheng-Bo Cai, Yu-Xiang Wang, Jing-Ran Xie et al. · 0 citations
#natural language process... Preprint Sep 2026

Lost with a Map: Conversational State and Behavioral Reliability in Language Models

Task-oriented dialogue requires maintaining and updating information across turns, yet language models expose no explicit belief-state object. We study how conversational state is represented, updated, and used inside eight instruction-tuned language models from four families on MultiWOZ and SGD. Structure and values s...

Atahan Dokme, Larry Heck · 0 citations
#artificial intelligence Preprint Sep 2026

VoiceLongMemEval: Do Assistants Remember How You Sounded?

VoiceLongMemEval (VLME) benchmark, where every answer depends on paralinguistic metadata attached to conversational turns, which is otherwise unrecoverable from the words alone, is presented.

Ramit Pahwa, Parivesh Priye, Apoorva Beedu · 1 citation
Preprint Aug 2026

When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue

A scalable framework that identifies text-biased surface interpretations and converts disagreement regions into conflict QA examples and develops an agentic-style reasoning framework that converts speech into an Audio Twin, a text-readable representation of localized acoustic cues that exposes acoustic evidence to the...

Yen-Ju Lu, Yu-Zhe Wang, Yao-Han Guan et al. · 1 citation
#natural language process... Preprint Sep 2026

Spoken Language Models that Think Aloud

While Chain-of-Thought (CoT) reasoning has improved the capability of language models, directly applying it to Spoken Language Models (SLMs) may introduce long silent intervals under the serial"think-then-speak"paradigm, disrupting real-time spoken interaction. To address this issue, we propose an asynchronous think-al...

Junyi Ao, Kainan Peng, Mingbo Ma et al. · 0 citations
Preprint Aug 2026

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

Hear2Act is introduced, a unified evaluation protocol for text and spoken assistants with 480 persona-grounded scenarios, hidden user concerns, and objectively verifiable outcomes that show that prosody matters when lexical evidence is insufficient, and that audio-capable LLMs can recover information from speech but do...

Xin-Yi Liu, H. Nayyeri, Dilek Hakkani-Tur et al. · 3 citations · ⚡1

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.