Skip to content

Latent-IM: Latent Interaction Management for Speech LLMs

Jul 2026 · arXiv.org · Vol abs/2607.26928 · 0 citations · 47 references
Computer Science

TL;DR

Latent-IM is introduced, an internal dialogue-management framework that provides a general interface for choosing and deploying conversational moves under different objectives and is used to reproduce human move choices, improving average end-to-end move accuracy by 12.5 points over the unsteered backbone while performing comparably to fine-tuning.

Abstract

Classical spoken dialogue systems often separated dialogue management from response realization: a policy selected the next dialogue action, and a generation component expressed that action. As dialogue systems shift toward LLMs, this decomposition has largely disappeared into the model's hidden representations. We ask whether an LLM-internal analogue of state estimation and action control can be recovered for conversational moves such as acknowledging, checking, querying, explaining, and replying. We formulate move control as two coupled problems: selection, predicting the appropriate next move from the dialogue context, and realization, causally producing a chosen move at generation time. We introduce Latent-IM, an internal dialogue-management framework that provides a general interface for choosing and deploying conversational moves under different objectives. Here, we use this control to reproduce human move choices, improving average end-to-end move accuracy by 12.5 points over the unsteered backbone while performing comparably to fine-tuning.

View source

Similar papers

#natural language process... Preprint Sep 2026

PragAlign: Feedback-Guided Pragmatic Alignment for Controlled Synthetic Dialogue Generation

The results support PragAlign as a quality-control framework for improving evaluator-defined communicative constraint satisfaction, while showing that affective realization and independent human-perceived quality remain open challenges.

Smitha Muthya Sudheendra, Jaideep Srivastava · 0 citations
#artificial intelligence Preprint Sep 2026

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon. This leaves open whether spoken dialogue models can sustain diverse roles over extended interactions, especially beyond predefined fictional characters....

Yu-Qi Wang, Feng-Yuan Liu, Hao-Chen Luo et al. · 0 citations
#natural language process... Preprint Sep 2026

AVERT: Audio-Verified Adjudication for Spoken Dialogue State Tracking

Spoken dialogue state tracking recovers slot-value pairs from speech, where ASR errors concentrate in entity values and persist across turns, making it both a generation and an editing problem. A strong per-turn text editor corrects much of this but, operating on the transcript alone, leaves three recoverable errors: a...

C. Lee, H. Pfister · 0 citations
Book Open access Sep 2026

Conversational Style in Open Domain Dialogue Systems: What Makes a Response Sound Natural

Conversational style in this setting is best understood as a pragmatic orientation toward the user rather than toward the system’s own content, and that the system must be mixed-initiative to manifest a conversational style.

Vrindavan Harrison, M. Walker · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.