Skip to content

ER-EDF: A Psychology-Grounded Emotion Regulation Framework for Speech Empathetic Dialogue Generation in Large Audio-Language Models

Sep 2026 · 0 citations · 48 references
Computer Science

TL;DR

Experiments show that ER-EDF consistently improves empathetic response quality in both automatic and human evaluations, highlighting the importance of jointly modeling emotion perception and regulation in spoken empathetic dialogue systems, paving a new direction for psychologically grounded empathetic AI.

Abstract

Empathetic response generation in spoken dialogue systems requires both accurate emotion perception and appropriate emotion regulation. Grounded in psychological theories such as the Perception-Action Model and emotion regulation theory, effective empathy depends not only on inferring a user's affective state but also on regulating how it is expressed in responses. However, recent large audio-language models (LALMs) largely treat emotion as a direct conditioning signal, lacking explicit regulatory mechanisms, which often leads to affect mirroring rather than calibrated support. We propose ER-EDF, a psychology-grounded framework that explicitly decouples emotion perception and emotion regulation in LALMs. Perception tracks the user's emotional state, while regulation determines how this state should guide empathetic response generation. The framework is model-agnostic and integrates seamlessly into existing LALMs. We further construct a spoken empathetic dialogue dataset and introduce empathy-aware evaluation metrics beyond lexical matching. Experiments across five LALMs and two datasets show that ER-EDF consistently improves empathetic response quality in both automatic and human evaluations, highlighting the importance of jointly modeling emotion perception and regulation in spoken empathetic dialogue systems, paving a new direction for psychologically grounded empathetic AI.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Exposing Weaknesses in Emotion Recognition in Conversations

An LLM-as-Judge framework is introduced that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision, suggesting that standard single-label evaluation is therefore insufficient.

Amir Ben Khalifa, Fanny Bezancon, B. Abdulrazak et al. · 0 citations
Preprint Oct 2026

EchoChat: Structured Cognitive Reasoning in Empathetic Spoken Dialogue

Empathetic spoken dialogue is a sophisticated cognitive process that requires not only recognizing emotions but also inferring a user's latent mental states to provide appropriate support. However, current SpeechLLMs often treat empathy as a direct input-to-response mapping, leading to"superficially warm"but emotionall...

Ding-Dong Wang, Shu-Jie Liu, Ya-Yue Deng et al. · 0 citations
Book Open access Oct 2026

Do Emotional Cues Matter? Exploring Support Strategy Selection in LLM-Based Supportive Conversations

Multimodal human–AI systems increasingly accept both text and speech, yet speech carries paralinguistic emotional cues that text does not. While prior work has evaluated the quality of large language model (LLM) responses, little is known about how vocal emotional cues reshape the support strategies an LLM selects and...

Wei-Yi Tian, Safak Dogan, Jie Meng · 0 citations
#artificial intelligence Preprint Sep 2026

EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision

Empathetic response generation requires models to decide not only what to say, but also how to respond to the previous speaker's affective situation, so multi-annotator emoji distributions are used as weak affective--attitudinal evidence to induce a latent control space that operationally approximates listener stance.

Zi-Yuan Jin, Yuxuan Ge, Zheng Tian · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.