Empathetic response generation requires models to decide not only what to say, but also how to respond to the previous speaker's affective situation, so multi-annotator emoji distributions are used as weak affective--attitudinal evidence to induce a latent control space that operationally approximates listener stance.
Abstract
Empathetic response generation requires models to decide not only what to say, but also how to respond to the previous speaker's affective situation. We formulate this as response-side affective-orientation control and use multi-annotator emoji distributions as weak affective--attitudinal evidence, rather than as output symbols or gold labels, to induce a latent control space that operationally approximates listener stance. We construct EmojiDialogue, an utterance-level extension of EmpatheticDialogues with emoji votes and confidence scores, and propose EmoStance, which models source-side affective expression, predicts a soft response-side orientation from dialogue context and speaker roles, and steers a frozen instruction-tuned LLM through continuous prefix embeddings. In blind pairwise evaluation with 20 annotators and 800 judgments, EmoStance achieves a 62.2% decisive win rate, with the clearest gains in contextual specificity and perceived responsiveness, while remaining complementary to external-knowledge methods. Code, annotation metadata, and reconstruction scripts are available in our GitHub repository: https://github.com/18277390221/EmoStance.
EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations, and EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning and Group Relative Policy Optimization, are introduced, establishing a foundational framework for advan...
Junyu Wang, Si-Yuan Zhang, Peiyuan Jiang et al.· 1 citation
Experiments show that ER-EDF consistently improves empathetic response quality in both automatic and human evaluations, highlighting the importance of jointly modeling emotion perception and regulation in spoken empathetic dialogue systems, paving a new direction for psychologically grounded empathetic AI.
Hong-Yu Jin, Wen-Da Zhang, Run-Qiu Fei et al.· 0 citations
Multimodal human–AI systems increasingly accept both text and speech, yet speech carries paralinguistic emotional cues that text does not. While prior work has evaluated the quality of large language model (LLM) responses, little is known about how vocal emotional cues reshape the support strategies an LLM selects and...
Wei-Yi Tian, Safak Dogan, Jie Meng· Companion Publication of the...· 0 citations
As conversational companions, large language models (LLMs) often have access to users'emotional states. We study how this affective context modulates LLM sycophancy in subjective, evaluative interactions, where users share actions or opinions that invite feedback. Drawing on ingratiation theory, we measure sycophancy a...
Jiayi Li, Sanjana Menon, Brett Frischmann et al.· 0 citations
Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions. Existing emotion benchmarks mainly annotate surface polarity or final emotion categories, while lacking a structured acc...
Zhen-Yan Zheng, Yun-Yao Zhang, Jun Sheng et al.· 0 citations
Detecting true felt emotions when speakers suppress or mask their internal state poses a fundamental challenge for affective computing systems. We study a specific form of affective dissonance in dyadic speech: utterances where a speaker’s self-reported emotion diverges from all external observer ratings, indicating th...
Cheng-Shiuan Lin, T. Eze, Da-Wei Xie et al.· Proceedings of the 28th Inte...· 0 citations