EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations, and EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning and Group Relative Policy Optimization, are introduced, establishing a foundational framework for advancing emotionally intelligent dialogue systems.
Abstract
Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a systematic, theory-driven cognitive framework. We introduce EmoSBench, the first comprehensive EI evaluation benchmark for SLMs constructed upon the four-branch theoretical model, covering Perceiving, Understanding, Using, and Managing Emotion across ten sub-tasks. Preliminary assessments on EmoSBench reveal a substantial gap: even leading proprietary models like GPT-4o-Audio achieve only 52.6%, significantly trailing human baselines. To bridge this gap, we develop EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To facilitate its effective training, we curate EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations. Concurrently, we introduce a reward mechanism integrating a Steep Exponential Accuracy Reward (SEAR) and a Rationale Fidelity Reward (RFR) to enforce precise ordinal scoring and valid reasoning. Experiments demonstrate that EmoS reaches 83.8% accuracy, approaching human-level performance. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust real-world generalization, establishing a foundational framework for advancing emotionally intelligent dialogue systems.
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld-130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human-in-the-loop workflow. Supervised fine-tuning on this corpus reveals significant mutual benefits derived from multi-task learning. Second, to fully unlock the latent reasoning potential, we propose Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation. Extensive experiments demonstrate that OneEmo achieves state-of-the-art performance against similarly sized baselines across most benchmarks. Notably, despite having significantly fewer parameters than commercial models, OneEmo delivers highly competitive results. This paper paves the way for more reliable and interpretable affective computing. The code is available at https://github.com/waHAHJIAHAO/OneEmo.
Jiahao Huang, Zheng Lian, Jingyi Zhang et al.· 0 citations
Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of their visual emotional intelligence remains largely absent from recent model releases. We attribute this gap to a structural mismatch between conventional AICA paradigms and the open-ended, instruction-driven nature of MLLMs, where further analysis reveals four major limitations: omission of plausible responses, limited emotion taxonomies, neglect of contextual factors, and labor-intensive annotation. To overcome these barriers, we introduce Emotion Statement Judgement (ESJ), a statement-verification formulation that preserves the expressiveness of the input space while constraining outputs to discriminative judgements. We further develop INSETS, a labor-efficient pipeline that instantiates ESJ at scale by constructing INSETS-462k and supporting MVEI, a rigorously refined benchmark spanning sentiment polarity, emotion interpretation, scene context, and perception subjectivity. Beyond evaluation, we build EmObserver, an emotion-oriented MLLM optimized on ESJ through an elaborate multi-stage recipe. Extensive evaluation of broad-spectrum MLLMs on MVEI reveals fine-grained insights into current artificial visual emotional intelligence, while experiments on multiple AICA benchmarks demonstrate the accuracy, generalization, and reasoning faithfulness of EmObserver. Collectively, these results establish ESJ as a practical formulation, MVEI as a comprehensive benchmark, and EmObserver as an advanced baseline for advancing MLLM-oriented visual emotional intelligence. Code will be released at: https://github.com/wdqqdw/EmObserver.
Daiqing Wu, Dongbao Yang, Jiashu Yao et al.· arXiv.org· 1 citation
EmoLASP's LLM pipeline demonstrates the potential advantages of using a reasoning approach to ensure emotion prediction consistency and to reduce both the cost of fine-tuning and the cost of prompting with long dialogue histories.
Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks, yet it remains unclear whether they encode emotion through language-specific correlations or language-agnostic representations. We present the first neuron-level interpretability study of this question. We define Multilingual Emotion Neurons (MLENs) as functional units exhibiting stable emotional selectivity and aligned causal effects across languages, and introduce Consistency-Regularized Fusion (CR-Fusion) to identify them. Across four modern LALMs and 12 typologically diverse languages, emotion-sensitive neurons identified independently per language show minimal overlap, and additional monolingual identification data saturates quickly without isolating more transferable units, motivating identification from pooled cross-lingual evidence. Causal interventions demonstrate that MLENs identified by CR-Fusion provide more precise and transferable affective control than monolingual neuron sets in both zero-shot and low-resource settings. Leave-one-out ablations further reveal asymmetric transfer: individual identification languages, including low-resource ones, contribute non-redundant evidence, while several low-resource languages benefit most from the resulting cross-lingual transfer. Together, our findings provide the first causal, neuron-level account of how LALMs encode emotion across languages, and establish multilingual neuron identification as an effective mechanism for understanding cross-lingual affective behavior.
Xiutian Zhao, Philipp Koehn, Björn W. Schuller et al.· 1 citation
The ubiquity of generative AI (GenAI) requires social computing scholarship to examine how such extensive AI mediation may (re)shape the emotional and analytic qualities of human communication, both of which are central to audience engagement, trust, and information integrity. This study conducts a large-scale evaluation of 11 custom and fine-tuned LLMs, including DeepSeek, Llama (Meta), Mistral, and Gemma (Google), by asking them to rewrite the complete set of content, published by a local news outlet over a 12-year span (2011–2023). Using transformer-based sentiment models (i.e., RoBERTa fine-tuned on GoEmotions) and lexicon-based measures (i.e., LIWC, NRCLex), we compare the emotional expressiveness and analytic style of human-written versus AI-adapted content across short-form (titles/headlines) and long-form communication (article bodies/content). Analyzing 3,623 original news articles, nearly 40,000 primary-prompt AI-generated rewrites, and an additional generic-prompt ablation set, we estimate linear mixed-effects models that account for article-level clustering and control for model and prompt heterogeneity, communication form, verbosity, and readability. Results show that, under the corpus and prompting conditions examined here, LLMs tend to intensify emotional expression relative to human-written texts. Meanwhile, LLM rewrites tend to preserve, and in many cases increase, linguistic markers associated with analytic style relative to human benchmarks. Interestingly, even under the simpler generic prompt, LLMs in our sample display a shift toward more positive emotional categories, including joy, excitement, admiration, optimism, and gratitude, whereas human writers in this corpus display a broader affective range and more nuanced expression. To assess risks of clickbait-like headline framing, we fine-tune a BERT model for clickbait detection and examine linguistic cues in titles and headlines, yet we find minimal evidence that LLMs’ emotional amplification devolves into sensational and hyperbolic framing. Lastly, to enhance validity, we conduct a blinded human evaluation study on a randomly sampled subset of the corpus. Across 17 small experiments, human judges (n = 170) independently evaluated analytic style, reasoning depth, and clickbait characteristics. These human evaluations provide convergent evidence that AI-generated texts were perceived as more analytically structured and slightly higher in perceived reasoning depth, while AI-generated headlines were also rated as less clickbait-oriented. Overall, our findings suggest that GenAI and LLMs have the potential to enhance the affective intensity of communication without necessarily compromising its informational quality. However, as human communication becomes increasingly and inevitably AI-mediated, current LLMs may shift the affective register, narrow the emotional spectrum, and introduce tonal biases in public communication that could subtly influence how information is disseminated, and how language is traditionally interpreted.
Samuel B. Mazzone, J. Harlan, K. M. Sajjadul Islam et al.· ACM Transactions on Social C...· 0 citations
Emotion Recognition in Conversations (ERC) aims to identify speakers'emotions in multi-turn dialogue. Accurate emotion recognition can support a wide range of applications, including empathetic conversational agents, mental health support, and educational technologies. While many recent approaches rely on task-specific fine-tuning, such models may exploit dataset-specific cues. A central yet rarely questioned assumption in ERC is that each utterance can be assigned a single unambiguous emotion label. To investigate this assumption, we study ERC using Large Language Models (LLMs) in a zero-shot setting while incorporating preceding conversational turns as context. We show that aggregate metrics mask systematic failures. Errors concentrate around utterances containing negations, exclamations, and interjections. This pattern is consistent across all evaluated models, suggesting limitations in the benchmarks rather than model-specific weaknesses. A controlled re-annotation study involving four human annotators supports this finding: strong agreement is observed in only 35 percent of cases, with neutral utterances dominating high-agreement instances, while many emotional categories fall into low-agreement regimes. These findings suggest that many apparent model errors reflect genuine annotation ambiguity rather than poor emotion understanding. Standard single-label evaluation is therefore insufficient. To address this limitation, we introduce an LLM-as-Judge framework that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision.
Amir Ben Khalifa, Fanny Bezancon, Bessam Abdulrazak et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.