Jul 2026· International Conference on Conversational User Interfaces· pp. 1-22· 0 citations· 150 references
Computer Science
TL;DR
A systematic review of empirical studies reveals that empathy is highly situated and driven by functional goals, like health and well-being, transactional service, social interaction, and learning support, and a disproportionate focus on text-based over voice-based interfaces.
Abstract
The advent of Large Language Models has accelerated interest in empathetic conversational agents. Despite a surge in empirical research, artificial empathy remains deeply fragmented, often serving as a catch-all term for diverse interactional phenomena. Addressing this conceptual gap, we systematically review 89 empirical studies to map how human-machine empathy is operationalized. Our synthesis reveals that empathy is highly situated and driven by functional goals, like health and well-being, transactional service, social interaction, and learning support. Within these contexts, we classify affective responsiveness by its directional flow, detailing how agents project, elicit, or mediate empathy. We structure the literature into a cohesive framework spanning linguistic, paralinguistic, identity, and architectural strategies. Furthermore, our methodological evaluation reveals a reliance on adapted clinical metrics, a scarcity of longitudinal studies, and a disproportionate focus on text-based over voice-based interfaces. Ultimately, this review equips researchers and practitioners with an actionable foundation for designing, measuring, and implementing contextually appropriate and empathetic agents.
Empathy is increasingly incorporated into large language model (LLM)–based conversational systems, particularly in healthcare settings. However, a persistent gap remains between the empathy expressed by these systems and the empathy expected and perceived by end users. This misalignment limits the effectiveness, trust, and acceptance of empathic AI, especially in emotionally sensitive domains such as cancer support. To address this challenge, this paper proposes an initial empathy-tuning for developers to systematically align system-delivered empathy with user expectations. Central to this approach is the assumption that empathy is not one-size-fits-all, as users require different forms and intensities of empathy depending on context and timing. We explore empathy tuning through a structured, developer-guided pipeline and demonstrate it via a prototypical implementation using the EPITOME framework. The approach is evaluated with prompt-based experiments and initial user studies in cancer-related scenarios. Our findings provide preliminary evidence that empathy in LLM-based systems can be tuned and that calibrated empathy improves alignment between system-generated and user-perceived empathy, while also highlighting the need for dynamic, runtime adaptation of empathic behaviour depending on conversation content.
Maryam R. Yeganeh, Samuel Fricker, Tamira Leber et al.· 2026 IEEE 34th International...· 0 citations
Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker's emotional experience. Emotional support conversation selects and sequences support for the seeker's current needs. Sustained use introduces a further goal. Effective support should sustain users'capacities for emotion regulation, coping, self-endorsed decisions, and social connection across the interaction lifecycle. We propose capability-sustaining emotional dialogue (CSED) as a longitudinal research paradigm that aligns supportive strategy with this goal and organizes data, models, system design, evaluation, and governance around repeated use, non-use, transition, and termination. A targeted literature-and-corpus audit motivates this position. In a PRISMA-ScR-guided sample, 95% of 60 system-building papers pursue relief-oriented goals. None evaluates capability or longitudinal outcomes, and only 1 considers dependency, autonomy, or termination risk. In 300 ESConv supporter turns, capability-relevant functions appear in 43.0%, while generic suggestions account for 22.0%, compared with 4.0% reappraisal, 6.7% self-efficacy support, and 0.3% boundary behavior. We release a protocol for extending the audit to model behavior. An illustrative process model connects latent user capability to six design commitments, four evaluation timescales, and lifecycle constraints. The resulting agenda makes CSED testable across data, policy design, training, evaluation, and governance.
Ming Wang, Jiaqi Wu Young, Wenfang Wu et al.· arXiv.org· 0 citations
Creating just educational systems requires disrupting how power operates in literacy teacher education programs and attention to preservice teacher (PST) agency. This study is grounded in literature concerned with PST agency, particularly in coaching interactions. Data consist of one empathy conversation, a dialogic coaching tool used to disrupt coaching hierarchies and center participant stories. We draw on positioning theory to ask: (1) How does a PST demonstrate agency in an empathy conversation? (2) How do empathy conversations disrupt traditional coaching norms? Utilizing conversation analysis and an analytic framework of agency, we show agency in a PST's telling of her own schooling history and current experiences. She co-constructs critiques of social structures with the field supervisor (FS), conveys key moments of choice in her story, and names systemic inequities. The structure and invitation of the empathy conversion allowed space for these agentic moves and extends beyond the limitations of lesson-focused coaching to foster dialogue about the dyad's lived experiences and educational systems. Implications include the potential to further explore how PST agency can emerge in coaching conversations.
Meagan Pike Dean, Melissa Mosley Wetzel· Literacy Research Theory Met...· 0 citations
This qualitative comparative study examines how ChatGPT-4o simulates empathy when responding to emotionally charged English in an EFL context. It aims to compare artificial intelligence and human responses in emotional recognition, pragmatic tone, empathetic support, and linguistic authenticity. The study is significant because EFL interaction requires learners to interpret affective cues while selecting socially and culturally appropriate language. Fifty prompts generated 50 AI responses and 1,000 human responses from 20 advanced-level EFL learners; the data were organized into 50 prompt-level comparison sets and analyzed through qualitative content analysis and comparative discourse analysis. Two trained coders applied a hybrid framework and achieved substantial agreement (Cohen’s κ = .86). ChatGPT recognized the intended emotion in 88% of its responses, used an appropriate tone in 84%, and displayed empathetic and pragmatically relevant support in 90%. Performance weakened with implicit, mixed, and culturally nuanced cues, while supportive language was sometimes formulaic or overly therapeutic. Human responses were more varied, culturally situated, and pragmatically flexible. The study recommends using ChatGPT as a teacher-mediated supplementary resource for emotional vocabulary and pragmatic practice, with explicit attention to cultural context, recurrent response formulas, and the distinction between simulated and human empathy.
Abdullah A. Al Fraidan, Jumana Waleed Buhaimed· Arab World English Journal· 0 citations
As online health information-seeking shifts to conversational AI, high-quality information retrieval increasingly relies on users'``communicative acts''(proactively sharing and seeking information)---similar to how effective diagnosis and personalized guidance are elicited in patient-clinician communication. Drawing on health communication research, this study examines how a chatbot's modality of empathetic expression (Verbal, Visual, Multimodal) and the conversational context (General, Sensitive, Mental Health) influence these acts through a 2 x 2 x 3 within-subjects experiment (N = 48). The results revealed that while verbal and multimodal empathy significantly increased reply length, communicative acts were largely shaped by conversational context, with Sensitive context triggering more question-asking and Mental Health context leading to heightened concerns, assertive responses, and unprompted information disclosure. Combined with qualitative findings, we discuss design implications for building context-sensitive AI health inquiry systems that can encourage active user participation.
Users increasingly turn to large language models for emotional support, yet little is known about how these models actually conduct a psychotherapy interaction. We introduce an ontology of ten therapeutic moves: compact, function-based categories grounded in the MULTI-60 inventory, validated through an annotation campaign with five licensed psychologists, and scaled with a judge-based approach that matches expert agreement. Applying it to real counseling transcripts and model-led sessions, we compare the move distributions between human clinicians and a panel of frontier models. Models over-use inquiry at up to three times the human rate, neglect psychoeducation, and are strongly context-anchored: they carry forward strategies initiated by a human clinician but rarely initiate them themselves. Exposing the ontology as a set of tools roughly halves the mean deviation from the human move distribution and improves turn-level alignment with human therapist by 7-9 percentage points, without any fine-tuning.
Afonso Baldo, Hugo Pitorro, Areti Vassilopoulos et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.