Skip to content
Preprint

Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration

Jul 2026 · 0 citations · 32 references
Computer Science

Abstract

Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languages remains largely unexplored. To address this gap, we curate 625 authentic mental health cases from three complementary sources: (1) publicly available Facebook posts discussing mental health concerns, (2) transcripts from the Bangladeshi television program"Ami Akhon Ki Korbo", and (3) anonymized student questionnaire responses covering diverse emotional and psychological challenges. Based on these cases, we build an evaluation corpus comprising advice written by licensed clinical psychologists and responses generated by three modern proprietary LLMs: GPT-4o Mini, Claude 4.5 Haiku, and Gemini 2.5 Pro. We further propose the Role-Playing Reflective Chain-of-Thought Advisory Framework (RP-RCAF), a task-specific prompting strategy that combines expert-authored few-shot examples with structured self-reflection to produce supportive, culturally aware, and ethically aligned counseling through a compassionate advisor persona. We also introduce the Grok 4-Based Response Evaluation and Scoring Framework (G-REFS), which integrates automated assessment with expert psychologist validation across emotional sensitivity, cultural appropriateness, linguistic clarity, and ethical soundness. Experimental results show that RP-RCAF consistently outperforms conventional prompting across all evaluated models and produces responses that more closely align with professional psychological counseling.

View source

Similar papers

Book Open access Jul 2026

From Scripted Responses To Therapeutic Dialogue: A Linguistic And Human Values Analysis Of Mental Health Chatbots

Mental health (MH) chatbots are increasingly used to provide accessible, on-demand emotional support, yet it remains unclear how these systems linguistically construct and communicate care. This work-in-progress examines whether MH chatbots produce responses that reflect supportive value orientations and counseling-adjacent tone. We conduct an observational analysis of responses from three widely used MH chatbots (Wysa, Sintelly, and Youper) across context-aware scenario prompts and a standardized-question session. Responses are analyzed using the SemEval’23 “Adam Smith” human value detection model and LIWC’22 psycholinguistic measures, including Language Style Matching (LSM), Clout, and Authenticity. Values such as “Security: Personal” and “Benevolence: Caring” appear consistently across systems, with contextual variation in secondary value emphasis. Linguistic patterns show moderate-to-high LSM and consistently high Clout, with Authenticity varying by scenario. These findings are exploratory signals intended to inform future evaluation and design of supportive conversational mental health systems.

Maleeha Sheikh, Chao Chen, Md. Romael Haque · 0 citations
Conference Jul 2026

Effectiveness of Student-Centered Fine-Tuning of Large Language Models for Mental Health Support

We developed a large language model designed to explore student mental well-being support in conversational settings, aimed at providing accessible, empathetic, and accurate responses to students facing challenges such anxiety, stress, loneliness, and academic pressure. Many students face barriers to seeking traditional counseling, such as stigma, scheduling constraints or discomfort with face-to-face interactions. The system addresses these challenges by offering a potential accessible conversational support channel through natural, conversational interactions with a large language model. Our approach focused on two strategies. Firstly, we enhanced the model's communication style to reflect counseling best practices such as empathy, active listening, and emotional validation. The second strategy is to enhance the model's understanding of mental health scenarios using realworld text sources and instructions. The model was trained on diverse, anonymized datasets from real counseling transcripts, emotional support dialogue corpora, and peer-support forums. We integrated prompt engineering, fine-tuning, and an iterative self-reflection loop to identify potentially unsupported or hallucination-prone responses, with the goal to improve factuality and safety in generated responses. We find that fine-tuning on student-centered data consistently outperforms both baseline and mixed-data approaches, emphasizing the importance of domainspecific adaptation. The model shows potential for confidential support, suggesting possible use as an early stage aid for coping strategies, and connects students to campus resources, reducing barriers to help seeking and supporting academic performance.

Sarthak Musmade, Lu Liu · 0 citations
Conference Jul 2026

Performative Empathy vs. Authentic Support: Benchmarking LLM Counselling Alignment for Digital Mental Health

Large Language Models (LLMs) are increasingly used for emotional support, yet their conversational behaviors often diverge from professional therapeutic standards. Rather than evaluating diagnostic accuracy, we assess how well these LLMs align with supportive conversational practices in digital mental well-being contexts. We present AuthenDia4MH, a transferable framework that transforms psychotherapy insights such as emotion consistency, sentiment dynamics, and linguistic simplicity into scalable quantitative metrics. Using a mental health Q&A dataset, we benchmark diverse frontier models against verified expert counsellors. Our results reveal distinct behavioral tradeoffs: proprietary reasoning models (e.g., GPT-4o, Claude) exhibit performative empathy characterized by hyper-agreeability and structural rigidity and suffer from a sophistication penalty, producing verbose responses that are significantly less accessible than human experts, while certain open-weight models (e.g., Ministral-8B) align more closely with the linguistic simplicity and naturalistic phrasing of professional counsellors. By quantifying these divergences, this work provides a benchmark for evaluating web-based mental health AI systems, providing transparent accountability mechanisms as these platforms become essential infrastructure for global mental health support.

Alexander Marrapese, Basem Suleiman, Jinglin Sun et al. · 0 citations
Preprint Jun 2026

How sensitive do we want AI to be? Socio-communicative competencies of large language models in healthcare

Background. Effective clinical practice relies heavily on the socio-communicative skills of medical professionals. Large language models (LLMs) have been proposed for tasks such as triaging patients, report drafting or translating medical jargon to support informed decision-making. These applications require both factual and social competence. This study evaluates dialogues between LLMs and participants to assess the current state of socio-communicative competencies displayed in LLM-generated texts. Methods. We extracted a subset of extended dialogues from the HELP-Med dataset, comprising 1800 conversation transcripts of interactions between human participants seeking medical information and three different LLMs, GPT 4o, Llama 3 and Command R+. Two experts coded the transcripts for demonstrations of socio-communicative behaviours (non-hostility, sensitivity, structuring, non-intrusiveness) using the IC-MD instrument, originally designed to evaluate interactional competencies in medical student admissions. Results. The LLMs in our study showed strength in non-hostility, mixed results in sensitivity and non-intrusiveness and performed poorly in structuring. Conclusion. Current LLMs lack the consistent and reliable socio-communicative skills needed for safe and effective use as healthcare advisors. While existing frameworks for assessing interactional competencies may support the development of more socially responsive LLMs, they will require adaptation to account for the differences in desirable behaviour between humans and LLMs.

Dorothee Amelung, Andrew M. Bean, Sabine C. Herpertz et al. · 0 citations
Book Open access Jul 2026

Large Language Models and the Evolution of Online Help-Seeking for Mental Health

Large language models (LLMs) are rapidly becoming embedded in everyday mental health help-seeking practices, particularly among young people who already turn to digital platforms as gateways to mental health support. While LLMs offer unprecedented immediacy and accessibility, their integration into help-seeking ecosystems raises important questions for digital health research. This opinion paper argues that LLMs fundamentally reshape the developmental processes underpinning online help-seeking. Traditional digital help-seeking requires active exploration, searching, comparing sources and reflecting on lived experience, processes that contribute to mental health literacy and resilience. In contrast, LLMs collapse informational plurality into singular, authoritative-sounding responses, potentially shifting users from active exploration toward passive consumption. We discuss the risks of sycophancy, and over-reliance on immediacy, and consider how these dynamics may alter developmental trajectories of coping and help-seeking agency. We argue that preserving agency, connectedness, and reflective engagement must be central to the design of conversational AI in health contexts.

Claudette Pretorius · 0 citations
Open access Jul 2026

Comparing Human and Large Language Model Responses to Patients Online Questions: Towards Multi-dimensional Patient-centered Support

Patients and caregivers seek informational and emotional support throughout medical care, especially when interpreting unfamiliar laboratory test results. Although resources such as patient portals and online health communities (OHCs) help address questions, gaps remain. The emergence of large language models (LLMs) offers the potential to be a complementary source of support to assist patients and caregivers in understanding and using their test results. The objective of our study is to empirically compare LLM responses to patients online questions containing their laboratory test results to responses written by peers in an OHC. We compared the 519 peer replies to 122 laboratory test-related posts from an OHC to 488 responses generated from four LLMs using mixed computational and qualitative methods. LLMs frequently provided clear explanations of medical terminology and structured interpretations of numeric results but were longer and less readable. Peers offered more personalized, context-specific emotional support. Overall, LLMs have the potential to complement peer responses in OHCs, but require greater emotional depth, reasoning transparency, and alignment with community norms.

M. Hussein, R. Doshi, L. He et al. · 0 citations