Skip to content
Conference

Performative Empathy vs. Authentic Support: Benchmarking LLM Counselling Alignment for Digital Mental Health

Jul 2026 · International Conference on Digital Health · pp. 237-246 · 0 citations · 24 references

Abstract

Large Language Models (LLMs) are increasingly used for emotional support, yet their conversational behaviors often diverge from professional therapeutic standards. Rather than evaluating diagnostic accuracy, we assess how well these LLMs align with supportive conversational practices in digital mental well-being contexts. We present AuthenDia4MH, a transferable framework that transforms psychotherapy insights such as emotion consistency, sentiment dynamics, and linguistic simplicity into scalable quantitative metrics. Using a mental health Q&A dataset, we benchmark diverse frontier models against verified expert counsellors. Our results reveal distinct behavioral tradeoffs: proprietary reasoning models (e.g., GPT-4o, Claude) exhibit performative empathy characterized by hyper-agreeability and structural rigidity and suffer from a sophistication penalty, producing verbose responses that are significantly less accessible than human experts, while certain open-weight models (e.g., Ministral-8B) align more closely with the linguistic simplicity and naturalistic phrasing of professional counsellors. By quantifying these divergences, this work provides a benchmark for evaluating web-based mental health AI systems, providing transparent accountability mechanisms as these platforms become essential infrastructure for global mental health support.

View source