Jul 2026· Journal of Information & Knowledge Management· 0 citations
Abstract
The demand for scalable and empathetic mental health support is driving increased interest in the use of large language models (LLMs) as advisory tools. Very few studies have been published that show how LLMs perform psychologically and demonstrate cross-model variation. We introduce DASS21-EvaLLM, a counsellor-in-the-loop evaluation system as an advisory appropriateness screening instrument for DASS-21 integration with four prominent LLMs (ChatGPT, Gemini, LLaMA and Mistral). The DASS21-EvaLLM provides the ability to rate, annotate and compare responses within a single interface. Using 65 simulated cases of clients and 13 licensed counsellors’ assessments, we considered the advisory quality of LLMs based upon each client’s profile for depression, anxiety and stress according to three specific criteria (accuracy, empathy and clarity), including a novel Weighted Score Index (WSI), for comprehensive and multi-dimensional comparison of advisory performance among LLMs. Overall results show that Gemini gives the highest quality overall as well as the highest level of empathy among LLMs while ChatGPT has the next highest level of advisory quality. Mistral and LLaMA both had specific strengths in certain scenarios, but both lacked emotional engagement and low levels of interpretability overall. Our contributions are: (i) a replicable evaluation protocol and workflow for evaluating LLM-based psychological advisories with counsellor oversight, (ii) a transparent WSI rubric and audit trail for per-criterion scoring and commentary, and (iii) evidence-based guidance for model selection and governance in digital mental health applications. DASS21-EvaLLM is an evaluation and training tool not a diagnostic system that supports safer deployment, improves counselling practice and supervision, and informs the design of responsible, human-centred advisory systems.
Large language models (LLMs) have expanded the potential of conversational AI in mental health support, yet counseling inherently relies on trust and relational aspects that may not transfer directly to these systems. We examine how users’ trust and experiences differ between human and LLM-based counseling, conducting a within-subjects study in which participants discussed their own concerns with both a licensed human counselor and an LLM-based counselor through text-based sessions. We assessed subjective distress and trust across five dimensions, and analyzed open-ended feedback across different trust profiles. The human counselor condition received higher overall trust, with the largest gaps in Faith and Personal Attachment, while Understandability remained comparable across conditions. Participants also reported greater reductions in subjective distress following human counselor sessions. Qualitative analysis further revealed that empathy, exploratory questioning, contextual understanding, and personalization were key factors shaping users’ trust, suggesting design implications for LLM-based counseling systems.
Sohhyung Park, Yein Song, Sungzoon Cho· IEEE Access· 0 citations
The results suggest that the chatbot is a usable alternative for administering screening instruments and that process-level behavioral traces may add contextual information beyond final questionnaire responses.
A. Netto, Dennis Paulino, Rafael Ris-Ala et al.· Information Hiding· 0 citations
Large Language Models (LLMs) are increasingly used for emotional support, yet their conversational behaviors often diverge from professional therapeutic standards. Rather than evaluating diagnostic accuracy, we assess how well these LLMs align with supportive conversational practices in digital mental well-being contexts. We present AuthenDia4MH, a transferable framework that transforms psychotherapy insights such as emotion consistency, sentiment dynamics, and linguistic simplicity into scalable quantitative metrics. Using a mental health Q&A dataset, we benchmark diverse frontier models against verified expert counsellors. Our results reveal distinct behavioral tradeoffs: proprietary reasoning models (e.g., GPT-4o, Claude) exhibit performative empathy characterized by hyper-agreeability and structural rigidity and suffer from a sophistication penalty, producing verbose responses that are significantly less accessible than human experts, while certain open-weight models (e.g., Ministral-8B) align more closely with the linguistic simplicity and naturalistic phrasing of professional counsellors. By quantifying these divergences, this work provides a benchmark for evaluating web-based mental health AI systems, providing transparent accountability mechanisms as these platforms become essential infrastructure for global mental health support.
Alexander Marrapese, Basem Suleiman, Jinglin Sun et al.· International Conference on...· 0 citations
Large language models are best understood as emerging assessment-support tools rather than replacements for clinical evaluation because the limited pace of academic validation means that, at present, LLMs are best understood as emerging assessment-support tools rather than replacements for clinical evaluation.
K. Aafjes-van Doorn, Francine Cheng Ty, A. Hua et al.· Journal of Psychopathology a...· 0 citations
We present LLM4SDM, the first study of open-source smaller language models (OS-sLLMs) for automated assessment of shared decision making (SDM) using the Observer OPTION12 framework. Unlike previous work that relies on large commercial models and the shorter OPTION5 instrument, our study focuses on privacy-preserving locally deployable models and Dutch melanoma consultation transcripts. Using expert-annotated clinical consultations, we evaluate three general-domain and two medical-domain OS-sLLMs during a development-phase pilot study. Results show that general-domain models outperform medical-domain models, which exhibit substantial hallucination and instruction-following failures. Gemma3:12b achieves the strongest agreement with human annotations (Pearson r=0.51, Spearman \r{ho}=0.59). Item-level and qualitative analyses reveal systematic challenges related to temporal discourse reasoning, conversational role attribution, and evidence grounding. We further introduce a Judge-LLM consensus framework designed to support disagreement resolution among multiple models. Our findings suggest that while current OS-sLLMs cannot replace human annotators, they offer a promising foundation for privacy-preserving human-in-the-loop SDM assessment.
Tamara Wit, Lifeng Han, C. Heipon et al.· 1 citation
UNSTRUCTURED
Inadequate post-care patient education contributes to preventable readmissions and adverse outcomes that disproportionately affect medically complex, high-need communities. Large language models (LLMs) show promise for generating personalized, plain-language patient education at scale. However, existing LLM evaluation frameworks prioritize technical accuracy over patient accessibility and alignment with health literacy, and few explicitly account for the attitudinal influences that clinician evaluators may introduce into the rating process. In this Viewpoint, we introduce an evaluation framework that pairs the Technology Acceptance Model (TAM) with the Medical Condition Regard Scale (MCRS). We call it the TAM-MCRS LLM Evaluation Framework, a novel clinician-centered approach for comparing which LLMs produce the highest-quality post-care patient education across accuracy, appropriateness, clarity, and completeness. We intend for this framework to be used to evaluate LLM-generated patient education outputs through a two-arm design that pairs an expert clinician panel with automated assessment methods, allowing for inter-arm comparison using clinical vignettes while accounting for measured evaluator attitudinal variance. The framework was developed through the National Institutes of Health Artificial Intelligence/Machine Learning Consortium to Advance Health Equity and Researcher Diversity (AIM-AHEAD) Clinicians Leading Ingenuity IN AI Quality (CLINAQ) fellowship program, in partnership with Ochsner Health and Xavier University of Louisiana. The TAM-MCRS framework integrates two theoretical lenses. TAM maps Perceived Usefulness onto accuracy and completeness, and Perceived Ease of Use onto clarity and appropriateness. Clinicians rate each output with a TAM-based questionnaire, and we then administer the MCRS as a post-scoring attitudinal covariate to see whether their regard for the conditions represented in the vignettes influences those ratings. Together, the two lenses are intended to produce evidence that is objective, theoretically grounded, clinically realistic, and disparity-responsive. Implications for clinician informaticists, health system governance, and responsible artificial intelligence (AI) deployment are discussed. This Viewpoint reflects the authors' position and is written for clinician informaticists, health system AI governance leaders, implementation scientists, and investigators evaluating LLM-generated patient education.
D. Austria, Grace Williams, Christopher Girardo et al.· JMIR Formative Research· 0 citations