Skip to content
Open access

PsyEval: a comprehensive large language model evaluation benchmark for mental health.

Jul 2026 · npj Mental Health Research · 0 citations
Medicine

TL;DR

This work introduces PsyEval, a benchmark specifically designed to evaluate LLMs in mental health-related tasks across three core dimensions: knowledge, diagnosis, and emotional support, and reveals considerable gaps in LLMs' current ability to reason accurately and respond appropriately in mental health contexts.

Abstract

Evaluating large language models (LLMs) in the mental health domain presents distinct challenges due to the subtle, context-dependent, and subjective nature of psychological symptoms. We introduce PsyEval, a benchmark specifically designed to evaluate LLMs in mental health-related tasks across three core dimensions: knowledge, diagnosis, and emotional support. PsyEval is constructed to reflect the complexity of mental health scenarios and provides a structured framework for assessing model performance within this sensitive domain. Using PsyEval, we evaluate eleven advanced LLMs with different prompting strategies to investigate how prompting affects their responses. The results reveal considerable gaps in LLMs' current ability to reason accurately and respond appropriately in mental health contexts, while also indicating promising directions for future model enhancement.

Read PDF

Similar papers

Review Aug 2026

Large language model applications for real-time clinical mental health assessment: Current potential and future directions.

Large language models are best understood as emerging assessment-support tools rather than replacements for clinical evaluation because the limited pace of academic validation means that, at present, LLMs are best understood as emerging assessment-support tools rather than replacements for clinical evaluation.

K. Aafjes-van Doorn, Francine Cheng Ty, A. Hua et al. · 0 citations
Conference Jul 2026

A Multi-Domain Human Expert Evaluation of Clinical and Behavioral Knowledge in Large Language Models

Large Language Models (LLMs) such as ChatGPT and Gemini are increasingly used to answer medical and psychological questions, yet systematic evaluations across domains with differing reasoning demands remain limited. We present a multi-domain expert evaluation of two state-of-the-art models, ChatGPT Pro (v5.2) and Gemini 3 Pro, across three healthcare domains: Gynecology, Pathology, and Psychology. We curated 300 open-ended, realistic questions, 100 per domain, designed to elicit clinical reasoning, mechanistic interpretation, and conceptual explanation. Responses were independently scored by domain experts using a standardized five-point rubric. Results reveal domain-dependent performance patterns, with both models performing well in guideline-aligned advisory tasks in Gynecology, lower in mechanistic diagnostic contexts in Pathology, and showing divergent strengths in conceptual psychology questions. Quantitative and qualitative analyses highlight recurring limitations in contextual nuance, mechanistic depth, and safety framing. These findings underscore the importance of domainstratified evaluation and expert oversight when deploying LLMs for healthcare information and provide a reproducible framework for future assessments.

Pragna Prahallad, Pranathi Prahallad, Dhrithi P. Desai · 0 citations
Conference Jul 2026

Trustworthy Mental Health Assessment via Confidence-Guided LLMs

Depression and anxiety disorders are among the most prevalent and debilitating mental health conditions worldwide, imposing substantial personal, social, and economic burdens. Although recent advances in Large Language Models (LLMs) have shown promise in supporting mental health assessment and intervention, existing approaches often lack contextual awareness, real-time adaptability, and privacy-preserving personalization. To address these limitations, we propose a novel, context-aware and privacy-preserving mental health evaluation architecture that synergistically integrates LLM-driven intelligence. The proposed system enables personalized, continuous, and stigma-free mental health support by combining structured multiple-choice questionnaires with advanced language models, including GPT-3.5-turbo and Groq, to analyze user inputs, identify behavioral patterns, and predict potential mental health conditions such as depression and anxiety. Furthermore, the platform provides individualized recommendations, including self-care strategies, lifestyle adjustments, mindfulness practices, and referrals to healthcare professionals when appropriate. Recognizing the critical importance of reliability in sensitive healthcare settings, we introduce an ensemble-based aggregation framework that explicitly incorporates classification confidence and uncertainty quantification across multiple LLMs. Experimental results demonstrate that the proposed approach outperforms existing LLM models. By prioritizing user anonymity and data privacy, the proposed system reduces psychological barriers to seeking mental health support and promotes early intervention.

Jashraj Jani, Sara Akif, Wassila Lalouani · 0 citations
Review Open access Jul 2026

Scalable, context-sensitive psychiatric assessment with large language models and brief diaries

Abstract Background Accurate psychiatric assessment requires understanding a person’s unique experience within their psychosocial context. Clinical interviews have been the gold standard for assessment as the only methods capable of this complex task, but they are time and resource-intensive. Consequently, psychiatric assessment typically relies on patient report surveys that are decontextualized and narrow in scope. This comprehensiveness-scalability tradeoff is a major bottleneck in studying and treating psychopathology. We propose using large language models (LLMs) to score psychopathology from brief personal narratives as a low-burden, context-sensitive solution. Methods Participants (N = 108) completed brief (~1 minute), freeform audio diaries daily for 2 weeks. We used six LLMs to score wide-ranging psychopathology (Internalizing, Detachment, Disinhibition, Antagonism, Anankastia) from the diary transcripts. Leveraging an array of self-report and clinical interview measures, we tested the convergent, discriminant, concurrent, and clinical validity of LLM ratings for between-person differences and within-person fluctuations in psychopathology. Results Supporting convergent and discriminant validity, LLM ratings correlated most strongly with corresponding self-report domains at the between (average convergent r = .42) and within-person (r = .28) levels. LLM and self-report ratings had similar patterns of associations with external variables, except for Anankastia and Antagonism. Further, every LLM-rated domain related to psychopathology ascertained by clinical interview. Conclusions Across multiple forms of validity, we showed that LLMs can assess most major forms of psychopathology from mere minutes of audio. These results support scoring open-ended narratives with LLMs as a scalable, portable method to translate idiographic diagnostic data into standardized psychiatric assessments.

Whitney R. Ringwald, Aman Taxali, Michael Angstadt et al. · 0 citations
Review Aug 2026

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate. Mental health is of acute public-health concern as millions of people turn to LLMs for psychological support, and most existing evaluations are bespoke academic benchmarks that are difficult to integrate into developer workflows. We introduce HealthBench-Psych and HealthBench-Psych-Hard. We screened HealthBench's 5,000 physician-rubric conversations for mental-health relevance with a transparent LLM-applied rubric, then validated the subset through two rounds of blinded clinician review with concealed known-exclude controls, yielding 610 conversations (12.2% of the corpus). Evaluating 20 frontier and open models under a cross-vendor panel of three LLM judges, we find a statistically tied frontier cluster, measurable refusal behavior in two models, and near-identical rankings across judges ($\tau \ge 0.92$). We release the subset, pipeline, model responses, grades, and analysis code as a reusable resource.

M. Flathers, Phuong Anh Nguyen, J. Noorily et al. · 0 citations
Conference Open access 2026

A Comprehensive Investigation of Empathetic Dialogue Systems for Mental Health Support Using Large Language Models

A multimodal emotion-aware architecture, which pays attention to memory-enhanced personalization and emotion-specific reinforcement learning, is introduced and hybrid human-AI approaches, which focus on safety and empathetic conversation to improve current mental health systems are recommended.

Rainn Ji · 0 citations