MentalHospital, a virtual evaluation environment for LLM-based psychiatric clinical encounters, and MentalEval, five domain-specific evaluators covering communication empathy, interviewing professionalism, clinical-note quality, diagnostic rigor, and treatment appropriateness, trained with rubric-grounded SFT and expert-guided DPO are introduced.
Abstract
Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters. We introduce $\textbf{MentalHospital}$, a virtual evaluation environment for LLM-based psychiatric clinical encounters. MentalHospital instantiates the Subjective Interviewing, Objective Examination, Diagnostic Assessment, and Treatment Planning (S.O.A.P.) workflow, using skill-augmented standardized patients constructed from 1,193 de-identified psychiatric electronic health record (EHR) cases spanning all major ICD-11 categories and 76 disorders. Each encounter is assessed through a dual-track protocol that combines objective comparison against EHR-derived references with subjective assessment of clinical process quality. To scale specialist judgment, we develop $\textbf{MentalEval}$, five domain-specific evaluators covering communication empathy, interviewing professionalism, clinical-note quality, diagnostic rigor, and treatment appropriateness, trained with rubric-grounded SFT and expert-guided DPO. Survey responses from 22 clinicians support MentalHospital's clinical fidelity (3.88/5), while MentalEval achieves strong expert alignment with an average QWK of 0.944. Benchmarking shows that even the strongest LLM trails clinicians by 37.28 percentage points in objective psychiatric competence, with mental status assessment as a key bottleneck.
Background
Psychiatric medicine presents unique diagnostic and therapeutic challenges, often involving multimorbidity, polypharmacy, and atypical presentations requiring complex reasoning. Artificial Intelligence, particularly Large Language Models (LLMs), is emerging as a support tool in these settings. However, the clinical validity, interpretability, and reliability of LLMs in psychiatry remain largely unexplored, particularly their ability to generate transparent, guideline-consistent reasoning.
Aims
This study evaluates the clinical reasoning capabilities of LLMs in complex psychiatric scenarios. The primary aim is to assess the validity of their long chain-of-thought (CoT) reasoning. A secondary aim is to determine whether LLM assistance improves clinicians’ diagnostic accuracy.
Method
Three LLMs, Gemini 2.5, Grok, and DeepSeek R1, will be assessed using ten complex psychiatric cases sourced from a non-public clinical manual and rated with the Amsterdam Clinical Challenge Scale. Each model's CoT response will be evaluated by blinded panels of psychiatrists, residents, and general practitioners using standardized metrics for factual accuracy, coherence, and medical plausibility. In a second phase, clinicians will answer thirty diagnostic questions with and without support from the best-performing LLM. The study uses step-by-step reasoning prompts and few-shot examples to elicit detailed responses and includes bias mitigation strategies such as randomization, blinding, and statistical controls.
Results
Analyses will assess inter-rater reliability, metric redundancy, and reasoning quality. Closed-source models are expected to outperform open-source ones. LLM assistance is anticipated to improve diagnostic accuracy, especially among non-specialists.
Conclusions
This study provides a framework for evaluating LLMs in psychiatry and supporting their safe, evidence-based integration into mental health practice.
Vittorio De Vita, Bianca Destro Castaniti, Antonio Cristiano et al.· Monolith alpha· 0 citations
Large language models are best understood as emerging assessment-support tools rather than replacements for clinical evaluation because the limited pace of academic validation means that, at present, LLMs are best understood as emerging assessment-support tools rather than replacements for clinical evaluation.
Katie Aafjes-van Doorn, Francine Cheng Ty, A. Hua et al.· Journal of Psychopathology a...· 0 citations
BACKGROUND
Large language model (LLM)-generated hospital courses are increasingly integrated into electronic health records (EHRs), yet their accuracy and safety in pediatric populations remain poorly characterized.
OBJECTIVE
To evaluate the accuracy, text quality, and perceived potential harm of EHR-integrated and LLM-generated hospital courses in pediatric inpatient care during early clinical implementation.
METHODS
We conducted a descriptive evaluation from June 10 to August 8, 2025, at an academic freestanding children's hospital using an Epic EHR with an integrated LLM tool (GPT-4o and GPT-4.1). Clinicians across multiple roles, including attending physicians, residents, and advanced practice providers, reviewed LLM-generated hospital courses for their own patients. Clinicians identified and categorized errors (hallucinations, inaccuracies, or omissions). They also rated text quality (comprehensiveness, conciseness, coherence) on a 5-point scale and perceived harm on an 8-point scale.
RESULTS
A total of 129 LLM-generated hospital courses were reviewed (median length of stay, 3 days; IQR, 2-7) by 50 involved clinicians. Hallucinations occurred in 21% (95% CI, 14%-29%) of the hospital courses, inaccuracies in 41% (53/129; 95% CI, 33%-50%), and omissions in 24% (31/129; 95% CI, 17%-32%). Overall, perceived harm ratings were low (median, 0; IQR, 0-1). Text quality ratings were high (median [IQR]: comprehensiveness, 4 [3-5]; conciseness, 4 [4-5]; coherence, 4 [4-5]) and comparable with prior literature.
CONCLUSION
In this pediatric evaluation of LLM-generated hospital courses reviewed by frontline clinicians, errors were common, but perceived potential harm was low, even assuming use without clinician correction. These findings support the use of LLM-generated hospital courses as starting drafts when paired with clinician review and institutional safeguards.
Jasmine E. Kim, J. Hron, Daniel J Kats et al.· Hospital Pediatrics· 0 citations
Current evidence supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.
Yun-Jia Wu, Qi Yan, Dingcheng Tian· AI Medicine· 0 citations
MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time, and represents a step towards building safe LLM systems to enhance patient education in psychiatry.
Alexander J. Hish, A. Nagendran, S. Compton· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.