Skip to content
Open access

Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage

Sep 2026 · Frontiers in Psychiatry · Vol 17 · 0 citations · 58 references
Medicine

TL;DR

The three models answered many late-life depression questions accurately and safely, but none demonstrated consistently reliable performance across complex or high-risk scenarios.

Abstract

Background General-purpose large language models are increasingly used by patients and caregivers to obtain mental health information and guidance about when professional care is required. In late-life depression, broadly accurate information may nevertheless be unsafe when cognitive change, multimorbidity, frailty, polypharmacy, self-neglect, caregiver dependence, or suicide risk is not adequately recognized. This study compared the clinical accuracy, safety, geriatric-specific appropriateness, triage performance, and response consistency of three large language models when answering patient- and caregiver-centered questions about late-life depression. Methods We conducted a blinded, paired benchmarking study using 90 questions covering six geriatric psychiatry domains and equally distributed across low-, moderate-, and high-risk strata. Each question was submitted independently to GPT-5.5 Instant via ChatGPT, Gemini 3.5 Flash via Gemini, and Seed2.0 Pro via Doubao, generating 270 primary-round responses. A stratified subset of 30 questions was resubmitted in separate conversations to assess test-retest consistency, yielding 360 responses overall. Two psychiatrists independently evaluated anonymized outputs against prespecified item-specific reference standards, with clinically important disagreements adjudicated by a third senior psychiatrist. The primary outcome was the proportion of clinically acceptable responses, defined using accuracy, clinical safety, geriatric appropriateness, warning-sign recognition, and triage criteria. Results Clinically acceptable responses were generated for 78.9% of questions by ChatGPT, 72.2% by Gemini, and 60.0% by Doubao (overall P<0.001). The difference between ChatGPT and Gemini was not statistically significant, whereas both outperformed Doubao. This model ranking remained unchanged under alternative core-safety and more stringent optimal-response definitions. Complete geriatric-specific appropriateness was achieved in 64.4%, 55.6%, and 43.3% of responses, respectively. Major safety errors occurred in 5.6% of ChatGPT responses, 10.0% of Gemini responses, and 16.7% of Doubao responses (raw P = 0.015; FDR-adjusted q=0.023). Performance declined substantially with increasing clinical risk. In exploratory post hoc caregiver-centered analyses, clinical acceptability was 76.7% for ChatGPT, 70.0% for Gemini, and 53.3% for Doubao, while explicit caregiver-directed action was present in 80.0%, 70.0%, and 53.3% of responses, respectively. Among high-risk questions, clinically acceptable response rates were 63.3%, 50.0%, and 36.7%, while under-triage occurred in 16.7%, 26.7%, and 36.7% of responses, respectively. Test-retest clinical consistency was highest for ChatGPT (90.0%), followed by Gemini (83.3%) and Doubao (73.3%). Conclusions The three models answered many late-life depression questions accurately and safely, but none demonstrated consistently reliable performance across complex or high-risk scenarios. Clinically important weaknesses involved geriatric-specific interpretation, recognition of indirect risk signals, crisis-response completeness, under-triage, and response stability. General-purpose large language models may support selected low-risk educational tasks but should not independently guide emergency triage, medication changes, suicide-risk management, or other safety-critical decisions in older adults.

Read PDF

Similar papers

#large language models Review Open access Sep 2026

Quality and safety of large language model–generated medication review outputs in geriatric pharmacotherapy: a two-stage comparative vignette-based benchmark evaluation

These findings support supervised use of large language models and evaluation approaches that assess reasoning and prioritisation as well as target detection and should not be interpreted as evidence that one model is clinically superior in real-world practice.

Kubra Cingar Alpay, D. Ozata, T. Gedik et al. · 0 citations
Open access Sep 2026

Applications of large language models in anxiety and depression patient care: a cross-model comparative analysis of dialogue quality

The disease burden of anxiety and depression is becoming increasingly severe, and the shortage of mental health clinical resources restricts patients’ access to care services. Large language models offer a new pathway for digital psychological care, yet standardized horizontal comparisons of response efficacy among dif...

Jie Min, Rong-Rong Jiang, Tao Yang et al. · 0 citations
#small language model Preprint Aug 2026

Performance of a domain-specific large language model in answering patient questions in psychiatry

MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time, and represents a step towards building safe LLM systems to enhance patient education in psychiatry.

Alexander J. Hish, A. Nagendran, S. Compton · 0 citations
Sep 2026

Use of multiple post-diagnostic dementia services across healthcare tiers in Beijing, China: A caregiver-reported cross-sectional study.

BackgroundPost-diagnostic dementia care, including for Alzheimer's disease, is multi-component and requires coordinated access to cognitive assessment, medication management, consultation for behavioral and psychological symptoms of dementia (BPSD), and non-pharmacological interventions. Evidence on service use and tie...

Chen-Li Zhu, Hai-Jin Li, Jing-Xuan Gao et al. · 0 citations
Open access Sep 2026

A structured exploratory multidisciplinary evaluation of four large language models responding to frequently asked patient questions about pregabalin

Significant limitations persist in safety warning compliance, response comprehensiveness, and hallucination rate for selected basic pharmaceutical questions, and significant limitations persist in safety warning compliance, response comprehensiveness, and hallucination rate.

Y. L. T. Bayala, I. A. Tinni, Thierry Boris Wend-Yam Yaméogo et al. · 0 citations
Open access Oct 2026

Large Language Models in Alzheimer’s Care: Clinical Use Cases, Safety Challenges, and Implementation Pathways

Background: Alzheimer’s disease (AD) care is increasingly communication-intensive, requiring sustained caregiver education, symptom monitoring, behavioral management, and coordination across fragmented clinical settings. Large language models (LLMs) can generate, summarize, and adapt natural language at scale, creating...

T. Alkam, E. Tarshizi, A. V. Van Benschoten · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.