Applications of large language models in anxiety and depression patient care: a cross-model comparative analysis of dialogue quality
Abstract
The disease burden of anxiety and depression is becoming increasingly severe, and the shortage of mental health clinical resources restricts patients’ access to care services. Large language models offer a new pathway for digital psychological care, yet standardized horizontal comparisons of response efficacy among different mainstream models in anxiety and depression care consultations are lacking. A cross-sectional comparative study design was adopted. Thirty standardized open-ended questions covering six themes (Diagnosis, Symptoms, Risk factors, Assessment, Management, and Outcome) were constructed. Uniform prompts were used to batch-query five tested models and collect responses. Three evaluators implemented two rounds of blind independent scoring with a 2-week interval, quantifying scores across five dimensions (Accuracy, Integrity, Relevance, Diversity, and Interpretability). Data processing and statistical testing were conducted using ICC, Spearman correlation, Wilcoxon signed-rank test, Kendall’s W, and PCA principal component analysis. Effect sizes and degrees of freedom were reported where applicable. GPT-5.3-Codex (21.35 points) and Claude-Opus-4.6 (21.25 points) showed the highest overall scores, whereas Gemini-3-Flash showed intermediate performance. The overall between-model difference was significant, F(4, 145) = 10.251, p < 0.001, η² = 0.220 (95% CI, 0.10–0.32), indicating a large overall effect. Significant pairwise contrasts were accompanied by moderate-to-large or large effects (Cohen’s d = 0.68–1.62). Test-retest ICCs were in the moderate-to-excellent range, and Spearman correlations between the two rounds were positive across all model-rater combinations (ρ = 0.549–0.864; df = 28; all p < 0.01). Wilcoxon signed-rank tests showed no systematic retest shift (all p > 0.05; effect-size r = 0.006–0.310). Within-model inter-rater correlations were generally strong. All models scored higher descriptively on Risk factors and Symptoms items, whereas Diagnosis showed the lowest overall mean score. Claude-Opus-4.6 had the highest mean scores for Accuracy and Interpretability, GPT-5.3-Codex for Integrity, and Gemini-3-Flash for Relevance and Diversity. PC1 and PC2 cumulatively explained 64.0% of the total variance. GPT-5.3-Codex and Claude-Opus-4.6 demonstrated the highest comprehensive efficacy under the standardized simulated consultation framework. Each model showed distinct performance patterns across themes and evaluation dimensions, which may inform future model optimization and prospective validation. Further real-world evaluation is required before these findings can support scenario-specific clinical deployment.