Evaluating Large Language Models for AI-Assisted Decision Support in Legal Capacity Assessment: A Comparative Study Using Interdisciplinary Medical Board Recommendations as the Expert Medical Reference Standard
Background: Legal capacity assessment requires multidisciplinary evaluation integrating cognitive, functional, neurological, psychiatric, and medico-legal information. Although large language models (LLMs) have shown promise in structured clinical reasoning, their role in supporting legal capacity assessment remains unclear. This study evaluated the performance of LLMs as AI-assisted decision-support tools using standardized medico-legal case vignettes, with interdisciplinary medical board (IMBD) recommendations under the Turkish Civil Code (TCC) as the expert medical reference standard; IMBD recommendations constitute an expert medical reference standard rather than final judicial determinations. Methods: We retrospectively analyzed 234 court-referred adult cases (2018–2024). Standardized, anonymized medico-legal case vignettes were independently evaluated by ChatGPT-5.2, Gemini 3 Pro, and Claude 4.5 Sonnet. Model outputs were compared with IMBD recommendations. The models were evaluated as AI-assisted decision-support tools and did not replace or influence clinical or judicial decision-making. Performance was assessed using accuracy, macro-F1, Cohen’s κ, AUC, calibration, decision-curve analysis, test–retest reliability, and human-factor outcomes; probability-based metrics were derived from secondary logistic models fitted to the categorical model outputs. Results: Article 405 was the most frequent outcome (65.8%). Dementia increased the likelihood of Article 405 recommendations (OR 6.5, 95% CI 1.2–35.2; p = 0.029), whereas higher Activities of Daily Living scores were protective (OR 0.96; p = 0.004). Gemini 3 Pro achieved the highest accuracy (86.3%), while ChatGPT-5.2 achieved the highest macro-F1 score (0.79) and AUC (0.91). Agreement with IMBD recommendations ranged from κ = 0.60 to 0.73, with high temporal stability (κ = 0.87–0.95); pairwise differences between the three models were not statistically significant after Holm correction. Performance was highest for cases in which Article 405 was recommended and for cases in which no guardianship was recommended, but remained limited for Article 408 (29.6% accuracy). The mean System Usability Scale score was 72.4, and safety flags were identified in 4 of 234 cases (1.7% of cases, corresponding to 4 of 702 individual model outputs). Conclusions: LLMs demonstrated agreement with IMBD recommendations on standardized medico-legal case vignettes, supporting further investigation of their potential role as AI-assisted decision-support tools under expert supervision. Because errors in this domain can directly affect fundamental rights, the use of AI in the legal system and in sensitive medical fields carries substantial risks and must remain strictly limited to expert-supervised decision support. Further prospective studies are needed to evaluate their safe integration into medico-legal practice.