From risk classification to clinical action: a prespecified paired pilot benchmark of public large language model interfaces for diabetes-related foot ulcer prevention
Abstract
Large language models (LLMs) may assist in the prevention of diabetes-related foot ulcers; however, their performance in classification may not translate effectively to context-dependent decisions. This study aimed to evaluate the accuracy, clinical actionability, reproducibility, and safety of five public LLM web interfaces in the context of the International Working Group on the Diabetic Foot (IWGDF) 2023 risk stratification and preventive management. A prespecified, paired, blinded, noninterventional pilot benchmark was conducted using 20 translated, de-identified, curated cases structured in a fixed clinical sequence and evenly distributed across the four IWGDF risk categories. These input conditions represent an idealized, best-case benchmark rather than routine clinical documentation. Each interface assessed all cases under standardized no-search conditions. The primary outcome was exact agreement with a frozen multidisciplinary consensus risk category. Additional outcomes included screening frequency, need for referral, specialty and urgency of referrals, completeness, interreviewer reliability, repeated-generation stability, and safety. Each interface classified 20/20 cases in concordance with the frozen IWGDF reference (100%; Wilson 95% CI 83.9%-100.0%). The 16.1-percentage-point interval below the observed ceiling indicates limited precision and remains compatible with clinically meaningful error in new cases. Of 100 primary outputs, agreement was observed in 89/100 (89.0%) for referral need, 68/75 (90.7%) for referral specialty, and 92/100 (92.0%) for urgency. Median information-item and mandatory-measure coverage was 100.0% for both measures; however, completeness scoring had limited inter-reviewer reliability and should be interpreted cautiously. In an exploratory, hypothesis-generating four-case repeated-generation substudy, all-three-generation consensus concordance was observed in 16/20 interface-case combinations for referral need, 14/15 eligible combinations for specialty, and 19/20 combinations for urgency. These descriptive counts are not estimates of failure probability or tail behavior. Under the prespecified curated no-search benchmark conditions, no major safety errors were observed among 140 outputs; this absence of observed events does not establish safety in routine clinical use. Performance was highest for structured guideline mapping, though reliability diminished in referral and individualized management across repeated generations. These findings highlight the necessity for auditable, clinician-supervised decision support instead of autonomous deployment.