Aug 2026· Bariatric Surgical Practice and Patient Care· 0 citations· 11 references
TL;DR
Strategic model selection can enhance preoperative workflows by balancing comprehensiveness with targeted, actionable assessment, however, deployment should remain clinician-supervised because outputs may vary with model updates, prompt phrasing, and local workflow constraints.
Abstract
Psychological assessment is essential for bariatric surgery candidacy but remains inconsistent and time-consuming. This study compared three large language models (LLMs) in generating structured psychological screening checklists to support standardized preoperative evaluation.
Models A, B, and C generated checklists for five standardized bariatric vignettes. Three blinded psychologists independently rated the outputs on 5-point Likert scales for clinical relevance, completeness, specificity, and usability. Statistical analysis included Jaccard similarity to quantify content overlap and analysis of variance with Tukey
post hoc
tests to compare expert ratings.
Model A produced the most extensive checklists (mean = 35.6 items, standard deviation = 4.2), while Models B and C were more concise (28.4 and 26.7 items, respectively). Model A achieved the highest completeness scores, whereas Model B was rated highest for specificity and usability. Interrater agreement was good to excellent (intraclass correlation coefficient = 0.79–0.87). Moderate content overlap (Jaccard similarity = 0.45–0.63) suggested complementary model strengths. Between-model differences were significant for completeness,
F
(2, 6) = 11.5,
p
= 0.008, and specificity,
F
(2, 6) = 9.8,
p
= 0.013.
LLMs differ significantly in checklist breadth and focus. Strategic model selection can enhance preoperative workflows by balancing comprehensiveness with targeted, actionable assessment. However, deployment should remain clinician-supervised because outputs may vary with model updates, prompt phrasing, and local workflow constraints.
LLMs demonstrate passing-level oral board performance and more consistent grading than human surgeons in this simulation context, suggesting potential value for further investigation of LLM integration into surgical training and assessment frameworks, pending replication in larger samples.
Kian A. Huang, Haris K. Choudhary, A-Lan Xu et al.· Cureus· 0 citations
Google Gemini outperformed ChatGPT-4.0 and CoPilot in accuracy and comprehensiveness for ICTEV management using the Ponseti method, whereas ChatGPT-4.0 offered more readable but less detailed answers.
A. Vescio, G. Testa, M. Sapienza et al.· Journal of Children's Orthop...· 0 citations
Although all three LLMs achieved favorable overall ratings, Claude performed significantly better than ChatGPT and DeepSeek for scoliosis FAQs taking into account its higher accuracy and consistency.
Yi-Chen Wang, Xuanyan Liu, Mengjie Chen et al.· Frontiers in Public Health· 0 citations
Instruments with the strongest psychometric profiles, such as QLICD-HY and DASI, are recommended for accurate patient assessment, yet future research must evaluate clinical responsiveness through longitudinal designs.
Afif Rakha Murtadha, T. Andayani, Dwi Endarti· Journal of Pharmacy and Scie...· 0 citations
The findings suggest that the KCL shows promise as a psychometrically evaluated instrument for frailty screening, especially in community settings, and further research is needed to address current evidence gaps and expand its applicability across diverse populations and clinical contexts.
Yang Zhao, Ting-Ting Wang, Ling-Na Kong et al.· Geriatric Nursing· 0 citations
The Two-Point Regression method offers an easily implementable and defensible approach for identifying excellence in numerically graded OSCEs, improving alignment between examiner global judgements and awarded grades.
A. Lunn, Christopher J. Harrison, J. McLachlan· Medical Teacher· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.