Skip to content

Comparative Performance and Utility of Large Language Models in Generating Psychological Screening Checklists for Bariatric Surgery

Aug 2026 · Bariatric Surgical Practice and Patient Care · 0 citations · 11 references

TL;DR

Strategic model selection can enhance preoperative workflows by balancing comprehensiveness with targeted, actionable assessment, however, deployment should remain clinician-supervised because outputs may vary with model updates, prompt phrasing, and local workflow constraints.

Abstract

Psychological assessment is essential for bariatric surgery candidacy but remains inconsistent and time-consuming. This study compared three large language models (LLMs) in generating structured psychological screening checklists to support standardized preoperative evaluation. Models A, B, and C generated checklists for five standardized bariatric vignettes. Three blinded psychologists independently rated the outputs on 5-point Likert scales for clinical relevance, completeness, specificity, and usability. Statistical analysis included Jaccard similarity to quantify content overlap and analysis of variance with Tukey post hoc tests to compare expert ratings. Model A produced the most extensive checklists (mean = 35.6 items, standard deviation = 4.2), while Models B and C were more concise (28.4 and 26.7 items, respectively). Model A achieved the highest completeness scores, whereas Model B was rated highest for specificity and usability. Interrater agreement was good to excellent (intraclass correlation coefficient = 0.79–0.87). Moderate content overlap (Jaccard similarity = 0.45–0.63) suggested complementary model strengths. Between-model differences were significant for completeness, F (2, 6) = 11.5, p = 0.008, and specificity, F (2, 6) = 9.8, p = 0.013. LLMs differ significantly in checklist breadth and focus. Strategic model selection can enhance preoperative workflows by balancing comprehensiveness with targeted, actionable assessment. However, deployment should remain clinician-supervised because outputs may vary with model updates, prompt phrasing, and local workflow constraints.

View source

Similar papers

Open access Aug 2026

Large Language Models as Examinees and Graders in Simulated General Surgery Oral Board-Style Cases: A Psychometric Pilot Study

LLMs demonstrate passing-level oral board performance and more consistent grading than human surgeons in this simulation context, suggesting potential value for further investigation of LLM integration into surgical training and assessment frameworks, pending replication in larger samples.

Kian A. Huang, Haris K. Choudhary, A-Lan Xu et al. · 0 citations
Open access Jul 2026

Comparison of responses from large language models using artificial intelligence for parent-focused inquiries on clubfoot treatment and Ponseti management

Google Gemini outperformed ChatGPT-4.0 and CoPilot in accuracy and comprehensiveness for ICTEV management using the Ponseti method, whereas ChatGPT-4.0 offered more readable but less detailed answers.

A. Vescio, G. Testa, M. Sapienza et al. · 0 citations
Review Open access Jul 2026

Evaluating large language models as tools for public health education on scoliosis

Although all three LLMs achieved favorable overall ratings, Claude performed significantly better than ChatGPT and DeepSeek for scoliosis FAQs taking into account its higher accuracy and consistency.

Yi-Chen Wang, Xuanyan Liu, Mengjie Chen et al. · 0 citations
Review Open access Aug 2026

Psychometric Properties of Quality of Life Instruments for Patients with Hypertension: A Systematic Review

Instruments with the strongest psychometric profiles, such as QLICD-HY and DASI, are recommended for accurate patient assessment, yet future research must evaluate clinical responsiveness through longitudinal designs.

Afif Rakha Murtadha, T. Andayani, Dwi Endarti · 0 citations
Review Open access Aug 2026

Psychometric properties of the Kihon checklist for frailty screening: A scoping review.

The findings suggest that the KCL shows promise as a psychometrically evaluated instrument for frailty screening, especially in community settings, and further research is needed to address current evidence gaps and expand its applicability across diverse populations and clinical contexts.

Yang Zhao, Ting-Ting Wang, Ling-Na Kong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.