Aug 2026· Cureus· Vol 18· 0 citations· 20 references
Medicine
TL;DR
LLMs demonstrate passing-level oral board performance and more consistent grading than human surgeons in this simulation context, suggesting potential value for further investigation of LLM integration into surgical training and assessment frameworks, pending replication in larger samples.
Abstract
Introduction: Surgical oral board examinations are vulnerable to examiner variability, potentially compromising scoring reliability. This study evaluated six public frontier large language models (LLMs) as both examinees and graders on general surgery oral board-style cases. Methods: Six commercially available LLMs were assessed across three standardized cases using a fully crossed psychometric design. Responses were graded on a 0-3 ordinal scale by six LLMs and three blinded senior surgeons using a structured rubric with anchored descriptors for each score level. Analyses included performance comparisons, reliability testing, and variance decomposition. Results: All LLMs achieved passing or near-passing performance. Significant differences existed between models (Friedman χ² = 12.51, p = 0.028). AI graders demonstrated greater internal consistency than the three-surgeon human panel in this pilot study (Cronbach's α = 0.697 vs. −0.923), though this comparison is limited by the small human panel and the absence of rater calibration for human graders. The dominant source of score variance was the Examinee × Rater interaction, suggesting that grader-driven disagreement accounted for more score variability than actual examinee performance differences in this dataset. Conclusions: LLMs demonstrate passing-level oral board performance and more consistent grading than human surgeons in this simulation context, suggesting potential value for further investigation of LLM integration into surgical training and assessment frameworks, pending replication in larger samples. Future work should evaluate whether these findings generalize to live examination settings, compare against calibrated human rater panels, and explore AI-augmented panel designs.
Strategic model selection can enhance preoperative workflows by balancing comprehensiveness with targeted, actionable assessment, however, deployment should remain clinician-supervised because outputs may vary with model updates, prompt phrasing, and local workflow constraints.
Y. K. Çalışkan, Fatih Başak· Bariatric Surgical Practice...· 0 citations
There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.
T. Davis, B. Guevel, K. Logishetty et al.· Annals of the Royal College...· 0 citations
This study compared the educational suitability, overall quality, translation-based readability, clinical intent alignment, and safety of frozen shoulder patient education materials generated by GPT-5, DeepSeek, Doubao, Wenxin Yiyan, and Tongyi Qianwen. Twenty clinician-curated questions covering five content categorie...
The evaluated LLMs showed small and method-dependent differences in DISCERN-based information quality and variable differences in readability, but the findings should not be interpreted as evidence of factual accuracy, clinical safety, or suitability for individualized decision-making.
Oktay Polat, Berk Koncalıoğlu, Mert Gündoğdu et al.· Healthcare· 0 citations
In this single-generation benchmark, the sampled outputs showed different profiles in specialist-rated quality, clinical content, potential clinical risk, readability, and externally observable response latency under the specific interface configurations tested.
I. Karaca, Esmanur Başer, Emre Ulubaş et al.· Journal of Clinical Medicine· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.