Skip to content
Open access

Large Language Models as Examinees and Graders in Simulated General Surgery Oral Board-Style Cases: A Psychometric Pilot Study

Aug 2026 · Cureus · Vol 18 · 0 citations · 20 references
Medicine

TL;DR

LLMs demonstrate passing-level oral board performance and more consistent grading than human surgeons in this simulation context, suggesting potential value for further investigation of LLM integration into surgical training and assessment frameworks, pending replication in larger samples.

Abstract

Introduction: Surgical oral board examinations are vulnerable to examiner variability, potentially compromising scoring reliability. This study evaluated six public frontier large language models (LLMs) as both examinees and graders on general surgery oral board-style cases. Methods: Six commercially available LLMs were assessed across three standardized cases using a fully crossed psychometric design. Responses were graded on a 0-3 ordinal scale by six LLMs and three blinded senior surgeons using a structured rubric with anchored descriptors for each score level. Analyses included performance comparisons, reliability testing, and variance decomposition. Results: All LLMs achieved passing or near-passing performance. Significant differences existed between models (Friedman χ² = 12.51, p = 0.028). AI graders demonstrated greater internal consistency than the three-surgeon human panel in this pilot study (Cronbach's α = 0.697 vs. −0.923), though this comparison is limited by the small human panel and the absence of rater calibration for human graders. The dominant source of score variance was the Examinee × Rater interaction, suggesting that grader-driven disagreement accounted for more score variability than actual examinee performance differences in this dataset. Conclusions: LLMs demonstrate passing-level oral board performance and more consistent grading than human surgeons in this simulation context, suggesting potential value for further investigation of LLM integration into surgical training and assessment frameworks, pending replication in larger samples. Future work should evaluate whether these findings generalize to live examination settings, compare against calibrated human rater panels, and explore AI-augmented panel designs.

Read PDF

Similar papers

Aug 2026

Comparative Performance and Utility of Large Language Models in Generating Psychological Screening Checklists for Bariatric Surgery

Strategic model selection can enhance preoperative workflows by balancing comprehensiveness with targeted, actionable assessment, however, deployment should remain clinician-supervised because outputs may vary with model updates, prompt phrasing, and local workflow constraints.

Y. K. Çalışkan, Fatih Başak · 0 citations
Open access Aug 2026

Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy.

There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.

T. Davis, B. Guevel, K. Logishetty et al. · 0 citations
Review Open access Oct 2026

Performance of five large language models in frozen shoulder patient education: a multidimensional evaluation of readability, quality, clinical alignment, and safety

This study compared the educational suitability, overall quality, translation-based readability, clinical intent alignment, and safety of frozen shoulder patient education materials generated by GPT-5, DeepSeek, Doubao, Wenxin Yiyan, and Tongyi Qianwen. Twenty clinician-curated questions covering five content categorie...

Jia-Chen Liang, Wan-Ying Gui, Hua-Nan Li · 0 citations
Review Open access Aug 2026

Evaluation and Comparison of Large Language Model Responses to Frequently Asked Questions Regarding Patellofemoral Pain Syndrome: A Quality and Readability Assessment Study

The evaluated LLMs showed small and method-dependent differences in DISCERN-based information quality and variable differences in readability, but the findings should not be interpreted as evidence of factual accuracy, clinical safety, or suitability for individualized decision-making.

Oktay Polat, Berk Koncalıoğlu, Mert Gündoğdu et al. · 0 citations
Open access Sep 2026

Comparative Evaluation of Large Language Model Interfaces in Third Molar Surgery Complication Scenarios: Response Quality, Clinical Content, Potential Clinical Risk, Readability, and Externally Observable Response Latency—A Cross-Sectional Comparative Benchmark Study

In this single-generation benchmark, the sampled outputs showed different profiles in specialist-rated quality, clinical content, potential clinical risk, readability, and externally observable response latency under the specific interface configurations tested.

I. Karaca, Esmanur Başer, Emre Ulubaş et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.