Benchmarking Large Language Model Performance in Generating and Assessing Radiology Objective Structured Clinical Examination
Abstract
High-quality radiology assessment questions are essential for education competency evaluation but labor-intensive to create. To compare four large language models (LLMs) in generating and evaluating radiology objective structured clinical examination (OSCE)–style questions and responses. Fifty Radiopaedia cases across 10 subspecialties were used to generate 200 radiology OSCE-style sessions (three questions per session) using four LLMs: GPT-4o (4O), Llama 3-70b (LM), Claude 3.5 Sonnet (CL), and Gemini 1.5 Flash (GM). Three expert radiologists blindly evaluated sessions for clarity, clinical relevance, difficulty, option accuracy, assessment accuracy, and feedback quality on 5-point Likert scale. Generalized estimating equations (GEE) compared model performance, accounting for rater correlation. The primary accuracy metric was Top Two Box Accuracy (2TBA; all raters ≥4); secondary metrics included Top Box Accuracy (TBA), Average Score ≥4 Accuracy (AS4A), and Perfect 5 Accuracy (P5A; all raters = 5). Final rankings were determined via Borda count. Inter-rater reliability ranged from fair to good agreement (Gwet's AC2: 0.39–0.73). GEE analysis showed significant performance variations (p < 0.05) across metrics. Final Borda scores: 4O (18.5), LM (13.5), CL (11.0), GM (7.0). TBA was high (70.7–100%). For the primary metric 2TBA, 4O showed significantly higher odds of achieving high-quality scores vs. GM for clarity (odds ratio [OR] 4.52, p < 0.001) and overall assessment accuracy (OR 3.78, p < 0.001). Under the strictest threshold (P5A), 4O led in composite Question General Accuracy (52.7% vs. 32.7% LM, 21.3% CL, 1.3% GM). Qualitative failure modes included distractor ambiguity, conceptual repetition, and information leakage. This study provides a benchmark of LLM performance for radiology OSCE-style content generation and evaluation during a specific snapshot of artificial intelligence development (August 2024). 4O demonstrated the highest relative performance, though none consistently achieved expert-level quality. While LLMs can augment radiology education, expert review remains mandatory for high-stakes assessment.