The future application of LLMs as decision assistance tools for modified Angoff standard setting while maintaining expert human oversight is suggested, suggesting greater consistency in the generated estimates.
Abstract
Abstract Background Standard setting is essential for a defensible assessment in medical education. The modified Angoff method requires several expert judges, and determining the minimally competent candidate is cognitively challenging. However, empirical evidence on the role of AI in standard setting is unclear. Objective This study aimed to examine the role of large language models (LLMs) in the modified Angoff standard setting method for multiple-choice questions in a basic medical science examination compared with faculty judges. Methods This study was conducted in year 2 of the MD program in the United Arab Emirates. Ten faculty judges and 5 LLMs (GPT-5.2, Grok, DeepSeek, MedGemma, and Claude 4.5 Sonnet) determined the Angoff cutoff scores for a summative examination (120 multiple-choice questions). Standardized prompts were used for the LLMs to mimic the same information provided for faculty judges. Generalizability (G) theory analyses were performed using a fully crossed item × rater design to estimate variance components, G and Φ coefficients, decision study, and root mean squared error (RMSE) of the Angoff cutoff scores. Results Faculty-generated Angoff estimates (mean 69.13, SD 10.92) were comparable to LLM-generated estimates (mean 68.94, SD 12.39). A 2-tailed paired-sample t test revealed no statistically significant difference between the 2 groups (95% CI −2.21 to 2.91; t119=0.27; P=.79). Generalizability theory analysis demonstrated moderate reliability for the faculty panel (G coefficient=0.738; Φ coefficient=0.692). Despite comprising only 5 LLMs, the LLM panel demonstrated higher reliability (G coefficient=0.823; Φ coefficient=0.815), lower RMSE (1.28 vs 2.33), higher item-related variance (46.85% vs 18.36%), and lower rater-related variance (2.64% vs 16.40%) than the faculty panel. Pass rates were similar using LLM- and faculty-derived cutoff scores (52/73, 71.2%). LLMs differed in their minimally competent candidate conceptualization and approaches to determining item-level percentages of correct responses. Furthermore, the correlations between Angoff estimates and item-related P values were larger in LLMs than in faculty judges (r=0.552 vs 0.437). Conclusions In this single-institution study, the evaluated LLMs generated modified Angoff estimates that were broadly comparable to those of faculty judges. Generalizability theory analyses demonstrated higher G and Φ coefficients, lower rater-related variance, and lower RMSE for the evaluated LLM outputs under standardized prompting conditions, indicating greater consistency in the generated estimates. Application of the resulting cutoff scores produced pass and fail rates similar to those derived from faculty judges. These findings suggest the future application of LLMs as decision assistance tools for modified Angoff standard setting while maintaining expert human oversight.
Modified Angoff scores generally do predict group-level "borderline-pass" students' performance adequately, however their accuracy varies by question type, content area and psychometric property.
Kelechi Nnaemeka Chukwudi, B. Hallahan, C. McDonald· Irish Journal of Psychologic...· 0 citations
Multiple-choice questions (MCQs) are widely used in written assessments, particularly in high-stakes medical examinations. Developing high-quality MCQs is time-consuming and requires subject matter expertise. Large language models (LLMs), such as ChatGPT-4o, have therefore been proposed as tools to support item g...
Wilma Anschuetz, Daniel Stricker, Claudia Canonica et al.· BMC Medical Education· 0 citations
Assessment format shapes how medical students study. While multiple-choice questions (MCQ) are widely used, they may incentivize pattern recognition and memorization. We developed and implemented a novel open-ended assessment format, the Diagnostic Reasoning (DxR) exam, designed to mirror the step-by-step, iterat...
Hope M. Cherian, Rachel Dockter, Rebecca L. Toonkel et al.· The journal of the Internati...· 0 citations
MCQs were easier and yielded higher scores, whereas SAQs provided a more challenging assessment with comparable discrimination, and using both formats together may enhance assessment quality in undergraduate medical education.
A. M. Ammar, H. E. El Naggar, M. Ahmed et al.· BMC Medical Education· 0 citations
Large language models (LLMs) are increasingly being considered for assessment support in health professions education; however, evidence of their performance in essay-style examinations remains limited. In particular, little is known about the reproducibility and operational stability of LLM-based grading under dif...
Asgeir Brevik, H. Jerpseth, S. Lafontan· Frontiers in Education· 0 citations
Findings underscore the rapid progress of these models, particularly open-weight systems, and the value of official German medical licensing examinations as a restricted-access benchmark with reduced public exposure, and carry implications for high-stakes assessment and AI-assisted medical education.
L. Cirkel, Johannes Knitza, Volker Schillings et al.· npj Digital Medicine· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.