Skip to content
Open access

AI-Assisted Angoff Standard Setting for Multiple-Choice Examinations in Medical Education: Comparative Study

Sep 2026 · JMIR Medical Education · Vol 12 · 0 citations · 25 references
Medicine

TL;DR

The future application of LLMs as decision assistance tools for modified Angoff standard setting while maintaining expert human oversight is suggested, suggesting greater consistency in the generated estimates.

Abstract

Abstract Background Standard setting is essential for a defensible assessment in medical education. The modified Angoff method requires several expert judges, and determining the minimally competent candidate is cognitively challenging. However, empirical evidence on the role of AI in standard setting is unclear. Objective This study aimed to examine the role of large language models (LLMs) in the modified Angoff standard setting method for multiple-choice questions in a basic medical science examination compared with faculty judges. Methods This study was conducted in year 2 of the MD program in the United Arab Emirates. Ten faculty judges and 5 LLMs (GPT-5.2, Grok, DeepSeek, MedGemma, and Claude 4.5 Sonnet) determined the Angoff cutoff scores for a summative examination (120 multiple-choice questions). Standardized prompts were used for the LLMs to mimic the same information provided for faculty judges. Generalizability (G) theory analyses were performed using a fully crossed item × rater design to estimate variance components, G and Φ coefficients, decision study, and root mean squared error (RMSE) of the Angoff cutoff scores. Results Faculty-generated Angoff estimates (mean 69.13, SD 10.92) were comparable to LLM-generated estimates (mean 68.94, SD 12.39). A 2-tailed paired-sample t test revealed no statistically significant difference between the 2 groups (95% CI −2.21 to 2.91; t119=0.27; P=.79). Generalizability theory analysis demonstrated moderate reliability for the faculty panel (G coefficient=0.738; Φ coefficient=0.692). Despite comprising only 5 LLMs, the LLM panel demonstrated higher reliability (G coefficient=0.823; Φ coefficient=0.815), lower RMSE (1.28 vs 2.33), higher item-related variance (46.85% vs 18.36%), and lower rater-related variance (2.64% vs 16.40%) than the faculty panel. Pass rates were similar using LLM- and faculty-derived cutoff scores (52/73, 71.2%). LLMs differed in their minimally competent candidate conceptualization and approaches to determining item-level percentages of correct responses. Furthermore, the correlations between Angoff estimates and item-related P values were larger in LLMs than in faculty judges (r=0.552 vs 0.437). Conclusions In this single-institution study, the evaluated LLMs generated modified Angoff estimates that were broadly comparable to those of faculty judges. Generalizability theory analyses demonstrated higher G and Φ coefficients, lower rater-related variance, and lower RMSE for the evaluated LLM outputs under standardized prompting conditions, indicating greater consistency in the generated estimates. Application of the resulting cutoff scores produced pass and fail rates similar to those derived from faculty judges. These findings suggest the future application of LLMs as decision assistance tools for modified Angoff standard setting while maintaining expert human oversight.

Read PDF

Similar papers

Open access Sep 2026

An exploration of the utility of the modified Angoff standard-setting method for multiple-choice examinations in determining the "just good enough" undergraduate medical student in psychiatry.

Modified Angoff scores generally do predict group-level "borderline-pass" students' performance adequately, however their accuracy varies by question type, content area and psychometric property.

Kelechi Nnaemeka Chukwudi, B. Hallahan, C. McDonald · 0 citations
Review Open access Aug 2026

AI-augmented versus expert-authored multiple-choice questions: a psychometric comparison in a high-stakes specialty examination

Multiple-choice questions (MCQs) are widely used in written assessments, particularly in high-stakes medical examinations. Developing high-quality MCQs is time-consuming and requires subject matter expertise. Large language models (LLMs), such as ChatGPT-4o, have therefore been proposed as tools to support item g...

Wilma Anschuetz, Daniel Stricker, Claudia Canonica et al. · 0 citations
Review Open access Aug 2026

Beyond Multiple Choice: The Association between Diagnostic Reasoning Exams and Medical Student Learning Strategies

Assessment format shapes how medical students study. While multiple-choice questions (MCQ) are widely used, they may incentivize pattern recognition and memorization. We developed and implemented a novel open-ended assessment format, the Diagnostic Reasoning (DxR) exam, designed to mirror the step-by-step, iterat...

Hope M. Cherian, Rachel Dockter, Rebecca L. Toonkel et al. · 0 citations
Open access Sep 2026

Comparative analysis of short answer questions and multiple-choice questions in formative assessment of first-year medical students

MCQs were easier and yielded higher scores, whereas SAQs provided a more challenging assessment with comparable discrimination, and using both formats together may enhance assessment quality in undergraduate medical education.

A. M. Ammar, H. E. El Naggar, M. Ahmed et al. · 0 citations
#large language models Open access Sep 2026

Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment

Large language models (LLMs) are increasingly being considered for assessment support in health professions education; however, evidence of their performance in essay-style examinations remains limited. In particular, little is known about the reproducibility and operational stability of LLM-based grading under dif...

Asgeir Brevik, H. Jerpseth, S. Lafontan · 0 citations
Open access Aug 2026

Evaluating foundation models on official German medical licensing examinations: Implications for high-stakes assessment and AI-assisted medical education

Findings underscore the rapid progress of these models, particularly open-weight systems, and the value of official German medical licensing examinations as a restricted-access benchmark with reduced public exposure, and carry implications for high-stakes assessment and AI-assisted medical education.

L. Cirkel, Johannes Knitza, Volker Schillings et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.