Skip to content
Open access

An exploration of the utility of the modified Angoff standard-setting method for multiple-choice examinations in determining the "just good enough" undergraduate medical student in psychiatry.

Sep 2026 · Irish Journal of Psychological Medicine · pp. 1-6 · 0 citations · 19 references
Medicine

TL;DR

Modified Angoff scores generally do predict group-level "borderline-pass" students' performance adequately, however their accuracy varies by question type, content area and psychometric property.

Abstract

Objectives

Multiple-choice questions (MCQs) are widely used in medical education, with the Angoff standard-setting method frequently used to determine competence. However, minimal research has determined the accuracy of item-level Angoff scores in predicting the actual performance of borderline-pass students.

Methods

This retrospective study analysed five years (2019-2023) of psychiatry MCQ data (350 questions) from 191 fourth-year borderline-pass undergraduate medical students (scoring 45-55% in their psychiatry module MCQ) at the University of Galway. Angoff-derived cut scores were compared with actual student scores across clinical and knowledge-based questions, diagnostic categories, and content domains.

Results

Angoff (mean = 53.0%) and actual student scores (mean = 54.2%) showed minimal overall difference and were moderately correlated (r = 0.57, p < 0.001). However, questions where clinical scenarios were presented were underestimated (-6.2%), and knowledge-based questions overestimated (+2.9%) by standard-setters. Question difficulty was not a significant predictor of score differences. Lower discrimination scores predicted standard setter marks being higher than student marks for knowledge-based (B = -31.96, β = -0.19, t = 2.61, p = 0.01) questions.

Conclusion

Modified Angoff scores generally do predict group-level "borderline-pass" students' performance adequately, however their accuracy varies by question type, content area and psychometric property. Further standard-setter training, and inclusion of more diverse specialist input may improve standard-setting accuracy across less familiar domains.

Read PDF

Similar papers

Open access Aug 2026

Evidence of validity of the 10-item Big Five Inventory (BFI-10) in Peruvian medical students

Introduction: The Big Five model is the most widely accepted theoretical framework for assessing personality traits, with applications in medical education, such as to predict academic performance, burnout, and student well-being. The BFI-10 is an ultra-short 10-item version designed for time-constrained contexts, alth...

J. Flores-Cohaila, Brayan Miranda-Chávez, J. Huarcaya-Victoria et al. · 0 citations
Open access Aug 2026

Diagnostic discriminative accuracy and cut-off values of the Arabic-language PHQ-9 in a treatment-seeking mental health sample.

THEORETICAL BACKGROUND The PHQ-9 is a widely used international screening instrument for depressive disorders, but optimal cut-off values vary depending on population and geographical region. RESEARCH QUESTION This study aims to evaluate the psychometric properties and diagnostic accuracy of the Arabic PHQ-9 and iden...

A. Geiling, Amelie Pettrich, Maya Böhm et al. · 0 citations
Open access Sep 2026

Comparative analysis of short answer questions and multiple-choice questions in formative assessment of first-year medical students

MCQs were easier and yielded higher scores, whereas SAQs provided a more challenging assessment with comparable discrimination, and using both formats together may enhance assessment quality in undergraduate medical education.

A. M. Ammar, H. E. El Naggar, M. Ahmed et al. · 0 citations
Review Open access Oct 2026

A Psychometric Evaluation of the Modified Self-Efficacy in Clinical Performance Scale with Undergraduate Nursing Students Across Year Levels: A Cross-Sectional Study

Self-efficacy in clinical performance reflects nursing students’ confidence in their ability to perform clinical activities and is an important contributor to competence and confidence at graduation. Understanding and evaluating this construct requires valid and reliable instruments that can be used across different st...

Beth Pierce, T. F. van de Mortel, Jeanne Allen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.