AI-augmented versus expert-authored multiple-choice questions: a psychometric comparison in a high-stakes specialty examination
Multiple-choice questions (MCQs) are widely used in written assessments, particularly in high-stakes medical examinations. Developing high-quality MCQs is time-consuming and requires subject matter expertise. Large language models (LLMs), such as ChatGPT-4o, have therefore been proposed as tools to support item g...