Skip to content
Open access

Quality of Large Language Model-Generated MCQs Across Three Medical Disciplines: An Expert Rater-Based Comparison of Gemini, GPT-4 and Perplexity Pro

Aug 2026 · Advances in Medical Education and Practice · Vol 17, pp. 1-14 · 0 citations · 31 references
Medicine

TL;DR

Assessing LLM-generated assessment items provides valuable insight into the quality of LLM-supported MCQ questions, and there is no significant difference in mean scores and ranks among the three LLMs across different quality criteria.

Abstract

Background The increasing number of student cohorts has compelled academics to create a larger number of Multiple-Choice Question (MCQ) items. Large language models (LLMs) can help educators generate assessment items across multiple disciplines. The study compared the ability of LLMs (Gemini Advanced, Perplexity Pro, and ChatGPT 4.0) to generate high-quality, clinical-scenario-based MCQ items across three disciplines in a medical program, using an 8-point quality rubric. Materials and Methods Learning Outcomes (LOs) from Anatomy and Pathology disciplines of a pre-clinical semester 4 module and Family Medicine discipline of a clinical semester 6 module of a medical undergraduate program were selected. Using a pre-determined descriptive prompt, 63 MCQ items were generated from three AI tools. The quality of item construction was assessed by external content experts who were blinded to item generation using an 8-criterion rubric and a 4-point Likert scale. Mean scores and ranks for MCQs under each LLM were analysed, and a Friedman test was conducted to compare them. Kendall’s W showed that all criteria except one demonstrated some effect and weak-to-fair inter-rater reliability. Results When measuring key problem-solving skills, Perplexity Pro-generated MCQ items received a higher percentage of “strongly agree” ratings in Anatomy (57.1%, n=63), Pathology (56.5%, n=63), and Family Medicine (71.4%, n=63). Perplexity Pro received a higher percentage of “strongly agree” ratings in Anatomy, at 47.6% (n = 63), when evaluating specific content. Gemini Advanced also scored highly, with 68.8% in Pathology and 81% in Family Medicine (n = 63). A comparative analysis of the higher mean scores and ranks across 24 quality criteria showed that Perplexity Pro, Gemini Advanced, and ChatGPT 4.0 achieved 13, 10, and 1 higher mean scores, respectively. There is no significant difference in mean scores and ranks among the three LLMs across different quality criteria. Conclusion MCQs created by Perplexity Pro and Gemini Advanced achieved comparatively higher percentages of “strongly agree” ratings across quality criteria for item construction. Assessing LLM-generated assessment items provides valuable insight into the quality of LLM-supported MCQ questions.

Read PDF

Similar papers

Review Open access Sep 2026

Benchmarking Large Language Model Performance in Generating and Assessing Radiology Objective Structured Clinical Examination

This study provides a benchmark of LLM performance for radiology OSCE-style content generation and evaluation during a specific snapshot of artificial intelligence development (August 2024).

Ankush Ankush, Samriddhi Burman, Sydney Smith et al. · 0 citations
Review Open access Aug 2026

Large Language Models as Simulated Candidates in Objective Structured Clinical Examinations: A Rubric-Mapping Proof-of-Concept Study

Background Large language models (LLMs) perform well on knowledge-based medical examinations, but evidence on their utility for performance-based assessments such as Objective Structured Clinical Examinations (OSCEs) remains limited. Beyond simulating candidate answers, LLMs could support quality assurance of OSCE stat...

É. Lupon, Alexandre O Gérard, Alexandre Destere et al. · 0 citations
#large language models Review Sep 2026

A Real-World Evaluation of Large Language Model-Generated Hospital Courses in Pediatrics.

BACKGROUND Large language model (LLM)-generated hospital courses are increasingly integrated into electronic health records (EHRs), yet their accuracy and safety in pediatric populations remain poorly characterized. OBJECTIVE To evaluate the accuracy, text quality, and perceived potential harm of EHR-integrated and L...

Jasmine E. Kim, J. Hron, Daniel J. Kats et al. · 0 citations
#large language models Open access Sep 2026

Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment

Large language models (LLMs) are increasingly being considered for assessment support in health professions education; however, evidence of their performance in essay-style examinations remains limited. In particular, little is known about the reproducibility and operational stability of LLM-based grading under dif...

Asgeir Brevik, H. Jerpseth, S. Lafontan · 0 citations
Review Open access Aug 2026

A Comparative Analysis of Large Language Models (LLMs) in Generating High-Quality Pharmacology Questions for Undergraduate Medical Education

The National Medical Commission embraced competency-based medical education (CBME) in 2019, which has a strong impact on higher-order cognitive abilities, professionalism, and clinical competence. Pharmacology in Phase II MBBS requires such competency-based evaluations. Large language models (LLMs) might help in...

Gurusamy Sivagnanam, Parasuraman Jayabharathy, Ganesan Justina Princess · 0 citations
Open access Aug 2026

Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting

How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next, and findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.

Nıyazı Çetın, A. Atılan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.