Skip to content
Review Open access

A Comparative Analysis of Large Language Models (LLMs) in Generating High-Quality Pharmacology Questions for Undergraduate Medical Education

Aug 2026 · The Journal of medical research · 0 citations · 12 references

Abstract

The National Medical Commission embraced competency-based medical education (CBME) in 2019, which has a strong impact on higher-order cognitive abilities, professionalism, and clinical competence. Pharmacology in Phase II MBBS requires such competency-based evaluations. Large language models (LLMs) might help in creating CBME-aligned essay and short-note question assessments. Developing effective assessment tools in medical education requires significant time and expertise. With CBME, there is an increasing emphasis on assessments to evaluate the clinical reasoning rather than simple recall. In this context, the present study examined the ability of three advanced LLMs to generate pharmacology assessment questions for Phase II MBBS students. Each model was provided with a standardized prompt to generate a 36-mark question paper consisting of one structured long-answer question and five structured short notes. The outputs were independently reviewed and evaluated by two senior pharmacologists and by cross-model assessment by the LLMs, using a 4-point Likert scale. All three models generated complete assessment questions. ChatGPT-5.2 has shown the highest mean score of 3.77 ± 0.27, followed by Claude Opus 4.5 and Gemini 3 Pro at 3.75 ± 0.34 and 3.58 ± 0.22, respectively. Although the clinical relevance of questions was strong, human reviewers identified minor factual inaccuracies. Though LLMs are supportive tools for the generation of assessment questions, they require human supervision for the maintenance of precision and appropriateness.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.