Generative artificial intelligence and the transformation of medical education: a scoping review
Abstract
Generative artificial intelligence (GenAI) is moving rapidly from demonstrations of model capability into medical-school teaching, yet it remains unclear how it is being used with students and whether reported outcomes support claims of educational transformation. We conducted a scoping review following JBI methodology and PRISMA-ScR. PubMed, Europe PMC, and ERIC were searched for English-language journal articles published from 01 January 2023 through 20 August 2026. A supplementary PubMed sensitivity search used broader learner terms; Crossref, citation searching, and recent reviews supported additional discovery and verification. Eligible reports involved students enrolled in entry-to-practice medical degrees, documented actual exposure to a GenAI-enabled teaching or learning activity, and reported a post-exposure learning, task-performance, educational-process, or learner-experience outcome. Model benchmarks, attitude-only surveys, and studies confined to residents or other health professions were excluded. Evidence was charted by study design, learner stage, pedagogical function, outcome type, and follow-up. 153 reports representing 149 linked project-level evidence units were included. Publications increased sharply after 2024 and covered simulation and clinical skills (47 reports), case-based learning and clinical reasoning (21), tutoring and personalized study (20), assessment and feedback (22), writing and professional formation (13), and other applications. OpenAI models predominated (102 reports). 75 reports combined objective with learner-reported outcomes, whereas 41 relied only on self-report or qualitative evidence. Controlled trials and comparative studies produced encouraging short-term signals for history taking, communication, clinical reasoning, and selected knowledge outcomes, but null or unfavorable findings were also common. Only 13 reports contained a follow-up or retention signal, and no patient-level outcome was identified. Evidence more clearly shows changes in the design and availability of learning activities than durable improvement in learning. Task design, grounding, feedback quality, and faculty oversight emerged as plausible conditions for benefit but were not tested as quantitative moderators. Future studies should prioritize active comparators, preregistration, model and prompt reporting, retention and transfer, adverse effects, and equitable implementation.