Evaluating foundation models on official German medical licensing examinations: Implications for high-stakes assessment and AI-assisted medical education
Aug 2026· npj Digital Medicine· Vol 9· 0 citations· 30 references
Medicine
TL;DR
Findings underscore the rapid progress of these models, particularly open-weight systems, and the value of official German medical licensing examinations as a restricted-access benchmark with reduced public exposure, and carry implications for high-stakes assessment and AI-assisted medical education.
Abstract
We evaluated proprietary and open-weight foundation models on 24 German medical licensing examinations (2019–2024), including 7485 items and response data from 119,878 sittings. For fair comparison, the eight vision-capable models were evaluated on the full benchmark and all thirteen on a shared text-only subset. On the full benchmark, Gemini 3.1 Pro achieved the highest overall accuracy, reaching 99.31% on the first (M1) and 98.37% on the second (M2) examination. On the shared text-only subset, proprietary frontier models performed at near-ceiling levels, several open-weight models (including GLM-5 and DeepSeek V3.2-Thinking) were highly competitive, and even compact ones exceeded mean student performance. Image-present items were more difficult for both students and models, but the associated decline was disproportionately larger for models than for students. Human- and model-defined difficulty subsets showed limited overlap, and model-hard subsets revealed residual differences among top systems. These findings underscore the rapid progress of these models, particularly open-weight systems, and the value of official German medical licensing examinations as a restricted-access benchmark with reduced public exposure. They carry implications for high-stakes assessment and AI-assisted medical education, notably multimodal assessment, human-aligned educational tools, and privacy-preserving local deployment.
The future application of LLMs as decision assistance tools for modified Angoff standard setting while maintaining expert human oversight is suggested, suggesting greater consistency in the generated estimates.
S. Kassab, Mariam Shadan, A. Ziganshina et al.· JMIR Medical Education· 0 citations
Background: Large language models (LLMs) have demonstrated strong performance on standardized medical examinations, with recent studies reporting performance approaching or exceeding that of senior medical residents. However, examination accuracy alone does not establish how models arrive at their answers or the relati...
F. Gafoor, M. Syed, M. Halai et al.· medRxiv· 0 citations
LLM performance in licensing examinations was strongly influenced by domain, search access, and visual characteristics, whereas associations with response format were less consistent and model-dependent.
Toshitsugu Sakurai, Daichi Aizawa, Kazuyoshi Okawa et al.· Journal of Medical Education...· 0 citations
Vision–language models (VLMs) are increasingly evaluated in medical education, yet their performance on visually intensive assessments remains incompletely understood. We compared four state-of-the-art VLMs, GPT-4o and GPT-5 (both accessed via ChatGPT), Gemini 2.5 Flash, and Gemini 3 Pro, with fifth-year medical studen...
O. Erdem, A. Yilmaz, A. Şahin et al.· Scientific Reports· 0 citations
The National Medical Commission embraced competency-based medical education (CBME) in 2019, which has a strong impact on higher-order cognitive abilities, professionalism, and clinical competence. Pharmacology in Phase II MBBS requires such competency-based evaluations. Large language models (LLMs) might help in...
Gurusamy Sivagnanam, Parasuraman Jayabharathy, Ganesan Justina Princess· The Journal of medical resea...· 0 citations
Diverse cloud-scale foundation/multimodal models and locally deployable models suitable for inference on consumer-grade GPUs are evaluated, achieving the highest accuracy of up to 90% on MCQs, similarity score of 0.72 on literature synthesis, and accuracy of 80% on a small sample of multimodal reasoning questions, with...
Shun Ye, V. C. Suja, Chen-Long Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.