Skip to content
Open access

Evaluating foundation models on official German medical licensing examinations: Implications for high-stakes assessment and AI-assisted medical education

Aug 2026 · npj Digital Medicine · Vol 9 · 0 citations · 30 references
Medicine

TL;DR

Findings underscore the rapid progress of these models, particularly open-weight systems, and the value of official German medical licensing examinations as a restricted-access benchmark with reduced public exposure, and carry implications for high-stakes assessment and AI-assisted medical education.

Abstract

We evaluated proprietary and open-weight foundation models on 24 German medical licensing examinations (2019–2024), including 7485 items and response data from 119,878 sittings. For fair comparison, the eight vision-capable models were evaluated on the full benchmark and all thirteen on a shared text-only subset. On the full benchmark, Gemini 3.1 Pro achieved the highest overall accuracy, reaching 99.31% on the first (M1) and 98.37% on the second (M2) examination. On the shared text-only subset, proprietary frontier models performed at near-ceiling levels, several open-weight models (including GLM-5 and DeepSeek V3.2-Thinking) were highly competitive, and even compact ones exceeded mean student performance. Image-present items were more difficult for both students and models, but the associated decline was disproportionately larger for models than for students. Human- and model-defined difficulty subsets showed limited overlap, and model-hard subsets revealed residual differences among top systems. These findings underscore the rapid progress of these models, particularly open-weight systems, and the value of official German medical licensing examinations as a restricted-access benchmark with reduced public exposure. They carry implications for high-stakes assessment and AI-assisted medical education, notably multimodal assessment, human-aligned educational tools, and privacy-preserving local deployment.

Read PDF

Similar papers

Open access Sep 2026

Do Large Language Models Use the Clinical Vignette? A Question Ablation Study on the Orthopaedic In-Training Examination

Background: Large language models (LLMs) have demonstrated strong performance on standardized medical examinations, with recent studies reporting performance approaching or exceeding that of senior medical residents. However, examination accuracy alone does not establish how models arrive at their answers or the relati...

F. Gafoor, M. Syed, M. Halai et al. · 0 citations
Open access Sep 2026

Assessment Design in the Era of Large Language Models: Evidence From Japanese Health Professions Licensing Examinations

LLM performance in licensing examinations was strongly influenced by domain, search access, and visual characteristics, whereas associations with response format were less consistent and model-dependent.

Toshitsugu Sakurai, Daichi Aizawa, Kazuyoshi Okawa et al. · 0 citations
Open access Sep 2026

Performance of vision-language models compared with 252 medical students on text-only and image-based dermatology examinations

Vision–language models (VLMs) are increasingly evaluated in medical education, yet their performance on visually intensive assessments remains incompletely understood. We compared four state-of-the-art VLMs, GPT-4o and GPT-5 (both accessed via ChatGPT), Gemini 2.5 Flash, and Gemini 3 Pro, with fifth-year medical studen...

O. Erdem, A. Yilmaz, A. Şahin et al. · 0 citations
Review Open access Aug 2026

A Comparative Analysis of Large Language Models (LLMs) in Generating High-Quality Pharmacology Questions for Undergraduate Medical Education

The National Medical Commission embraced competency-based medical education (CBME) in 2019, which has a strong impact on higher-order cognitive abilities, professionalism, and clinical competence. Pharmacology in Phase II MBBS requires such competency-based evaluations. Large language models (LLMs) might help in...

Gurusamy Sivagnanam, Parasuraman Jayabharathy, Ganesan Justina Princess · 0 citations
#artificial intelligence Review Sep 2026

BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering

Diverse cloud-scale foundation/multimodal models and locally deployable models suitable for inference on consumer-grade GPUs are evaluated, achieving the highest accuracy of up to 90% on MCQs, similarity score of 0.72 on literature synthesis, and accuracy of 80% on a small sample of multimodal reasoning questions, with...

Shun Ye, V. C. Suja, Chen-Long Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.