Performance of vision-language models compared with 252 medical students on text-only and image-based dermatology examinations
Abstract
Vision–language models (VLMs) are increasingly evaluated in medical education, yet their performance on visually intensive assessments remains incompletely understood. We compared four state-of-the-art VLMs, GPT-4o and GPT-5 (both accessed via ChatGPT), Gemini 2.5 Flash, and Gemini 3 Pro, with fifth-year medical students on ten consecutive dermatology clerkship examinations administered between September 2023 and January 2025. Examinations combined text-only questions (multiple-choice, multiple-select, and matching; 60% of total score) with image-based, structured open-ended questions (40%) spanning seven dermatologic sub-domains. Model outputs were evaluated using expert-validated answer keys and grading rubrics, with repeated runs to assess output variability. All VLMs significantly outperformed medical students on text-only examinations (mean scores > 95 vs. 84.9; p < 0.001), showing minimal sensitivity to exam difficulty. In contrast, image-based performance was heterogeneous: Gemini 3 Pro and GPT-5 achieved higher scores than students, whereas students significantly outperformed GPT-4o and Gemini 2.5 Flash. Medical students showed the smallest difference between text-only and image-based scores within this examination setting. Among VLMs, sub-domain analyses suggested that strong visual description and diagnostic performance did not uniformly translate into etiological, differential diagnostic, or treatment-related reasoning. Within the constraints of this single-center, examination-based benchmark, these findings indicate that VLMs demonstrate strong text-based dermatologic knowledge, whereas multimodal performance remains uneven and model-dependent. These results should not be interpreted as evidence of standalone clinical competence, but rather as support for further controlled evaluation of VLMs as complementary, human-supervised tools in dermatology education.