Skip to content
Open access

Comparative Expert Evaluation of Multimodal Large Language Models for Pediatric Rash Diagnosis: Clinical Utility, Safety, Information Quality, and Readability

Sep 2026 · Children · 0 citations · 30 references

Abstract

Background/Objectives: Multimodal large language models (LLMs) can interpret clinical text and images, but their performance in pediatric rash assessment remains uncertain. This study compared the clinical utility, safety, information quality, diagnostic correctness, and readability of ChatGPT, Gemini and Grok. Methods: Fifteen content-validated pediatric rash vignettes with brief histories and anonymized photographs were submitted once to each platform using a standardized zero-shot prompt. Three pediatricians blinded to platform identity independently rated the 45 responses using a five-point Clinical Utility and Safety (CUS) scale and a five-item modified DISCERN instrument. Diagnostic correctness was assessed descriptively; platform comparisons used Friedman tests with Bonferroni-adjusted Wilcoxon tests when appropriate. Results: Overall, 82.2% of CUS ratings were in categories 4–5 and 83.0% of modified DISCERN scores were ≥20/25; no rating was assigned to CUS category 1. Gemini and Grok had descriptively higher expert ratings than ChatGPT, but CUS did not differ significantly across platforms (p = 0.157), and although modified DISCERN differed globally (p = 0.038), no pairwise comparison remained significant after adjustment. In the single-query diagnostic assessment, at least one platform missed the reference diagnosis in 9/15 vignettes, and all three missed porphyria. Gemini generated the longest responses, whereas Grok produced the most linguistically complex text; neither response length nor readability was associated with expert ratings. Conclusions: The three multimodal LLMs produced predominantly clinically acceptable responses, but performance varied by vignette and platform. Because each vignette–platform combination was sampled once, diagnostic findings represent single-response observations rather than stable platform accuracy estimates. Clinical verification remains necessary.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.