Domain-specific multimodal AI model evaluated in this study demonstrated higher diagnostic consistency than the general-purpose LLM evaluated under the study conditions, suggesting the potential of specialized AI models as viable assistive tools, while highlighting the complementary utility of general-purpose LLMs in broader clinical contexts.
Abstract
Background
Chest radiography remains the most widely used imaging modality worldwide; however, its interpretation is inherently challenging because of overlapping anatomical structures and subtle findings. Recent advances in multimodal large language models (LLMs) have enabled automated radiology report generation, yet their clinical performance relative to domain-specific medical AI systems remains insufficiently validated.
Objectives
This study aimed to evaluate the performance and clinical applicability of a domain-specific multimodal AI model (M4CXR) compared with a general-purpose LLM (ChatGPT-4o) for chest radiograph interpretation.
Methods
In this retrospective study, 500 anonymized chest radiographs from a single tertiary care center were analyzed. Four board-certified radiologists independently evaluated AI-generated reports from both models. Key outcomes included key finding detection (categorized as complete, partial, or inconsistent), report generation time, and report discrepancies assessed using the RADPEER scoring system. Agreement between original and M4CXR-assisted RADPEER scores was assessed using intraclass correlation coefficients and weighted Cohen's kappa. Statistical analyses included paired t-tests, and chi-square tests.
Results
M4CXR demonstrated significantly higher report consistency than GPT-4o, with complete concordance observed in 55.8% versus 19.8% of cases, and lower inconsistency rates (25.2% vs. 46.4%, P<.001). The use of M4CXR significantly reduced report generation time compared with unaided interpretation (16.3 ± 12.9 s vs. 179.2 ± 50.4 s, P<.001). RADPEER-based discrepancy analysis revealed no significant differences between original and AI-assisted interpretations. Agreement between original and M4CXR-assisted RADPEER scores showed good reliability (ICC = 0.701), and weighted kappa analysis showed substantial agreement (κw = 0.652).
Conclusions
Domain-specific multimodal AI model evaluated in this study demonstrated higher diagnostic consistency than the general-purpose LLM evaluated under the study conditions. These findings suggest the potential of specialized AI models as viable assistive tools, while highlighting the complementary utility of general-purpose LLMs in broader clinical contexts. Future integration should prioritize human-AI collaboration and prospective multi-center validation.
Accurate interpretation of spine imaging is essential for clinical decision-making, yet the diagnostic potential of large language models (LLMs) for radiological report analysis remains inadequately evaluated in terms of sample size, multi-model comparison, reproducibility, and cross-institutional generalisability. Her...
Hao-Lai Liu, Hao Zhang, Hai-Xin Wei et al.· npj Digital Medicine· 0 citations
LLMs generated radiology-relevant indications from clinical notes that were more comprehensive and factual than clinician indications, and when generated by the proprietary LLM, were ranked most useful in protocoling and imaging interpretation.
A. Serapio, Timothy L. Chen, Brian Tangsombatvisit et al.· Radiology· 1 citation
Implementation of commercial AI-assisted pulmonary nodule assessment on chest CT scans reduced radiologist reporting time in a real-world clinical setting within a real-world clinical setting.
Jasika Paramasamy, Arlette E Odink, T. Mulders et al.· Radiology· 0 citations
Foundation models have recently demonstrated strong capabilities across a wide range of medical imaging tasks. However, their performance in structured clinical interpretation settings remains insufficiently explored. In lung cancer screening, interpretative variability persists despite standardized frameworks such as...
B. Renoust, P. Baudot, T. Foriel et al.· 0 citations
The addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, but this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.
Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al.· Journal of Medical Internet...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.