Skip to content
Open access

Performance evaluation of domain-specific and general-purpose AI models for chest radiograph interpretation: a comparative study.

Jul 2026 · BMC Medical Imaging · 0 citations
Medicine

TL;DR

Domain-specific multimodal AI model evaluated in this study demonstrated higher diagnostic consistency than the general-purpose LLM evaluated under the study conditions, suggesting the potential of specialized AI models as viable assistive tools, while highlighting the complementary utility of general-purpose LLMs in broader clinical contexts.

Abstract

Background

Chest radiography remains the most widely used imaging modality worldwide; however, its interpretation is inherently challenging because of overlapping anatomical structures and subtle findings. Recent advances in multimodal large language models (LLMs) have enabled automated radiology report generation, yet their clinical performance relative to domain-specific medical AI systems remains insufficiently validated.

Objectives

This study aimed to evaluate the performance and clinical applicability of a domain-specific multimodal AI model (M4CXR) compared with a general-purpose LLM (ChatGPT-4o) for chest radiograph interpretation.

Methods

In this retrospective study, 500 anonymized chest radiographs from a single tertiary care center were analyzed. Four board-certified radiologists independently evaluated AI-generated reports from both models. Key outcomes included key finding detection (categorized as complete, partial, or inconsistent), report generation time, and report discrepancies assessed using the RADPEER scoring system. Agreement between original and M4CXR-assisted RADPEER scores was assessed using intraclass correlation coefficients and weighted Cohen's kappa. Statistical analyses included paired t-tests, and chi-square tests.

Results

M4CXR demonstrated significantly higher report consistency than GPT-4o, with complete concordance observed in 55.8% versus 19.8% of cases, and lower inconsistency rates (25.2% vs. 46.4%, P<.001). The use of M4CXR significantly reduced report generation time compared with unaided interpretation (16.3 ± 12.9 s vs. 179.2 ± 50.4 s, P<.001). RADPEER-based discrepancy analysis revealed no significant differences between original and AI-assisted interpretations. Agreement between original and M4CXR-assisted RADPEER scores showed good reliability (ICC = 0.701), and weighted kappa analysis showed substantial agreement (κw = 0.652).

Conclusions

Domain-specific multimodal AI model evaluated in this study demonstrated higher diagnostic consistency than the general-purpose LLM evaluated under the study conditions. These findings suggest the potential of specialized AI models as viable assistive tools, while highlighting the complementary utility of general-purpose LLMs in broader clinical contexts. Future integration should prioritize human-AI collaboration and prospective multi-center validation.

Read PDF

Similar papers

Open access Aug 2026

Multicenter evaluation of four large language models for automated spine imaging diagnosis

Accurate interpretation of spine imaging is essential for clinical decision-making, yet the diagnostic potential of large language models (LLMs) for radiological report analysis remains inadequately evaluated in terms of sample size, multi-model comparison, reproducibility, and cross-institutional generalisability. Her...

Hao-Lai Liu, Hao Zhang, Hai-Xin Wei et al. · 0 citations
Aug 2026

Radiologically Relevant Clinical History Summarization with Large Language Models: A Multireader Performance Study.

LLMs generated radiology-relevant indications from clinical notes that were more comprehensive and factual than clinician indications, and when generated by the proprietary LLM, were ranked most useful in protocoling and imaging interpretation.

A. Serapio, Timothy L. Chen, Brian Tangsombatvisit et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Performance vs Consistency: Evaluating a Foundation Model in Lung-RADS Screening

Foundation models have recently demonstrated strong capabilities across a wide range of medical imaging tasks. However, their performance in structured clinical interpretation settings remains insufficiently explored. In lung cancer screening, interpretative variability persists despite standardized frameworks such as...

B. Renoust, P. Baudot, T. Foriel et al. · 0 citations
Open access Aug 2026

Incremental Diagnostic Value of Clinical Information for Large Language Models Across Multiple Organs: Retrospective Study

The addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, but this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.

Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.