Overall, radiology-specific VFMs show promising transferability, but clinical translation remains constrained by limited data representativeness, heterogeneous benchmarks, incomplete reporting and insufficient deployment-oriented evaluation.
Abstract
Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation remain heterogeneous. We conducted a PRISMAScR scoping review of peer-reviewed studies published between January 2017 and March 2026 describing foundation models trained exclusively on radiological imaging data. Sixty-seven studies were included and mapped across three pillars: data scale and heterogeneity, architectural and pretraining scalability, and downstream transferability and generalization. Datasets primarily covered brain MRI, thoracoabdominal CT, and chest X-ray, ranging from fewer than 100,000 samples to multi-million-image cohorts. Transformer-based architectures and self-supervised pretraining predominated, particularly masked image modeling, contrastive learning and multi-stage approaches. Evaluation focused mainly on segmentation and classification, whereas cross-center, cross-scanner, anatomical and modality-shift validation was inconsistently reported. Alignment with FUTURE-AI principles was uneven. Overall, radiology-specific VFMs show promising transferability, but clinical translation remains constrained by limited data representativeness, heterogeneous benchmarks, incomplete reporting and insufficient deployment-oriented evaluation.
Findings across the reviewed studies suggest that pretraining with RadImageNet produces feature representations that are more aligned with radiological image characteristics than those derived from general-purpose datasets.
V. P. Hara Gopal, S. Nagaraju· Neural computing & applicati...· 0 citations
This review provides a comprehensive and structured synthesis of FMs in medical image analysis by systematically organizing studies into two primary categories: vision-only foundation models (VFMs) and vision-language foundation models (VLFMs), based on their architectural foundations, training strategies, and downstre...
P. Rajendran, M. Safari, Wen-Feng He et al.· Medical Image Analysis· 0 citations
As artificial intelligence becomes increasingly integrated into medical imaging practice, its robustness across heterogeneous real-world settings remains a major challenge. We quantified the effect of real-world distribution shifts on three-dimensional AI models for lung nodule analysis on CT and, motivated by these sh...
B. Bercean, Rafael Medelean, A. Tenescu et al.· Journal of imaging informati...· 0 citations
Current evidence suggests that foundation models are most promising when they are deployed as interactive, auditable components within human-in-the-loop workflows, where they can reduce annotation burden, support rapid draft segmentation, and improve consistency across large imaging studies.
Juntao Wei· Scientific Journal of Techno...· 0 citations
The findings indicate that the most convincing gains arise from task-adapted hybrid designs that combine local feature extraction with global context modeling, rather than from an unconditional superiority of transformers over convolutional networks.
Sam Ansari, Nastaran Faraji, Luke K. Topham et al.· Frontiers in Artificial Inte...· 0 citations
AI-enhanced neuroimaging is rapidly developing, with applications in tumor segmentation, hemorrhage detection, vascular lesion mapping, tractography, and augmented reality. Despite high algorithmic performance in experimental studies, few of these tools are integrated or validated into routine neurosurgeon practice. Th...