Skip to content
Review Open access

From 2D Vision–Language Models to Volumetric Medical AI: Large Language Models and Foundation Models for 3D Medical Imaging

Sep 2026 · Computation · 0 citations · 26 references

Abstract

Multimodal large language models (MLLMs) and vision–language models (VLMs) have rapidly entered medicine, demonstrating promising performance in clinical reasoning, radiology report generation, and visual question answering (VQA). However, many current multimodal architectures and pretrained visual backbones remain fundamentally rooted in two-dimensional (2D) image processing, even though major clinical imaging modalities, including computed tomography (CT), magnetic resonance imaging (MRI), optical coherence tomography (OCT), and echocardiography, are inherently volumetric or temporal. This narrative review examines the transition from 2D vision–language systems to volumetric multimodal AI, tracing the evolution from 2D and slice- or projection-based approaches through sequential and video-like methods to three-dimensional (3D) vision foundation models and native 3D VLMs/MLLMs. We examine their representational and computational trade-offs, evaluation gaps, and clinically grounded benchmarks. Approaches differ substantially in how they represent and preserve 3D information. Slice- and projection-based methods offer computational efficiency but may discard spatial context, whereas sequential and native volumetric approaches increasingly model relationships across the full imaging study. Recent 3D foundation models and multimodal systems demonstrate the feasibility of reusable volumetric representations and language-enabled 3D image interpretation, but face barriers in computational cost, training-data scale, evaluation methodology, and clinical reliability. Only 53% of Med-Gemini-3D reports were judged clinically acceptable, and natural language processing (NLP) metrics such as BLEU and ROUGE correlate poorly with diagnostic correctness. True 3D multimodal medical intelligence remains in its early stages. Future progress requires efficient volumetric representation strategies, clinically grounded evaluation frameworks, standardized benchmarks, and robust cross-institution validation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.