Generalizable CT vision-language modeling for population health and disease risk
Abstract
Vision-language foundation models (VLMs) for computed tomography (CT) are emerging tools that learn generalizable representations from large-scale clinical imaging data. While these models can predict task-specific labels, the extent to which their representations capture the clinical, physiological, and longitudinal variation of real-world patient populations remains unclear. We introduce Percival, a CT-native VLM trained on more than 400,000 CT-report pairs from the Penn Medicine BioBank using a dual-encoder symmetric contrastive framework. Across over 20,000 held-out participants, Percival’s latent space aligns with demographic, physiological, and laboratory variation, and supports phenome-wide associations across the electronic health record. We evaluate Percival against alternative foundation-model paradigms, including vision-only contrastive and multi-organ segmentation, demonstrating that vision-language pretraining captures clinical information not fully accessible to vision-only alternatives across disease classification and longitudinal risk modeling. Together, these findings indicate that CT-VLMs uncover latent structure aligned with clinical, physiological, and prognostic variation across the disease-prevalence spectrum.