Skip to content
Preprint

Can Protein-Derived Knowledge Improve Pathology Foundation Models?

Sep 2026 · 0 citations · 37 references
Computer Science

Abstract

Molecularly guided pathology foundation models (PFMs) exploit transcriptomic or proteomic information to enrich whole-slide image (WSI) representations, yet effectively leveraging large standalone molecular corpora remains challenging. First, existing molecular foundation models encode protein sequences or single-cell states, not the patient-level bulk expression profiles paired with WSIs. Second, because cross-modal supervision is restricted to paired WSI-omics samples, knowledge from standalone molecular corpora reaches the pathology encoder only indirectly, creating a paired-support bottleneck. To address these challenges, we propose a three-stage framework that decouples proteomic knowledge acquisition from cross-modal transfer, yielding ProSlide, a slide-level hierarchical pathology foundation model. First, to close the modality gap, we pretrain a Proteomic Foundation Encoder (PFE) on 12,695 sample-level bulk protein profiles using virtual profile generation and expression-space multi-view pretraining. Second, we pretrain ProSlide, a patch-region-slide encoder, to predict protein expression from paired WSI-protein samples. Third, to relax the paired-support bottleneck, we introduce Prot2Path, a cross-modal relational distillation objective. For each paired sample, it aligns the similarity distributions of the WSI and its protein profile over a shared, frozen bank of PFE-encoded paired and standalone profiles. We evaluate ProSlide on 12 downstream tasks across breast, lung, and renal cancers. Despite being pretrained with only 2,229 WSIs and 12,695 sample-level protein profiles, ProSlide achieves the highest mean accuracy and AUC within each cancer group.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.