While foundation models have been shown to learn biological representations from large transcriptomic atlases, it remained unknown whether proteomics data allow the same. We here therefore introduce OmicsFM, a modality-agnostic transformer pretrained through masked abundance reconstruction on an unprecedented proteomics data corpus of 48,837 quality-filtered proteomics profiles from 1,397 reprocessed PRIDE projects. Interestingly, despite training on 14- to 93-fold fewer profiles than matched bulk- and single-cell transcriptomic models, respectively, our proteomics model rivals both. On held-out projects, OmicsFM attention networks recovered more molecular relationships than co-expression methods and existing single-cell foundation models across nine reference databases that reveal pathway-level organization. Sample-level embeddings preserved biological structure across independent studies, and its representations transferred successfully to cell-type classification, gene-essentiality prediction, and perturbation-response prediction, while consistently outperforming task-specific models. Moreover, our results show that proteomics and transcriptomics representations capture complementary biology. OmicsFM thus firmly establishes the possibility of training highly performant proteomics-based foundation models, and their importance in modelling and uncovering fundamental biology.
Sander Heyndrickx, R. Gabriels, Harikrishnan Ramadasan et al.· bioRxiv· 0 citations
The potential for open modification searching to correct potential mistakes in non-canonical proteins detection by preventing modified canonical peptides or variants from being incorrectly identified as non-canonical peptides is shown.
V. Vasylieva, Enrico Massignani, Tine Claeys et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.