Skip to content
Review Open access

Foundation models in omics research: a comprehensive survey

Jul 2026 · Briefings in Bioinformatics · Vol 27 · 0 citations · 165 references
Medicine

TL;DR

This review systematically analyzes the emerging landscape of FMs in omics research, spanning sequence modeling, cell state characterization, and multimodal integration, and proposes a roadmap for the next generation of FMs, advocating for architectures that move beyond statistical correlation to incorporate causal reasoning, temporal dynamics, and autonomous experimental validation.

Abstract

Abstract The rapid expansion of high-throughput omics has created molecular datasets of unprecedented scale and complexity. These data are rich in biological information yet inherently sparse and high-dimensional, often limiting the effectiveness of conventional machine learning techniques. Foundation models (FMs), built on large-scale self-supervised pretraining, offer a robust alternative by learning generalizable representations directly from raw biological data. This review systematically analyzes the emerging landscape of FMs in omics research, spanning sequence modeling, cell state characterization, and multimodal integration. We organize the current literature into three distinct paradigms—sequence-centric, cell-centric, and multi-omics—to clarify a field currently fragmented by diverse tokenization strategies and architectural choices. Beyond methodology, we evaluate the practical utility of these models in tasks ranging from biomarker discovery to perturbation response prediction. We also identify critical barriers to adoption, including high computational costs, interpretability challenges, and the lack of standardized benchmarks. To support reproducible research, we provide a curated catalog of essential datasets and evaluation frameworks. Finally, we propose a roadmap for the next generation of FMs, advocating for architectures that move beyond statistical correlation to incorporate causal reasoning, temporal dynamics, and autonomous experimental validation.

Read PDF

Similar papers

Open access Aug 2026

A generalized supervised contrastive learning framework for integrative multi-omics prediction models

Advancements in multi-omics research have demonstrated the potential of integrating human microbiome and metabolomics data to better understand physiological processes and improve prediction accuracy in studies of human health. While conventional models utilizing single-omics data provide valuable perspectives, they often fail to capture the complexity of biological systems. Recent developments in supervised contrastive learning frameworks have enhanced predictive performance for categorical responses, yet limitations persist in extending these methods to continuous outcomes. A robust model capable of addressing these gaps could significantly enhance multi-omics predictions and provide new insights into complex biological interactions. Here, we present MB-SupCon-cont, a novel supervised contrastive learning framework designed for both categorical and continuous responses in multi-omics data. MB-SupCon-cont improves prediction accuracy by incorporating a generalized contrastive loss function that defines similarity and dissimilarity for continuous responses using three distance-based weighting methods. Through simulation studies and two real-world datasets for Type 2 Diabetes (T2D) and High-Fat Diet (HFD), we demonstrate that MB-SupCon-cont consistently achieves lower prediction errors than tuned conventional models, canonical correlation analysis, and autoencoder baselines, with most reaching statistical significance. We further provide a validation-based rule for selecting the weighting method and show that the learned embeddings align more closely with the response and recover known microbe and metabolite associations. The framework also provides superior representation learning and improves data visualization in lower-dimensional spaces. These findings suggest that MB-SupCon-cont is a powerful tool for general multi-omics prediction and may have broad applicability in biomedical research.

Sen Yang, Shidan Wang, Yiqing Wang et al. · 0 citations
Open access Aug 2026

Self-Attention over Parallel Dense Embeddings for High-Dimensional Omic Data

High-dimensional omic datasets present major challenges for machine learning due to their sparse biological signal, strong feature heterogeneity, and high dimensionality. In this work, we propose PLAT (Parallel Latent Attention Transformer), a neural architecture for high-dimensional tabular transcriptomic data. The model projects input gene expression features into multiple parallel latent representations, each processed independently through self-attention to capture complementary feature interactions while maintaining moderate model complexity. The proposed architecture was evaluated using both controlled Negative Binomial simulations designed to reproduce RNA-seq overdispersion and the TCGA-BRCA breast cancer dataset comprising 499 patients and 4376 gene expression variables for ER+/ER− classification. Comparative analyses against a baseline multilayer perceptron and a lightweight FT-Transformer showed that PLAT achieves competitive predictive performance while maintaining a comparable number of trainable parameters. Simulation experiments further indicate that its main advantage is concentrated in specific high-dimensional settings with an intermediate proportion of informative features. To assess model interpretability, we additionally performed a SHAP-based analysis of the baseline MLP and compared it with the attention-derived gene rankings. Although both models identified largely different sets of predictive genes, functional enrichment analyses consistently highlighted biological processes and disease pathways associated with breast cancer, supporting the biological relevance of the learned latent representations. These results suggest that PLAT provides an effective and interpretable framework for high-dimensional transcriptomic classification.

Kamal Elatifi, Nicolas Jäger Gallego, Á. Sánchez-Pla et al. · 0 citations
Open access Aug 2026

A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings

In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime.

Erik Vonkaenel, Lisa M. Bramer, Javier E. Flores et al. · 0 citations

Explainable AI for analyzing cancer outcomes using large-scale genome sequencing data

Metastatic cancer remains a leading cause of global mortality, yet accurate prognosis is frequently hampered by high-dimensional molecular features and heterogeneous clinical presentations. While traditional staging systems and linear models provide a foundational risk assessment, they often fail to capture the complex, nonlinear interactions between metastatic topology, genomic burden, and functional sequence variation. To address this, recent advances in machine learning and genomic foundation models present a transformative opportunity to integrate diverse data types into an explainable predictive framework. Consequently, this research developed a multi-tier, explainable AI framework designed to risk-stratify patients and predict overall survival using clinical and genomic covariates. Additionally, the framework aimed to surface sequence-level disease drivers by implementing joint variant calling from RNA-seq data and leveraging transformer-based architectures. The study employed a two-track methodological approach encompassing populationscale modeling and sequence-level deep learning. For the population-scale aim, a retrospective analysis was conducted on the Memorial Sloan Kettering-Metastatic cohort, consisting of 25,775 patients. Five distinct classifiers XGBoost, Logistic Regression, Random Forest, Decision Tree, and Naive Bayes were trained on a balanced subset of 20,338 patients utilizing an 80/20 stratified split. Model explainability was established through Shapley Additive Explanations (SHAP), while survival dynamics were evaluated using Kaplan-Meier estimates, Cox proportional hazards models, and an XGBoost-Cox variant. Concurrently, a pilot study involving 60 individuals, comprising 30 breast cancer cases and 30 controls, investigated sequence-level drivers using RNAseq data. A joint variant calling pipeline generated a unified genomic variant call format for association testing, and three genomic foundation models DNABERT-2, HyenaDNA, and Nucleotide Transformer were fine-tuned for 50 epochs on variantcentered windows spanning 100 base pairs in either direction to classify case versus control status. The results revealed stark contrasts in performance between the clinical and genomic modeling tracks. In survivability predictions, XGBoost emerged as the superior classifier, achieving an accuracy of 0.74 and an AUC of 0.82, while the XGBoost-Cox model outperformed the traditional Cox model with a C-index of 0.70 compared to 0.66. Through explainability and hazard-based analyses, metastatic site count, tumor mutational burden, the fraction of the genome altered, and the presence of liver and bone metastases were identified as the most potent prognostic indicators across pan-cancer and cancer-specific models. Conversely, the sequence-level transformer models exhibited severe overfitting, with test performance remaining near stochastic levels between 49 percent and 51 percent accuracy. Although DNABERT-2 achieved the highest nominal accuracy at 50.63 percent and HyenaDNA showed superior computational efficiency, the pilot ultimately indicated that fine-tuning transformers on raw sequences in small cohorts is heavily limited by a high signal-to-noise ratio and the polygenic complexity of cancer. Ultimately, this research demonstrates that explainable machine learning models can robustly predict survivability and highlight actionable features for oncology dashboards. However, future sequence-level deep learning efforts must pivot toward using frozen transformer embEd. D.ings or larger, multi-center cohorts to ensure equitable and generalizable clinical adoption.

P. Nalela · 0 citations
Preprint Aug 2026

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Bioassay activity prediction is often data-limited because drug-discovery datasets rely on time-consuming and expensive wet-lab experiments for data generation and evaluation. This challenge has inspired recent research into molecular foundation models (MFMs), which aim to encode general-purpose chemical knowledge into molecular representations that generalize well in data-constrained scenarios. This paper presents Monroe, a new MFM with several innovations over the existing state of the art: increased scale allowing pre-training on over 81 million molecules from the PM6 quantum chemistry dataset; improved graph representation of stereochemistry; improved training losses including conformer denoising and embedding decorrelation; improved multi-task learning; and the use of a prior-data-fitted model (TabPFN) for downstream in-context prediction. Our evaluations use a principled pairwise comparison framework that measures statistically significant performance differences. Across established Polaris benchmarks, Monroe matches or exceeds existing MFMs, while on activity cliff benchmarks, designed to assess utility for molecular discovery, it achieves significant improvements over prior methods. Finally, ablation and transfer experiments show that PFN-based downstream predictors also substantially improve two leading existing models, MiniMol and CheMeleon, yielding new state-of-the-art variants we call MiniMol_PFN and CheMeleon_PFN, suggesting that our downstream adaptation strategy generalizes beyond Monroe. Source code is at github.com/blazejba/monroe.

Blazej Banaszewski, Andrew W. Fitzgibbon · 0 citations