Skip to content
Open access

Scaling recipes for single-cell RNA sequencing foundation models: when do scaling laws hold?

Sep 2026 · bioRxiv · 0 citations · 36 references
Biology

TL;DR

Empirical relationships linking the optimal learning rate and depth-to-width ratio to model size and depth or compute provide a quantitative framework for estimating the expected returns from additional resources and selecting suitable hyperparameters and architectures, thereby supporting the development of increasingly capable foundation models for omics data.

Abstract

Deep learning models exhibit empirical scaling laws whereby performance changes predictably with model size, dataset size, and training compute. Although these relationships are well established in domains such as language and image modelling, their applicability to biological data remains unclear. Here, we investigate scaling behaviour in foundation models trained on large collections of single-cell transcriptomes. We show that pre-training loss decreases systematically with model capacity and training compute, exhibiting a power-law dependence on model size. The strength and regularity of these trends differ between model formulations. We identify and quantify empirical relationships linking the optimal learning rate and depth-to-width ratio to model size and depth or compute. These results demonstrate that scaling principles extend to transcriptomic modelling. More broadly, they provide a quantitative framework for estimating the expected returns from additional resources and selecting suitable hyperparameters and architectures, thereby supporting the development of increasingly capable foundation models for omics data.

Read PDF

Similar papers

Open access Aug 2026

Single-cell foundation models benefit from cross-modal training: adding proteomics data beats parameter scaling

It is shown that with the right training recipe, heterogeneous proteomics data can improve the learned representations of single-cell RNAseq samples, demonstrating strong out-of-distribution generalization and suggesting that multimodal pretraining is a promising path toward more informative biological foundation model...

Maximilien Burq, Peter Cimermancic, Charlie Kim et al. · 1 citation
Open access Aug 2026

Benchmarking single-cell foundation models in a zero-shot setting

Evaluating four single-cell foundation models suggests that current single-cell foundation models provide useful representations for some downstream tasks in zero-shot conditions but do not yet offer a universal replacement for task-specific methods.

Yasmine Gaballa, Somaia K. Ahmed, T. Abdelaal · 0 citations
Open access Aug 2026

FP8 Inference in Genomic Foundation Models: Theoretical vs. Realized Speedups on GenomeOcean

Genomic Foundation Models (GFMs) are increasingly used for large-scale sequence analysis and generation. Compared with frontier language models, GFMs are typically smaller and frequently operate on long genomic sequences, with evaluation often requiring preservation of biologically meaningful structure and sequence-lev...

Mutian Yu, Robert Egan, Feng-Chen Liu et al. · 0 citations
Preprint Aug 2026

CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models

This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression values, potentially encouraging the reproduction of assay-specific techni...

Hai-Ping Liu, Qian Zhao, Li-Jing Lin et al. · 0 citations
Open access Sep 2026

Global tree encoding of atlas-scale single-cell genomics

MILK is presented, a scalable computational framework that organizes high-dimensional single-cell populations into unified tree representations that establish the hierarchical organization of biological data as a scalable and unifying representation of cellular identity, enabling integrative analysis of single-cell gen...

Brett Kiyota, Chaehyeon Lee, Hao-Yang Yao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Towards a knowledge-enhanced single-cell foundation model

Single-cell foundation models (scFMs) increasingly rely on large-scale transcriptomic pretraining, yet expanding pretraining data can yield diminishing gains while substantially increasing computational cost. Our data scaling analyses showed that incorporating biological knowledge, including cell-level text annotation...

Han-Qing Zhang, Jie Bao, Mei Ma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.