Empirical relationships linking the optimal learning rate and depth-to-width ratio to model size and depth or compute provide a quantitative framework for estimating the expected returns from additional resources and selecting suitable hyperparameters and architectures, thereby supporting the development of increasingly capable foundation models for omics data.
Abstract
Deep learning models exhibit empirical scaling laws whereby performance changes predictably with model size, dataset size, and training compute. Although these relationships are well established in domains such as language and image modelling, their applicability to biological data remains unclear. Here, we investigate scaling behaviour in foundation models trained on large collections of single-cell transcriptomes. We show that pre-training loss decreases systematically with model capacity and training compute, exhibiting a power-law dependence on model size. The strength and regularity of these trends differ between model formulations. We identify and quantify empirical relationships linking the optimal learning rate and depth-to-width ratio to model size and depth or compute. These results demonstrate that scaling principles extend to transcriptomic modelling. More broadly, they provide a quantitative framework for estimating the expected returns from additional resources and selecting suitable hyperparameters and architectures, thereby supporting the development of increasingly capable foundation models for omics data.
It is shown that with the right training recipe, heterogeneous proteomics data can improve the learned representations of single-cell RNAseq samples, demonstrating strong out-of-distribution generalization and suggesting that multimodal pretraining is a promising path toward more informative biological foundation model...
Maximilien Burq, Peter Cimermancic, Charlie Kim et al.· bioRxiv· 1 citation
Evaluating four single-cell foundation models suggests that current single-cell foundation models provide useful representations for some downstream tasks in zero-shot conditions but do not yet offer a universal replacement for task-specific methods.
Yasmine Gaballa, Somaia K. Ahmed, T. Abdelaal· bioRxiv· 0 citations
Genomic Foundation Models (GFMs) are increasingly used for large-scale sequence analysis and generation. Compared with frontier language models, GFMs are typically smaller and frequently operate on long genomic sequences, with evaluation often requiring preservation of biologically meaningful structure and sequence-lev...
Mutian Yu, Robert Egan, Feng-Chen Liu et al.· bioRxiv· 0 citations
This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression values, potentially encouraging the reproduction of assay-specific techni...
Hai-Ping Liu, Qian Zhao, Li-Jing Lin et al.· 0 citations
MILK is presented, a scalable computational framework that organizes high-dimensional single-cell populations into unified tree representations that establish the hierarchical organization of biological data as a scalable and unifying representation of cellular identity, enabling integrative analysis of single-cell gen...
Brett Kiyota, Chaehyeon Lee, Hao-Yang Yao et al.· bioRxiv· 0 citations
Single-cell foundation models (scFMs) increasingly rely on large-scale transcriptomic pretraining, yet expanding pretraining data can yield diminishing gains while substantially increasing computational cost. Our data scaling analyses showed that incorporating biological knowledge, including cell-level text annotation...
Han-Qing Zhang, Jie Bao, Mei Ma et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.