Skip to content

Scaling an Autoregressive Transformer for Single-Cell Generation

Aug 2026 · 0 citations · 18 references
Computer Science Biology

TL;DR

The first jointly-fit two-exponent scaling law and compute-optimal frontier for a single-cell foundation model is found, finding the first jointly-fit two-exponent scaling law and compute-optimal frontier for a single-cell foundation model.

Abstract

We study a self-supervised generation task for single-cell gene expression vectors: given a set of vectors from a cell type, we aim to generate additional gene expression vectors of that cell type. For this task we characterize both the biological fidelity of the generated gene expression vectors and the scaling behavior of the pretraining loss. The model is a causal transformer paired with a learned quantized VAE tokenizer, trained with a cross-entropy loss. To evaluate the model, we condition it on held-out gene expression vectors of a cell type and generate vectors of gene expression, comparing the resulting distribution over gene expression vectors to the ground truth distribution of that cell type. We study the scaling properties of the proposed architecture by varying the number of trained parameters and the amount of training data. To our knowledge, we find the first jointly-fit two-exponent scaling law and compute-optimal frontier for a single-cell foundation model. Finally, we discuss how this pretrained model could be finetuned for perturbation response prediction.

View source

Similar papers

Open access Sep 2026

scRep: A Latent-Space Self-Distilled Foundation Model for Single-Cell Representation Learning

Single-cell foundation models have shown strong potential for learning transferable representations from large-scale transcriptomic data. However, many existing approaches rely on reconstructing masked gene expression values, creating a potential mismatch between observation-space reconstruction and the goal of learnin...

Sheng-Jie Wang, Zong-Yong Hu, Yunlong Bie et al. · 0 citations
Open access Aug 2026

Learning Discrete Cell and Niche Codes from Spatial Transcriptomics Using Dual Residual Vector Quantization

SQUINT is presented, a graph vector-quantized variational autoencoder that learns two disjoint codebooks per cell from a shared architecture that outperforms or is competitive with strong baselines on identification and achieves the most faithful cross-section integration.

Sebastian Birk, Arpit Merchant, Amirhossein Vahidi et al. · 0 citations
Open access Aug 2026

AdaGeneBudget: Cell-Adaptive Gene-Token Allocation for Efficient Single-Cell Foundation Models

AdaGeneBudget is introduced, a training-free gene-token selection method that combines each gene’s expression with reference-derived inverse detection frequency and retains the shortest ranked prefix that captures a target fraction of the cell’s expression-specificity score mass, which establishes biologically informed...

Dohee Kim, Uiwon Hwang · 0 citations
#artificial intelligence Preprint Sep 2026

P2P: Cross-View Population Denoising for Unpaired Single-Cell Perturbation Response Prediction

AIVC (AI Virtual Cell) is a learned simulator of cellular behavior across conditions. Predicting how a cell population responds transcriptionally to a genetic perturbation is a core task. Perturb-seq records that response by destructive sequencing, so a control cell and a perturbed cell are never observed as a pair, an...

Hao-Jie Yang, Ran Su · 0 citations
#small language model Preprint Aug 2026

The Von-Neumann State-Space Transformer for neural decoding

A von-Neumann inspired hypothesis of efficient computation as an alternative for neural decoding, a memory-augmented Transformer whose feed-forward block is a low-rank instruction bank: a shared base operator plus a small set of learned low-rank instructions, from which a per-token code synthesizes the weight matrix ac...

Morteza Sarafyazd · 0 citations
Open access Aug 2026

A platform for automated training of mammalian cell physiology

Controlling cell physiology is difficult, not only because of cells’ complexity, but also their capacity for real-time adaptation to interventions, leading to challenges such as drug resistance and transgene silencing. Accumulating evidence suggests that this adaptivity resembles classical forms of learning defined in...

Patrick Erickson, Douglas Hazel, Ramses Martinez et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.