Skip to content
Open access

AdaGeneBudget: Cell-Adaptive Gene-Token Allocation for Efficient Single-Cell Foundation Models

Aug 2026 · bioRxiv · 0 citations · 16 references
Biology

TL;DR

AdaGeneBudget is introduced, a training-free gene-token selection method that combines each gene’s expression with reference-derived inverse detection frequency and retains the shortest ranked prefix that captures a target fraction of the cell’s expression-specificity score mass, which establishes biologically informed, cell-adaptive gene-token allocation as a practical complement to architectural and systems-level efficiency methods.

Abstract

Single-cell foundation models (scFMs) represent each cell using sequences of gene-associated tokens, making embedding extraction increasingly costly as the number of cells and expressed genes grows. Existing input policies typically rely on fixed input budgets, with retained genes determined by random subsampling, model-native ranking, or a fixed dataset-level highly variable gene (HVG) panel. However, they do not jointly determine, for each cell, which genes to retain and how many tokens to allocate. We introduce AdaGeneBudget, a training-free gene-token selection method that combines each gene’s expression with reference-derived inverse detection frequency and retains the shortest ranked prefix that captures a target fraction of the cell’s expression-specificity score mass. The resulting cell-specific budget is bounded by predefined minimum and maximum lengths, requires no cell-type labels, and leaves the pretrained backbone unchanged. We evaluated AdaGeneBudget in a frozen-backbone inference setting using pretrained scGPT and Geneformer models on Kang and PBMC reference-mapping tasks, with an additional scPRINT comparison against its official HVG policy and an expressed-only HVG control. Across four scGPT and Geneformer backbone–dataset pairs, AdaGeneBudget substantially reduced mean gene-token counts and peak GPU memory while increasing embedding-extraction throughput by up to 4.63×. Despite this compression, it preserved native-level aggregate annotation utility and consistently outperformed token-matched random selection. AdaGeneBudget also preserved fine-grained and low-support cell identities and retained lineage-marker programs and stimulation-associated pathway genes under compression. In scPRINT, both HVG controls achieved higher annotation macro-F1, whereas AdaGeneBudget more faithfully preserved the stimulation-induced embedding direction. These results establish biologically informed, cell-adaptive gene-token allocation as a practical complement to architectural and systems-level efficiency methods for applying existing scFMs to new datasets. They also suggest a cell-adaptive input-allocation principle for future models operating under finite token budgets.

Read PDF

Similar papers

Open access Sep 2026

LucaCell: a sequence-centric foundation model for cross-species single-cell analysis

LucaCell, a sequence-centric foundation model that represents genes through pre-trained mRNA sequence embeddings rather than static gene annotations, is presented, showing that sequence-informed gene representation can improve the generalization of single-cell foundation models across species, data types, and predictiv...

Yan Sun, Yong He, Min-Si Ren et al. · 0 citations
Open access Sep 2026

Cell-level random splits leak group-owned answers in single-cell benchmarks

Machine learning models in single-cell biology increasingly forecast differentiation, reprogramming and therapeutic response from early transcriptomic profiles. Testing whether a model has learned real biology requires held-out cells. Single-cell data, however, are grouped: cells from the same clone, patient or batch s...

Cang Hu, Sha Sun · 0 citations
Open access Aug 2026

scFair: Geometry-Aware Gene Budgets and Same-Rank Extension for Highly Variable Gene Selection

Background Highly variable gene (HVG) selection begins almost every single-cell RNA-seq analysis. While ranking formulas have been compared extensively, the integer gene budget at which any ranking must be truncated is typically left to the user and habitually fixed near 2,000. Relying on such a convention carries hidd...

Zhao Li, A. James, Sheng-Xuan Li · 0 citations
#artificial intelligence Preprint Sep 2026

ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing

Proteins perform diverse cellular functions, and even single amino-acid substitutions can alter stability, activity, or molecular interactions. Protein language models (PLMs) provide a scalable approach for modeling such sequence--function relationships from unlabeled sequences, but increasing the size of dense Transfo...

Ming-Rui Li, Si-Xian Shen, Min-Zhang Li et al. · 0 citations
Open access Aug 2026

Pretraining Enhances Megabase-Scale Gene Expression Prediction with GeneUnet

GB.GeneUnet, an 837M-parameter transformer-based U-Net pretrained on 6 trillion tokens from multi-species genomes in OpenGenome2 is introduced, extending genomic context to 1 Mb with up to 100× inference speedup over GeneMoE, a preliminary MoE transformer baseline of similar model size pretrained on the same data.

Ning Sun, William de Vazelhes, Pan Li et al. · 0 citations
Sep 2026

scGFormer: A Multi-Scale Graph-Transformer for Cell Type Annotation in Single-Cell RNA Sequencing.

ScGFormer is equipped with a biology-guided adaptive contrastive learning strategy, which is designed to account for zero inflation, balance class distributions, and refine dynamic graphs during training, thereby facilitating robustness and adaptability.

Ziqi Yuan, Hong-Wei Zhang, Cheng Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.