This work systematically evaluates a diverse set of deep learning and non-deep-learning methods for their ability to predict differential expression outcomes under two generalization regimes: unseen perturbations within the same cell line, and unseen cellular contexts across cell lines.
Abstract
Accurate predictions of transcriptomic responses to genetic perturbations could unlock our understanding of gene functions and regulatory networks. While a growing number of methods and benchmarks target this task, existing evaluations focus on mean expression accuracy alone. This overlooks differential expression (DE), which captures both mean and variance and forms the basis for biological interpretation and experimental follow-up. Here, we systematically evaluate a diverse set of deep learning and non-deep-learning methods for their ability to predict DE outcomes under two generalization regimes: unseen perturbations within the same cell line, and unseen cellular contexts across cell lines. We find that simple baselines, such as embedding-based nearest neighbors, are competitive and often outperform specialized deep learning models for DE classification across datasets and evaluation metrics. We further show that sparsity calibration, motivated by the structure of single-cell data, substantially improves DE classification for deep learning models that do not explicitly account for sparsity. Together, our findings establish practical baselines and evaluation principles for benchmarking perturbation models on DE prediction.
The results show that compact biological representations can support accurate and computationally efficient perturbation prediction, and highlight the importance of perturbation representations and population-construction procedures in low-data benchmarks.
Dewei Hu, Marc Pielies Avellí, L. J. Jensen et al.· bioRxiv· 0 citations
The performance and flexibility of State set the stage for scaling the development of AI models of cell state, a machine learning model that predicts perturbation effects while accounting for cellular heterogeneity within and across experiments.
Abhinav K. Adduri, Dhruv Gautam, Beatrice Bevilacqua et al.· Cell· 4 citations
Evaluating four single-cell foundation models suggests that current single-cell foundation models provide useful representations for some downstream tasks in zero-shot conditions but do not yet offer a universal replacement for task-specific methods.
Yasmine Gaballa, Somaia K. Ahmed, T. Abdelaal· bioRxiv· 0 citations
TranScouter is introduced, a lightweight encoderdecoder framework that represents perturbed genes using LLM-derived embeddings of their text summaries and represents biological conditions using transcriptomic profiles of control cells from the target condition.
This work shows that gene function is predictable from coexpression substantially because it reflects differences in expression between cell types, and these differences are also intrinsic to the ground truth labels, and indicates that function prediction models trained on bulk coexpression are largely limited to cell-...
Alex Adrian-Hamazaki, P. Pavlidis· bioRxiv· 0 citations
PerturbLDM, a latent-diffusion framework for conditional generation of single-cell transcriptional responses, is introduced, showing support for conditional response generation across data scales and biological settings.
Li-Shan Yu, Kang-Lin Hsieh, Y. Chu et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.