Skip to content
Open access

Method Choice, Not Biology, Determines In Silico Perturbation Results: A Systematic Evaluation of Eight Methods Across Four Datasets

Aug 2026 · bioRxiv · 0 citations · 36 references
Biology

TL;DR

Computer methods that simulate gene knockout experiments from single-cell RNA sequencing data are increasingly popular, but researchers lack guidance on which method to choose, so preliminary guidance for method selection is established, including cross-pathway validation, direction-aware benchmarking, and minimum data requirements.

Abstract

Most in silico perturbation methods for single-cell transcriptomics have been validated only on individual datasets, leaving their reliability and generalizability unknown. Through systematic cross-method, cross-dataset benchmarking of eight methods spanning six mathematical frameworks across four datasets, we find that six of eight methods—including widely used VAE-based and tensor decomposition approaches—fail to produce detectable transcription factor (TF)-to-pathway signals. Only CellOracle and DDIM consistently detected TF-to-glycolysis directional regulation. Cross-pathway analysis in PBMC monocytes revealed biologically coherent TF-pathway associations beyond glycolysis (SPI1→glycolysis 4.4× enrichment, FOS→AP-1 targets 4.4×), with SOX9 serving as a biological specificity control (no pathway enrichment). Method choice alone could reverse biological conclusions: DDIM and scTenifoldKnk rankings were significantly anti-correlated (ρ=-0.811, p=0.027). CRISPRi Perturb-seq validation in K562 cells confirmed TF knockdown suppresses glycolysis gene expression (JUN δ=-1.72, CEBPB δ=-1.59, SPI1 δ=-1.57, FOS δ=-0.70), but CellOracle-predicted perturbation directions did not match experimental directions (40.9% agreement, not different from chance), revealing a fundamental gap between steady-state correlation and causal perturbation. Diagnostic analyses using VAE latent space profiling, correlation distribution comparison, and gene-gene graph analysis identified distinct failure modes in unsuccessful methods: VAE latent space competition (STAT3 signal-to-noise 0.44 vs. SPI1 4.25), correlation noise (TF-glycolysis |r|=0.038 indistinguishable from background |r|=0.047), and graph non-specificity (0.84× enrichment). A controlled ablation experiment showed that adding a GRN prior to DDIM did not improve target recall (delta=0 for all TFs), confirming that performance differences are multi-factorial. These findings establish preliminary guidance for method selection, including cross-pathway validation, direction-aware benchmarking, and minimum data requirements (≥500 cells, ≥1,000 HVGs). Author Summary Computational methods that simulate gene knockout experiments from single-cell RNA sequencing data are increasingly popular, but researchers lack guidance on which method to choose. We systematically tested eight such methods across four different cell types, including macrophages from osteoarthritis and rheumatoid arthritis patients, blood monocytes, and leukemia cells. We found that only two methods—CellOracle and DDIM—reliably detected how transcription factors control metabolism. These two methods also detected biologically coherent signals across multiple pathways, not just metabolism. Worryingly, two different methods applied to the same data could produce opposite conclusions about which genes regulate which pathways. We also observed that detecting a perturbation signal does not guarantee predicting its direction correctly: CRISPR-based experimental validation showed that computational methods captured which genes respond to TF perturbation but not whether they are upregulated or downregulated. Through systematic diagnostic analysis, we identified why unsuccessful methods failed: the key TF signals are too weak relative to background variation for purely data-driven approaches to detect. Based on our results, we recommend CellOracle for initial screening (requiring at least 500 cells and 1,000 highly variable genes), cross-pathway validation for any TF→target inference, and orthogonal experimental validation when perturbation direction matters. Our evaluation framework and practical guidelines help researchers choose perturbation methods appropriate for their specific biological questions.

Read PDF

Similar papers

Open access Sep 2026

Benchmarking methods for inferring single-cell transcription factor activity using large-scale perturbation sequencing data

A comprehensive evaluation of eight mainstream TFA inference methods using large-scale, high-quality single-cell perturbation sequencing (Perturb-seq) datasets shows that metaTF, which employs an integrated GRN, achieves the best performance across multiple metrics, including TF coverage, predictive accuracy for pertur...

Yue-Hui Zhu, Dong-Mei Han, Zhen Wang · 0 citations
Open access Sep 2026

Systematic benchmarking and optimal strategy selection of cross-species integration methods

Abstract Single-cell RNA sequencing provides an unprecedented resolution for cellular heterogeneity and gene regulation, fostering cross-species comparative analyses with increasing interspecies data. However, integrating single-cell transcriptomic data faces challenges, including gene selection, evolutionary distance,...

Ruo-Lin Wang, Jun-Juan Zheng, Chu-Ning Mao et al. · 0 citations
Review Open access Aug 2026

Revisiting differential expression analysis: An updated six-dimensional comparative study

A systematic and updated benchmarking framework for DE analysis is outlined, emphasizing a balance between accuracy and consistency, and a “BaGua (eight trigrams)” map summarizing the multi-dimensional performances of methods is provided.

Jian-Xiong Wu, Shaolei Lu, Hui Yao et al. · 0 citations
Open access Oct 2026

Dataset structure outweighs method choice in single-cell cell-type annotation

Automated cell-type annotation is a prerequisite for nearly all single-cell RNA-sequencing (scRNA-seq) analysis, and the proliferation of annotation tools — spanning marker-based, similarity-based, classical machine-learning, deep-learning, semi-supervised, large language model (LLM), and transformer foundation-model f...

Oliver Wardhana · 0 citations
Open access Sep 2026

Regulon-informed cellular representations reveal task-dependent generalization in drug combination prediction

Drug combination models must represent cellular context in a form that remains informative when the tested cells or compounds differ from those used for training. Transcription factor regulons offer a biologically structured representation, but their contribution can depend on the accompanying features and the intended...

Elizaveta Ignatova, M. Likhter · 0 citations
#artificial intelligence Preprint Sep 2026

AmbiModBench: Benchmarking Gene Perturbation Prediction Beyond Shared Responses

Predicting cellular responses to genetic perturbations helps prioritize experiments in single-cell genomics, where exhaustive measurement is infeasible. While computational models increasingly predict these responses, three evaluation deficiencies obscure what their scores demonstrate. First, absolute metrics cannot se...

Si-Kai Huang, Zhi-Wen Yang, Kai Yu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.