This work introduces ensemble refinement for scRNA-seq and scATAC-seq embeddings, inspired by ensemble methods from statistical machine learning, and implements BatchRefiner, a fast post-processing tool to enhance batch integration.
Abstract
Data from single-cell RNA sequencing (scRNA-seq) and the Assay for Transposase-Accessible Chromatin (scATAC-seq) are high-dimensional, sparse, and undesirably capture technical variability between experiments or batches. Many analysis methods thus seek to produce a low-dimensional cell-by-feature embedding space that groups together biologically similar cells across batches while distancing dissimilar cells. Here, we introduce ensemble refinement for scRNA-seq and scATAC-seq embeddings, inspired by ensemble methods from statistical machine learning, and implement BatchRefiner, a fast post-processing tool to enhance batch integration. We extensively benchmark widely-used scRNA-seq embedding methods on both batch integration and biological conservation over a wide range of datasets, before and after the addition of BatchRefiner. We extend these benchmarking approaches to provide the first comprehensive benchmark of batch integration for scATAC-seq embedding methods, including BatchRefiner. Importantly, we formalize a significance statistic, which we use to demonstrate BatchRefiner’s significant improvement in batch integration across a wide range of embedding methods, atlas-scale datasets, and established metrics.
Batch integration is a central preprocessing step in single-cell genomics, where datasets collected across experiments, donors, and protocols must be combined despite pervasive technical batch effects. The leading integration methods produce a shared low-dimensional embedding, which discards the corrected gene-expressi...
Single-cell RNA sequencing (scRNA-seq) technology has rapidly advanced in recent years, driving significant breakthroughs in developmental biology, cancer research, immunology, and other related fields. However, existing clustering methods still face performance bottlenecks when handling large-scale scRNA-seq data. To...
Hong-Yi Yuan, Chun-Yan Wang, Qiu-Cheng Sun et al.· PLoS ONE· 0 citations
Results indicate that discrete, anchor-constrained latent modeling provides a powerful and biologically coherent solution for unpaired single-cell multi-omics integration, and scCoA-VQA can accurately capture the meaningful regulatory structure.
Han Peng, Wu-Chao Liu, Yifang Cai et al.· Proceedings of the 32nd ACM...· 0 citations
Advances in single-cell RNA sequencing (scRNA-seq) enable the exploration of cellular heterogeneity at single-cell resolution. Accurate cell clustering is a prerequisite for delineating these identities; however, the inherent sparsity and prevalent dropout noise of scRNA-seq data pose major algorithmic challenges. Exis...
Xiang-Zhen Kong, Hui-Bo Tian, J. Shang et al.· IEEE transactions on computa...· 0 citations
ScGFormer is equipped with a biology-guided adaptive contrastive learning strategy, which is designed to account for zero inflation, balance class distributions, and refine dynamic graphs during training, thereby facilitating robustness and adaptability.
Ziqi Yuan, Hong-Wei Zhang, Cheng Liu et al.· IEEE transactions on computa...· 0 citations
Correcting batch effects and integrating single-cell sequencing datasets has been a crucial step in large-scale biological studies. Many methods have been published for this task, with complementary strengths in various scenarios. Our previous work, LIGER, leveraging integrative non-negative matrix factorization (iNMF)...
Yi-Chen Wang, Andrew Robbins, Gaurav T. Gadhvi et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.