Aug 2026· Genome Research· pp. gr.281981.126· 0 citations
Medicine
TL;DR
CARA is introduced, a cross-omics Bayesian framework that transfers cell type knowledge from scRNA-seq to scATAC-seq and detects novel cell types in single-cell DNA methylation data, demonstrating seamless extensibility to new modalities.
Abstract
Accurate cell type annotation is essential for revealing the dynamic, cell type-specific accessibility of regulatory elements from single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data. However, unlike the more mature single-cell RNA-sequencing (scRNA-seq) cell type annotation workflows, scATAC-seq cell type annotation remains challenging due to extreme sparsity, high dimensionality, the scarcity of labelled scATAC references, and pronounced batch effects across datasets. To enable annotation without relying on extensive scATAC labels, we introduce CARA, a cross-omics Bayesian framework that transfers cell type knowledge from scRNA-seq to scATAC-seq. CARA simultaneously enables cell type annotation, batch correction, and detection of cell types absent from the reference by combining pretraining on scRNA-seq data and semisupervised learning on target scATAC-seq data, along with distribution alignment, dynamic class weighting, and novel cell type detection. Across diverse benchmark datasets, CARA consistently outperforms baseline methods, remaining robust to batch effects. CARA's cross-modal alignment and robust annotation strategy preserve fine-grained lineage structure, enabling reconstruction of the hematopoietic differentiation trajectory. Through multidimensional uncertainty and latent-space clustering, CARA identifies novel, rare, or aberrant populations outside the reference cell type space, providing candidates for further biological validation and perturbation. Using an omics-specific generative framework and distribution alignment, CARA delivers accurate knowledge transfer and detects novel cell types in single-cell DNA methylation data, demonstrating seamless extensibility to new modalities. Ultimately, CARA offers a powerful and flexible solution for cross-modal cell type annotation in complex single-cell settings, facilitating the discovery of novel cell types and mechanistic insight into cell type-specific regulation across diverse analyses.
Single-cell chromatin accessibility (scATAC-seq) profiles genome-wide regulatory elements that shape immune cell identity and function, but its interpretation is currently limited by low cell type resolution and small reference datasets. Existing datasets annotate fewer than 20 immune cell types and are too coarse to resolve heterogeneity and characterize cell type-specific gene regulatory programs and functions. Here, we present a large-scale scATAC-seq resource that substantially improves immune cell annotation and regulatory inference.
By integrating matched-donor scRNA-seq and scATAC-seq data from human peripheral blood mononuclear cells (PBMCs) with trimodal TEA-seq (single-cell ATAC, RNA, and surface protein), we classified 36 immune cell types, including 4 myeloid, 6 B cell, 5 NK cell, 6 CD4 T cell, and 15 CD8 T cell subtypes. Cell frequencies from published scRNA-seq and new scATAC-seq labels were highly correlated (median ρ = 0.84). Labels were applied to our longitudinal multi-modal dataset of 206 samples spanning over 3 million PBMCs from 78 healthy human donors.
We used these annotations to define baseline epigenetic states, age-associated differences, and epigenetic changes following influenza vaccination. Our analysis revealed extensive sets of differentially accessible tiles and enriched transcription factor motifs that define cell type-specific regulatory identities. Linking these chromatin regions and transcription factors to differentially expressed target genes enabled the construction of gene regulatory circuits associated with cell type, aging, and vaccination. Additionally, we trained a classification model for high resolution cell type labeling and doublet detection in new scATAC-seq datasets.
Together, this multi-modal atlas and associated cell type-labeling model provide an unprecedented reference for immune cell gene regulatory circuits and a valuable resource for exploring the epigenome of human immune cells.
n/a
Computational and Systems Immunology (COMP)
Sydney Kuhl, Upaasana Krishnan, A. Tjaernberg et al.· Journal of Immunology· 0 citations
Single-cell and single-nucleus RNA sequencing (scRNA-seq and snRNA-seq) have transformed cardiovascular biology by resolving cellular heterogeneity and disease-specific cell states. The interpretive power of these technologies, however, hinges critically on accurate cell-type annotation, the assignment of biologically meaningful labels to clusters or individual cells. Mis-annotation risks systematic bias in mechanistic inference, obscures rare populations, and undermines cross-study reproducibility. This Review systematically integrates current conceptual frameworks, computational methodologies, and domain-specific best practices for cell-type annotation in cardiovascular single-cell research. We first examine the biological principles that underpin annotation, including the definition of cell types, the distinction between cell states and transitional phenotypes, considerations of granularity, and ontology mapping. We then delineate comprehensive analytical workflows spanning preprocessing and quality control (empty droplet filtering, doublet detection, normalization, and batch correction), dimensionality reduction (PCA, UMAP, t-SNE), and graph-based clustering. Turning to annotation strategies, we survey manual marker-based curation with canonical cardiovascular markers, reference-correlation methods (SingleR), supervised machine-learning classifiers (CellTypist, SingleCellNet, Garnett, scBERT), reference-mapping and label-transfer platforms (Azimuth, scArches, Symphony), and hybrid probabilistic frameworks (scANVI). We also summarize the software ecosystems of R/Seurat and Python/Scanpy/scvi-tools. Benchmarking evidence indicates that annotation accuracy depends more on reference data quality than on algorithmic sophistication. SingleR, CellTypist, and Azimuth have emerged as leading performers, whereas ensemble consensus approaches that integrate two or more independent methods enhance robustness. We further highlight cardiovascular-specific challenges, including modality-dependent differences in cardiomyocyte representation between scRNA-seq and snRNA-seq, vascular and stromal heterogeneity, immune cell tissue adaptation, disease-induced transitional cell states, and barriers to cross-species translation. Finally, we present a ten-step best-practice workflow, a comprehensive reporting and reproducibility checklist, and illustrative case studies drawn from the Adult Human Heart Atlas and CardioAtlas. We offer practical recommendations and discuss future directions encompassing multimodal integration, spatial transcriptomics, and AI-assisted annotation, all aimed at ensuring reproducible and interpretable annotations across laboratories [1-5].
Lu Sun, Li Ma, Li-Zhi Chen et al.· International Journal of Bio...· 0 citations
Cell type annotation is essential for gaining biological insight from single-cell RNA sequencing data, yet manual labeling remains time-consuming and difficult to reproduce. Various computational approaches have been developed to automate this process, and recent studies suggest that large language models can infer cell types with promising accuracy in single-cell analysis. However, most workflows still rely on cluster-specific markers derived from gene expression alone or manual curation. As a result, marker selection can be sensitive to statistical criteria and dataset-dependent bias, which may lead to the selection of less informative genes or missing important markers, while providing limited biological context. To address this limitation, we introduce CELLIA, an LLM-based workflow for automated and robust cell type annotation. CELLIA employs an integrative evidence-knowledge marker selection strategy that combines statistical differential expression criteria with curated tissue-specific marker resources to identify informative marker genes. In benchmarking analyses of 102 cell types, this approach improved agreement with manual annotations. In addition, CELLIA achieved higher agreement in subtype-level analyses of closely related immune populations and was further evaluated in a non-immune stromal subtype setting, covering 25 cell types in total. By integrating evidence-knowledge from gene expression with curated biological prior knowledge, CELLIA provides a more stable marker selection and improves the reliability of LLM-cell type annotation.
Jiyeong Shin, Soyoung Jeong, Hyun Je Kim et al.· BioData Mining· 0 citations
Accurate identification of cell types and states is essential for reliable single-cell RNA-sequencing analyses, yet current methods remain sensitive to continuous biological states, data preprocessing choices, and reference selection. Here we present ClustoCell, a reference-free method that resolves cell identity using within-cell transcriptional architecture. By stratifying gene expression of each cell into high and medium tiers, ClustoCell constructs cell-cell similarity graphs that prioritize intrinsic expression structure over global variance. Across 450 datasets spanning over 24 million cells, ClustoCell recovered expert annotations with high concordance (92%). Benchmarked against state-of-the-art methods, ClustoCell identifies more stable and coherent cell types and states, avoids excessive partitioning of closely related cells, and improves the identification of cell type-specific markers. From transcriptional structure alone, ClustoCell resolves rare and transitional cell states, distinguishes malignant from non-malignant cells, and refines expert cell annotations. Applied to immunotherapy datasets, ClustoCell uncovered coordinated pre-treatment immune circuits linking T cell states to PD-1 responsiveness in a tumour-type-specific manner. ClustoCell provides an interpretable and scalable foundation for single-cell analysis and translational profiling.
Abbas Salavaty, M. Foroutan, N. Pretel et al.· bioRxiv· 0 citations
Single-cell RNA sequencing profiles cellular heterogeneity at atlas scale, making automated annotation essential. However, target datasets often contain novel cell types missing from incomplete references. We present scOLAR, an ontology-guided open-set framework that learns prototypes over the Cell Ontology and uses both reference and target expression to annotate known classes while detecting unfamiliar populations. Guided by ontology hierar-chies and decision-boundary regularization, scOLAR penalizes coarse-lineage misclassification and groups novel cells without requiring predefined cluster counts. Across benchmarks, scOLAR achieves a novelty-detection AUROC of 0.9726 and an average precision of 0.9871, enabling structured post-hoc lineage-level interpretation of populations absent from the reference.