Skip to content
Open access

Annotation-free phenotype prediction using knowledge-augmented clustering from single-cell RNA sequencing data

Jul 2026 · Briefings in Bioinformatics · Vol 27 · 0 citations · 37 references
Medicine

TL;DR

Evaluated across three public scRNA-seq datasets, scCap consistently outperforms baseline models in predictive accuracy and identifies disease-associated subpopulations previously reported in the literature without relying on predefined cell-type annotations.

Abstract

Abstract Single-cell RNA sequencing has emerged as a transformative tool, enabling precise phenotype prediction and the detailed identification of disease-associated cell subpopulations. However, many existing computational approaches still rely on predefined cell-type annotations during model training. This dependence makes their predictive performance highly sensitive to subjective annotation quality, labeling inconsistencies, and dataset-specific biases, ultimately hindering their generalizability across diverse patient cohorts. To address these challenges, we propose scCap, an annotation-free framework that leverages knowledge-augmented clustering for robust phenotype prediction. Specifically, the framework first constructs initial clusters from raw gene expression profiles and subsequently refines them within the embedding space of a pretrained single-cell foundation model, allowing the clusters to better reflect broader biological organization while preserving fine-grained cellular heterogeneity. The resulting knowledge-augmented clusters are then integrated into a hierarchical multiple instance learning framework with dual-level attention, enabling interpretable predictions at both the cell and cluster levels. Evaluated across three public scRNA-seq datasets, scCap consistently outperforms baseline models in predictive accuracy. Furthermore, scCap identifies disease-associated subpopulations previously reported in the literature without relying on predefined cell-type annotations. These results demonstrate that scCap provides a robust and interpretable framework for annotation-free phenotype prediction.

Read PDF

Similar papers

Jul 2026

Functionally Guided Graph Learning for Robust Cross-Patient Cell-Type Annotation in Single-Cell RNA Sequencing.

Cross-patient cell-type annotation in single-cell RNA sequencing (scRNA-seq) remains challenging due to pronounced interpatient heterogeneity and distribution shifts across patient-specific cellular contexts. Conventional annotation approaches often rely on proximity-driven graph construction or expression similarity, which may introduce spurious cell-cell connections and lead to unstable knowledge transfer across patients. To address this limitation, we propose PathoGraph, a functionally guided graph learning framework for robust cross-patient cell-type annotation. The proposed method integrates KEGG-7-based biosemantic graph structure learning with cross-patient representation adaptation. Specifically, pathway-derived functional semantic profiles are incorporated to refine patient-specific cell graphs, encouraging biologically coherent neighborhoods and suppressing noise introduced by purely expression-based similarity. Based on the refined graphs, a cross-patient representation adaptation mechanism further aligns embeddings between labeled reference patients and unlabeled query patients to facilitate reliable annotation transfer. Experiments on three cross-patient scRNA-seq data sets, including leukemia, breast invasive carcinoma, and colorectal cancer data sets, demonstrate that PathoGraph achieves stable annotation performance across 32 directed reference-to-query transfer tasks. Across all tasks, PathoGraph obtained an average ACC of 84.28% and an F1-score of 84.08%, showing competitive and stable performance compared with representative marker-based, correlation-based, and model-based annotation methods. Ablation studies further show that removing the biosemantic graph learning module reduces the average accuracy to 83.48%, highlighting the importance of functional-guided graph refinement. In addition, post hoc functional relevance analyses in immune-cell and cancer-associated contexts suggest that the learned cell-cell graphs capture biologically relevant neighborhood structures beyond expression-driven proximity. The source code and processed data are publicly available at: https://github.com/LiYuechao1998/PathoGraph.

Yue C. Li, Mengmeng Wei, Xinfei Wang et al. · 0 citations
Open access Jul 2026

PRISM: Prior-enhanced Inference for Spatial Transcriptomic Cell Type Mapping

Abstract Motivation Cell type annotation in spatial transcriptomics (ST) is fundamental for deciphering complex tissue organization and spatially resolved biological processes. Most existing methods perform ST cell type annotation by transferring labels from single-cell RNA-seq (scRNA) data to ST data, but typically rely on weakly constrained representations that neglect structured spatial dependencies and treat marker gene selection as an isolated preprocessing step. This renders them vulnerable to substantial domain gaps as well as platform-specific noise, resulting in unstable predictions and limited biological interpretability. Results To address these issues, we propose Prior-enhanced Inference for Spatial Transcriptomic Cell Type Mapping (PRISM), a novel three-stage framework integrating biological prior construction, pseudo-label generation, and multi-level ST refinement. First, PRISM constructs a cross-domain biological prior to explicitly extract marker genes to enforce positive biological discriminability. Next, it adopts a prior-enhanced self-training strategy, where scRNA-trained ensembles generate reliable pseudo-label candidates for ST data, serving as a robust anchor for cross-domain adaptation. Finally, the framework consolidates high-quality ensemble predictions selected via metric-guided evaluation, encodes spatial information, and optimizes the model under dual-directional biological constraints. Extensive experiments on eleven ST datasets across six platforms, two species, and multiple tissue contexts validate PRISM. Specifically, on the five labeled benchmarks, PRISM shows strong overall performance under both Accuracy and Macro-F1 evaluation across brain and non-brain tissues. Moreover, under fully label-free settings, PRISM achieves the best overall composite rank across all datasets, demonstrating strong robustness to domain shift and platform heterogeneity. Availability and implementation PRISM is available at https://github.com/lilab-ai4s/PRISM and https://doi.org/10.5281/zenodo.20529683.

Yiheng Xu, Xuehao Wang, Shuqi Liu et al. · 0 citations
Open access Jul 2026

Deep interpretable learning of sample representations for characterizing disease states in single-cell transcriptomics

Single-cell transcriptomics technology offers unprecedented insights into molecular heterogeneity. However, capturing sample-level representations that reflect both systemic and cellular states remains challenging, especially when disease annotations are mostly available as coarse sample-level labels. Here, we introduce Phenoverse, an interpretable deep learning framework that learns sample-level disease state representations through cell type-aware residual encoding, prototype learning, and Perceiver-based aggregation. Applied to independent single-cell transcriptomic cohorts of COVID-19, Alzheimer’s disease, and systemic lupus erythematosus, totaling over 5 million cells, we demonstrate that learned sample representations enable disease state prediction and encode a continuous spectrum of disease severity on unseen data that correlate with multiple clinical and pathological measures, despite being trained solely on binary phenotype labels. Further, we demonstrate that trajectory-derived genes reveal cross-cohort molecular programs and show consistently higher reproducibility than traditional case-control comparisons. Finally, prototype learning provides intrinsic model interpretability and enables the characterization of cell type-specific disease states. Taken together, Phenoverse offers an interpretable disease-phenotyping approach to dissecting sample heterogeneity, and our results highlight its utility in translating complex single-cell transcriptomic data into patient-level biological insights.

Manoj M Wagle, Yongheng Wang, Soham Samanta et al. · 0 citations
Open access Jul 2026

FloREN: Decoding Immune Regulatory Networks through Interpretable Graph Transformer Patient Representations

A Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method that enables improved sample stratification and biomarker discovery and supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).

Iñigo Clemente‐Larramendi, S. Hillion, D. Cornec et al. · 0 citations