A Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method that enables improved sample stratification and biomarker discovery and supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).
Abstract
Single-cell RNA sequencing (scRNA-seq) enables detailed characterization of cellular heterogeneity, yet understanding the full cellular and regulatory environment of complex tissues remains challenging. In the era of large single-cell atlases, this technology has become increasingly accessible, and datasets have grown in scale and statistical power. As a result, sample representation methods have emerged as a promising strategy to summarize patient-level biological variation. However, most existing approaches rely on unsupervised learning frameworks with ambiguous biological interpretability. Here we present a Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method. FloREN models single-cell data as a heterogeneous network integrating cells and genes together with gene regulatory and cell-cell communication relationships. Through condition-aware embeddings and interpretable attention networks, FloREN enables improved sample stratification and biomarker discovery. In addition, the framework supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).
Single-cell transcriptomics technology offers unprecedented insights into molecular heterogeneity. However, capturing sample-level representations that reflect both systemic and cellular states remains challenging, especially when disease annotations are mostly available as coarse sample-level labels. Here, we introduce Phenoverse, an interpretable deep learning framework that learns sample-level disease state representations through cell type-aware residual encoding, prototype learning, and Perceiver-based aggregation. Applied to independent single-cell transcriptomic cohorts of COVID-19, Alzheimer’s disease, and systemic lupus erythematosus, totaling over 5 million cells, we demonstrate that learned sample representations enable disease state prediction and encode a continuous spectrum of disease severity on unseen data that correlate with multiple clinical and pathological measures, despite being trained solely on binary phenotype labels. Further, we demonstrate that trajectory-derived genes reveal cross-cohort molecular programs and show consistently higher reproducibility than traditional case-control comparisons. Finally, prototype learning provides intrinsic model interpretability and enables the characterization of cell type-specific disease states. Taken together, Phenoverse offers an interpretable disease-phenotyping approach to dissecting sample heterogeneity, and our results highlight its utility in translating complex single-cell transcriptomic data into patient-level biological insights.
Manoj M Wagle, Yongheng Wang, Soham Samanta et al.· bioRxiv· 0 citations
Precise resolution of cellular heterogeneity within complex tissues is fundamental to deciphering disease etiologies from bulk transcriptomic profiles. While computational deconvolution offers a scalable alternative, current deep learning methods predominantly operate as "black boxes," neglecting the structural constraints of biological laws. This reliance on purely data-driven feature extraction often yields biologically incoherent predictions and limited mechanistic interpretability. iDCF (Interpretable Deconvolution of Cell Fractions) is a novel framework that enforces biological topology onto deep neural networks. The iDCF architecture employs a dual-stream design, synergizing a standard deep network with a knowledge-based sparse neural network (KSNN) explicitly masked by pathway definitions and protein-protein interaction (PPI) networks. In comprehensive benchmarks, iDCF achieves top-tier performance, consistently ranking among state-of-the-art methods in accuracy and robustness. iDCF integrates the SHapley Additive exPlanations (SHAP) framework, bridging the gap between computational inference and biological intuition. The model's decision logic is governed by established biological mechanisms rather than spurious statistical correlations, validating its reliability. Validations across clinical contexts, including Alzheimer's disease, ovarian cancer, and diabetes, demonstrate iDCF's ability to recover disease-relevant cellular dynamics. iDCF offers a high-performance, interpretable, and biologically grounded tool for deconvolving cell-type proportions, facilitating deeper insights into tissue heterogeneity in health and disease.
Hongming Guo, Tingfang Wu, Wen-Zheng Wang et al.· PLoS Computational Biology· 0 citations
In the tumor microenvironment, cell's state is influenced by cell-cell interactions (CCIs) with neighboring cells in its niches. Identifying dysregulated CCIs that are associated with pathogenic process pinpoints targets for drug discovery. Imaging-based spatial transcriptomics and single-cell RNA sequencing provide, respectively, single-cell spatial information and transcriptome-wide measurements needed to study CCIs, but neither modality provides both. Existing spatial transcriptomics foundation models also cannot effectively learn from spatially resolved single-cell data with full-transcriptome coverage, explicitly infer the CCI mechanisms driving cell state-niche associations, or interpretable enough to support direct biological interpretations. Here, we present GITIII-scale, a hierarchical, interpretable pan-cancer spatial transcriptomics foundation model for TME representation learning that investigates cell state-niche associations and their underlying ligand-receptor (LR) signaling pathways. GITIII-scale uses transformers to model interactions between pairs of cells at defined spatial distances, an interpretable single-layer graph transformer without a feed-forward network to decompose how each gene in a receiver cell is influenced by each neighboring sender cell, and a graph transformer to generate cellular-neighborhood embeddings. Trained on our assembled pan-cancer database of specimen-matched scRNA-seq and imaging-based spatial transcriptomics datasets, GITIII-scale generated TME embeddings that recovered niche-associated state changes more accurately than existing spatial transcriptomics foundation models in cancer types unseen during training. A case study of an unseen breast cancer dataset further demonstrated the model's interpretability by identifying potentially drug-targetable LR pathways associated with endothelial overgrowth and tumorigenesis.
Xiaohui Xiao, Jia-Shu He, Shiyang Zhang et al.· 0 citations
Gene regulatory networks (GRNs) describe regulatory interactions between transcription factors and their target genes and are essential for understanding cellular processes and disease mechanisms. Recent advances in single-cell RNA sequencing (scRNA-seq) have enabled data-driven GRN inference at single-cell resolution. However, the high sparsity and noise inherent in scRNA-seq data pose substantial challenges for accurately recovering regulatory relationships. Existing graph neural network (GNN)-based approaches often rely on localized message passing, which can lead to over-smoothing and limited modeling of long-range regulatory dependencies. To address these limitations, a structure-aware interleaved-attention graph learning framework, termed IAGRN, is proposed for GRN inference from scRNA-seq data. Specifically, it interleaves topology-constrained local attention with distance-aware global attention, enabling effective integration of structural priors and long-range regulatory signals. Graph Laplacian positional encoding is further incorporated to preserve topological information and enhance node representations. Evaluations on seven public benchmark datasets demonstrate that IAGRN consistently improves GRN reconstruction under highly sparse conditions and achieves competitive performance compared with existing approaches.
Yue Wang, Sicheng Tian, Dan Li· International Journal of Mol...· 0 citations
Motivation Single-cell RNA sequencing (scRNA-seq) has become an attractive tool for studying complex diseases, in which transient cell states affecting diverse cell populations characterise disease development and progression. However, due to data sparsity and disease heterogeneity analysis is often challenging. With recent advances in machine learning, two widely used approaches have emerged for learning cellular representations: large-scale foundation models and biological knowledge-guided methods. Despite their complementary strengths, there is currently no unified workflow for systematically comparing and integrating these approaches. Results Here, we present scRepresenter, an open-source workflow for computing, integrating, and validating cellular embeddings derived from foundation models and biological knowledge-guided methods in the context of complex diseases. It consists of two components: a command-line workflow that computes cellular embeddings and performs downstream analyses, and an interactive Shiny application for visualizing and comparing the computed embeddings. scRepresenter supports four categories of cellular representations: (1) expression-based, (2) knowledge-guided, (3) foundation model-derived, and (4) hybrid embeddings that combine foundation model-derived representations with knowledge-guided representations. This approach takes a cell-by-gene count matrix as input and outputs an integrated object containing the computed embeddings. Then, this object can be uploaded into our interactive Shiny application to compare different embeddings. Availability The workflow is available at https://github.com/GuilhermePocas/scRepresenter Contact AL291@cam.ac.uk; MA2129@cam.ac.uk
Guilherme Pocas, Muhammad Umar, Oliver Davis et al.· bioRxiv· 0 citations
Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes. Here we present BioM-JEPA, a joint-embedding predictive architecture that instead predicts aggregate representations of graph-connected gene blocks defined by protein-association and corpus-derived coexpression evidence. A student network infers each target-block representation from the remaining genes in a cell, while a slowly updated teacher supplies the corresponding target from the full observed gene set. Under the reported extraction procedure, block-level prediction produced embeddings with higher effective rank and weaker association with detected-gene depth in the tested diagnostics than token-prediction, random-block and reconstruction controls. Across CellBench tasks, frozen BioM-JEPA embeddings retained expression, pathway and neighbourhood information and achieved the lowest aggregate perturbation-response error among the evaluated models. Representation diagnostics were also consistent with canonical pancreatic programmes and compositional relationships between genetic perturbations. Linear attention avoids constructing a quadratic gene-by-gene attention matrix; in a matched one-epoch hPancreas experiment at batch size 8, BioM-JEPA provided 5.75-fold higher fine-tuning throughput and 3.76-fold higher held-out embedding throughput than scFoundation. Together, these results support graph-connected gene blocks as useful prediction units for JEPA-style representation learning in single-cell biology.
Yuhao Wang, Zelin Zang, Yuxuan Liu et al.· 0 citations