Skip to content

Author

Xiaobo Sun

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

TCellAlign: Cross-study T-cell Populations Alignment with Nomenclature-Guided Multi-Agent Workflow

Cell type standardization plays a central role in integrating biological knowledge across single-cell studies. While standardized resources (e.g., Cell Ontology, Nomenclature Frameworks) provide unified vocabularies of cell populations, scientific publications and public datasets continue to use heterogeneous study-specific labels, making cross-study comparison difficult even when biologically equivalent cell populations are described. In this work, we are the first to formulate this challenge as an evidence-grounded cell population alignment problem and propose TCellAlign, a multi-agent framework that includes literature retrieval, information extraction, nomenclature-guided label alignment, and evidence-based adjudication. This modular design preserves the original terminology and supporting evidence reported by each study while producing standardized labels that can be compared across studies. We further construct a manually validated benchmark dataset linking study-specific labels, CZ CELLxGENE annotations, and standardized T-cell nomenclature across 44 manually curated, published studies (including over seven million cells) spanning four biological categories: healthy, cancer, infectious disease and inflammatory diseases. Across the evaluated tasks, TCellAlign achieves stronger semantic agreement than ontology-based baselines and maintains transcriptomic coherence with both open-source and closed-source large language models (LLM) backbones. By connecting literature, datasets, and expert's nomenclature, TCellAlign enables consistent interpretation of T-cell subtypes and states across studies, facilitating biological knowledge integration and the development of future foundation models built upon standardized cellular representations.

Peng Xie, Rongjia Zhou, Zhi-Li Ou et al. · 0 citations
Book Open access Aug 2026

DiffPro: Decoupled Generative Prior with Diffusion Models for Efficient Few-Shot Drug Synergy Prediction

Predicting drug synergy is essential for optimizing combination therapies in cancer treatment. Under extreme data scarcity, existing computational methods struggle to generalize to new cell lines. Although meta-learning approaches have shown promise, a critical limitation lies in their reliance on a unimodal Gaussian prior for task representation. In extreme few-shot settings, this assumption can overly pull the task posterior toward the prior mean, reducing the discriminability of task representations and pushing the model toward a generic mean-value predictor. To overcome this, we propose DiffPro, a novel framework that integrates a Latent Diffusion Model (LDM) as a structural prior. Unlike Gaussian-based methods, our diffusion module learns a flexible, data-driven prior that can model highly complex task distributions. This learned prior better reflects task heterogeneity in few-shot drug synergy prediction under extreme data scarcity, yielding more discriminative task representations. Comprehensive experiments on the DrugComb dataset demonstrate that DiffPro achieves a significant improvement in performance. Notably, in the challenging 5-shot setting, it outperforms the best-performing baseline, delivering approximately a 17.3% relative gain in R2 (0.264 vs. 0.225) and a 4.6% relative reduction in MSE. These results confirm that the diffusion prior successfully regularizes the latent space, mitigating task representation collapse and enabling robust, high-precision predictions for combination therapy design.

Shuting Jin, Xu Guo, Anqi Huang et al. · 0 citations
Book Open access Aug 2026

Concord: Building Consensus Representations for Single Cells with Collaborative Random Projection

Foundation models (FMs) have recently transformed single-cell genomics by learning transferable representations from large-scale single-cell data, enabling a wide range of downstream biomedical applications. Inspired by natural language processing, existing single-cell FMs adapt transformer architectures by treating genes as tokens and cells as sequences. However, transformers are inherently agnostic to input order, while genes lack a natural sequential structure. Current approaches rely on heuristic strategies, such as expression-based gene sorting, to impose positional information, which often fail to capture relative relationships among collectively expressed genes and between gene identities and their expression levels, leading to information loss and limited generalization. In this work, we propose Concord, a novel single-cell FM that explicitly models relationships between gene identities and expression levels through a collaborative rotary attention (CRA) mechanism. Specifically, Concord employs two collaborative attention modes: a gene-to-expression rotational attention that produces gene-enhanced expression representations, and an expression-to-gene rotational attention that yields expression-enhanced gene representations. These two processes provide distinct yet complementary views of the same cell-level expression profile; accordingly, we employ contrastive learning to align their semantic representations. Our theoretical analysis demonstrates that CRA effectively leverages the relative distances in one embedding space to establish stable and consistent dependencies in the other modality. Extensive experiments on various single-cell datasets demonstrate that Concord outperforms existing FMs across various downstream tasks and provides more informative and transferable gene and cell representations. Our code is available at https://github.com/Catchxu/Concord.

Kaichen Xu, Mianpeng Liu, Hao Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.