Skip to content
Open access

Deconvolution-derived cell-type expression targets for personal genome sequence-to-expression prediction

Jul 2026 · bioRxiv · 0 citations · 31 references
Medicine Biology

TL;DR

Results support deconvolution as a feasible approach for generating cell-type-resolved training targets, while showing that target quality and limited cohort size remain important constraints.

Abstract

Sequence-to-function models learn regulatory features from genomic sequence, but they remain limited in their ability to predict gene-expression differences among individuals. Cell-type-specific regulatory effects may be obscured in bulk RNA sequencing, whereas paired genotype and single-cell expression cohorts remain small. We evaluated whether deconvolution of bulk RNA-seq could provide scalable cell-type-specific targets for personal-genome expression prediction. GTEx v8 bulk RNA-seq from six tissues was deconvolved with BayesPrism using single-nucleus reference profiles, producing targets across 83 tissue–cell-type contexts. Deconvolved expression agreed with matched pseudobulked GTEx single-nucleus RNA-seq, with median donor-level Pearson correlations across genes ranging from 0.53 to 0.73 by tissue. We compared genotype-feature models, regressors trained on frozen Enformer representations, and fine-tuned Enformer and Borzoi models. Across random and nonlinear-enriched gene sets, sequence-derived approaches generally outperformed genotype-feature baselines, while frozen Enformer features were competitive with end-to-end fine-tuning. For the random gene set, Fisher-averaged Pearson correlations were 0.122–0.142 for sequence-derived approaches and 0.081–0.086 for genotype-feature baselines in a coverage-aware sensitivity analysis. Model performance was positively associated with deconvolution–pseudobulk agreement for sequence-derived models (r = 0.35–0.43 across tissue–cell-type contexts), suggesting that target reliability may constrain downstream prediction. Context-specific Enformer fine-tuning did not materially out-perform a shared, combined-context strategy. These results support deconvolution as a feasible approach for generating cell-type-resolved training targets, while showing that target quality and limited cohort size remain important constraints. Frozen pretrained representations provide a computationally efficient and competitive baseline for personal sequence-to-expression modeling.

Read PDF

Similar papers

Open access Aug 2026

Mapping enhancer–gene regulatory interactions from single-cell data

A family of classification models, scE2G, is introduced that predict enhancer–gene regulatory interactions from single-cell datasets and enable mapping of these interactions across diverse cell types and tissues and will enable accurate mapping of enhancer–gene regulatory interactions across thousands of human cell types.

Maya U. Sheth, Wei-Lin Qiu, X. Ma et al. · 1 citation
Open access Sep 2026

CellMAGE: cell-type deconvolution for multi-parent population analysis of gene expression

Single-cell RNA-sequencing remains prohibitively expensive for multiparental population (MPP) studies. Existing deconvolution methods treat bulk RNA-seq as genetically anonymous mixtures, but in MPPs, the proportional contribution of each parental strain to each progeny’s transcriptome is already known. CellMAGE (Cell-type deconvolution for Multi-parent Analysis of Gene Expression) weights parental cell-type profiles by each progeny’s known genetic composition, requiring no model training and no minimum sample size. Validated in 16 Diversity Outbred mice across 12 prefrontal cortex cell types and 23,116 genes, predicted and measured gene expression were statistically equivalent (±0.05) in all cell types (pooled Spearman ρ = 0.923, 95% CI: 0.910, 0.934). CIBERSORTx required 96 additional samples to resolve at most 12.4% of genes and only 3 cell types; CellMAGE outperformed it even within this restricted comparison (per-cell-type median ρ: 0.908-0.974 vs. 0.353-0.641). CellMAGE is applicable to any MPP with parental single-cell data, including diploid crop MAGIC populations. Article summary Identifying which cell types manifest genetic effects on complex traits requires cell-type-specific gene expression data, but single-cell sequencing is prohibitively expensive for population genetic sample sizes. Ball et al. present CellMAGE, a computational method that delivers single-cell fidelity from bulk sequencing in multiparental populations, at bulk sequencing cost. Because each individual’s genotypes are known, CellMAGE weights parental cell-type profiles by genetic composition directly, requiring no model fitting. Validated in Diversity Outbred mice across 12 brain cell types, CellMAGE predictions were highly accurate, outperformed an existing method, and apply to any multiparental population with available parental single-cell data, including crop species.

Robyn L. Ball, Alyssa Klein, Ashley A. Auth et al. · 0 citations
Open access Aug 2026

A statistical framework for disease classification with scRNA-Seq Data

This work introduces a two-stage statistical framework for interpretable patient-level disease classification from single-cell data, and recovered biologically coherent, cell-type specific gene signatures consistent with known disease mechanisms, demonstrating improved interpretability without sacrificing predictive accuracy.

Zhi-Wei Xiao, William Torous, Jeffrey B. Cheng et al. · 0 citations
Open access Aug 2026

Integrated inference of cellular compositions and gene expression programs by deconvolution

While computational deconvolution is routinely used to estimate cell-type proportions from tissue mixtures, reconstructing cell-type-specific transcriptomes at single-sample resolution remains a fundamentally underdetermined algorithmic challenge. Consequently, accurate single-sample, gene-level inference is rarely achieved by existing tools. Here, we systematically benchmarked multiple deconvolution approaches across diverse biological contexts using both pseudo-bulk mixtures and real bulk RNA-seq datasets derived from multiple tissues. Evaluating the critical computational limitations in these models, we developed BayesPrism-DWLS, a framework that enables integrated inference of cell-type proportions and cell-type-specific expression at single-sample resolution. Applied to mouse colon bulk RNA-seq and spatial transcriptomics, BayesPrism-DWLS revealed cell-type-specific genes and pathways that were undetectable at the bulk or spot level. Therefore, this framework provides a robust, high-resolution tool for dissecting cell heterogeneity and supports mechanistic studies informed by cell-type-specific transcriptional programs using cost-effective sequencing data.

Ze Zhang, Xu Wang, Fan Hong et al. · 0 citations
Open access Aug 2026

Systematic assessment of the biological impact of cellular deconvolution on downstream analyses of disease transcriptomes

Background Cellular deconvolution methods estimate cell-type proportions from bulk RNA-seq data, typically using single-cell RNA-seq–derived signatures, enabling separation of disease-associated transcriptional changes into composition-driven and cell-intrinsic effects. However, these approaches depend on model assumptions and the stability of cell-type signatures, and it remains unclear how deconvolution-related uncertainties influence downstream analyses and biological conclusions. Results We systematically evaluated the effect of cell-type correction on disease-relevant transcriptomic insights, using Alzheimer’s disease (AD) as a model and the Mount Sinai Brain Bank cohort as a primary dataset. Applying dtangle, selected after comparison with another deconvolution approach, we estimated cell-type proportions across four brain regions and assessed how correction reshaped differential gene expression and pathway enrichment. Cell-type correction (CTC) markedly altered differentially expressed gene (DEG) profiles in a region-dependent manner: the superior temporal gyrus lost all significant signals, while the frontal pole gained DEGs with improved cross-region concordance. At the pathway level, correction shifted enrichment from synaptic loss and immune activation toward suppression of stress-response and immune regulatory programs, suggesting that composition changes partly obscure cell-intrinsic regulatory signals. Overlap with AD genome-wide association study loci and replication in an independent cohort indicated that cell-intrinsic changes are more consistently validated than composition-driven changes. Notably, KCNN2 and RIMS1, not currently recognized as canonical AD biomarkers, emerged as robust transcriptional signatures, potentially reflecting both composition-driven and cell-intrinsic dysregulation and warranting further investigation. Conclusions Parallel evaluation of uncorrected and CTC analyses distinguishes composition-driven from cell-intrinsic transcriptional effects and highlights robust disease signatures in heterogeneous tissues such as the brain.

Sanga Mitra, Maziya Ibrahim, Manikandan Narayanan · 0 citations
Open access Sep 2026

Network-informed deconvolution of bulk immune gene co-expression reveals single-cell programs and spatial organization

Introduction Single-cell RNA sequencing (scRNA-seq) has opened unprecedented possibilities to explore the complexity of the immune system. However, existing methods primarily rely on expression-based clustering analysis, which lacks mechanistic explanations for immune cell states and encounters challenges in integrating multi-scale data. Methods We developed a network-informed deconvolution framework that constructs Bayesian network-derived regulatory structures using immune-related genes from context-matched bulk RNA-seq datasets. Network markers were extracted from these structures and projected onto peripheral blood mononuclear cell (PBMC) and lung adenocarcinoma (LUAD) scRNA-seq datasets to identify network biomarkers and define immune cell states. Spatial transcriptomic analysis was further used to evaluate the spatial coherence of network-defined cell states. The scRNA-seq and spatial transcriptomic datasets analyzed in this study were generated from prospectively collected samples by our team, while context-matched bulk RNA-seq cohorts were used to derive population-level immune gene network structures. Results The framework identified structure-defined immune subpopulations in both PBMC and LUAD datasets and revealed functional heterogeneity across multiple immune lineages. Spatial transcriptomic analysis further showed that network-associated immune clusters exhibited closer spatial proximity than non-associated clusters, supporting the spatial coherence of network-defined cell states. Discussion This framework provides a network-informed representation for immune cell subpopulation identification and functional characterization. By linking bulk immune gene co-expression, single-cell programs, and spatial organization, this approach offers an additional perspective for understanding immune dynamics in both normal and pathological states and may provide an analytical basis for more precise immunotherapy-related studies.

Yi-Ming Li, Yu-Jie You, Rui-Xian Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.