PIANO: Probabilistic Inference Autoencoder Networks for multi-Omics enables robust generative modeling of gene expression and scales single-cell integration to 100 million cells
PIANO accurately integrates single-cell data across species and across single-cell and spatial transcriptomics modalities and models spatially-resolved gene expression during Alzheimer’s disease progression in human brains and integrates over 100 million cancer cells to model drug perturbations.
Abstract
Single-cell RNA technologies enable the routine acquisition of transcriptomic atlases. However, these molecular profiles are influenced by overlapping sources of variation. Since these covariates confound comparisons, data integration is the first step in most analyses. Three challenges remain: correcting strong batch effects, scaling to millions of cells, and modeling how covariates influence gene expression. To address these challenges, we developed PIANO: Probabilistic Inference Autoencoder Networks for multi-Omics, a deep learning framework whose central feature is a generative model of gene expression data. Additionally, PIANO achieves robust integrations and trains 10x faster than previous methods. PIANO accurately integrates single-cell data across species and across single-cell and spatial transcriptomics modalities. As practical applications, PIANO models spatially-resolved gene expression during Alzheimer’s disease progression in human brains and integrates over 100 million cancer cells to model drug perturbations. In summary, PIANO’s integration and generative modeling capabilities will empower novel insights for countless future studies.
Interindividual heterogeneity in Alzheimer’s disease (AD) remains poorly understood, as disparate single-cell studies leave it unclear whether findings reflect shared architecture or dataset-specific idiosyncrasies. Here, we present panAD, a transcriptomic atlas of >3 million nuclei from 791 individuals across 13 studi...
Negin Rahimzadeh, S. Morabito, S. Khullar et al.· bioRxiv· 0 citations
Gene co-expression maps transcriptome-wide gene-gene relationships, yet high-quality estimates cover less than half the genome. Meanwhile, spatial omics either profiles restricted in situ panels or lacks cellular resolution. Extending co-expression transcriptome-wide could overcome these limitations by inferring unassa...
Gene co-expression networks reflect transcription factor-associated regulation together with epigenetic influences such as DNA methylation, chromatin states, and histone modifications. To distinguish gene-gene dependencies that persist after accounting for methylation covariation, large-scale statistical inference cond...
Sung-Dong Lee, Qing Zhao, Dongwon Kim et al.· bioRxiv· 0 citations
Systematic identification of transcriptional and epigenetic regulators (TERs) remains a challenge in myeloid leukemia. Current methods for TER identification typically rely on single data types and show limited power for long-range regulatory interactions. Here we present TERfinder, a deep learning framework that...
Single-cell DNA methylation profiling technology captures novel epigenetic data modality but are challenging to analyze because of their heterogeneity, high dimensionality, and ultra-sparsity. Here we present SMORE (Single-cell MethylOme Reduction and Embedding), a computational method for joint dimensionality reductio...
Joint single-cell transcriptomic–metabolomic profiling remains technically intractable. Here we present CHIMERA (Cell-level Hybrid Inference of Metabolome Embedded on RNA Atlas), a data-driven framework that learns transcriptome-to-metabolome mappings from spatially paired multi-omics data and transfers them to unpaire...
Lu-Yu Yang, Jing Xu, Tian-Lun Wang et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.