MILK is presented, a scalable computational framework that organizes high-dimensional single-cell populations into unified tree representations that establish the hierarchical organization of biological data as a scalable and unifying representation of cellular identity, enabling integrative analysis of single-cell genomic data across diverse contexts.
Abstract
The rapid expansion of single-cell genomic datasets has led to the compilation of biological resources comprising hundreds of millions of cells across tissues, developmental stages, and disease states. This has underscored the need for scalable and interpretable data representations that preserve the complex relationships and multi-scale organization of cellular states, while remaining computationally tractable at atlas scale. Existing approaches based on discrete abstractions have enabled cell annotation, clustering, and trajectory inference, but are often optimized for local inference tasks and may obscure continuous cellular relationships and multi-resolution structure within complex transcriptional and other genomic landscapes. Moreover, increasing dataset sizes often require information-reduction strategies such as random downsampling, limiting the resolution of rare cell populations and heterogeneous cellular states. Here, we present MILK, a scalable computational framework that organizes high-dimensional single-cell populations into unified tree representations. Across large-scale transcriptomic atlases, MILK enables representative subsampling with preserved information, supporting the tractable application of existing algorithms for tasks including deep generative model training and foundation model benchmarking. Additionally, MILK enables holistic, multi-resolution analyses that capture global developmental trajectories, characterize disease-associated cellular perturbations across tissues, and facilitate comparison of transcriptional programs across species within a coherent hierarchical framework. Together, these results establish the hierarchical organization of biological data as a scalable and unifying representation of cellular identity, enabling integrative analysis of single-cell genomic data across diverse contexts.
Abstract Motivation Resolving cellular heterogeneity requires methods that capture both discrete cell types and their functional relationships. Existing single-cell clustering approaches often produce flat partitions, limiting their ability to reveal rare cell states and continuous biological transitions. Here, we intr...
Cellular diversity in multicellular organisms arises from the functional specialization of individual cells and the influence of both the local tissue microenvironment and external stimuli. Understanding this heterogeneity requires accurate characterization of cell types and the molecular dynamics that define them. In...
J. López-Castiblanco, L. López-Kleine, Yesid Cuesta-Astroz· Journal of Investigative Med...· 0 citations
The results indicate that GmGM provides a unified, reproducible framework for joint cell clustering and gene-network inference, capable of revealing cellular structure beyond that captured by conventional pipelines.
O. Lanzetta, L. Cutillo, Bailey Andrew et al.· 0 citations
Spatial transcriptomics maps gene expression at cellular resolution, revealing how cells organize into multicellular niches. Yet computational analyses remain dataset-specific, without a transferable representation of tissue organization that generalizes across datasets, tasks and tissues or predicts how tissues behave...
Sebastian Birk, M. V. Sanian, Amirhossein Vahidi et al.· bioRxiv· 4 citations· ⚡1
A comprehensive, context-specific guide to current annotation strategies for spatial transcriptomics is presented and open-set recognition of reference-absent cell states, adaptive incorporation of spatial context, and improved resolution of rare and transitional cell identities are identified as central priorities for...
Yu-Ling Zhu, Yunfei Hu, M. Xie et al.· Research Square· 0 citations
Malva is presented, a computational platform that enables ultrafast, species-agnostic and reference-free interrogation of the raw sequence space, enabling searching for any sequence, mutation, splice junction or pathogen, or spatial location of arbitrary transcripts.
D. León-Periñán, Nikos Karaiskos, N. Rajewsky· Nature· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.