Skip to content
Open access

A generative model for dimensionality reduction with millions of features and few samples

Aug 2026 · bioRxiv · 0 citations · 15 references
Biology

TL;DR

It is demonstrated that it is feasible to train a deep generative model for dimensionality reduction with millions of features using few samples, which makes this type of generative model a more versatile alternative to standard methods for dimensionality reduction.

Abstract

Motivation In this paper, we demonstrate that it is feasible to train a deep generative model for dimensionality reduction with millions of features using few samples, which makes this type of generative model a more versatile alternative to standard methods for dimensionality reduction. Specifically, we hypothesize that for a decoder-only model, the number of training samples required is almost independent of the feature dimensionality in most network architectures. Results Through an extensive set of experiments on synthetic non-linear data, we validate this hypothesis. We also train the model on a downsampled version of the 1000 Genomes Project (1KGP) dataset to further assess its behavior under controlled reductions in sample size. Furthermore, we train a deep generative decoder (DGD) on a curated dataset from the International Cancer Genome Consortium (ICGC), which contains 4.4 million features. It is trained on approximately 4,000 samples and tested on 1,000 samples. The resulting latent representation exhibits clear clustering, and when methods are reduced to the same number of dimensions, it outperforms PCA and VAE for tumor type classification. Additionally, the DGD is computationally efficient and can be trained on a 16GB GPU. Availability and implementation Code is available at https://github.com/cpancott/ReceptiveDGD. Contact corrado.pancotti@helmholtz-munich.de; akrogh@di.ku.dk Supplementary information Supplementary data are available with this preprint.

Read PDF

Similar papers

Open access Aug 2026

Self-Attention over Parallel Dense Embeddings for High-Dimensional Omic Data

Results suggest that PLAT provides an effective and interpretable framework for high-dimensional transcriptomic classification and functional enrichment analyses consistently highlighted biological processes and disease pathways associated with breast cancer, supporting the biological relevance of the learned latent re...

Kamal Elatifi, Nicolas Jäger Gallego, Á. Sánchez-Pla et al. · 0 citations

A Unified Toolkit for Evaluating Nonlinear Dimensionality Reduction Techniques

This thesis builds on an existing diagnostics toolkit mainly for t-SNE and UMAP and turns it into a more accessible package for interested practitioners, while also extending it with diagnostics tools.

Kasra Amirani, S. Huisman, E. V. van Nieuwenburg · 0 citations
#machine learning Preprint Sep 2026

Selective Inference for Deep Clustering in Latent Spaces

An SI framework for deep clustering with a fixed pretrained encoder that provides a principled approach to quantifying the statistical reliability of structures discovered by deep clustering and enables valid statistical testing of differences between clusters identified in the latent space.

Eina Mizui, Tomohiro Shiraishi, S. Nishino et al. · 0 citations
Preprint Aug 2026

Accelerated Learning of High Dimensional Functions with a Tensor-Featured Training Network

This work presents a method to accelerate the optimization of learning high dimensional functions using deep neural network (DNN), and studies the effect of adding features which distill pretrained DNN into TNs using a discretize and decompose strategy.

Karl Pierce, Y. Khoo, Haizhao Yang · 0 citations
Review Open access Sep 2026

An Introduction to Stochastic Deep Learning

Deep neural networks (DNNs) have achieved remarkable success in prediction, but their deterministic formulation makes many statistical inference tasks difficult. StoNet, short for stochastic neural network, addresses this limitation by reformulating a DNN as a probabilistic latent‐variable model, in which the outputs o...

Fa-Ming Liang · 0 citations
#machine learning Preprint Aug 2026

Boosting Data Augmentation with Stochastic Weight Averaging

Stochastic weight averaging applied to classification as an alternative ensembling technique that does not require repeated training runs and provides an equivariance boost that goes beyond what could be expected from the performance increase due to SWA alone.

Longde Huang, Axel Flinth, Jan E. Gerken · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.