Skip to content

FloDR: An invertible dimensionality reduction method based on a normalising flow

Jul 2026 · arXiv.org · Vol abs/2607.26278 · 0 citations · 43 references
Computer Science

TL;DR

While FloDR only uses the first two output coordinates to create a two-dimensional embedding, it retains the remaining coordinates rather than discarding them, which enable diagnostic visualisations that are computed from the exact inverse of the model that drew the layout rather than from an approximate one.

Abstract

It is common for two-dimensional embeddings of high-dimensional data to be read far beyond what they can support. Distances in and between clusters, the meaning behind empty spaces, and the amount of structure hidden at each point are generally invisible in the output of methods such as t-SNE and UMAP. This is because the information that could support the meaning of these properties is discarded during the optimisation process. Here, we present FloDR, a dimensionality reduction method that embeds data through an invertible normalising flow. While FloDR only uses the first two output coordinates to create a two-dimensional embedding, it retains the remaining coordinates rather than discarding them. In addition to the embedding, an exact inverse and an exact density are properties of a trained mapping, which enable diagnostic visualisations that are computed from the exact inverse of the model that drew the layout rather than from an approximate one. Specifically, we draw two fields, the conditional spread, which measures how much of the original data remains undetermined at each embedding position in input units, and the hidden contrast, which measures how much information about a labelled contrast the two plotted coordinates discard. Both fields are rendered with a prespecified test against a held out portion of the input data and a bootstrap confidence. A field that fails the test is reported as refused.

View source

Similar papers

A Unified Toolkit for Evaluating Nonlinear Dimensionality Reduction Techniques

This thesis builds on an existing diagnostics toolkit mainly for t-SNE and UMAP and turns it into a more accessible package for interested practitioners, while also extending it with diagnostics tools.

Kasra Amirani, S. Huisman, E. V. van Nieuwenburg · 0 citations
Preprint Aug 2026

Efficient Estimation of High Information Projections using Nearest Neighbours

An intuitive method for dimensionality reduction is proposed, which is highly effective for finding interesting projections of multivariate data and aiding the downstream tasks of cluster analysis and outlier detection.

David P. Hofmeyr · 0 citations
Preprint Sep 2026

Cursive: The Trace from the Curse of Dimensionality

Modern data are increasingly high-dimensional or non-Euclidean. As dimension grows, new statistical patterns can emerge in the relations among observations, while a conventional statistical summary may fail to retain the signal they carry. This paper names and organizes a research program around this observation, calli...

Hao Chen · 0 citations
Preprint Aug 2026

Gromov-Wasserstein Quantization and Clustering: Structure, Rates, and Algorithms

Numerical experiments show that GW quantization opens up many modeling possibilities beyond normal clustering methods and that the introduced algorithm leads to useful numerical solutions with approximation quality often in line with theoretically optimal rates.

F. Beier, S. Eckstein · 0 citations
Jul 2026

FADEx: Feature Attribution and Distortion-based Explanation of Dimensionality Reduction

Dimensionality Reduction (DR) is a fundamental tool for high-dimensional data exploration, reducing the complexity of latent spaces of machine learning models, and assisting in the explanation of complex opaque models. However, non-linear DR techniques often function as opaque transformations themselves, making it chal...

Lucas Greff Meneses, Evandro S. Ortigossa, Cláudio T. Silva et al. · 0 citations
#natural language process... Preprint Aug 2026

When Can We Work in Embedding Space? What Text Embeddings Preserve

In an application to 363 U.S. metropolitan areas, embedding-based clusters of LLM-generated economic descriptions recover interpretable economic archetypes and separate local employment dynamics more sharply than clustering on model residuals, or on a curated set of industry and demographic covariates.

S. Freyaldenhoven · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.