Skip to content
Preprint

Efficient Estimation of High Information Projections using Nearest Neighbours

Aug 2026 · 0 citations · 30 references
Mathematics Computer Science

TL;DR

An intuitive method for dimensionality reduction is proposed, which is highly effective for finding interesting projections of multivariate data and aiding the downstream tasks of cluster analysis and outlier detection.

Abstract

An intuitive method for dimensionality reduction is proposed, which is highly effective for finding interesting projections of multivariate data. Following similar intuitive motivation to a number of existing techniques, the proposed method is based on enhancing the nearest neighbour relationships in the data. The proposed projection arises from the spectral decomposition of a matrix designed to encode the local covariance structure in the data, where the local covariance at a point is captured by pairs of its nearest neighbours. We show that under standard regularity conditions this matrix is a consistent estimator of the so-called ``Density Information Matrix''(DIM); a non-parametric analogue of the Fisher Information Matrix. Spectral decompositions of DIMs have been shown to be connected with the important problems of Independent Components Analysis and, in the supervised context, Sufficient Dimension Reduction. However, existing estimators of the DIM are computationally expensive to compute and only target the DIM of a surrogate density, which is proportional to the square of the true underlying density. In addition, we go on to explore the practical utility of our method in aiding the downstream tasks of cluster analysis and outlier detection.

View source

Similar papers

2026

Probabilistic Nested Homogeneous Spaces for Dimensionality Reduction

A novel probabilistic version of the NHS model (PNHS) is presented for dimensionality reduction of high dimensional manifold-valued data in Riemannian homogeneous spaces and has several advantages over its deterministic counterpart namely, the NHS model.

Xi-Ran Fan, B. Vemuri · 0 citations
Preprint Aug 2026

Tensor Covariance Estimation via Kronecker-Structured Sparse Inverse Cholesky

A unified framework for scalable estimation of tensor covariances based on a Kronecker-structured sparse inverse Cholesky (KSIC) projection, proving that the KSIC estimator gainfully exploits cross-mode information and is robust to data scarcity.

Wentao Zhan, Matthias Katzfuss · 0 citations
Preprint Aug 2026

High-dimensional nonparametric changepoint detection via low-rank degree-two density projection

Detecting distributional changes in high dimension is difficult when neither the pre-change nor post-change density is parametrically specified. We introduce a representation-based approach that retains all degree-at-most-two density information while replacing density estimation by matrix mean estimation. For observat...

Guoqing Zhang, Zhaixin Chen · 0 citations
Preprint Aug 2026

Approximating the null distribution of generalized distance covariance

The null distribution of distance covariance is usually approximated by permutation, which is prohibitive when very small p-values are needed, or by matching a few moments to a parametric family, which is inaccurate in the tails. A third option is to approximate the limiting distribution, a weighted sum of chi-square v...

D. Edelmann · 0 citations
Preprint Sep 2026

Cluster-Based Dimensionality Reduction by Nonparametric Distributional Screening

The objective is not to construct a low-rank projection, but to retain an interpretable subset of the original coordinates that preserves the distributional information distinguishing the clusters that preserves the distributional information distinguishing the clusters.

S. Jha, Rishikesh Muralimohan, Praveen Athauda Arachchi et al. · 0 citations
Preprint Oct 2026

High-Dimensional Regularization of the Spatial Sign Covariance Matrix for Robust Shape Estimation

Covariance estimation is a key component of many applications in system identification and data-driven control. Although heavy-tailed distributions may lack a covariance matrix to estimate, the shape matrix provides a well-defined, scale-free generalization for the broad family of elliptical distributions. In this sett...

Jonas Elmerraji, J. Spall, Mateo Díaz · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.