An intuitive method for dimensionality reduction is proposed, which is highly effective for finding interesting projections of multivariate data and aiding the downstream tasks of cluster analysis and outlier detection.
Abstract
An intuitive method for dimensionality reduction is proposed, which is highly effective for finding interesting projections of multivariate data. Following similar intuitive motivation to a number of existing techniques, the proposed method is based on enhancing the nearest neighbour relationships in the data. The proposed projection arises from the spectral decomposition of a matrix designed to encode the local covariance structure in the data, where the local covariance at a point is captured by pairs of its nearest neighbours. We show that under standard regularity conditions this matrix is a consistent estimator of the so-called ``Density Information Matrix''(DIM); a non-parametric analogue of the Fisher Information Matrix. Spectral decompositions of DIMs have been shown to be connected with the important problems of Independent Components Analysis and, in the supervised context, Sufficient Dimension Reduction. However, existing estimators of the DIM are computationally expensive to compute and only target the DIM of a surrogate density, which is proportional to the square of the true underlying density. In addition, we go on to explore the practical utility of our method in aiding the downstream tasks of cluster analysis and outlier detection.
A novel probabilistic version of the NHS model (PNHS) is presented for dimensionality reduction of high dimensional manifold-valued data in Riemannian homogeneous spaces and has several advantages over its deterministic counterpart namely, the NHS model.
Xi-Ran Fan, B. Vemuri· Proceedings of machine learn...· 0 citations
A unified framework for scalable estimation of tensor covariances based on a Kronecker-structured sparse inverse Cholesky (KSIC) projection, proving that the KSIC estimator gainfully exploits cross-mode information and is robust to data scarcity.
Detecting distributional changes in high dimension is difficult when neither the pre-change nor post-change density is parametrically specified. We introduce a representation-based approach that retains all degree-at-most-two density information while replacing density estimation by matrix mean estimation. For observat...
The null distribution of distance covariance is usually approximated by permutation, which is prohibitive when very small p-values are needed, or by matching a few moments to a parametric family, which is inaccurate in the tails. A third option is to approximate the limiting distribution, a weighted sum of chi-square v...
The objective is not to construct a low-rank projection, but to retain an interpretable subset of the original coordinates that preserves the distributional information distinguishing the clusters that preserves the distributional information distinguishing the clusters.
S. Jha, Rishikesh Muralimohan, Praveen Athauda Arachchi et al.· 0 citations
Covariance estimation is a key component of many applications in system identification and data-driven control. Although heavy-tailed distributions may lack a covariance matrix to estimate, the shape matrix provides a well-defined, scale-free generalization for the broad family of elliptical distributions. In this sett...
Jonas Elmerraji, J. Spall, Mateo Díaz· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.