This paper proposes visualizing the median of multiple NLDR outputs rather than relying on individual projections and finds that taking the median of multiple projections performs comparably to individual runs on multiple quality metrics, while increasing perturbation emphasizes global over local structure.
Abstract
Widely used non-linear dimensionality reduction (NLDR) methods such as UMAP and t-SNE are stochastic--repeated runs on the same data can produce different low-dimensional projections. In this paper, we explore two problems related to projection variability: on some datasets clusters, structure, and outliers may change run-to-run, and on others projections can be extremely stable when overfitting noise. To address the first problem, we propose visualizing the median of multiple NLDR outputs rather than relying on individual projections. To address the second, we perturb input data before creating consensus embeddings. We find that taking the median of multiple projections performs comparably to individual runs on multiple quality metrics, while increasing perturbation emphasizes global over local structure. We show through a set of exploratory visualizations that even relatively simple ensemble presentations can be used to better communicate the reliability of projection patterns.
This thesis builds on an existing diagnostics toolkit mainly for t-SNE and UMAP and turns it into a more accessible package for interested practitioners, while also extending it with diagnostics tools.
Kasra Amirani, S. Huisman, E. V. van Nieuwenburg· 0 citations
Modern data are increasingly high-dimensional or non-Euclidean. As dimension grows, new statistical patterns can emerge in the relations among observations, while a conventional statistical summary may fail to retain the signal they carry. This paper names and organizes a research program around this observation, calli...
This paper reformulates SN parameter estimation as a convex optimization problem over a positive semidefinite matrix, replacing the original nonconvex likelihood search with a formulation amenable to standard optimization tools, and clarify the expressive power of the SN class by connecting polynomial log-density model...
Arindam Roychowdhury, Luis G. Crespo, H. Lam· 0 citations
While FloDR only uses the first two output coordinates to create a two-dimensional embedding, it retains the remaining coordinates rather than discarding them, which enable diagnostic visualisations that are computed from the exact inverse of the model that drew the layout rather than from an approximate one.
Abdallah Baraka, Daniel Probst· arXiv.org· 0 citations
Multi-label data often contain high-dimensional features, outlier instances, and noisy labels, all of which can lead to the curse of dimensionality and decreased performance in downstream tasks. Although numerous data reduction methods have been developed, existing approaches face two major limitations: 1) existing met...
Li Yang, Yan-Yong Huang, Jin-Yuan Chang et al.· Proceedings of the Thirty-Fi...· 0 citations
The objective is not to construct a low-rank projection, but to retain an interpretable subset of the original coordinates that preserves the distributional information distinguishing the clusters that preserves the distributional information distinguishing the clusters.
S. Jha, Rishikesh Muralimohan, Praveen Athauda Arachchi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.