An SI framework for deep clustering with a fixed pretrained encoder that provides a principled approach to quantifying the statistical reliability of structures discovered by deep clustering and enables valid statistical testing of differences between clusters identified in the latent space.
Abstract
Deep clustering is a powerful approach for discovering meaningful structures in high-dimensional data by learning a low-dimensional latent representation prior to clustering. Despite its empirical success, assessing the statistical reliability of the resulting clusters remains challenging. Testing discovered clusters on the same data induces selection bias and invalidates classical $p$-values. Selective inference (SI) provides a principled framework for correcting this bias, but existing methods focus on clustering performed directly on the observed features. In this work, we develop an SI framework for deep clustering with a fixed pretrained encoder. The key challenge is that cluster assignments are determined through a nonlinear transformation from the original data space to the latent space, resulting in a substantially more complex selection process than in conventional clustering. Our method provides a computationally tractable way to account for this process and enables valid statistical testing of differences between clusters identified in the latent space. Synthetic experiments demonstrate that the proposed method controls the Type I error rate while achieving higher power than valid but conservative baselines, and genomic applications show that it can identify significant cluster differences while appropriately accounting for selection bias. Our framework provides a principled approach to quantifying the statistical reliability of structures discovered by deep clustering.
A Reliability-Based Deep Embedded Clustering (RDEC) approach that improves clustering reliability through a novel distance-aware weighting strategy, coupled with a novel Kullback–Leibler divergence objective function to focus on the most representative data instances and guide the clustering process.
Meaad Altwaimi, M. B. Ben Ismail, Ouiem Bchir· Algorithms· 0 citations
Subspace clustering methods have been widely used in high-dimensional data analysis due to their excellent capability in processing high-dimensional data. However, traditional methods struggle with nonlinear data, whereas kernel mapping-based methods are limited by kernel function design. Although deep subspace cluster...
Li Guo, Qian Wang· Journal of King Saud Univers...· 0 citations
Multi-view clustering aims to utilize information from multiple feature representations to uncover underlying data structures. Most existing methods emphasize learning a consensus representation by enforcing consistency across views. However, those structures that cannot be directly incorporated into the clustering spa...
Gao-Kai Wang, Yazhou Ren, Feng-Yu Zhang et al.· Proceedings of the Thirty-Fi...· 0 citations
It is demonstrated that it is feasible to train a deep generative model for dimensionality reduction with millions of features using few samples, which makes this type of generative model a more versatile alternative to standard methods for dimensionality reduction.
C. Pancotti, P. Fariselli, J. Meisner et al.· bioRxiv· 0 citations
It is shown that layer-wise generative learning can spontaneously uncover and progressively amplify class-related structure in unlabeled data and improve average clustering can coexist with reduced accessibility for a few difficult class pairs.
Patrick Krauss, Achim Schilling, Andreas K. Maier et al.· 0 citations
Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.