The results show that KMHC can improve clustering consistency when the two base clusterers provide complementary partitions, particularly in datasets where both local compactness and hierarchical structure are informative, however, the method remains limited by the quadratic cost of HAC and by its current batch-mode design.
Abstract
This paper proposes KMHC, a parallel hybrid clustering algorithm that combines K-means and Hierarchi- cal Agglomerative Clustering (HAC) through an overlap-based label alignment and deterministic voting consensus mechanism. The objective of KMHC is to exploit the complementarity between centroid-based compactness and hierarchical structural information in order to improve clustering consistency and reduce assignment errors. Unlike conventional ensemble clustering methods that rely on repeated resampling, large pools of base partitions, or co-association matrices, KMHC uses a lightweight two-clusterer architec- ture with explicit cluster correspondence estimation. The proposed method is formally defined through an overlap matrix, a cluster mapping function, and a deterministic tie-breaking rule. Its computational com- plexity is also analyzed, showing that the main bottleneck is the HAC component, while the alignment and voting stages introduce limited additional cost. KMHC is evaluated on benchmark and synthetic datasets with different dimensionalities, noise levels, and cluster distributions. Its performance is compared with K-means, HAC, DBSCAN, Spectral Clustering, and Gaussian Mixture Models using precision, normal- ized mutual information, adjusted Rand index, error rate, runtime, and stability over multiple independent runs. Statistical significance tests and an ablation study are also conducted to assess the contribution of K-means, HAC, label alignment, and voting. The results show that KMHC can improve clustering perfor- mance when the two base clusterers provide complementary partitions, particularly in datasets where both local compactness and hierarchical structure are informative. However, the method remains limited by the quadratic cost of HAC and by its current batch-mode design. Future work will investigate approximate hierarchical clustering, incremental extensions, and streaming-data scenarios.
Partial clustering ensemble integrates multiple incomplete base clusterings to generate a consensus one. The key challenge of partial clustering ensemble is to infer the consensus clustering from the incomplete and unreliable base clusterings without access to the original data. To this end, we propose a novel partial...
Teng-Sheng Qi, Hao-Ye Qiu, Yuheng Jia et al.· IEEE Transactions on Neural...· 0 citations
We introduce Round-Trip KNN Clustering (RTKNNC), a graph-based method for finding cluster structure at several neighbourhood scales without requiring the number of clusters in advance. Unlike approaches that first make a $k$-nearest-neighbour (KNN) graph undirected, RTKNNC keeps both directions of the neighbour relatio...
E. P. Marinho, C. Ranieri, Fabricio Aparecido Breve· 0 citations
A robust clustering method that estimates cluster centers and covariance matrices using density power divergence measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm is introduced.
Anirban Mondal, Paromita Banerjee, A. Mandal· 0 citations
This work proposed to use Regularised Multidimensional Scaling using Radial Basis Function (RBF-MDS) for dimension reduction, a multidimensional scaling that mitigates the impact of irrelevant or redundant features, enabling 2D/3D visualisation of clusters for interpretability.
Afsana Akter Setu, J. Singha, Sohana Jahan· Dhaka University Journal of...· 0 citations
Clustering is an unsupervised learning technique that partitions unlabeled data into groups. Most existing methods require user-specified parameters, such as the number of clusters or neighborhood size. Conversely, we propose automatic depth-based local center clustering (A-DLCC), a fully data-driven method that elimin...
Si-Yi Wang, Alexandre Leblanc, P. McNicholas· 0 citations
In unsupervised learning, Clustering is a core method used to determine unseen arrangements and structures within datasets by grouping similar instances together. Among the many clustering algorithms, Among clustering techniques, K-Means continues to be one of the most popular owing to its ease of implementation, fast...
I. Khan, H. Daud, Rajalingam Sokkalingam et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.