Skip to content
Open access

KMHC: A Parallel Hybrid Clustering Algorithm Integrating K-Means and Hierarchical Agglomerative Clustering with Voting-Based Consensus

Aug 2026 · Informatica · 0 citations · 25 references

TL;DR

The results show that KMHC can improve clustering consistency when the two base clusterers provide complementary partitions, particularly in datasets where both local compactness and hierarchical structure are informative, however, the method remains limited by the quadratic cost of HAC and by its current batch-mode design.

Abstract

This paper proposes KMHC, a parallel hybrid clustering algorithm that combines K-means and Hierarchi- cal Agglomerative Clustering (HAC) through an overlap-based label alignment and deterministic voting consensus mechanism. The objective of KMHC is to exploit the complementarity between centroid-based compactness and hierarchical structural information in order to improve clustering consistency and reduce assignment errors. Unlike conventional ensemble clustering methods that rely on repeated resampling, large pools of base partitions, or co-association matrices, KMHC uses a lightweight two-clusterer architec- ture with explicit cluster correspondence estimation. The proposed method is formally defined through an overlap matrix, a cluster mapping function, and a deterministic tie-breaking rule. Its computational com- plexity is also analyzed, showing that the main bottleneck is the HAC component, while the alignment and voting stages introduce limited additional cost. KMHC is evaluated on benchmark and synthetic datasets with different dimensionalities, noise levels, and cluster distributions. Its performance is compared with K-means, HAC, DBSCAN, Spectral Clustering, and Gaussian Mixture Models using precision, normal- ized mutual information, adjusted Rand index, error rate, runtime, and stability over multiple independent runs. Statistical significance tests and an ablation study are also conducted to assess the contribution of K-means, HAC, label alignment, and voting. The results show that KMHC can improve clustering perfor- mance when the two base clusterers provide complementary partitions, particularly in datasets where both local compactness and hierarchical structure are informative. However, the method remains limited by the quadratic cost of HAC and by its current batch-mode design. Future work will investigate approximate hierarchical clustering, incremental extensions, and streaming-data scenarios.

Read PDF

Similar papers

Sep 2026

Reliable Co-Occurrence Guided Partial Clustering Ensemble.

Partial clustering ensemble integrates multiple incomplete base clusterings to generate a consensus one. The key challenge of partial clustering ensemble is to infer the consensus clustering from the incomplete and unreliable base clusterings without access to the original data. To this end, we propose a novel partial...

Teng-Sheng Qi, Hao-Ye Qiu, Yuheng Jia et al. · 0 citations
#machine learning Preprint Oct 2026

Round-Trip KNN Clustering: multiscale hierarchical cluster detection on directed nearest-neighbour graphs

We introduce Round-Trip KNN Clustering (RTKNNC), a graph-based method for finding cluster structure at several neighbourhood scales without requiring the number of clusters in advance. Unlike approaches that first make a $k$-nearest-neighbour (KNN) graph undirected, RTKNNC keeps both directions of the neighbour relatio...

E. P. Marinho, C. Ranieri, Fabricio Aparecido Breve · 0 citations
Preprint Aug 2026

Robust K-means Clustering using the Density Power Divergence Measure

A robust clustering method that estimates cluster centers and covariance matrices using density power divergence measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm is introduced.

Anirban Mondal, Paromita Banerjee, A. Mandal · 0 citations
Open access Aug 2026

Density-Peak-Based Clustering in Reduced Feature Spaces

This work proposed to use Regularised Multidimensional Scaling using Radial Basis Function (RBF-MDS) for dimension reduction, a multidimensional scaling that mitigates the impact of irrelevant or redundant features, enabling 2D/3D visualisation of clusters for interpretability.

Afsana Akter Setu, J. Singha, Sohana Jahan · 0 citations
#machine learning Preprint Sep 2026

Automatic depth-based local center clustering via $\beta$-integrated local depth and adaptive grouping

Clustering is an unsupervised learning technique that partitions unlabeled data into groups. Most existing methods require user-specified parameters, such as the number of clusters or neighborhood size. Conversely, we propose automatic depth-based local center clustering (A-DLCC), a fully data-driven method that elimin...

Si-Yi Wang, Alexandre Leblanc, P. McNicholas · 0 citations
Conference Aug 2026

Optimal Cluster Selection in Unsupervised Machine Learning Using K-Means Clustering

In unsupervised learning, Clustering is a core method used to determine unseen arrangements and structures within datasets by grouping similar instances together. Among the many clustering algorithms, Among clustering techniques, K-Means continues to be one of the most popular owing to its ease of implementation, fast...

I. Khan, H. Daud, Rajalingam Sokkalingam et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.