Jul 2026· International Conference on Machine Learning and Embedded Systems· Vol 14295, pp. 1429502 - 1429502-5· 0 citations· 12 references
Engineering
TL;DR
Improved K-means method (MD-Kmeans) is proposed, which integrates K-nearest-neighbor– based density estimation with a maximum-dispersion strategy to ensure representative and well-distributed initial centers, and employs a balanced objective that jointly enhances intra-cluster compactness and inter-cluster separability.
Abstract
Clustering analysis is an essential task in data mining and machine learning, and the classical K-means algorithm is widely used due to its efficiency. However, its random initialization often leads to unstable results, especially on complex or nonuniform datasets, where it easily falls into local optima. Moreover, its objective function focuses solely on intra-cluster compactness while overlooking inter-cluster separability, thus limiting global clustering performance. To address these issues, this paper proposes an improved K-means method (MD-Kmeans). The algorithm integrates K-nearest-neighbor– based density estimation with a maximum-dispersion strategy to ensure representative and well-distributed initial centers, and employs a balanced objective that jointly enhances intra-cluster compactness and inter-cluster separability. Experimental results show that MD-Kmeans achieves notable improvements in Adjusted Rand Index (ARI), Silhouette Coefficient (SC), and Davies–Bouldin Index (DBI), outperforming traditional K-means and recent variants, particularly on non-uniform datasets.
This paper presents a comparative analysis of optimization techniques for the minimum sum-of-squares clustering (MSSC) problem—widely known in applied research as the K-means clustering problem—in the context of big data. K-means is the most widely used algorithmic framework for solving this problem, but MSSC methods c...
Ravil Mussabayev, R. Mussabayev· Symmetry· 0 citations
K-SCAN is presented -- a novel hybrid algorithm that optimizes this trade-off between robustness to noise and ability to detect non-linear clusters while maintaining structural stability, and achieves more than a 3-fold speed-up over the hierarchical BIRCH algorithm.
F. Kosiorowski, Grzegorz Sroka· arXiv.org· 0 citations
This work proposed to use Regularised Multidimensional Scaling using Radial Basis Function (RBF-MDS) for dimension reduction, a multidimensional scaling that mitigates the impact of irrelevant or redundant features, enabling 2D/3D visualisation of clusters for interpretability.
Afsana Akter Setu, J. Singha, Sohana Jahan· Dhaka University Journal of...· 0 citations
Sector-Mean Initialization is proposed, a deterministic initialization strategy with O(N) time complexity that partitions the two-dimensional data space into angular sectors around the global centroid and initializes centroids using sector-wise means.
A robust clustering method that estimates cluster centers and covariance matrices using density power divergence measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm is introduced.
Anirban Mondal, Paromita Banerjee, A. Mandal· 0 citations
The results show that KMHC can improve clustering consistency when the two base clusterers provide complementary partitions, particularly in datasets where both local compactness and hierarchical structure are informative, however, the method remains limited by the quadratic cost of HAC and by its current batch-mode de...