2026· E3S Web of Conferences· Vol 723, pp. 01008· 0 citations· 5 references
TL;DR
The proposed method uses the L2 norm to directly map each multi-dimensional data point into a one-dimensional scalar value, which serves as a sorting criterion and is stored in a dictionary data structure to deterministically extract initial centroids.
Abstract
The K-Means clustering algorithm is a fundamental tool in machine learning, but its performance is often strongly affected by the instability of traditional random initialization methods, which can lead to convergence to poor local optima. Although many studies have proposed deterministic models to address this issue, they often involve high computational cost with
O
(
n
2
) complexity. This paper introduces a new engineering approach that is lightweight and efficient. Specifically, the method uses the L2 norm to directly map each multi-dimensional data point into a one-dimensional scalar value. This value serves as a sorting criterion and is stored in a dictionary data structure to deterministically extract
K
initial centroids. Experimental evaluation shows that the proposed method completely eliminates variability, fixing the Adjusted Rand Index (ARI) at stable values of 0.6345 for the Digits dataset and 0.7302 for the Iris dataset across all runs. These results demonstrate that a practical data structure-based approach can achieve high reliability with very low computational cost.
K-SCAN is presented -- a novel hybrid algorithm that optimizes this trade-off between robustness to noise and ability to detect non-linear clusters while maintaining structural stability, and achieves more than a 3-fold speed-up over the hierarchical BIRCH algorithm.
F. Kosiorowski, Grzegorz Sroka· arXiv.org· 0 citations
Sector-Mean Initialization is proposed, a deterministic initialization strategy with O(N) time complexity that partitions the two-dimensional data space into angular sectors around the global centroid and initializes centroids using sector-wise means.
A robust clustering method that estimates cluster centers and covariance matrices using density power divergence measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm is introduced.
Anirban Mondal, Paromita Banerjee, A. Mandal· 0 citations
This paper presents a comparative analysis of optimization techniques for the minimum sum-of-squares clustering (MSSC) problem—widely known in applied research as the K-means clustering problem—in the context of big data. K-means is the most widely used algorithmic framework for solving this problem, but MSSC methods c...
Ravil Mussabayev, R. Mussabayev· Symmetry· 0 citations
This work proposed to use Regularised Multidimensional Scaling using Radial Basis Function (RBF-MDS) for dimension reduction, a multidimensional scaling that mitigates the impact of irrelevant or redundant features, enabling 2D/3D visualisation of clusters for interpretability.
Afsana Akter Setu, J. Singha, Sohana Jahan· Dhaka University Journal of...· 0 citations
A proportional capture relation is introduced that links optimal and current centers based the assignment proportions of lines, enabling a refined analysis that bypasses the triangle inequality barrier.
Ting Liang, Xiao-Liang Wu, Junyu Huang et al.· Neural Information Processin...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.