Skip to content
Review Open access

Optimizing K-Means Clustering for Big Data: A Review

Sep 2026 · Symmetry · 0 citations · 63 references

Abstract

This paper presents a comparative analysis of optimization techniques for the minimum sum-of-squares clustering (MSSC) problem—widely known in applied research as the K-means clustering problem—in the context of big data. K-means is the most widely used algorithmic framework for solving this problem, but MSSC methods can suffer from scalability issues when dealing with large datasets. The paper reviews approaches for overcoming these issues, including decomposition, sampling, initialization, preprocessing, parallel and distributed computation, data summarization, acceleration of distance computations, and hybridization with various metaheuristic frameworks. The experimental evaluation compares selected big data MSSC algorithms under a common benchmark protocol and assesses them according to the dominance criterion provided by the “less is more” approach (LIMA), i.e., simultaneously along the dimensions of clustering quality, speed, and simplicity. The results reveal distinct accuracy–time regimes and a multi-algorithm Pareto front under LIMA dominance, indicating that no method is universally preferable and that greater algorithmic complexity does not by itself ensure a better practical trade-off. Lightweight simplest methods generally favor speed but may sacrifice accuracy, while more complex hybrid methods reach stricter accuracy levels at substantially greater computational cost; competitive stochastic-sampling methods occupy intermediate accuracy–time–simplicity trade-offs.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.