Skip to content

Sector-Mean: Deterministic Initialization of K-Means Centroids via Angular Sector Partitioning

Sep 2026 · 0 citations · 39 references
Computer Science

TL;DR

Sector-Mean Initialization is proposed, a deterministic initialization strategy with O(N) time complexity that partitions the two-dimensional data space into angular sectors around the global centroid and initializes centroids using sector-wise means.

Abstract

K-Means is one of the most widely used clustering algorithms, but its susceptibility to initial centroid selection remains a primary bottleneck for its convergence speed and clustering accuracy. This paper proposes Sector-Mean Initialization, a deterministic initialization strategy with O(N) time complexity that partitions the two-dimensional data space into angular sectors around the global centroid and initializes centroids using sector-wise means. We evaluate the method on established two-dimensional benchmarks (SIPU, Birch) and multiple real-world datasets, comparing against random, K-Means++, and Max-Min initialization under identical Lloyd iterations. The statistical analysis of Friedman's test (p<0.05) and Nemenyi post-hoc comparison indicates that, while delivering equivalent clustering quality as K-Means++ and Max-Min, Sector-Mean offers significant computational efficiency. Experimental results show that Sector-Mean reduces the initialization time by 74.9% and 59.8% in comparison to K-Means++ and max-min, respectively. And, it yields the lowest average number of iterations, achieving approximately 5% fewer iterations than K-Means++ and 16% fewer than max-min. These results highlight that Sector-Mean initialization offers a deterministic and computationally efficient initialization strategy while preserving cluster quality.

View source

Similar papers

Review Open access Sep 2026

Optimizing K-Means Clustering for Big Data: A Review

This paper presents a comparative analysis of optimization techniques for the minimum sum-of-squares clustering (MSSC) problem—widely known in applied research as the K-means clustering problem—in the context of big data. K-means is the most widely used algorithmic framework for solving this problem, but MSSC methods c...

Ravil Mussabayev, R. Mussabayev · 0 citations
Conference Aug 2026

Optimal Cluster Selection in Unsupervised Machine Learning Using K-Means Clustering

In unsupervised learning, Clustering is a core method used to determine unseen arrangements and structures within datasets by grouping similar instances together. Among the many clustering algorithms, Among clustering techniques, K-Means continues to be one of the most popular owing to its ease of implementation, fast...

I. Khan, H. Daud, Rajalingam Sokkalingam et al. · 0 citations
Preprint Aug 2026

Robust K-means Clustering using the Density Power Divergence Measure

A robust clustering method that estimates cluster centers and covariance matrices using density power divergence measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm is introduced.

Anirban Mondal, Paromita Banerjee, A. Mandal · 0 citations
Open access Aug 2026

KMHC: A Parallel Hybrid Clustering Algorithm Integrating K-Means and Hierarchical Agglomerative Clustering with Voting-Based Consensus

The results show that KMHC can improve clustering consistency when the two base clusterers provide complementary partitions, particularly in datasets where both local compactness and hierarchical structure are informative, however, the method remains limited by the quadratic cost of HAC and by its current batch-mode de...

Aida Chefrour, Soufiane Khedairia · 0 citations
Open access Aug 2026

Density-Peak-Based Clustering in Reduced Feature Spaces

This work proposed to use Regularised Multidimensional Scaling using Radial Basis Function (RBF-MDS) for dimension reduction, a multidimensional scaling that mitigates the impact of irrelevant or redundant features, enabling 2D/3D visualisation of clusters for interpretability.

Afsana Akter Setu, J. Singha, Sohana Jahan · 0 citations
Preprint Sep 2026

A Geometry-Aware Framework for Clustering Cylindrical Data

Cylindrical data pair an angle with a linear measurement. Clustering that ignores the periodicity of the angle breaks up groups lying across its origin. We formulate the K-means algorithm for a generic distance on the cylinder and instantiate it with the chordal distance of the ambient space and the geodesic distance a...

Giuseppe Pandolfo, Luca Coraggio, Antonio D'Ambrosio · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.