Skip to content
Preprint

Robust K-means Clustering using the Density Power Divergence Measure

Aug 2026 · 0 citations · 33 references
Mathematics

TL;DR

A robust clustering method that estimates cluster centers and covariance matrices using density power divergence measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm is introduced.

Abstract

We introduce a robust clustering method, MK-means DPD, that estimates cluster centers and covariance matrices using density power divergence (DPD) measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm. Since Mahalanobis distance-based K-means lacks a general convergence guarantee, we further introduce a convergent variant, Density-Consistent MK-means DPD (DC-MK-means DPD), which redefines the cluster assignment step in terms of a pointwise DPD loss. We prove a formal theorem establishing that the resulting algorithm converges in a finite number of steps. We also propose two new robust internal evaluation indices, a Median Davies-Bouldin Index and a Trimmed Calinski-Harabasz Index, to ensure that performance comparisons are not themselves distorted by outliers. The efficacy of the proposed methods is demonstrated on simulated data, showing superiority over existing methods, and on two real datasets: Iris data, to identify similar species, and COVID-19 case fatality rate and infection rate data for countries worldwide, examining the resulting clusters'geographic and socio-economic patterns.

View source

Similar papers

Open access Aug 2026

Density-Peak-Based Clustering in Reduced Feature Spaces

This work proposed to use Regularised Multidimensional Scaling using Radial Basis Function (RBF-MDS) for dimension reduction, a multidimensional scaling that mitigates the impact of irrelevant or redundant features, enabling 2D/3D visualisation of clusters for interpretability.

Afsana Akter Setu, J. Singha, Sohana Jahan · 0 citations
Open access Sep 2026

Robust ℓp-Norm Two-Dimensional Discriminative Clustering for Image Data

For image data, matrix-based clustering methods are gaining popularity because they can directly process two-dimensional (2D) data structures without vectorization. However, most existing approaches rely on squared Frobenius norms in their objective functions, making them sensitive to outliers and noise commonly encoun...

Yan-Ru Guo, Xiang-Yu Hua · 0 citations
Open access Sep 2026

Sparse k-Means Clustering with Lasso-Based Feature Selection

Clustering high-dimensional data is challenging when irrelevant or weakly informative features obscure the underlying cluster structure. To address this issue, we propose two sparse k-means clustering algorithms, S-KM1 and S-KM2, that perform cluster estimation and feature selection by sparse feature weighting simultan...

Miin-Shen Yang, Shazia Parveen · 0 citations
#machine learning Preprint Sep 2026

Riemannian Difference-of-Convex Optimization for K-Means Clustering

This paper replaces the cardinality constraint with a difference-of-convex (DC) penalty and establishes a global error bound to prove that the penalized and constrained formulations share the same global minimizers whenever the penalty parameter exceeds a finite threshold.

Meng Xu, Bo Jiang, Han-Fu Zhang et al. · 0 citations
Open access Aug 2026

Conservative entropy-regularized feature-weighted fuzzy C-means clustering method

This paper proposes the CERFW-FCM (Conservative Entropy-Regularized Feature-Weighted Fuzzy C-Means) method as a conservatively stabilized extension of the classical fuzzy C-means algorithm for fuzzy clustering of multidimensional numerical data with heterogeneous feature informativeness. The study is motivated by the o...

D. Symonov, Y. Symonov, Bohdan Zaika · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.