A robust clustering method that estimates cluster centers and covariance matrices using density power divergence measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm is introduced.
Abstract
We introduce a robust clustering method, MK-means DPD, that estimates cluster centers and covariance matrices using density power divergence (DPD) measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm. Since Mahalanobis distance-based K-means lacks a general convergence guarantee, we further introduce a convergent variant, Density-Consistent MK-means DPD (DC-MK-means DPD), which redefines the cluster assignment step in terms of a pointwise DPD loss. We prove a formal theorem establishing that the resulting algorithm converges in a finite number of steps. We also propose two new robust internal evaluation indices, a Median Davies-Bouldin Index and a Trimmed Calinski-Harabasz Index, to ensure that performance comparisons are not themselves distorted by outliers. The efficacy of the proposed methods is demonstrated on simulated data, showing superiority over existing methods, and on two real datasets: Iris data, to identify similar species, and COVID-19 case fatality rate and infection rate data for countries worldwide, examining the resulting clusters'geographic and socio-economic patterns.
This work proposed to use Regularised Multidimensional Scaling using Radial Basis Function (RBF-MDS) for dimension reduction, a multidimensional scaling that mitigates the impact of irrelevant or redundant features, enabling 2D/3D visualisation of clusters for interpretability.
Afsana Akter Setu, J. Singha, Sohana Jahan· Dhaka University Journal of...· 0 citations
For image data, matrix-based clustering methods are gaining popularity because they can directly process two-dimensional (2D) data structures without vectorization. However, most existing approaches rely on squared Frobenius norms in their objective functions, making them sensitive to outliers and noise commonly encoun...
Clustering high-dimensional data is challenging when irrelevant or weakly informative features obscure the underlying cluster structure. To address this issue, we propose two sparse k-means clustering algorithms, S-KM1 and S-KM2, that perform cluster estimation and feature selection by sparse feature weighting simultan...
This paper replaces the cardinality constraint with a difference-of-convex (DC) penalty and establishes a global error bound to prove that the penalized and constrained formulations share the same global minimizers whenever the penalty parameter exceeds a finite threshold.
Meng Xu, Bo Jiang, Han-Fu Zhang et al.· 0 citations
This paper proposes the CERFW-FCM (Conservative Entropy-Regularized Feature-Weighted Fuzzy C-Means) method as a conservatively stabilized extension of the classical fuzzy C-means algorithm for fuzzy clustering of multidimensional numerical data with heterogeneous feature informativeness. The study is motivated by the o...
D. Symonov, Y. Symonov, Bohdan Zaika· International Scientific Tec...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.