Aug 2026· Humanities and Social Sciences Communications· 0 citations
TL;DR
A structured overview of classical and advanced clustering approaches, including hierarchical, partition-based, density-based, density-based, model-based, subspace, grid-based, and search-based metaheuristic techniques are provided.
Abstract
Clustering is a core technique in unsupervised learning that organizes unlabeled data into meaningful groups based on similarity. It has wide applications in domains such as bioinformatics, pattern recognition, social network analysis, computer vision, and artificial intelligence. Owing to the diversity and complexity of real-world datasets, numerous clustering paradigms have been developed, each with specific advantages and limitations. This survey provides a structured overview of classical and advanced clustering approaches, including hierarchical, partition-based, density-based, model-based, subspace, grid-based, and search-based metaheuristic techniques. We further examine commonly used similarity measures and validation metrics, including internal and external evaluation criteria, to highlight their role in assessing clustering quality. A comparative taxonomy is presented to clarify algorithmic characteristics, scalability, robustness, and parameter sensitivity under varying data conditions. Despite significant progress, challenges remain in handling noisy and high-dimensional data, determining the optimal number of clusters, and ensuring computational efficiency. Emerging directions such as hybrid frameworks, self-supervised learning, and multi-view clustering offer promising avenues for developing more adaptive and scalable clustering solutions.
This paper presents a comparative analysis of optimization techniques for the minimum sum-of-squares clustering (MSSC) problem—widely known in applied research as the K-means clustering problem—in the context of big data. K-means is the most widely used algorithmic framework for solving this problem, but MSSC methods c...
Ravil Mussabayev, R. Mussabayev· Symmetry· 0 citations
Community detection in networks is a crucial task across diverse fields. While modularity maximization is a widely used approach, it suffers from limitations such as resolution limits and sensitivity to noise. Modularity density, an alternative measure, addresses some of these issues by focusing on minimizing out-of-cl...
S. Gamal· IISE Annual Conference &...· 0 citations
In unsupervised learning, Clustering is a core method used to determine unseen arrangements and structures within datasets by grouping similar instances together. Among the many clustering algorithms, Among clustering techniques, K-Means continues to be one of the most popular owing to its ease of implementation, fast...
I. Khan, H. Daud, Rajalingam Sokkalingam et al.· International Conference on...· 0 citations
The rapid growth of high-dimensional and heterogeneous datasets has intensified the demand for
scalable and robust optimization approaches in data analytics. Traditional deterministic
algorithms frequently suffer from premature convergence, initialization sensitivity, and reduced
performance in nonlinear or noisy en...
Biralatei Fawei· International Journal of Com...· 0 citations
This work proposed to use Regularised Multidimensional Scaling using Radial Basis Function (RBF-MDS) for dimension reduction, a multidimensional scaling that mitigates the impact of irrelevant or redundant features, enabling 2D/3D visualisation of clusters for interpretability.
Afsana Akter Setu, J. Singha, Sohana Jahan· Dhaka University Journal of...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.