Jul 2026· ACM Transactions on Intelligent Systems and Technology· Vol 17, pp. 1-25· 0 citations· 38 references
TL;DR
Overall, the results position MILES as a robust and generalizable alternative to conventional boosting, bagging, and cost-sensitive ensembles for multiclass imbalanced learning.
Abstract
Imbalanced multiclass learning remains challenging due to skewed class distributions, class overlap, and heterogeneous within-class structure. We revisit the Multiclass Imbalance Learning in Ensembles through Selective Sampling (MILES) framework and study two clustering-based variants: MILES \({}^{k}\) , which uses \(k\) -means with an SSE-based heuristic for selecting the number of clusters, and MILES \({}^{FF}\) , which uses FarthestFirst to explore an alternative centroid-based partitioning strategy. Both variants combine clustering-guided selective sampling with resampling to construct diverse and more balanced training subsets, improving representation of difficult classes while preserving local decision structure. We evaluate MILES on eight real-world multiclass datasets spanning different imbalance regimes, overlap levels, and feature complexities. Across these experiments, MILES is consistently competitive in Multiclass Area Under the Curve ( \(MAUC\) ) and achieves its strongest gains in Macro-F1, with statistically significant improvements over several strong ensemble baselines. MILES also achieves strong Micro-F1 on multiple datasets, while its performance on Geometric Mean ( \(G\) -Mean) is competitive but more dataset-dependent. Class-wise analysis further shows that the improvement arises from better recovery of hard classes and systematic reduction of dominant baseline confusion patterns, especially on Page and Satellite, while Hyperspectral highlights a limitation case where high dimensionality and stronger overlap reduce the benefits of centroid-based selective sampling. A glaciology case study on glacier algae prediction further demonstrates the practical utility of the framework. Overall, the results position MILES as a robust and generalizable alternative to conventional boosting, bagging, and cost-sensitive ensembles for multiclass imbalanced learning.
This study proposes complexity-guided cluster-based adaptive oversampling (CCAS), a novel resampling technique for low- to mid-sized imbalanced datasets that outperforms the benchmark methods across metric-classifier combinations.
Abhilash Panda, U. Chattaraj· Sādhanā· 0 citations
Multi-class imbalanced classification remains difficult because minority classes can be poorly recognised even when aggregate performance appears acceptable. Many imbalance-handling methods still rely on a fixed technique, a single selected technique, or one level of adaptation, despite the fact that technique suitabil...
S. Obe, D. Matthias, E. O. Bennett· Journal of Artificial Intell...· 0 citations
Evaluating the performance of Greedy K-center across a variety of metric spaces shows that mapping unlabeled instances into a predictive probability space and weighting the result by entropy often dominates the other options for active learning selection with Greedy K-center.
Data-Informed Centroid Splitting (DICS), a clustering-based framework that constructs a compact and informative set of candidate splits using data-driven priors, significantly reduces the split search space for classification tasks while preserving predictive performance.
Class imbalance remains a critical challenge in supervised learning, often biasing classifiers toward majority classes. While resampling techniques like Synthetic Minority Oversampling Technique (SMOTE) are widely used, the combined effect of data balancing and hyperparameter optimization across diverse datasets is rar...
Necati Vardar, Mehmet Fatih Ören· Sakarya University Journal o...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.