Skip to content
Open access

MILES++: A Generalizable Clustering-Based Ensemble Framework for Multiclass Imbalanced Learning

Jul 2026 · ACM Transactions on Intelligent Systems and Technology · Vol 17, pp. 1-25 · 0 citations · 38 references

TL;DR

Overall, the results position MILES as a robust and generalizable alternative to conventional boosting, bagging, and cost-sensitive ensembles for multiclass imbalanced learning.

Abstract

Imbalanced multiclass learning remains challenging due to skewed class distributions, class overlap, and heterogeneous within-class structure. We revisit the Multiclass Imbalance Learning in Ensembles through Selective Sampling (MILES) framework and study two clustering-based variants: MILES \({}^{k}\) , which uses \(k\) -means with an SSE-based heuristic for selecting the number of clusters, and MILES \({}^{FF}\) , which uses FarthestFirst to explore an alternative centroid-based partitioning strategy. Both variants combine clustering-guided selective sampling with resampling to construct diverse and more balanced training subsets, improving representation of difficult classes while preserving local decision structure. We evaluate MILES on eight real-world multiclass datasets spanning different imbalance regimes, overlap levels, and feature complexities. Across these experiments, MILES is consistently competitive in Multiclass Area Under the Curve ( \(MAUC\) ) and achieves its strongest gains in Macro-F1, with statistically significant improvements over several strong ensemble baselines. MILES also achieves strong Micro-F1 on multiple datasets, while its performance on Geometric Mean ( \(G\) -Mean) is competitive but more dataset-dependent. Class-wise analysis further shows that the improvement arises from better recovery of hard classes and systematic reduction of dominant baseline confusion patterns, especially on Page and Satellite, while Hyperspectral highlights a limitation case where high dimensionality and stronger overlap reduce the benefits of centroid-based selective sampling. A glaciology case study on glacier algae prediction further demonstrates the practical utility of the framework. Overall, the results position MILES as a robust and generalizable alternative to conventional boosting, bagging, and cost-sensitive ensembles for multiclass imbalanced learning.

Read PDF

Similar papers

Open access 2026

Adaptive Ensemble Weighting with Local Competence for Multi-Class Imbalanced Classification

Multi-class imbalanced classification remains difficult because minority classes can be poorly recognised even when aggregate performance appears acceptable. Many imbalance-handling methods still rely on a fixed technique, a single selected technique, or one level of adaptation, despite the fact that technique suitabil...

S. Obe, D. Matthias, E. O. Bennett · 0 citations
Preprint Aug 2026

Diversity-Based Active Learning: An Evaluation of Metric Spaces for Active Learning Selection

Evaluating the performance of Greedy K-center across a variety of metric spaces shows that mapping unlabeled instances into a predictive probability space and weighting the result by entropy often dominates the other options for active learning selection with Greedy K-center.

Siddharth Chilamkur, D. Hochbaum · 0 citations
Preprint Aug 2026

DICS: Data-Informed Centroid Splitting for Decision Tree Classifiers

Data-Informed Centroid Splitting (DICS), a clustering-based framework that constructs a compact and informative set of candidate splits using data-driven priors, significantly reduces the split search space for classification tasks while preserving predictive performance.

Saifur Rahman Mazumder, Feng Yu · 0 citations
Open access Sep 2026

A Comparative Performance Evaluation of Classification Algorithms on Imbalanced Datasets

Class imbalance remains a critical challenge in supervised learning, often biasing classifiers toward majority classes. While resampling techniques like Synthetic Minority Oversampling Technique (SMOTE) are widely used, the combined effect of data balancing and hyperparameter optimization across diverse datasets is rar...

Necati Vardar, Mehmet Fatih Ören · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.