Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were properties of the method, and it is shown this practice is unsafe and, for a well-calibrated ensemble, imbalance handling is unnecessary.
Abstract
Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were properties of the method. We show this practice is unsafe. On the public Kaggle credit-card fraud dataset, under a leakage-free nested cross-validation protocol in which the decision threshold is selected on a held-out inner validation fold, a plain Random Forest at the default 0.5 threshold attains F1 = 0.861 +/- 0.021, and threshold tuning yields it no benefit (delta-F1 = -0.002). Read alone, this supports an appealing conclusion: for a well-calibrated ensemble, imbalance handling is unnecessary. We then apply the identical protocol to 45 binary tasks spanning imbalance ratios from 1:1.5 to 1:178 (2,025 model fits, four model families). The conclusion reverses. Random Forest benefits most from threshold tuning across the suite (delta-F1 = +0.101 +/- 0.134), not least, while three other families replicate their fraud-dataset behaviour almost exactly. SMOTE likewise harms the fraud dataset but helps across the suite (mean delta-F1 = +0.076; 138 wins, 39 losses; Wilcoxon p = 2.7e-17). Two further results. Threshold-tuning benefit is non-monotonic in the imbalance ratio: near zero below 1:5, peaking at +0.120 in the 1:15-1:40 band, declining to +0.045 beyond 1:100 - explaining why the fraud dataset, at 1:577, is an unrepresentative place to study the question. And we reject an intuitive heuristic: validation-set calibration error does not predict tuning benefit (expected calibration error r = -0.087; Brier r = +0.137), so calibration diagnostics cannot tell a practitioner whether tuning is worthwhile. We release the protocol, the 45-task harness, and all per-run metrics.
Class weighting is the most widely used cost-sensitive remedy for class imbalance, and inverse-frequency “balanced” weights are often applied as an unquestioned default. This paper asks whether the choice among competing weighting mechanisms actually matters, and for which classifiers. Seven mechanisms, namely Balanced...
Swati Satpute, Ajit More· International Research Journ...· 0 citations
A systematic comparative framework that integrates data-level resampling, cost-sensitive learning, and hybrid approaches to evaluate their performance under varying imbalance ratios and noise levels indicates that hybrid approaches consistently outperform standalone methods, achieving the most stable and balanced perfo...
Tamsir Ariyadi, E. Noche, Nisha Pandey et al.· Journal of Data Science· 0 citations
SoftMCC is a calibration-sensitive MCC-family selector with bounded stability and utility evidence, coupling an MCC-specific calibrated identity with a tie-aware, shared-pool selection protocol.
Detecting a few hundred fraudulent transactions among hundreds of thousands is an extreme class-imbalance problem where one miss can cost a full transaction value. Stacking heterogeneous classifiers is the standard recipe, yet under a leakage-free, precision–recall evaluation, its ranking gain over the best single mode...
Xinyao Liu, Guixiang Zhu· Machine Learning and Knowled...· 0 citations
SMOTE-Stack-XAI, a stacked ensemble framework that uses the Synthetic Minority Oversampling Technique (SMOTE) to address class imbalance, outperforms the RF, hybrid RF-SVM, and hybrid SVM-LR baselines with an accuracy of 99.3%, precision of 99.2%, recall of 99.4%, and F1-score of 99.3%.
Bhukya Dharma, D. Latha· International journal of com...· 0 citations
A novel oversampling algorithm: the adaptive weighting–synthetic minority oversampling technique (AW-SMOTE), which combines the two perspectives of boundary tightness and local density and provides global sample enhancement support.