Skip to content
Preprint

When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification

Aug 2026 · 0 citations · 15 references
Computer Science

TL;DR

Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were properties of the method, and it is shown this practice is unsafe and, for a well-calibrated ensemble, imbalance handling is unnecessary.

Abstract

Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were properties of the method. We show this practice is unsafe. On the public Kaggle credit-card fraud dataset, under a leakage-free nested cross-validation protocol in which the decision threshold is selected on a held-out inner validation fold, a plain Random Forest at the default 0.5 threshold attains F1 = 0.861 +/- 0.021, and threshold tuning yields it no benefit (delta-F1 = -0.002). Read alone, this supports an appealing conclusion: for a well-calibrated ensemble, imbalance handling is unnecessary. We then apply the identical protocol to 45 binary tasks spanning imbalance ratios from 1:1.5 to 1:178 (2,025 model fits, four model families). The conclusion reverses. Random Forest benefits most from threshold tuning across the suite (delta-F1 = +0.101 +/- 0.134), not least, while three other families replicate their fraud-dataset behaviour almost exactly. SMOTE likewise harms the fraud dataset but helps across the suite (mean delta-F1 = +0.076; 138 wins, 39 losses; Wilcoxon p = 2.7e-17). Two further results. Threshold-tuning benefit is non-monotonic in the imbalance ratio: near zero below 1:5, peaking at +0.120 in the 1:15-1:40 band, declining to +0.045 beyond 1:100 - explaining why the fraud dataset, at 1:577, is an unrepresentative place to study the question. And we reject an intuitive heuristic: validation-set calibration error does not predict tuning benefit (expected calibration error r = -0.087; Brier r = +0.137), so calibration diagnostics cannot tell a practitioner whether tuning is worthwhile. We release the protocol, the 45-task harness, and all per-run metrics.

View source

Similar papers

Open access Sep 2026

When Does Cost-Sensitive Weighting Matter? A Classifier-Capacity Anal-ysis for Imbalanced Classification

Class weighting is the most widely used cost-sensitive remedy for class imbalance, and inverse-frequency “balanced” weights are often applied as an unquestioned default. This paper asks whether the choice among competing weighting mechanisms actually matters, and for which classifiers. Seven mechanisms, namely Balanced...

Swati Satpute, Ajit More · 0 citations
Open access Sep 2026

Learning Under Extreme Class Imbalance: A Comparative Study of Algorithmic and Data-Level Solutions

A systematic comparative framework that integrates data-level resampling, cost-sensitive learning, and hybrid approaches to evaluate their performance under varying imbalance ratios and noise levels indicates that hybrid approaches consistently outperform standalone methods, achieving the most stable and balanced perfo...

Tamsir Ariyadi, E. Noche, Nisha Pandey et al. · 0 citations
Review Open access Aug 2026

DAUNT: Ensemble Disagreement as Actionable Uncertainty for Imbalanced Fraud Detection

Detecting a few hundred fraudulent transactions among hundreds of thousands is an extreme class-imbalance problem where one miss can cost a full transaction value. Stacking heterogeneous classifiers is the standard recipe, yet under a leakage-free, precision–recall evaluation, its ranking gain over the best single mode...

Xinyao Liu, Guixiang Zhu · 0 citations
Open access Aug 2026

SMOTE-Stack-XAI: An Explainable Stacked Ensemble Learning Framework Integrating Random Forest, XGBoost, SVM and Deep Neural Networks for Real-Time Credit Card Fraud Detection

SMOTE-Stack-XAI, a stacked ensemble framework that uses the Synthetic Minority Oversampling Technique (SMOTE) to address class imbalance, outperforms the RF, hybrid RF-SVM, and hybrid SVM-LR baselines with an accuracy of 99.3%, precision of 99.2%, recall of 99.4%, and F1-score of 99.3%.

Bhukya Dharma, D. Latha · 0 citations
Open access Sep 2026

Adaptive Weighting–Synthetic Minority Oversampling Technique

A novel oversampling algorithm: the adaptive weighting–synthetic minority oversampling technique (AW-SMOTE), which combines the two perspectives of boundary tightness and local density and provides global sample enhancement support.

Shen Yan, Hai-Feng Guo, Xiao-Ming Su · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.