This work introduces RUBRIC, a generator-agnostic filtering framework that formulates synthetic sample selection as a quality-over-quantity optimization problem and shows that RUBRIC improves F1-macro and recall while maintaining comparable ROC-AUC across several generators.
Abstract
Class imbalance poses a fundamental challenge in risk-sensitive applications such as fraud detection and medical diagnosis, where minority-class samples are scarce yet critical for accurate classification. Existing oversampling methods generate synthetic samples to rebalance class distributions; however, they often produce large numbers of low-quality candidates that distort decision boundaries or introduce artifacts, leading to overfitting and degraded generalization. In this work, we introduce RUBRIC, a generator-agnostic filtering framework that formulates synthetic sample selection as a quality-over-quantity optimization problem. RUBRIC ranks candidates using a realism-utility trade-off: realism is quantified by a learned discriminator that distinguishes real samples from synthetic samples, while utility captures proximity to the decision boundary through a concave margin-based scoring function. We show that, under mild regularity conditions, the proposed filtering strategy monotonically tightens the generalization bound for margin-based classifiers by jointly reducing distribution shift and suppressing near-negative tail contributions. Through extensive experiments on credit-card fraud detection and other imbalanced benchmarks, we demonstrate that RUBRIC improves F1-macro and recall while maintaining comparable ROC-AUC across several generators. We also provide explicit lambda-sensitivity analysis to show how users can recover AUPRC when ranking quality is prioritized.
Noise-Aware Adaptive-Difficulty Oversampling (NADOS) is proposed, which separates the assessment of seed trustworthiness from the allocation of synthesis effort and provides a practical strategy for noisy imbalanced classification.
Ednel Ashraff Misran, Syahid Anuar, Adam Mohd Khairuddin· International Journal of Adv...· 0 citations
Class imbalance is prevalent in real-world datasets. Minority samples are far fewer than majority samples. Traditional classifier design typically assumes balanced data, which causes classifiers to favor the majority class when faced with imbalanced datasets. Thus, there are high misclassification costs for minority cl...
Bounded predictive influence and reliability-guided geometry as complementary mechanisms for imbalanced learning with uncertain labels are supported as complementary mechanisms for imbalanced learning with uncertain labels.
M. Akhtar, J. Akarsh, M. Tanveer et al.· 0 citations
This study introduces a novel approach, named SDCGAN, which oversamples with a Conditional Generative Adversarial Network (CGAN) where the generator is fed with strengthened distribution information extracted from a curated set of minority samples.
Yu-Lin Zhang, Xiao-Zhe Wang, Kai-Wen Xue et al.· International Journal of Dat...· 0 citations
Multi-class imbalanced datasets are ubiquitous in domains like medical diagnostics, fraud detection, and learning performance classification, where minority classes are critical but underrepresented, and class overlap introduces noise and ambiguous boundaries. Prior research has explored oversampling and undersamplin...
Nhiem Ba Nguyen, Sinh Van Nguyen, B. Nguyễn· Vietnam Journal of Computer...· 0 citations
RoBell-RVFL is proposed, a robust and lightweight generalized bell random vector functional link network that redefines how randomized models handle class imbalance and noisy data and achieves adaptive control over sample contributions without sacrificing the closed-form learning efficiency of RVFL networks.
A. Rahaman, A. Quadir, M. Tanveer· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.